Skip to main content
The Metric Library is organized by how each metric works. This page flips that around: start from what you want to measure and jump straight to the metrics that measure it.
Just starting out? A solid baseline for most agents is Latency, End Reason, and one custom Binary LLM Judge for your specific success criteria. Add more as you learn where your agent struggles.

Did the agent do its job? (task resolution & correctness)

Was it fast and responsive? (latency & reliability)

Does it sound natural? (voice quality — voice agents)

See the Statistical metrics page for the full set of acoustic and prosody checks (background noise, artifacts, vocal fry, and more).

How did the customer feel? (sentiment & experience)

Did it follow the rules? (compliance & scripting)

Did it do the right things behind the scenes? (tools & traces)

Was the audio transcribed accurately? (STT accuracy)

What did it cost? (usage)


Once you know which metrics you want, add them to a run. For anything custom, see Write judge prompts and Configure metrics.