Did the agent do its job? (task resolution & correctness)
Was it fast and responsive? (latency & reliability)
Does it sound natural? (voice quality — voice agents)
See the Statistical metrics page for the full set of acoustic and prosody checks (background noise, artifacts, vocal fry, and more).
How did the customer feel? (sentiment & experience)
Did it follow the rules? (compliance & scripting)
Did it do the right things behind the scenes? (tools & traces)
Was the audio transcribed accurately? (STT accuracy)
What did it cost? (usage)
Once you know which metrics you want, add them to a run. For anything custom, see Write judge prompts and Configure metrics.