Both modes use the question, categories, or score range you configure. The difference is how they gather the evidence needed to answer.
Standard mode and Coval’s Judge
Standard mode is the recommended choice for most evaluations. It works well when the answer can be determined directly from the conversation, optionally with trace context. Coval’s Judge is Coval’s proprietary evaluator for Standard text LLM Judge metrics. It is built and optimized specifically for agent evaluations and is the default for new Binary, Categorical, and Numerical text judges. Coval’s Judge is designed to provide more robust and consistent evaluation performance across varied conversations than relying on a general-purpose default. Coval manages and continuously improves the evaluator, so its underlying implementation may evolve without requiring changes to your metrics. Each evaluation returns one result and one explanation. Where available, a High, Medium, or Low confidence indicator helps identify results that may benefit from closer review.Coval’s Judge applies to text-based Standard LLM Judges. Audio judges, Composite Evaluation, and Agentic mode use their own evaluation paths.
Agentic mode
Agentic mode is designed for questions that cannot be answered reliably from a single pass over the transcript. Instead of receiving all context at once, the judge determines what evidence it needs, investigates the relevant parts of the evaluation, and continues until it has enough information to answer—or determines that the evidence is insufficient. The investigation is read-only and limited to the current evaluation. It cannot call your agent’s tools, contact external systems, or change any data.Evidence Agentic mode can investigate
Agentic mode can reason across several kinds of evidence:- Conversation content — Read the complete conversation or locate specific turns, phrases, and patterns.
- Agent execution — Inspect traces, timing, statuses, and selected attributes describing what happened behind the scenes.
- Structured interaction data — Examine agent actions and their results, event sequences, transcript records, and speaking activity.
- Run and scenario context — Understand the test-case description and expected behaviors while keeping expectations separate from evidence of what actually happened.
- Focused analysis — Apply a narrower structured assessment to a relevant piece of evidence when the top-level question requires classification or scoring.
UNKNOWN instead of making an unsupported judgment.
When to use Agentic mode
Agentic mode is especially useful for:- Verifying actions, not claims — Confirm that an agent actually completed an action rather than merely saying it did.
- Checking order of operations — Determine whether identity verification, consent, or another required step happened before a sensitive action.
- Finding sparse evidence — Locate one decisive event in a long conversation without depending on a fixed context summary.
- Evaluating structured behavior — Reason about action inputs, outputs, retries, failures, and event sequences.
- Computing evidence-based outcomes — Count or categorize events whose meaning requires interpretation before aggregation.
- Diagnosing hidden failures — Connect what the customer heard with what occurred inside the agent’s execution.
- “Did the assistant verify the customer’s identity before sharing account-specific information?”
- “Did the assistant successfully submit the refund, rather than only promise to submit it?”
- “How many times did the customer need to repeat the same request before the assistant responded appropriately?”
- “Classify the outcome as Resolved, Escalated, Abandoned, or Unresolved using both the conversation and available execution evidence.”
When Standard mode is better
Use Standard mode when the answer is directly visible in the transcript, such as tone, acknowledgement, disclosure wording, or overall resolution. Standard mode is generally faster and more economical. Agentic mode may require several investigation steps, so reserve it for evaluations where adaptive evidence gathering materially improves the decision. For exact checks that require no interpretation—such as matching a required phrase or counting a known trace event—prefer a deterministic, regex, SQL, or trace metric.Availability and limitations
- Agentic mode supports text-based Binary, Categorical, and Numerical LLM Judges.
- It does not evaluate raw audio.
- Trace-based investigations require your agent to send OpenTelemetry data to Coval.
- Results depend on the evidence available for that evaluation.
- Missing or unavailable evidence may produce
UNKNOWN. - Agentic mode does not use Coval’s Judge; it has its own model selection.
Configure a judge mode
When creating or editing a text LLM Judge:- Choose Standard or Agentic under Mode.
- Choose a Binary, Categorical, or Numerical output.
- Write the evaluation question and define the result criteria.
- For Agentic mode, select the evidence categories the judge may investigate.
- Use Test Metric to evaluate the configuration against representative conversations before adding it to a run.