Skip to main content
Use human review to check a metric against labeled conversations before and after changing it. Start with disagreements, determine whether the metric or the human answer needs correction, and test a specific revision on the same examples. This guide assumes you have collected human labels. For setup, see Manage Projects. For labeling instructions, see Label Conversations.
Human review is supported for a subset of metric types. See the metric types reviewers can label.

Establishing a comparison sample

Open your project’s Insights and choose the date range, conversation source, and review scope. Record those settings and the number of comparable pairs so you can use the same sample after revising the metric. Check whether reviewers agree with each other before using their answers as the reference. Overlapping reviews in an Individual project can reveal ambiguous criteria or inconsistent labeling. Resolve those cases with the reviewers before changing the metric to match them. Also check the distribution of human answers. High agreement on a sample containing almost no failures does not show how well a metric detects failures. For binary metrics, read agreement alongside catch rate, class balance, and the constant-answer baseline. See Insights and Annotations for definitions and sample limits.

Diagnosing a disagreement

Follow a linked disagreement into the Annotations table. Compare the human answer and note with the automated answer and explanation, then read or listen to the conversation. Decide which kind of correction the evidence supports: Look for a pattern across several cases. For example, a resolution metric may accept an explanation of the cancellation policy as evidence that a cancellation was completed. If the intended criterion requires confirmation of cancellation, state that requirement explicitly in the metric.

Revising and testing the metric

Make a focused change that addresses the pattern you found. Open the metric, edit its prompt or relevant configuration, and use Test Metric to inspect its behavior on example conversations. See Write judge prompts for guidance on defining criteria and exceptions. Check both kinds of cases: ones the metric previously got wrong and ones it previously got right. A revision that fixes missed failures may also introduce false alarms. Keep a record of the previous results or export the annotations before evaluating the changed metric. Metric Details

Comparing results after the change

In project Insights, use Rerun Metrics to evaluate the revised metric on project conversations. Select the metric and choose Only labeled conversations when you want to compare its outputs with existing human answers. Set the time window and date basis to match your comparison sample, then check the scope before launching. When results arrive, return to Insights using the same filters. Check whether the original disagreements were resolved, whether new disagreements appeared, and whether the number of comparable pairs changed. A change in sample coverage can change the agreement rate even when the metric itself has not improved. Review remaining disagreements individually. Then evaluate a fresh labeled sample as well as the examples used to revise the metric, so the decision is not based only on cases you already inspected. Repeat this check when changes to your agent, evaluation criteria, or conversation mix affect what the metric needs to judge.