Skip to main content
Open Insights from a human review project to examine triage categories, reviewer activity, and agreement. Use the Annotations table when you need individual human answers, notes, and automated outputs behind a summary.
Choose a date range, conversation source, and review scope before comparing results. Conversation time filters by when calls occurred; Label time filters by when they were labeled. Comparison tiles support up to 30 days. Check each insight’s counts and coverage, since conversations, annotations, and comparable pairs are different totals.

Understanding patterns in conversations

Label triage counts conversations in each category reviewers applied. Each conversation is counted once per active category. In a project that allows multiple categories, one conversation can therefore contribute to several category totals. Start with a category you want to investigate and follow its linked count to the matching conversations. Read several examples to check whether reviewers are applying the category consistently. Project scope and reviewer activity show who has contributed and which conversations are covered. Assignee summaries, conversations labeled over time, and conversation-level progress help distinguish a change in review activity from a change in the conversations themselves. Metric annotation counts do not represent all triage work or completed reviews.

Understanding whether reviewers agree

Human-to-human agreement compares reviewers’ answers to the same metric on the same conversation. It requires overlapping reviews; reviewers who worked on different conversations do not form a comparison pair. The Reviewer agreement view shows agreement alongside the number of comparable reviewer pairs and overlapping conversations. Individual projects are useful for this analysis because they intentionally collect independent reviews of the same items. A shared queue with little overlapping review may not have enough evidence to measure reviewer consistency. Low agreement is a reason to revisit examples together. Reviewers may be interpreting an ambiguous criterion differently, missing context, or applying different thresholds. Clarify those differences before changing the automated metric to match one person’s answer. Always read the pair count alongside the rate: a high percentage based on very few overlapping reviews provides limited evidence.

Understanding how well metrics match human judgment

Human-to-machine agreement compares collected human answers with automated metric outputs. The metric agreement summary distinguishes agreements, disagreements, conversations that have not been reviewed, and labels that cannot be compared. For binary metrics, read several signals together: For example, suppose humans mark 95 of 100 conversations as successful. A metric that marks all 100 as successful has 95% agreement but catches none of the five failures. Its agreement is no better than the constant-answer baseline. Check catch rate and the disagreements alongside the overall agreement rate.

Separating missed problems from false alarms

Where available, Error direction shows a heatmap of human and machine answer combinations. It separates cases where the metric misses a problem identified by a human from cases where it flags a problem the human did not find. These patterns suggest different revisions: missed problems may need clearer failure criteria, while false alarms may need better exceptions or boundaries. Directional labels depend on the metric’s configured success condition. When that condition is unavailable, use the human and machine answer combinations without assuming that Yes always means success.

Checking balance, coverage, and priorities

Additional diagnostic views help put the agreement score in context:
  • Human class balance shows how binary and categorical human answers are distributed. Numerical metrics are excluded from this view.
  • Sample power shows human labels, machine outputs, comparable pairs, and gaps in coverage. Not reviewed and Not comparable represent missing comparison evidence, not additional disagreements.
  • Where κ sits helps interpret chance-corrected agreement. Its interpretation bands are guidance; class imbalance can affect κ, and some answer distributions make it undefined.
  • Fix priority ranks binary metrics using disagreement, evidence volume, and error direction to help choose what to investigate first.
Binary comparison diagnostics mark fewer than 100 comparable pairs as a limited sample. Priority and κ summaries require sufficient evidence, and κ must be computable where used. This threshold does not guarantee that a sample represents all your traffic. If a tile reports that only some metrics were returned, its ranking covers those metrics only.

Investigating labels and disagreements

An unexpected result may reflect a metric problem, an unclear review criterion, or a human labeling mistake. Suspect labels, marked as beta, surfaces binary labels worth a second look: conflicting reviewers, lone positive answers contradicted by a peer and the machine, and human answers that disagree with the latest automated output. These are investigation candidates, not proof that a human answer is wrong. Use the reviewer notes and machine explanation to understand the disagreement, then read or listen to the conversation. Linked counts and evidence links take you from summary insights to the relevant cases.

Working with the underlying annotations

Open the Annotations table from Insights to inspect individual reviews. Filter by source, review progress, reviewer, triage labels, metric comparisons, or time, and choose the metric columns relevant to your investigation. For example, focus on completed reviews where a resolution metric disagrees with humans, then inspect the shared failure pattern. Expand a conversation to examine reviewer rows, or open its details to compare machine outputs with individual human answers and notes. The details also provide conversation context and, for supported composite evaluations, criterion-level comparisons. Return to the assignment when a label needs another review. Filters and displayed metric columns are reflected in the page URL, so you can share the investigation with a collaborator who has access to the project.

Re-evaluating metrics after a change

Before rerunning a metric, review the disagreements and decide whether the metric, the human labels, or the review instructions need to change. Rerunning produces new automated results to compare with the collected human answers. After updating a metric, Rerun Metrics lets you evaluate it again on project conversations. Select the metrics, then choose all linked conversations, only labeled conversations, or conversations missing the selected metrics. Set the time window and whether it follows label date or conversation creation date, and review the scope before launching. Return to Insights after the results arrive to compare the new outputs with the human labels. See Improving metrics with human review for the full iteration workflow.

Sharing review findings

Use Download CSV from the Annotations table when you need the underlying evidence for analysis or discussion outside Coval. The export follows the current filters and selected columns, can include matching rows beyond those already loaded, and lets you set conversation dates and a maximum row count. Conversation and reviewer rows stay together, so the export may stop below that limit. For a report in Coval, Generate report creates and opens a report based on the project’s linked runs. It requires linked runs and is separate from exporting a filtered annotations view. When sharing a result, include the date range, conversation source, review scope, and sample size. Those details help collaborators understand what the findings support and whether a follow-up sample is needed.