Choose a date range, conversation source, and review scope before comparing results. Conversation time filters by when calls occurred; Label time filters by when they were labeled. Comparison tiles support up to 30 days. Check each insight’s counts and coverage, since conversations, annotations, and comparable pairs are different totals.
Understanding patterns in conversations
Label triage counts conversations in each category reviewers applied. Each conversation is counted once per active category. In a project that allows multiple categories, one conversation can therefore contribute to several category totals. Start with a category you want to investigate and follow its linked count to the matching conversations. Read several examples to check whether reviewers are applying the category consistently. Project scope and reviewer activity show who has contributed and which conversations are covered. Assignee summaries, conversations labeled over time, and conversation-level progress help distinguish a change in review activity from a change in the conversations themselves. Metric annotation counts do not represent all triage work or completed reviews.Understanding whether reviewers agree
Human-to-human agreement compares reviewers’ answers to the same metric on the same conversation. It requires overlapping reviews; reviewers who worked on different conversations do not form a comparison pair. The Reviewer agreement view shows agreement alongside the number of comparable reviewer pairs and overlapping conversations. Individual projects are useful for this analysis because they intentionally collect independent reviews of the same items. A shared queue with little overlapping review may not have enough evidence to measure reviewer consistency. Low agreement is a reason to revisit examples together. Reviewers may be interpreting an ambiguous criterion differently, missing context, or applying different thresholds. Clarify those differences before changing the automated metric to match one person’s answer. Always read the pair count alongside the rate: a high percentage based on very few overlapping reviews provides limited evidence.Understanding how well metrics match human judgment
Human-to-machine agreement compares collected human answers with automated metric outputs. The metric agreement summary distinguishes agreements, disagreements, conversations that have not been reviewed, and labels that cannot be compared. For binary metrics, read several signals together:
For example, suppose humans mark 95 of 100 conversations as successful. A metric that marks all 100 as successful has 95% agreement but catches none of the five failures. Its agreement is no better than the constant-answer baseline. Check catch rate and the disagreements alongside the overall agreement rate.
Separating missed problems from false alarms
Where available, Error direction shows a heatmap of human and machine answer combinations. It separates cases where the metric misses a problem identified by a human from cases where it flags a problem the human did not find. These patterns suggest different revisions: missed problems may need clearer failure criteria, while false alarms may need better exceptions or boundaries. Directional labels depend on the metric’s configured success condition. When that condition is unavailable, use the human and machine answer combinations without assuming that Yes always means success.Checking balance, coverage, and priorities
Additional diagnostic views help put the agreement score in context:- Human class balance shows how binary and categorical human answers are distributed. Numerical metrics are excluded from this view.
- Sample power shows human labels, machine outputs, comparable pairs, and gaps in coverage. Not reviewed and Not comparable represent missing comparison evidence, not additional disagreements.
- Where κ sits helps interpret chance-corrected agreement. Its interpretation bands are guidance; class imbalance can affect κ, and some answer distributions make it undefined.
- Fix priority ranks binary metrics using disagreement, evidence volume, and error direction to help choose what to investigate first.
Binary comparison diagnostics mark fewer than 100 comparable pairs as a limited sample. Priority and κ summaries require sufficient evidence, and κ must be computable where used. This threshold does not guarantee that a sample represents all your traffic. If a tile reports that only some metrics were returned, its ranking covers those metrics only.