Skip to main content
Sep 16
2026
3 updates
Highlights
  • Sofia investigations on Home
  • Reports and scheduling tools in the Coval connector
  • Sofia can configure alerts, personas, and dashboards
Sofia investigations on Home

Sofia now runs scheduled investigations over your workspace and reports its findings in a Recent Investigations section on Home, with evidence for each one. Acknowledge, resolve, or dismiss findings, flag the ones Sofia got wrong, and review the full history on the Investigations page.

Reports and scheduling tools in the Coval connector

The Coval connector now exposes tools to list, read, and create reports, and to list run templates and create, update, and inspect recurring evaluation schedules. New schedules stay disabled until you enable them. Test case tools also accept an input type and scripted turns, including keypad presses.

Sofia can configure alerts, personas, and dashboards

In the app, Sofia can now create and update alerts, configure a persona including its voice settings, build a complete dashboard in one action, update an existing run, and stage a batch of test case changes in one reply. Every change still waits for your confirmation.

Sep 14
2026
3 updates
Highlights
  • Voice designer, from prompt to voice
  • Sofia Notebooks in beta
  • Built-in baseline agent
Voice designer (prompt to voice)

Create a new voice by describing it in a prompt, then use it for personas in your simulations.

Sofia Notebooks (beta)

Save, schedule, and share Sofia analyses as notebooks, so a recurring investigation and its findings stay in one place. In beta.

Built-in baseline agent

Every organization now comes with a ready-made sample agent, so you can run a baseline evaluation without building an agent first.

Sep 7
2026
10 updates
Highlights
  • Load testing up to 1,000 concurrent calls
  • Median (P50) aggregation on dashboards
  • Manager insights in Human Review
Load testing up to 1,000 concurrent calls

Load testing now runs up to 1,000 concurrent simulated calls, so you can see how an agent holds up under production-scale traffic.

Median (P50) aggregation

Dashboard metrics can now aggregate by median (P50), alongside the existing aggregations.

Composite scores shown as percentages

Composite evaluation scores now display as percentages in runs tables, result badges, and Human Review labels.

Test SQL metrics against recent simulations

The SQL metric editor now shows results from recent simulations so you can check a query before saving, and surfaces clearer validation errors.

Metadata suggestions in filters

Conversation filters now suggest metadata keys and values as you type.

Open a simulation from a report row

Click any row in a report to open the underlying simulation.

Toggle chart series from the legend

While editing a dashboard, you can show or hide a chart’s series by clicking its legend.

Faster simulation and conversation lists

Simulation and conversation lists load faster and no longer pull full transcripts just to render the list.

Sofia can set up and troubleshoot alerts

Sofia can now help you create, edit, and troubleshoot alerts.

Manager insights

Human Review insights give managers a fuller picture, with reviewer agreement, per-labeler leaderboards and activity over time, suspect-label detection, and outlier views.

Aug 31
2026
1 update
Highlights
  • Workspaces, in beta
Workspaces (beta)

Split your organization into separate workspaces so agents, test sets, personas, metrics, and tags can be scoped to a team. In beta.

Aug 24
2026
1 update
Highlights
  • More reliable conversations listing
More reliable conversations listing

The public v1 conversations listing no longer times out on large default result sets.

Aug 17
2026
1 update
Highlights
  • Scheduled investigations on Home
Scheduled investigations on Home

Sofia’s scheduled investigations now surface findings on your Home page, where you can acknowledge, dismiss, or flag them.

Aug 10
2026
4 updates
Highlights
  • Web search for OpenAI Realtime agents
  • Set run metadata from the launch page
  • Template variables in more fields
OpenAI Realtime agents can use web search

OpenAI Realtime agents can now use web search during simulations, available in the app and through the API.

Set run metadata from the launch page

You can now set run metadata directly from the launch page when you start a run, not just through the API.

Template variables in more fields

Template variables now resolve in more places, including run metadata and test case fields, so you can drive more per-scenario values into a simulation.

Order the metrics reviewers see

You can now set the order in which metrics appear to reviewers in Human Review.

Aug 3
2026
7 updates
Highlights
  • Sofia in Claude
  • Sofia home briefs
  • New simulation voice languages
Sofia in Claude

Sofia is now available through the Coval connector in Claude, so you can reach your Coval data and runs from the assistant you already work in.

Sofia home briefs

Your home page now opens with a Sofia-generated brief of recent activity and findings across your runs.

New simulation voice languages

New voices add Japanese, Mandarin, Cantonese, and Bengali for simulations, and voice previews now show each voice’s language and accent.

Create LiveKit and Pipecat agents via API and CLI

The API and CLI can now create LiveKit and Pipecat agents directly, in SDK 0.6.0.

Merge reports through the API and CLI

You can now merge reports together through the API and CLI, matching what the app already offers.

Outlier scores highlighted

Outlier metric scores are now highlighted in runs and reports, so unusual results stand out at a glance.

Claim-based review workflow

Reviewers can claim a conversation when they start, which locks it and its labels so two people do not review the same one.

Jul 31
2026
1 update
Highlights
  • Coval connector for ChatGPT and Claude
Coval connector for ChatGPT and Claude

Install Coval from the ChatGPT plugin listing or Claude connector directory to inspect evaluation data, manage supported resources, launch supported runs, and consult Sofia from the assistant where you already work.

Jul 27
2026
3 updates
Highlights
  • Org-wide trace search
  • Reviewer agreement and insights
  • Playback markers in review
Org-wide trace search

Search traces across your whole organization from the app, and through the API and CLI, with configurable monitoring reports.

Reviewer agreement and insights

Human Review insights add reviewer agreement, a reviewer leaderboard, suspect-label detection, and outlier views, so you can see how reviewers compare and spot labels worth a second look.

Playback markers in review

Call playback now shows speech-activity markers, so you can see who was speaking and when while reviewing a conversation.

Jul 20
2026
5 updates
Highlights
  • Rerun metrics in bulk through the API
  • Review project membership controls
  • Review complete action on assignments
Bulk rerun metrics through the API

You can now rerun metrics across many simulations at once through the v1 API, with a new rerun-metrics endpoint, so you do not have to trigger them one at a time.

Review project membership controls

Human Review projects now support a required manager, membership history, and a membership audit log, so you can see who has access to a project and when it changed.

Review complete action on assignments

Reviewers can now mark an assignment review complete directly, making it clearer when a set of labels is finished.

Streaming responses in Sofia

Sofia, the in-app agent, now streams its responses as they generate, with speed controls, so answers start appearing right away.

Run names in the runs list

Test run names now display correctly in the runs list.

Jul 13
2026
5 updates
Highlights
  • Blind review and per-metric blind labeling
  • Per-labeler insights
  • Metric configuration validation
Blind review and per-metric blind labeling

Reviewers can now label conversations blind, without seeing existing or AI-generated labels, so human judgments stay independent. Blind mode can be set per metric.

Per-labeler insights

Human Review project insights add a per-labeler assignee tab and a per-labeler-per-day activity plot, so you can see who labeled what and how labeling is progressing.

Metric configuration validation

Metric configurations are now validated as you set them up, so misconfigured metrics are caught early instead of failing at run time.

Simulation scope persistence

Reports now remember your simulation-level scope, so a report keeps the exact set of simulations you selected.

Sofia runs Human Review actions and summarizes outcomes

Sofia, the in-app agent, can now take Human Review actions and summarize behavioral outcomes from your runs.

Jul 6
2026
6 updates
Highlights
  • Restore deleted test sets, personas, and agents
  • Sofia builds charts from your results
Restore deleted test sets, personas, and agents

Deleted test sets, personas, and agents can now be restored from a Recently Deleted view, so an accidental delete is no longer permanent.


Deletes are now soft deletes. Open Recently Deleted to restore a test set, persona, or agent. Metric restore is following shortly.
Human Review report filtering

Reports built from a Human Review project can now filter to just the human-labeled simulations, so the report reflects only the conversations your team reviewed.

Multiple conditions on metric rules

Default metric rules can now match more than one metadata condition at once, combined with AND, so you can target metrics more precisely.

Choose a type when creating an agent

Creating a new agent now opens a type picker for voice, text, and advanced setups instead of defaulting straight to voice.

Sofia builds charts from your results

Sofia, the in-app agent, can now generate charts and visualizations in its answers when you ask about your results.

Add tags from inside a metric

You can now add tags to a metric directly from the metric page, without going back to the list.

Jun 29
2026
4 updates
Highlights
  • Caller phone numbers in simulation results
  • Scatter plot outlier and legend controls
Caller phone numbers in simulation results

Simulation results now include the caller’s source and destination phone numbers, in the app and through the API.

Scatter plot outlier and legend controls

Scatter plots in Multi-Run Analysis add an exclude-outliers toggle that dims Tukey IQR outliers and drops them from the trend line, plus a legend toggle and a point-count caption for shared reports.

Accurate test set edit timestamp

A test set’s last-edited time now updates correctly after edits made through the API.

Deleted items excluded from reports

Deleted simulations and runs no longer appear in report rows.

Jun 22
2026
3 updates
Highlights
  • Sofia, your in-app agent, is now in beta
  • Template variables in chat agent setup
  • Clearer metric errors and smarter skipping
Sofia, your in-app agent (beta)

Sofia is an in-app agent that helps you set up tests, create metrics, and dig into results through chat, without leaving Coval.


In beta and rolling out to self-serve accounts. Ask Sofia to build a test set, configure a metric, or explain a run, and rate each reply with a thumbs up or down. See the Sofia overview for current capabilities and usage guidance.
Template variables in chat agent setup

Chat agents now resolve test case and agent template variables in custom headers, request payloads, and initialization payloads, so you can drive per-scenario values into your endpoint.

Clearer metric errors and smarter skipping

Metrics that do not apply to a call now skip with a clear reason instead of a generic failure, and LLM judge metrics retry automatically on truncated or off-list responses, so you see fewer spurious failures.

Jun 15
2026
7 updates
Highlights
  • Metric versioning with full change history
  • Standardized LLM-judge metric templates
  • AI-generated run summaries (beta)
  • Project-scoped Human Review pages
Metric versioning

Metrics now keep a version history, so you can see how a metric changed over time and review prior versions.


Each metric keeps a record of its prior states, newest first. You can review the history in the app or pull it through the v1 API.View in docs ↗
Standardized metric templates

A library of ready-made LLM judge metric templates is now available when you create a metric, and the metric gallery is reorganized into clearer categories.


Pick a standardized judge prompt as a starting point instead of writing one from scratch.
Run summaries (beta)

Runs now include an AI-generated summary of what happened, available in beta. Mark a summary helpful or not to help it improve.


The summary renders with clean formatting and a thumbs up or down control. As a beta feature, it should improve over time as we tune it.
Human Review project pages

Human Review now has project-scoped pages, so you can open a project’s overview and assignments on their own pages and share a direct link to them.


Each project gets its own overview and assignments view with breadcrumbs, so you can navigate to and share specific review work without losing context.View in docs ↗
Connection validation for Pipecat, LiveKit, and WebSocket agents

Coval now validates Pipecat, LiveKit, and WebSocket agent connections when you set them up, so configuration problems surface before you run.

Bring your own background sounds

You can now upload your own background sounds for simulations, so agents can be tested against the exact ambient conditions they will face in production.

Telephony recording uploads

Telephony call recordings in MP3 format now upload reliably, including low-sample-rate recordings that were previously rejected.

Jun 6
2026
2 updates
Highlights
  • New pitch variability metric
  • New perceived loudness (LUFS) metric
Pitch variability metric

A new metric that flags whether an agent sounds monotone or expressive across a call.

Perceived loudness (LUFS)

A new metric that measures perceived loudness across a call, so you can catch audio that is too quiet or too loud.

Jun 1
2026
5 updates
Highlights
  • Create and delete dashboards via the API and CLI
  • Broader chat-agent connectivity with SSE streaming
  • Per-conversation metadata in metric prompts
Dashboards via API and CLI

Create and delete dashboards programmatically through the API and CLI.


Manage dashboards as part of your own workflows and scripts, without setting each one up by hand in the app. Useful for spinning up consistent dashboards per project or per environment.View in docs ↗
Broader chat-agent connectivity

Chat agents now support SSE streaming and a configurable response format.

Metadata in metric prompts

Dynamic metrics can now reference per-conversation metadata directly in the prompt template.

Clearer concurrency limits

When your organization is running evaluations beyond your concurrency limits, we will store your data but flag the evaluations as an error. This lets you rerun them later, once you are operating within your concurrency limits.

Smoother metric authoring

Metric descriptions auto-populate and the name auto-fills when you create a metric.

May 26
2026
1 update
Highlights
  • More reliable categorical audio metrics
More reliable categorical audio metrics

Categorical audio metrics now flag a clear, descriptive error when no categories are configured, so misconfigurations surface right away.

May 19
2026
7 updates
Highlights
  • Visual IVR Tree Builder
  • IVR Flow Adherence metric
  • Custom trace aggregations
  • Tags for metrics, templates, and test sets
IVR Flow Adherence metric

A built-in metric that checks whether a call follows the intended IVR navigation path, with no custom scoring logic required.


The metric scores each call against the IVR path you define, so you can see where calls deviate from the intended route. Results are reported per call and roll up across the run.View in docs ↗
IVR Tree Builder

Define and simulate branching IVR call flows visually, directly in Coval.


Lay out the call paths a caller can take, then run simulations against the whole tree to see where agents take the wrong branch. No external diagramming or scripting needed.View in docs ↗
Custom trace aggregations

Choose how per-turn scores roll up across a multi-turn trace, for example worst-case or first-occurrence.

Tags

Apply custom tags to metrics, run templates, and test sets to organize and filter your library.

Expanded voice catalog

New voices are available in the persona picker, ready to use with no setup.

Flexible sign-in

Custom sign-in now works across all configured authentication methods.

Shared run access

Password-protected shared runs now open reliably across different saved password formats.

May 13
2026
6 updates
Highlights
  • Faster, more reliable reports
  • Easier metric selection
  • More resilient metric scoring
Faster, more reliable reports

Large reports that used to struggle to load now open reliably even at scale, thanks to a change in how we load them.

Easier metric selection

The metric picker now uses grouped, nested categories, so the right metric is faster to find and apply.

Larger conversation uploads

There is no longer a 1000-character limit on metadata for uploaded conversations.

Consistent LLM Judge labeling

Auto-generated metrics are now reliably tagged LLM Judge, so AI-scored and rule-based metrics are easy to tell apart.

Accurate audio metric scoring

Speech anomaly, volume variance, and audio sentiment metrics now scope to the channel and time range you configure.

More resilient metric scoring

Metrics now do a better job of handling a wider variety of metadata inputs.