Skip to main content
A simulated conversation is an interaction between a Coval-driven user and your voice or chat agent. Instead of testing your agent by hand, Coval plays the role of a realistic user, talks to your agent end to end, and evaluates the conversation against criteria you define. You run simulated conversations to catch regressions before they reach production, measure how your agent performs across many different users and scenarios, and understand where a conversation goes wrong. Coval supports two kinds of simulated conversations:
  • Text-based — for chat agents using text inputs and outputs.
  • Voice-based — for voice agents with audio inputs and outputs.

What you need to set up a simulated conversation

A simulated conversation brings together four building blocks. Each answers a different question about the test, and each lives in its own section of Coval so you can reuse it across many runs.

Agent

Who is being tested. A reusable connection to your voice or chat agent. Without an agent, there’s nothing to talk to — this is how Coval reaches your system.

Persona

Who is doing the talking. The simulated user — their voice, accent, and behavior. Personas let you test how your agent holds up against the range of real people it will actually meet.

Test Set

What the conversation is about. A collection of test cases (scenarios, goals, scripts) that tell the simulated user what to do. This defines the situations you want to cover.

Metrics

How you measure success. The pass/fail criteria applied to each conversation — from latency and sentiment to custom LLM-judged questions. Metrics turn a transcript into an answer about whether your agent did its job.
Once you have these, you don’t have to wire them together by hand every time. A Template saves a combination of agent, persona, test set, and metrics so you can launch the same evaluation consistently — or schedule it to run on a recurring basis.

Launching simulated conversations

When you launch an evaluation, Coval starts one or more simulated conversations and evaluates their results. A single launch is called a run. A run can contain many simulated conversations, for example one per test case in your test set. Open the Simulate page and follow these steps:
1

Open the Simulate page

Click + beside Simulated in the sidebar, or New Simulation on the Simulated Conversations page.
2

Pick a template, or configure manually

Start from a saved template for a one-click launch, or set it up yourself:
  • Select an agent to test
  • Select a persona
  • Choose a test set
  • Set your simulation parameters
  • Choose the metrics to track
  • (Optional) Add tags to label this run
3

Launch

Click Launch — Coval runs the simulated conversations and evaluates them against your metrics.
Once launched, the run appears on the Simulated Conversations page. Open it to inspect each simulated conversation and its metric results.

Re-running a simulation

If a single simulation looks off, select it in the run’s results table and click Resimulate to run it again in place against your latest configuration. This overwrites the previous result, so launch a new run instead if you want to keep the old one for comparison.

Reading run results

Open a run to see its results table — one row per simulated conversation, with a column for each metric.
  • Identifier columns — The Simulation ID column is shown by default. Use the column toggle to add or remove identifiers.
  • Tool calls in transcripts — When a transcript message contains several tool calls, the transcript shows the first ones and a +N more control. Click it to reveal the remaining calls without leaving the message.
  • Tags — Add or edit tags from the results header of the run or of an individual simulation result. Tags are stored on the run, so tagging one simulation result tags the parent run and every simulation in it. See Tagging runs.

Time limits

Every simulated conversation has a maximum duration. It ends when the conversation wraps up naturally, the objective is met, or the limit is reached — whichever comes first. The default limit depends on your plan: The limit can be set anywhere from 1 minute to 60 minutes for your organization — contact support@coval.dev to change it. Voice-to-voice agents (OpenAI Realtime, Gemini Live, Grok Realtime, and Deepgram Voice Agent) also carry a per-agent Simulation Timeout in their connection settings: 900 seconds (15 minutes) by default, up to 1800 seconds (30 minutes). When a voice simulation reaches its limit, Coval lets the agent’s final reply finish transcribing before ending the call, so the transcript includes what the agent was saying at the cutoff. The simulation’s end reason is recorded as DURATION_LIMIT — see the End Reason metric.

Next steps

Connect your agent

Set up the connection to your voice or chat agent.

Launch your first run

Put the pieces together and run your first evaluation.