- Text-based — for chat agents using text inputs and outputs.
- Voice-based — for voice agents with audio inputs and outputs.
What you need to set up a simulated conversation
A simulated conversation brings together four building blocks. Each answers a different question about the test, and each lives in its own section of Coval so you can reuse it across many runs.Agent
Who is being tested. A reusable connection to your voice or chat agent. Without an agent, there’s nothing to talk to — this is how Coval reaches your system.
Persona
Who is doing the talking. The simulated user — their voice, accent, and behavior. Personas let you test how your agent holds up against the range of real people it will actually meet.
Test Set
What the conversation is about. A collection of test cases (scenarios, goals, scripts) that tell the simulated user what to do. This defines the situations you want to cover.
Metrics
How you measure success. The pass/fail criteria applied to each conversation — from latency and sentiment to custom LLM-judged questions. Metrics turn a transcript into an answer about whether your agent did its job.
Launching simulated conversations
When you launch an evaluation, Coval starts one or more simulated conversations and evaluates their results. A single launch is called a run. A run can contain many simulated conversations, for example one per test case in your test set. Open the Simulate page and follow these steps:1
Open the Simulate page
Click + beside Simulated in the sidebar, or New Simulation on the Simulated Conversations page.
2
3
Launch
Click Launch — Coval runs the simulated conversations and evaluates them against your metrics.
Re-running a simulation
If a single simulation looks off, select it in the run’s results table and click Resimulate to run it again in place against your latest configuration. This overwrites the previous result, so launch a new run instead if you want to keep the old one for comparison.Reading run results
Open a run to see its results table — one row per simulated conversation, with a column for each metric.- Identifier columns — The Simulation ID column is shown by default. Use the column toggle to add or remove identifiers.
- Tool calls in transcripts — When a transcript message contains several tool calls, the transcript shows the first ones and a +N more control. Click it to reveal the remaining calls without leaving the message.
- Tags — Add or edit tags from the results header of the run or of an individual simulation result. Tags are stored on the run, so tagging one simulation result tags the parent run and every simulation in it. See Tagging runs.
Time limits
Every simulated conversation has a maximum duration. It ends when the conversation wraps up naturally, the objective is met, or the limit is reached — whichever comes first. The default limit depends on your plan:
The limit can be set anywhere from 1 minute to 60 minutes for your organization — contact support@coval.dev to change it.
Voice-to-voice agents (OpenAI Realtime, Gemini Live, Grok Realtime, and Deepgram Voice Agent) also carry a per-agent Simulation Timeout in their connection settings: 900 seconds (15 minutes) by default, up to 1800 seconds (30 minutes).
When a voice simulation reaches its limit, Coval lets the agent’s final reply finish transcribing before ending the call, so the transcript includes what the agent was saying at the cutoff. The simulation’s end reason is recorded as
DURATION_LIMIT — see the End Reason metric.
Next steps
Connect your agent
Set up the connection to your voice or chat agent.
Launch your first run
Put the pieces together and run your first evaluation.