Most Used
Runs
Launch evaluation runs and track results across your agents
Agents
Connect and configure your AI agents for testing
Simulated Conversations
View simulated conversation results with transcripts and metric scores
Test Sets
Create and manage test cases for your simulations
Getting Started
1
Get your API key
Obtain your API key from the Coval Dashboard. See the API Keys guide for detailed setup instructions.
2
Authenticate your requests
Include your API key in the
X-API-Key header:3
Create your first agent
Set up an agent configuration for testing:
4
Launch simulated conversations
Launch a run and inspect its results with the Runs and Simulated Conversations APIs.
Base URL
Authentication
Every request except the public OpenAPI specification endpoints requires an API key in theX-API-Key header. The header name is matched case-insensitively, so x-api-key works too.
A missing or invalid key returns 401 with the code UNAUTHENTICATED. A key that lacks the permission scope an endpoint requires returns 403 with PERMISSION_DENIED. See the API Keys guide for creating keys and choosing scopes.
Workspaces
Resources such as agents, test sets, runs, and conversations belong to a workspace within your organization. Send the optionalX-Coval-Workspace-Id header (also matched case-insensitively) with a workspace id to act in that workspace:
- Omit the header and the request uses your organization’s default workspace. Resources created before your organization had workspaces belong to the default workspace.
- Listings such as
GET /conversations/simulatedandGET /runsreturn only the resolved workspace’s resources, and resources you create belong to that workspace. - A header naming a workspace that does not exist, is archived or deleted, or belongs to another organization is rejected with
401and the codeUNAUTHENTICATED. - API keys and the Workspaces endpoints themselves are organization-level and are not filtered by the header.
Pagination
List endpoints page withpage_size and page_token:
page_sizeis the maximum number of results per page. It defaults to 50; the upper bound varies by endpoint and is shown on each endpoint’s page (for example 1000 for runs and simulated conversations, 250 for uploaded conversations).page_tokenis the opaque cursor returned asnext_page_tokenin the previous response. Omit it for the first page. Anullnext_page_tokenmeans there are no more results.
metric.{metric_id} predicate in filter) or embed them (include=metric_values for simulated and uploaded conversations, include=metric_averages for runs) are served newest-first and cap page_size at 100. See List simulations and List runs.
Filtering and sorting
List endpoints accept afilter expression and an order_by field.
filter is one or more field operator value predicates joined with AND or OR. Operators are =, !=, >, <, >=, and <=. Values may be unquoted or double-quoted; quote any value containing spaces, such as status="IN PROGRESS". The supported fields differ per endpoint and are listed on each endpoint’s page. For example, simulated conversations can be filtered on status, agent_id, persona_id, test_set_id, test_case_id, run_id, external_conversation_id, mutation_id, mutation_name, create_time, and metric.{metric_id}.
order_by takes a field name for ascending order or -field for descending. Most lists default to -create_time (newest first); uploaded conversations default to -occurred_at. The sortable fields are listed on each endpoint’s page.
A metric.{metric_id} predicate is more restrictive: it supports only AND alongside a small set of fields, and any non-default order_by is rejected with 400. See List simulations and List conversations for the details.
Response detail
Metric output lists on simulated and uploaded conversations accept two parameters that control how much comes back:viewselects the detail level.FULL(the default) includes every field, including the structuredresultandruntime_metadata.BASIComits those heavy fields while retaining the value, status, explanation, and bounded subvalues.include_supersededcontrols re-score history. Re-scoring a metric appends a new output rather than replacing the previous one. With the defaultfalse, each metric appears at most once with its most recent output; set it totrueto return every output, including those a later re-score has replaced.
Errors
Errors return anerror object with a machine-readable code, a human-readable message, and, where applicable, details naming the field at fault:
Every response carries an
X-Request-Id header. Include it when you contact support about a failed request.
Rate limits
Only Batch rerun metrics is rate limited today. Reruns count against an organization-level hourly limit on simulations queued (5000 per hour by default). When the limit is exceeded the API returns429 with the code RESOURCE_EXHAUSTED and a Retry-After header giving the number of seconds until the window rolls over. No other endpoint has a published quota.
OpenAPI Specification
We publish our OpenAPI specifications at public endpoints (no authentication required).List available specs
spec_name values, their URLs, and when each spec was last published (last_modified).
Fetch a specific spec
- Default response: YAML (
application/yaml) - JSON response: set
Accept: application/json