Skip to main content

Sofia

consult_sofia

Delegate a read-only evaluation question to Sofia. Sofia can use Coval product knowledge and inspect the authenticated organization’s runs, simulations, conversations, metrics, agents, personas, test sets, and dashboards. It cannot create, update, run, or delete resources.
In ChatGPT, consult_sofia takes a single standalone prompt of up to 4,000 characters and no conversation context. Include the exact facts the question needs in the prompt itself. Other MCP clients use the full parameters below.
Each conversation item must contain exactly the message content needed for the current request: Do not include API keys, access tokens, or other secrets in conversation context or session_id. For service reliability and evaluation, Coval records the prompt, the optional conversation messages supplied to this tool, and Sofia’s response. Messages elsewhere in the parent Claude, Codex, or other client session are not sent to Coval unless the client includes them in this tool call. Use direct tools for deterministic reads and writes. Use consult_sofia when the task benefits from evaluation expertise or synthesis across several Coval resources.

Run Management

list_runs

List evaluation runs with filtering and pagination.

get_run

Get detailed information about a specific run. Returns run details including status, progress, and metrics (if completed).

create_run

Launch a new evaluation run.

Agent Management

list_agents

List all configured agents.

get_agent

Get detailed configuration for a specific agent.

create_agent

Create a new agent configuration. Model Types:
  • MODEL_TYPE_VOICE - Inbound voice
  • MODEL_TYPE_OUTBOUND_VOICE - Outbound voice
  • MODEL_TYPE_CHAT - Chat/text
  • MODEL_TYPE_SMS - SMS messaging
  • MODEL_TYPE_WEBSOCKET - WebSocket

update_agent

Update an existing agent configuration.

Test Set Management

list_test_sets

List all test sets available for evaluation.

get_test_set

Get detailed information about a test set.

create_test_set

Create a new test set.

Test Case Management

list_test_cases

List test cases with optional filtering by test set.

get_test_case

Get detailed information about a test case.

create_test_case

Create a new test case in a test set. Scripted test cases. With input_type set to SCRIPT, the simulated user follows an exact sequence of turns instead of improvising from a scenario. Put the turns, in order, in simulation_metadata_input.script_turns. Each entry is either a plain string the simulated user speaks verbatim, a keypad press {"type": "dtmf", "digits": "1"} (digits may contain 0-9, *, #, and phone punctuation for full numbers), or {"type": "skip"} to stay silent for one turn.
Never write a keypad press as spoken text such as dtmf:1 or “press one”: the simulated user would read it aloud. Use a {"type": "dtmf", "digits": "..."} turn. For SCENARIO cases against a phone tree, instruct the simulated user in input_str to press keys instead.

update_test_case

Update an existing test case. To change the turns of a scripted case, prefer the top-level script_turns field. If you send both script_turns and simulation_metadata_input, the object is replaced first and the turns are then merged into the result, so script_turns always wins for that key.

Metrics

list_metrics

List available evaluation metrics.

get_metric

Get detailed configuration for a specific metric.

Reports

list_reports

List saved reports in the organization.

get_report

Get one saved report with a bounded page of its result rows. Row data is trimmed to fit the response budget. When the response flags that output was trimmed, narrow it with metric_ids and a smaller page_size, then continue with next_page_token. Check scope.has_more before treating the report as complete.

create_report

Save a report over a set of runs. Reports created through the connector are private to your organization.

Scheduling

list_run_templates

List reusable run configurations that a schedule can execute.

list_scheduled_runs

List recurring evaluation schedules.

get_scheduled_run

Get one schedule and a page of the runs it has triggered. History is served from the schedule’s most recent 500 runs. history_scope.upstream_capped is set when older runs exist beyond that window.

create_scheduled_run

Create a recurring evaluation schedule from a run template.
New schedules are created disabled. Enabling a schedule starts recurring evaluation work, which consumes simulation capacity on every trigger until you disable it.

update_scheduled_run

Update selected fields on an existing schedule. Provide at least one field besides the ID.

Personas

list_personas

List available simulated personas for testing.

get_persona

Get detailed configuration for a specific persona. Returns persona configuration including voice settings, language, and behavior.