Evaluation Datasets & Runs

The Eval API provides organization-scoped scaffolding for defining evaluation datasets, scheduling runs against an agent and rubric, and recording per-case results and aggregate scores. Automated execution is driven by the agent runtime worker; this API owns the durable state. All eval endpoints require the caller to be an organization owner or admin.

Base URL

For local development:

Endpoints

Datasets

A dataset is a named collection of test cases. Each case is a free-form JSON object; common fields are input and expected_output.

Create a dataset

Request body

Response

Update a dataset

Delete a dataset

Response

Runs

A run binds a dataset to an agent and optional rubric. The runtime worker executes the run and records results.

Create a run

Request body

Response

Record scores

After a run completes, record per-case results and aggregate scores.

Request body

Response

Valid statuses: pending, running, completed, failed.

Grade a run against a rubric

If rubric_id is omitted, the rubric set on the run is used.

Response

Error codes