Datasets API
The Datasets API manages project-scoped Datasets and their directly owned Examples.
Base Path#
/v1/projects/:projectId/datasetsAPI-key callers need workspace:read for reads and models:manage for writes. Signed-in Workbench
users need the corresponding project data permissions.
Dataset Endpoints#
GET /v1/projects/:projectId/datasets
POST /v1/projects/:projectId/datasets
GET /v1/projects/:projectId/datasets/:datasetId
PATCH /v1/projects/:projectId/datasets/:datasetId
DELETE /v1/projects/:projectId/datasets/:datasetIdCreate a Dataset:
curl -X POST "$PROVON_API_URL/projects/$PROJECT_ID/datasets" \
-H "Authorization: Bearer $PROVON_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Correct support failures",
"description": "Reviewed corrections from production conversations",
"schema": { "type": "chat", "version": 1 },
"collectionObjective": "correct_failed_conversations",
"collectionModel": "openai/gpt-4o-mini",
"collectionEnabled": true
}'Supported collection objectives are preserve_successful_behavior,
correct_failed_conversations, learn_from_user_corrections, and
teach_project_knowledge. The last objective can be filtered with
collectionFilters.knowledgeGoalIds.
preserve_successful_behavior copies verified outputs and does not require a model. The other
objectives use collectionModel, which defaults to provon/auto. Dataset responses include
collectionStatus (off, ready, or setup_required) and blockedReason; model-dependent
collection is not scheduled while its model needs setup.
Dataset schemas cannot be changed after creation.
Dataset Fields#
| Field | Type | Notes |
|---|---|---|
id |
string | Generated by the server. |
projectId |
string | Inherited from the URL. |
name |
string | Required. Non-empty. |
description |
string | null | Optional. |
schema |
object | Required. Currently only { "type": "chat", "version": 1 } is supported and it is immutable. |
collectionPolicy |
object | null | Required for automatic collection. Contains objective, revision, optional model, and optional filters. |
collectionPolicy.objective |
string | One of the supported collection objectives. |
collectionPolicy.revision |
string | Server-managed revision of the policy. |
collectionPolicy.model |
string | Model used by objectives that synthesize an expected response. |
collectionPolicy.filters |
object | Optional. knowledgeGoalIds filters teach_project_knowledge; categories is reserved. |
collectionEnabled |
boolean | Defaults to true on creation. Disable to freeze the working set. |
collectionStatus |
string | Effective state: off, ready, or setup_required. |
blockedReason |
string | null | Actionable reason automatic collection cannot currently run. |
exampleCount |
number | Counter maintained in the same transaction as Example create/delete. |
createdAt |
number | Unix epoch milliseconds. |
updatedAt |
number | Unix epoch milliseconds. |
Example Endpoints#
GET /v1/projects/:projectId/datasets/:datasetId/examples
POST /v1/projects/:projectId/datasets/:datasetId/examples
GET /v1/projects/:projectId/datasets/:datasetId/examples/:exampleId
PATCH /v1/projects/:projectId/datasets/:datasetId/examples/:exampleId
DELETE /v1/projects/:projectId/datasets/:datasetId/examples/:exampleIdCreate a manual Example:
curl -X POST \
"$PROVON_API_URL/projects/$PROJECT_ID/datasets/$DATASET_ID/examples" \
-H "Authorization: Bearer $PROVON_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"source": { "kind": "manual", "externalRef": "case-123" },
"tags": ["support"],
"payload": {
"input": {
"type": "chat",
"messages": [{ "role": "user", "content": "Question" }]
},
"expectedOutput": {
"type": "chat_message",
"message": { "role": "assistant", "content": "Answer" }
}
}
}'PATCH accepts source, tags, or payload. Legacy top-level input, expectedOutput, and
metadata fields are rejected.
Example Fields#
| Field | Type | Notes |
|---|---|---|
id |
string | Generated by the server. |
projectId |
string | Inherited from the Dataset. |
datasetId |
string | Inherited from the URL. |
source |
object | Required. See source kinds below. |
tags |
string[] | Optional. Non-empty strings, no duplicates. |
payload |
object | Required. Contains input and optional expectedOutput, rejectedOutput, rubric. |
createdAt |
number | Unix epoch milliseconds. |
updatedAt |
number | Unix epoch milliseconds. |
Source Kinds#
| Kind | Required Fields | Use When |
|---|---|---|
manual |
externalRef or traceIds |
The example is authored directly, possibly with a reference to an external case or trace. |
conversation |
conversationId, traceIds |
The example was derived from a Provon conversation. |
import |
externalRef |
The example was imported from an external system or file. |
synthetic |
externalRef or traceIds |
The example was generated synthetically and needs an audit reference. |
All source kinds may include derivationKey for idempotency or grouping.
Payload Shapes#
input supports:
{ "type": "chat", "messages": [...] }{ "type": "prompt", "prompt": "...", "system?": "..." }{ "type": "instruction", "instruction": "...", "context?": "..." }{ "type": "text", "text": "..." }{ "type": "json", "value": ... }
expectedOutput / rejectedOutput support:
{ "type": "chat_message", "message": { "role": "assistant", "content": "..." } }{ "type": "text", "text": "..." }{ "type": "json", "value": ... }{ "type": "classification", "label": "..." }
rubric is an object with criteria and optional scoring entries (label, description, score).
Errors#
Creating a Dataset with an unsupported schema:
{
"error": "Bad Request",
"message": "schema must be chat/v1"
}Creating an Example with a missing required source field:
{
"error": "Bad Request",
"message": "source.conversationId is required for conversation sources"
}Using legacy top-level fields on PATCH:
{
"error": "Bad Request",
"message": "Legacy top-level input/expectedOutput/metadata fields are not accepted"
}Pagination And Consistency#
List endpoints accept limit and cursor:
| Parameter | Type | Notes |
|---|---|---|
limit |
integer | Page size from 1 to 1000; default 100 |
cursor |
string | Opaque cursor from the previous response |
Dataset list responses expose nextCursor. Example list responses expose the cursor returned by the store.
Creating or deleting an Example updates the Dataset's exampleCount in the same transaction.
Automatic collection uses deterministic provenance keys to avoid duplicate Examples.
See Datasets for concepts and the Dataset quickstart for the user workflow.