> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/ai-evals/get-started/quickstart-api.md).

# REST API Quickstart

Integrate AI Evals into any tool or CI/CD pipeline using direct HTTP requests.

Integrate AI Evals into any tool or CI/CD pipeline using raw HTTP. Best for non-Harness environments, custom tooling, and language flexibility.

***

## The four-resource model

Every eval boils down to four resources plus a run.

```mermaid
flowchart LR
    T[Target] --> E[Evaluation]
    D[Dataset] --> E
    M[Metric Set] --> E
    E -->|run| R[Eval Run]
```

* **Target**: what to test. An HTTP agent endpoint, a prompt with a system message, or precomputed outputs.
* **Dataset**: the test cases. Each row has an `input`, and depending on scoring, optional `expected_output`, `expected_tools`, or `context`.
* **Metric Set**: how to score. A list of metrics with thresholds and weights.
* **Evaluation**: the wiring. Ties the three together plus run-time settings.
* **Eval Run**: the execution. Per-item scoring and aggregate results.

***

## Authentication

1. **Get API token:** Harness > Account Settings > Personal Access Tokens > New Token
2. **Base URL:** `https://app.harness.io/gateway/ai-evals/api/v1/orgs/{org}/projects/{project}/`
3. **Required headers:**
   * `x-api-key: {your-token}`
   * `Harness-Account: {your-account-id}`
   * `Content-Type: application/json`

Test authentication:

```bash
curl -X GET \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/targets?page=0&size=5" \
  -H "x-api-key: pat.YOUR_ACCOUNT.YOUR_TOKEN" \
  -H "Harness-Account: YOUR_ACCOUNT_ID"
```

If you get JSON with `{ "data": [...], "page": 0, "total_elements": N }`, authentication works.

***

## Build your first evaluation

The example evaluates a flight support agent. Substitute your own endpoint everywhere it appears below.

### 1. Create the target

```bash
curl -X POST \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/targets" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "SkyBuddy",
    "type": "agent",
    "description": "Flight support agent",
    "config": {
      "endpoint": "https://your-agent.example.com/api/eval",
      "method": "POST",
      "request_template": {"input": "{{input.content}}"},
      "response_path": "output",
      "timeout_ms": 60000
    },
    "tags": ["skybuddy", "demo"]
  }'
```

Response:

```json
{
  "id": "target-uuid",
  "name": "SkyBuddy",
  "type": "agent",
  "config": {},
  "created_at": "2026-07-16T10:00:00Z"
}
```

### 2. Create the dataset

```bash
curl -X POST \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/dataset" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "SkyBuddy Starter",
    "identifier": "skybuddy_starter",
    "description": "Basic flight queries and off-topic prompts"
  }'
```

### 3. Add dataset items

```bash
curl -X POST \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/dataset/{dataset-uuid}/items" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID" \
  -H "Content-Type: application/json" \
  -d '{
    "item_identifier": "status-ha482",
    "input": {"content": "What is the status of flight HA482?"},
    "expected_tools": ["get_flight_status"]
  }'
```

Response:

```json
{
  "id": "item-uuid-1",
  "item_identifier": "status-ha482",
  "input": {"content": "What is the status of flight HA482?"},
  "expected_tools": ["get_flight_status"],
  "created_at": "2026-07-16T10:02:00Z"
}
```

Add more items:

```bash
curl -X POST \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/dataset/{dataset-uuid}/items" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID" \
  -H "Content-Type: application/json" \
  -d '{
    "item_identifier": "offtopic-joke",
    "input": {"content": "Tell me a joke"},
    "expected_output": {"text": "help"}
  }'
```

### 4. List available metrics

```bash
curl -X GET \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/metrics?page=0&size=50" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID" \
  | jq '.data[] | select(.name | contains("ToolSelectionAccuracy") or contains("Contains"))'
```

Note the metric UUIDs from the response.

### 5. Create the metric set

```bash
curl -X POST \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/metric-sets" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "SkyBuddy Heuristics",
    "entries": [
      {
        "metric_id": "<ToolSelectionAccuracy UUID>",
        "threshold": 0.7,
        "weight": 1.0,
        "position": 0
      },
      {
        "metric_id": "<Contains UUID>",
        "threshold": 1.0,
        "weight": 1.0,
        "position": 1,
        "config": {"options": {"case_sensitive": false}}
      }
    ]
  }'
```

### 6. Create the evaluation

```bash
curl -X POST \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/evals" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "SkyBuddy Starter Eval",
    "target_id": "<target-uuid>",
    "dataset_id": "<dataset-uuid>",
    "metric_set_id": "<metric-set-uuid>",
    "concurrency": 3,
    "timeout_per_item_ms": 60000
  }'
```

### 7. Run it

```bash
curl -X POST \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/evals/{eval-uuid}/runs" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID"
```

Response:

```json
{
  "run_id": "run-uuid",
  "status": "running",
  "started_at": "2026-07-16T10:05:00Z"
}
```

### 8. Poll for completion

```bash
while true; do
  STATUS=$(curl -s \
    "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/runs/{run-uuid}" \
    -H "x-api-key: $TOKEN" \
    -H "Harness-Account: $ACCOUNT_ID" | jq -r '.status')
  
  echo "Status: $STATUS"
  
  if [ "$STATUS" = "completed" ] || [ "$STATUS" = "failed" ]; then
    break
  fi
  
  sleep 30
done
```

### 9. Fetch results

```bash
# Get summary scores
curl -X GET \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/runs/{run-uuid}" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID" \
  | jq '.summary_scores'

# Get per-item details
curl -X GET \
  "https://app.harness.io/gateway/ai-evals/api/v1/orgs/default/projects/my_project/runs/{run-uuid}/items?page=0&size=50" \
  -H "x-api-key: $TOKEN" \
  -H "Harness-Account: $ACCOUNT_ID"
```

***

## What success looks like

A completed run returns aggregate scores plus per-item results:

```json
{
  "run_id": "run-uuid",
  "status": "completed",
  "summary_scores": {
    "pass_rate": 0.83,
    "passed": 5,
    "failed": 1,
    "errored": 0
  },
  "dimension_scores": {
    "trajectory": 0.83,
    "correctness": 1.0
  }
}
```

Open the **Evaluations** tab in the AI Evals UI to see the same run with the full trace side by side.

***

## Troubleshooting

### Agent target returns HTTP 400 error with message about unexpected input structure

Without `request_template` on an agent target, the runner double-nests row input. Your agent receives `{input: {input: ...}}` instead of the expected structure and rejects with HTTP 400. Always specify the outgoing body shape explicitly on `type: agent` targets using the `request_template` field.

### API returns 401 Unauthorized even with valid token

Verify both `x-api-key` and `Harness-Account` headers are set. The token format is `pat.ACCOUNT_ID.TOKEN_STRING`. Also check that the token has not expired in Harness Account Settings.

### Dataset item creation fails with "input must be a JSON object" error

The `input` field must always be a JSON object, not a string. Use `{"content": "your-text-here"}` instead of passing the text directly. For example: `{"input": {"content": "What is 2+2?"}}` not `{"input": "What is 2+2?"}`.
