> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/ai-evals/get-started/quickstart-mcp.md).

# MCP Quickstart

Drive evals from an MCP-capable AI assistant like Claude Code or Cursor. Create, run, and calibrate evals in plain language.

Drive evals from an MCP-capable AI assistant (Claude Code, Cursor). Best for iterating with an agent that can create, run, and calibrate evals for you. Fastest path today.

***

## The four-resource model

Every eval boils down to four resources plus a run.

```mermaid
flowchart LR
    T[Target] --> E[Evaluation]
    D[Dataset] --> E
    M[Metric Set] --> E
    E -->|run| R[Eval Run]
```

* **Target**: what to test. An HTTP agent endpoint, a prompt with a system message, or precomputed outputs.
* **Dataset**: the test cases. Each row has an `input`, and depending on scoring, optional `expected_output`, `expected_tools`, or `context`.
* **Metric Set**: how to score. A list of metrics with thresholds and weights.
* **Evaluation**: the wiring. Ties the three together plus run-time settings.
* **Eval Run**: the execution. Per-item scoring and aggregate results.

***

## Install

The Harness MCP server ships as an npm package. The AI Evals toolset requires `harness-mcp-v2` **2.9.7 or later**.

```bash
npm i -g harness-mcp-v2 --registry=https://registry.npmjs.org
harness-mcp-v2 --version
```

***

## Configure

Claude Code reads MCP config from `~/.claude.json`. Add a server entry with your Harness credentials and toolsets.

```json
{
  "mcpServers": {
    "harness": {
      "type": "stdio",
      "command": "/opt/homebrew/bin/harness-mcp-v2",
      "args": [],
      "env": {
        "HARNESS_API_KEY": "pat.<account-id>....",
        "HARNESS_ACCOUNT_ID": "<account-id>",
        "HARNESS_DEFAULT_ORG_ID": "<org>",
        "HARNESS_DEFAULT_PROJECT_ID": "<project>",
        "HARNESS_ORG": "<org>",
        "HARNESS_PROJECT": "<project>",
        "HARNESS_BASE_URL": "https://app.harness.io/",
        "HARNESS_TOOLSETS": "ai-evals,pipelines,services,secrets,connectors"
      }
    }
  }
}
```

{% hint style="warning" %}
**Gotchas:**

* `ai-evals` must be in `HARNESS_TOOLSETS`, or leave the variable unset to enable defaults.
* The AI Evals toolset key is `ai-evals` (hyphenated). Some other toolsets use underscores (`file_store`, not `file-store`); check each toolset's documentation for the exact key.
* Both `HARNESS_ORG` / `HARNESS_PROJECT` and `HARNESS_DEFAULT_ORG_ID` / `HARNESS_DEFAULT_PROJECT_ID` are read by different code paths. Set both.
* After editing, kill running processes and reconnect: `pkill -9 -f harness-mcp-v2`, then `/mcp` in your client.
  {% endhint %}

***

## Build your first evaluation

The example evaluates a flight support agent. Substitute your own endpoint everywhere it appears below.

Six MCP calls take you from zero to a scored run.

### 1. Create the target

```json
{
  "name": "SkyBuddy",
  "type": "agent",
  "description": "Flight support agent",
  "config": {
    "method": "POST",
    "endpoint": "https://your-agent.example.com/api/eval",
    "timeout_ms": 60000,
    "request_template": { "input": "{{input.input}}" },
    "response_path": "output"
  },
  "tags": ["skybuddy", "demo"]
}
```

{% hint style="success" %}
**Sanity-check the target first.** Run `harness_execute resource_type=eval_target action=test` with a real input. If the agent does not return a real answer here, none of your evals will run.
{% endhint %}

### 2. Create the dataset

`expected_tools` drives trajectory metrics. `expected_output` drives answer metrics.

```json
{
  "name": "SkyBuddy Starter",
  "identifier": "skybuddy_starter",
  "items": [
    {"id": "status-ha482",   "input": {"input": "What's the status of flight HA482?"}, "expected_tools": ["get_flight_status"]},
    {"id": "search-blr-sin", "input": {"input": "Flights from BLR to SIN on 2026-07-12"}, "expected_tools": ["search_flights"]},
    {"id": "offtopic-joke",  "input": {"input": "Tell me a joke"}, "expected_output": {"text": "help"}},
    {"id": "write-cancel",   "input": {"input": "Cancel my booking HX7K2P"}, "expected_output": {"text": "airline"}}
  ]
}
```

### 3. Find metric UUIDs

Metrics are pre-seeded per account. Search for the ones you want:

```bash
harness_list resource_type=eval_metric filters={"search": "ToolSelectionAccuracy"}
harness_list resource_type=eval_metric filters={"search": "ContainsMetric"}
```

### 4. Create the metric set

```json
{
  "name": "SkyBuddy Heuristics",
  "entries": [
    { "metric_id": "<ToolSelectionAccuracy UUID>", "threshold": 0.7, "weight": 1, "position": 0 },
    { "metric_id": "<Contains UUID>",              "threshold": 1,   "weight": 1, "position": 1,
      "config": { "options": { "case_sensitive": false } } }
  ]
}
```

### 5. Create the evaluation

```json
{
  "name": "SkyBuddy Starter Eval",
  "target_id":     "<TARGET_UUID>",
  "dataset_id":    "<DATASET_UUID>",
  "metric_set_id": "<METRIC_SET_UUID>",
  "concurrency": 3,
  "timeout_per_item_ms": 60000
}
```

### 6. Run it

```bash
harness_execute resource_type=evaluation action=run \
  resource_id=<EVAL_UUID> \
  body={"trigger_type": "manual"}
```

The response includes a run UUID and a `pipeline_execution_id`. The runner is a Harness pipeline using `harnessdev/ai-evals-runner`.

### 7. Read results

```bash
harness_get  resource_type=eval_run      resource_id=<RUN_UUID>
harness_list resource_type=eval_run_item filters={"run_id": "<RUN_UUID>"} compact=false
```

The run returns `summary_scores`, `dimension_scores`, and pass/fail counts. Items return per-row input, output, per-metric scores, and any error messages.

{% hint style="info" %}
**Or just ask.** Once the MCP server is connected, describe the goal in plain language: "Register SkyBuddy at `<endpoint>`, create a starter dataset with four rows covering flight status and off-topic prompts, build a heuristic metric set with ToolSelectionAccuracy at 0.7 and Contains at 1.0, wire them into an evaluation and run it." Your assistant translates that into the six calls above.
{% endhint %}

***

## What success looks like

A working MCP run returns aggregate scores plus per-item results:

```json
{
  "run_id": "run_01K8Z...",
  "status": "completed",
  "eval_id": "eval_abc123",
  "summary_scores": {
    "pass_rate": 0.83,
    "passed": 5,
    "failed": 1,
    "errored": 0
  },
  "dimension_scores": {
    "trajectory": 0.83,
    "correctness": 1.0
  },
  "started_at": "2026-07-16T10:05:00Z",
  "completed_at": "2026-07-16T10:06:32Z"
}
```

**What the scores mean:**

* **pass\_rate: 0.83**: 83% of dataset rows passed all metrics
* **dimension\_scores**: Average scores per metric category (trajectory = tool metrics, correctness = answer metrics)
* **Status: completed**: Run finished successfully (other states: `running`, `failed`, `cancelled`)

Open the **Evaluations** tab in the AI Evals UI to see the same run with the full trace side by side.

### Natural language alternative

Once MCP is connected, you can also describe what you want in plain language:

```
Tell Claude: "Register SkyBuddy at https://my-agent.example.com/api/eval, 
create a starter dataset with 4 rows covering flight status and off-topic prompts, 
build a heuristic metric set with ToolSelectionAccuracy at 0.7 and Contains at 1.0, 
wire them into an evaluation and run it."
```

Claude translates that into the six MCP calls above and returns the run results.

***

## Troubleshooting

### Agent target returns HTTP 400 error with message about unexpected input structure

Without `request_template` on an agent target, the runner double-nests row input. Your agent receives `{input: {input: ...}}` instead of the expected structure and rejects with HTTP 400. Always specify the outgoing body shape explicitly on `type: agent` targets using the `request_template` field.

### MCP server changes in \~/.claude.json are not taking effect after restart

Claude Code caches MCP server processes. Editing `~/.claude.json` does not automatically re-spawn them. Run `pkill -9 -f harness-mcp-v2` to kill the cached process, then run `/mcp` in the client to reconnect. Restarting the terminal alone is not enough.
