> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/ai-evals/get-started/quickstart-ui.md).

# UI Quickstart

Follow the New Evaluation wizard to run your first AI evaluation in Harness AI Evals. No code required.

The fastest way to your first passing evaluation is the **Create Evaluation** wizard. It walks you through five short screens, then offers to auto-create a pipeline and run your eval on the spot.

The wizard covers each of the five steps in order.

```mermaid
flowchart LR
    B[1 Basic info] --> T[2 Target] --> D[3 Dataset] --> M[4 Metric set] --> C[5 Configuration] --> R[Run] --> V[Review]
```

The example evaluates a customer support agent that answers ecommerce questions. Substitute your own target everywhere it appears.

***

## Open the wizard

Complete these three actions to reach the wizard.

1. In your Harness project, select **AI Evals** in the left nav.
2. Select **Evaluations** > **Create Evaluation**.
3. The wizard opens on the **Basic info** step.

***

## Step 1: Basic info

Give the evaluation a name and a short description.

| Field           | Value                                                       |
| --------------- | ----------------------------------------------------------- |
| **Name**        | `Support Agent Quickstart`                                  |
| **Description** | (optional) e.g. `First eval for the customer support agent` |

Select **Continue**.

***

## Step 2: Target

The target is the AI system you are evaluating. You can pick an existing one or create a new one inline.

Select **Select Existing** and choose your registered target from the dropdown. If this is your first target, select **Create New** and fill in the target form:

| Field                | Value                                                                                                                                     |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- |
| **Name**             | `Support Agent`                                                                                                                           |
| **Type**             | **Agent** — called App (`agent`) in the API. Other options: **Prompt** (`prompt`), **Precomputed** — called Static (`static`) in the API. |
| **Endpoint**         | `https://your-agent.example.com/api/eval`                                                                                                 |
| **Method**           | `POST`                                                                                                                                    |
| **Request template** | `{ "input": "{{input.content}}" }`                                                                                                        |
| **Response path**    | `output`                                                                                                                                  |

Use the **Test Target** panel to send a sample input and confirm the agent returns a real answer. Select **Continue**.

{% hint style="warning" %}
**`request_template` is not optional.** Without it, the runner double-wraps `row.input`. Your agent sees `{input: {content: "..."}}` instead of `{input: "..."}`, and every row fails with HTTP 400.
{% endhint %}

***

## Step 3: Dataset

The dataset is your ground truth. Each row is a scenario the agent will be tested against.

Select **Select Existing** to pick a dataset, or **Create New** to add one inline. You can add rows three ways:

* **Paste JSONL** in the inline editor.
* **Upload a `.jsonl` file** for bulk imports.
* **Generate synthetic** rows with an LLM. Pick a strategy (`adversarial`, `rephrase`, `complexity_ladder`, or `use_case`), describe the use case, and review the generated rows before saving.

For this quickstart, paste this in the JSONL editor:

```json
{"id": "iphone-stock",   "input": {"content": "Do you have iPhone 15 Pro in stock?"}, "expected_output": {"text": "yes"}}
{"id": "macbook-price",  "input": {"content": "How much does the MacBook Pro 14-inch cost?"}, "expected_output": {"text": "$1,999"}}
{"id": "shipping",       "input": {"content": "Do you offer free shipping?"}, "expected_output": {"text": "shipping"}}
```

Select **Continue**.

{% hint style="success" %}
**Field rules.** `input` must be a JSON object, not a bare string. `expected_output` for `ContainsMetric` must be `{"text": "..."}`. A bare string silently fails.
{% endhint %}

***

## Step 4: Metric set

The metric set defines what "good" means for your agent. Pick a few metrics from the built-in catalog. Filter by type (**heuristic**, **LLM judge**, **embedding**, **AI judge**) to narrow the list.

Select **Create New** and name the metric set (for example, `Support Quality`). Add three metrics that cover writing quality, tone, and outcome.

| Metric                 | Threshold | What it checks                                                    |
| ---------------------- | --------- | ----------------------------------------------------------------- |
| `NoTyposMetric`        | `0.8`     | The output is grammatically clean.                                |
| `PolitenessMetric`     | `0.8`     | The response is professional and courteous.                       |
| `TaskCompletionMetric` | `0.7`     | The agent actually answered the question, not something adjacent. |

These are LLM-as-a-judge metrics, so select a **Judge LLM Connector** at the top of the form. Any OpenAI or Anthropic connector configured in your project works.

Select **Continue**.

{% hint style="success" %}
**Prefer no judge model?** Start with heuristic metrics like `ContainsMetric` and `ExactMatchMetric`. They score for free with no connector needed. Layer in LLM-judge metrics once your first run works.
{% endhint %}

***

## Step 5: Configuration

The last screen controls how the run behaves. The defaults work for most first runs.

| Setting               | Default     | What it does                                                                                 |
| --------------------- | ----------- | -------------------------------------------------------------------------------------------- |
| **Sampling strategy** | `All Items` | Run every row in the dataset. Use `first N` or `random N` for smoke tests on large datasets. |
| **Concurrency**       | `5`         | How many rows run against your agent in parallel.                                            |
| **Timeout per item**  | `30000` ms  | How long a single row can take before failing.                                               |
| **Cost limit**        | none        | Optional hard stop if judge tokens exceed a dollar amount.                                   |

Select **Save & Continue**.

***

## Create the pipeline and run

Once the eval is saved, a modal appears with two options.

* **Create Pipeline & Run**. Harness bundles the evaluation into a **Suite**, generates a default CI/CD pipeline, attaches your eval as an **AI Eval** step, and starts the first run right away.
* **Skip for now**. The eval is saved for later. You can trigger runs manually from the eval detail page.

Select **Create Pipeline & Run**. The run kicks off in a minute or two and forwards results to the Results page as it completes.

***

**You just shipped your first AI quality gate.** Every time you push code, this eval runs automatically and blocks deployment if quality drops. No more guessing whether your agent still works after a prompt tweak or a model swap.

{% hint style="info" %}
**Prefer to build resources one at a time?** You can also register a target, dataset, and metric set from their own dedicated tabs (**Targets**, **Datasets**, **Scorers**) and then wire them together in **Evaluations**. Useful when you want to reuse existing resources across many evaluations.
{% endhint %}
