> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/ai-evals/get-started/overview.md).

# AI Evals Overview

Harness AI Evals is the quality gate for AI in your CI/CD pipeline. Evaluate datasets against your AI systems on every change and catch regressions before they reach production.

Harness AI Evals is the quality gate for AI in your CI/CD pipeline. It runs versioned evaluations against your AI systems on every change, so you catch regressions before they reach production.

{% embed url="<https://youtu.be/VXUS-HnFXPE>" %}
Introducing Harness AI Evals
{% endembed %}

***

## What you will learn

* **How AI Evals works**: targets, datasets, and metric sets combine to score outputs.
* **Why LLM outputs need evaluation** and the five failure modes it catches.
* **The platform loop** that connects observation, evaluation, analytics, and CI/CD gates.
* **The architecture** and how the SDK and control plane split responsibilities.
* **The product surfaces** in the UI and what each one is for.
* **The core concepts** you use to build evaluations.

***

## How it works

Every evaluation binds three inputs together: a **target** (the AI system under test), a **dataset** (versioned golden test cases), and a **metric set** (the scoring rubric). AI Evals runs the target against each dataset row, scores the output with every metric in the set, and aggregates the results.

Group evaluations into eval suites, gate a pipeline stage on the suite result, and pipe production traces from Harness AgentTrace back into the dataset to keep coverage current. The same target, dataset, and metric set drive both a developer running `pytest` locally and a pipeline gating a deployment.

***

## Why LLM outputs need evals

**Traditional testing falls short for AI systems.** You cannot write an `assert` statement that checks if an answer is "relevant", "factually grounded", or "helpful". These qualities are subjective, context-dependent, and cannot be reduced to a boolean check. When an LLM returns a fluent but wrong answer, there is no exception to catch and no non-zero exit code to fail the build.

AI Evals exists because evaluating AI outputs requires a different approach: **score outputs against a rubric** rather than assert they match an exact value. Metrics define what "good" means across dimensions like correctness, groundedness, safety, tone, and task completion. Some metrics are deterministic (exact match, contains, regex). Others use LLMs as judges to grade subjective qualities. All of them produce a numeric score that you can threshold, weight, aggregate, and track over time.

Traditional unit and integration tests do not catch the failure modes that matter most for AI systems. Five gaps show up again and again in production:

| Challenge                   | What happens in production                                                                                         | How AI Evals addresses it                                                                                                                   |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------- |
| **Silent failures**         | An LLM returns a fluent, confident answer that is wrong. There is no exception to catch and no non-zero exit code. | Heuristic and LLM-as-a-judge metrics grade every output against a golden expected value, so wrong-but-fluent answers surface as low scores. |
| **Non-determinism**         | The same prompt with the same context returns a different answer across runs.                                      | Datasets run at configurable sampling and concurrency. Trend charts show pass rate across runs, not just one shot.                          |
| **No regression baseline**  | A prompt change, model swap, or retrieval tweak improves one query and regresses another.                          | Every eval run is versioned. Calibration re-derives thresholds from a baseline so regressions are visible on the trend view.                |
| **Context drift**           | Production inputs shift over time. What passed at launch stops passing months later.                               | Observe production traces via AgentTrace, promote real interactions to the dataset, and re-evaluate on the updated corpus.                  |
| **Cost and latency spikes** | A safer or more accurate model doubles token spend or triples response time.                                       | Runs report per-item duration and judge cost. Cost limits hard-stop a run before the bill runs away.                                        |

***

## The platform loop

AI Evals sits inside a four-step loop that connects observation to enforcement.

```mermaid
flowchart LR
    A["Observe<br/><small>AgentTrace captures<br/>production traces</small>"] --> B["Judge<br/><small>AI Evals runs datasets<br/>through metrics</small>"]
    B --> C["Understand<br/><small>Analytics turns runs<br/>into pass-rate trends</small>"]
    C --> D["Act<br/><small>AI Eval step<br/>gates pipelines</small>"]
    D -.-> A
```

* **Observe.** Harness AgentTrace captures production traces of every LLM call, tool invocation, and agent step. You inspect them, annotate them, and promote interesting cases to an eval dataset.
* **Judge.** AI Evals runs those datasets against your AI systems with configurable metrics, both offline in the SDK and online in the control plane.
* **Understand.** Analytics turns run results into pass-rate trends, regression alerts, cost curves, and latency histograms across evaluations, suites, and metrics.
* **Act.** AI Eval steps in Harness pipelines gate deployments on suite results. A failing domain suite blocks a stage; a master suite feeds an advisory signal without blocking.

The loop keeps evaluation continuous. Every change in the target, dataset, or scoring rubric flows back through it.

***

## Architecture at a glance

AI Evals is a two-layer system: an SDK you run anywhere Python runs, and a control plane hosted in Harness.

```mermaid
flowchart TB
    subgraph SDK["harness-evals SDK (pure Python)"]
        S1["Define datasets"]
        S2["Define metric sets"]
        S3["Run locally with pytest"]
    end
    subgraph CP["AI Evals control plane (Harness)"]
        C1["Targets, Datasets, Metric Sets"]
        C2["Evaluations & Runs"]
        C3["Suites"]
        C4["Observe, Analytics, Registry"]
    end
    S3 -->|forward results| C2
    C2 --> C4
    C3 --> C2
```

* **`harness-evals` SDK.** A pure-Python package that defines datasets, targets, metric sets, and evaluations, and runs them locally or in CI. Install with `pip install harness-evals`. Import as `harness_evals`. Source lives at [github.com/harness/harness-evals](https://github.com/harness/harness-evals) under Apache 2.0.
* **AI Evals control plane.** A hosted service in Harness that manages targets, datasets, metric sets, evaluations, runs, suites, and traces. Exposes the AI Evals product surfaces in the Harness UI: Targets, Datasets, Scorers, Evaluations, Suites, Observe, Analytics, and Registry.

Local runs from the SDK forward their results to the control plane, so a single definition of a dataset, evaluation, and metric set serves both a developer running `pytest` at their desk and a pipeline gating a deployment.

***

## Product surfaces

The AI Evals UI groups its surfaces into three stages of the evaluation workflow: build the pieces, run and gate on them, and observe how they perform over time.

### Build

Register the AI system, curate its test cases, and define the scoring rubric.

* **Targets.** Register the AI systems you want to evaluate. A target has a name, a type (Prompt, App, or Static — shown as Agent, Prompt, or Precomputed in the UI), and connection details. Targets are reusable across evaluations.
* **Datasets.** Create and version datasets of golden test cases. Add rows by pasting JSONL, uploading a file, or generating synthetic rows with an LLM. Each row has an input and optional expected values.
* **Scorers.** The Scorers tab is where you build metric sets by picking from the built-in metric library or writing custom metrics. Set thresholds and weights per metric. Attach a Judge LLM Connector when using LLM-based metrics.

### Run and gate

Wire targets, datasets, and metric sets into evaluations, then group them into suites for pipeline gating.

* **Evaluations.** Wire a target, a dataset, and a metric set together. Configure sampling, concurrency, timeout, and cost limits. Evaluations activate as soon as all three pieces are set.
* **Suites.** Group evaluations that share a pass strategy. Run a suite with one click and get a combined pass or fail across every evaluation inside it.

### Observe and iterate

Inspect production traces, track pass-rate trends, and version the prompts and agents that make up your AI system.

* **Observe.** Inspect production traces from AgentTrace. Filter by session, agent, or time range. Annotate traces and promote interesting cases to a dataset.
* **Analytics.** Track pass-rate trends across evaluations, suites, and metrics. Spot regression alerts, cost curves, and latency histograms. Compare runs side by side.
* **Registry.** Version prompts, agents, tools, and skills. Pin a specific version to an evaluation so it reproduces exactly, even if the source moves on.

***

## Key concepts map

Five building blocks compose every evaluation in AI Evals.

```mermaid
flowchart LR
    T[Target] --> E[Evaluation]
    D[Dataset] --> E
    M[Metric Set] --> E
    E --> R[Run]
    R --> S[Eval Suite]
    S --> P[Pipeline Step]
```

| Concept           | What it is                                                                                                                                                                                                    |
| ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Target**        | The AI system under test. An HTTP agent endpoint, an LLM prompt with a system message, or a precomputed output file.                                                                                          |
| **Dataset**       | A versioned collection of golden test cases. Each row is an input and, depending on scoring, optional `expected_tools`, `expected_output`, or `context`.                                                      |
| **Metric Set**    | The scoring rubric. A list of metrics with per-metric thresholds and weights. Metric types: heuristic, LLM-judged, embedding, or pre-built AI judge.                                                          |
| **Evaluation**    | The wiring. Binds a target, a dataset, and a metric set together, plus run-time settings like sampling, concurrency, timeout, and cost limits.                                                                |
| **Eval Suite**    | A group of evaluations with a shared pass strategy. Domain suites use `all_must_pass` and block a pipeline stage on failure. Master suites use `weighted_threshold` and feed advisory signals into Analytics. |
| **Pipeline Step** | The AI Eval step in a Harness pipeline. Runs an evaluation or a suite as part of CI/CD and gates the stage on the result.                                                                                     |

Go to [Core concepts](/ai-evals/core-concepts/targets.md) to review each building block in depth.

***

## What you can use today

AI Evals is generally available in Harness. Every project with AI Evals access can use the full product surface end to end: registering targets, versioning datasets, building metric sets, running evaluations, grouping them into suites, gating pipelines on suite results, and observing production traces. Both the UI and the `harness-evals` Python SDK are supported paths, and the AI Eval pipeline step is available in every Harness pipeline. No feature flag is required, and no separate enablement request is needed if your account already has AI Evals in the left navigation.

***

## Track what changes

AI Evals ships new metrics, target types, and UI capabilities on the same progressive-deployment cadence as the rest of the Harness Platform. New metric additions, changes to the AI Eval pipeline step, SDK releases, and UI-surface updates are documented in the AI Evals release notes. Check the release notes when calibrating existing metric sets, planning an SDK upgrade, or investigating a change in evaluation behavior between runs.

***

## Frequently asked questions

<details>

<summary>Do I need Python to use Harness AI Evals?</summary>

No. The AI Evals UI supports the full flow, including creating targets, datasets, metric sets, and evaluations, running them, and reviewing results, without writing code. Python is only required if you want to use the harness-evals SDK to define evaluations in code or run them locally with pytest.

</details>

<details>

<summary>Can I use my own judge model for LLM-as-a-judge metrics?</summary>

Yes. LLM-based metrics reference a Judge LLM Connector configured in your Harness project. You can use OpenAI, Anthropic, or any provider supported by Harness connectors. Heuristic and embedding metrics do not require a judge model.

</details>

<details>

<summary>Does AI Evals replace unit and integration tests?</summary>

No. AI Evals complements traditional tests. Unit and integration tests still cover deterministic code paths and system boundaries. AI Evals covers the non-deterministic quality of AI outputs, such as correctness, groundedness, safety, trajectory, and performance, which traditional tests cannot express.

</details>

<details>

<summary>What is the difference between a domain suite and a master suite?</summary>

A domain suite uses the all\_must\_pass pass strategy and blocks a pipeline stage if any single evaluation in it fails. A master suite uses weighted\_threshold, aggregates scores across domain suites, and feeds an advisory signal into Analytics without blocking deployments.

</details>

<details>

<summary>What is the difference between heuristic, LLM-as-a-judge, and embedding metrics?</summary>

Heuristic metrics are deterministic and free to run. They include exact match, contains, tool selection, regex, and JSON diff. LLM-as-a-judge metrics use a judge model to grade outputs on qualitative dimensions like faithfulness, toxicity, and hallucination. Embedding metrics compute cosine similarity between the output and the expected value. Heuristics run for free; LLM and AI-judge metrics require a Judge LLM Connector.

</details>

<details>

<summary>Can I evaluate a system that is not hosted in Harness?</summary>

Yes. Targets can point to external HTTP endpoints, external LLM providers, or precomputed output files that you upload. AI Evals does not require your AI system to run inside Harness. For private endpoints behind a VPC or firewall, contact Harness Support to review your networking setup.

</details>

***

## Ready to build your first eval?

[Get Started](/ai-evals/get-started/get-started.md)

***

## Related concepts

* Go to [Core concepts](/ai-evals/core-concepts/targets.md) to review targets, datasets, metric sets, evaluations, and suites in depth.
* Go to [Pipeline integration](/ai-evals/guides/overview-2.md) to gate a Harness pipeline on an eval suite.
* Go to [Multi-turn and agentic evals](/ai-evals/guides/overview.md) to evaluate conversations and tool-use trajectories.
* Go to [Observability and analytics](/ai-evals/guides/overview-1.md) to inspect production traces and track pass-rate trends.
* Go to [SDK reference](/ai-evals/guides/overview-3.md) to install `harness-evals` and define evaluations in code.
