> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/ai-evals/whats-supported.md).

# What's Supported

Platforms, targets, metrics, and integrations supported by Harness AI Evals

Harness AI Evals integrates with a wide range of target types, metrics, LLM providers, and CI/CD platforms. This page outlines all supported integrations, providers, and features available in Harness AI Evals.

For information about what is supported for other Harness modules and the Harness Platform overall, go to [Supported platforms and technologies](https://developer.harness.io/docs/platform/platform-whats-supported).

***

## 1. Target types

AI Evals evaluates three types of AI systems:

{% tabs %}
{% tab title="Prompt targets" %}
**Single LLM call with a prompt template**

Harness calls the LLM with your system message and user input.

**Supported providers:**

* OpenAI (GPT-4, GPT-3.5 Turbo)
* Anthropic (Claude 3 Opus, Sonnet, Haiku)
* Azure OpenAI
* Google Vertex AI (Gemini)

**Use case:** Testing prompt templates before deployment, rapid iteration on system messages
{% endtab %}

{% tab title="App targets" %}
**Any HTTP endpoint (black box evaluation)**

Harness sends HTTP requests to your agent or RAG pipeline.

**Supported configurations:**

* Any publicly accessible HTTPS endpoint
* Custom request templates (map dataset input to your API format)
* Private endpoints behind VPC/firewall (contact support)

**Use case:** Evaluating deployed agents, RAG pipelines, multi-step systems
{% endtab %}

{% tab title="Static targets" %}
**Precomputed outputs from file**

Upload outputs that were generated externally.

**Supported format:**

* JSONL (one JSON object per line)
* Uploaded via UI or API

**Use case:** Comparing model versions offline, scoring outputs without re-running inference
{% endtab %}
{% endtabs %}

***

## 2. Metrics (47+ built-in)

### By category

<details>

<summary><strong>Deterministic (free, no LLM judge)</strong></summary>

* **ExactMatch**: Output exactly matches expected
* **Contains**: Output contains expected substring
* **Regex**: Output matches regular expression
* **NumericDiff**: Numeric accuracy within threshold
* **JsonDiff**: JSON structure and values match
* **SchemaValidation**: JSON validates against schema

</details>

<details>

<summary><strong>Similarity (embedding-based)</strong></summary>

* **BLEU**: Translation quality metric
* **Levenshtein**: Edit distance
* **EmbeddingSimilarity**: Cosine similarity between embeddings

</details>

<details>

<summary><strong>LLM-as-a-judge (requires Judge connector)</strong></summary>

* **GEval**: General evaluation with custom rubric
* **RubricJudge**: Multi-criteria rubric scoring
* **Pairwise**: Compare two outputs side by side
* **PromptAlignment**: Output matches prompt intent
* **TaskCompletion**: Task was completed successfully
* **PolitenessMetric**: Professional and courteous tone
* **ToneMetric**: Tone matches target (formal, casual, etc.)
* **ClarityMetric**: Clear and easy to understand

</details>

<details>

<summary><strong>RAG (Retrieval-Augmented Generation)</strong></summary>

* **Faithfulness**: Answer is grounded in retrieved context
* **AnswerRelevancy**: Answer addresses the question
* **ContextPrecision**: Retrieved context is relevant
* **ContextRecall**: Retrieved context contains the answer

</details>

<details>

<summary><strong>Safety</strong></summary>

* **PII**: Detects personally identifiable information leakage
* **Toxicity**: Offensive or harmful content
* **PromptInjection**: Attempts to manipulate system behavior
* **Hallucination**: Confident but factually wrong statements
* **Bias**: Unfair treatment of groups

</details>

<details>

<summary><strong>Agent / tool use</strong></summary>

* **ToolCorrectness**: Called tools with correct arguments
* **ToolSelectionAccuracy**: Selected the right tools
* **PlanAdherence**: Followed the expected plan

</details>

<details>

<summary><strong>Conversation (multi-turn)</strong></summary>

* **GoalAccuracy**: Achieved conversation goal
* **RoleAdherence**: Stayed in character
* **KnowledgeRetention**: Remembered context from earlier turns
* **TopicAdherence**: Stayed on topic

</details>

<details>

<summary><strong>Operational</strong></summary>

* **LatencyMetric**: Response time within threshold
* **TokenCostMetric**: Token usage within budget
* **CostEfficiencyMetric**: Cost per quality point

</details>

Go to [Metrics](/ai-evals/core-concepts/metrics.md) for the complete catalog with examples.

***

## 3. LLM providers (judge models)

For LLM-as-a-judge metrics, configure a Judge LLM Connector in your Harness project:

| Provider             | Models                              | Setup Guide                                                                                            |
| -------------------- | ----------------------------------- | ------------------------------------------------------------------------------------------------------ |
| **OpenAI**           | GPT-4, GPT-4 Turbo, GPT-3.5 Turbo   | [OpenAI connector](https://developer.harness.io/docs/platform/harness-ai/openai-model-connector)       |
| **Anthropic**        | Claude 3 Opus, Sonnet, Haiku        | [Anthropic connector](https://developer.harness.io/docs/platform/harness-ai/anthropic-model-connector) |
| **Azure OpenAI**     | GPT-4, GPT-3.5 Turbo (Azure-hosted) | [Azure OpenAI connector](https://developer.harness.io/docs/platform/harness-ai/openai-model-connector) |
| **Google Vertex AI** | Gemini Pro, Gemini Ultra            | Contact support                                                                                        |

***

## 4. Dataset formats

### Upload methods

| Method                   | Format                   | Use Case                              |
| ------------------------ | ------------------------ | ------------------------------------- |
| **JSONL upload**         | One JSON object per line | Bulk import from existing test suites |
| **Inline creation**      | Add rows via UI form     | Quick testing with 5-10 cases         |
| **Synthetic generation** | LLM-powered strategies   | Generate edge cases and variations    |
| **API creation**         | REST API or MCP          | Programmatic dataset management       |

### Synthetic generation strategies

* **Adversarial**: Inputs designed to trick or confuse the agent
* **Rephrase**: Rephrase existing rows in different ways
* **Complexity ladder**: Progressively harder versions of a base case
* **Use case expansion**: Generate cases covering a range of scenarios

### Required fields

```json
{
  "input": {"content": "What is 2+2?"},
  "expected_output": {"text": "4"},
  "expected_tools": ["calculator"],
  "context": ["Math textbook page 42..."]
}
```

***

## 5. CI/CD integrations

### Harness Pipelines (native)

* **AI Eval step**: Native step type in CI/CD pipelines
* **Suite-based gating**: Block deployments on eval suite results
* **Pass strategies**: `all_must_pass` (blocks on any failure), `weighted_threshold` (advisory signal)

### Python SDK (any CI system)

The `harness-evals` SDK runs in any environment that supports Python 3.10+:

{% tabs %}
{% tab title="pytest" %}

```python
from harness_evals import Evaluation

def test_my_agent():
    eval = Evaluation.from_registry("my-eval")
    results = eval.run()
    assert results.pass_rate >= 0.8
```

{% endtab %}

{% tab title="GitHub Actions" %}

```yaml
- name: Run AI Evals
  run: |
    pip install harness-evals
    harness-evals run --suite my-suite
```

{% endtab %}

{% tab title="GitLab CI" %}

```yaml
test:
  script:
    - pip install harness-evals
    - harness-evals run --suite my-suite
```

{% endtab %}
{% endtabs %}

**Works in:** GitHub Actions, GitLab CI, Jenkins, CircleCI, Travis CI, Bitbucket Pipelines, any system that runs Python

***

## 6. Observability integrations

### Harness AgentTrace

Native integration for production trace ingestion:

* **Trace ingestion**: OpenTelemetry spans via OTLP protocol
* **Trace viewer**: Inspect LLM calls, tool invocations, costs, latency
* **Promote to dataset**: Convert production traces into test cases
* **Online evaluation**: Score sampled production traffic

### Third-party tracing

AI Evals ingests traces from:

* **Langfuse** (via OTLP export)
* **Custom OTLP exporters** (any OpenTelemetry-compatible system)

***

## 7. API and tooling access

### REST API

Full CRUD operations for all resources:

* **Targets, Datasets, Metrics, Metric Sets**: Create and manage evaluation components
* **Evaluations, Runs, Suites**: Execute and monitor evaluation runs
* **Authentication**: Via Harness PAT (Personal Access Token)

Go to [REST API Quickstart](/ai-evals/get-started/quickstart-api.md) for examples.

### MCP (Model Context Protocol)

Drive evals from AI assistants:

* **Supported clients**: Claude Code, Cursor, any MCP client
* **Requirements**: `harness-mcp-v2` v2.9.7+ with `ai-evals` toolset enabled

Go to [MCP Quickstart](/ai-evals/get-started/quickstart-mcp.md) for setup.

***

## 8. Beta features and roadmap

### Beta features

The following features are in beta. Contact [Harness Support](mailto:support@harness.io) to enable them for your account:

* **Private endpoint targets**: Evaluate agents behind VPC or firewall
* **Custom metric SDK**: Write custom metrics in Python
* **Multi-region datasets**: Store datasets in EU or APAC regions

### Recently released

Go to the AI Evals release notes for the latest updates.

### Roadmap

Go to the AI Evals roadmap for planned features.

***

## Next steps

* Go to [Get Started](/ai-evals/get-started/get-started.md) to build your first evaluation.
* Go to [Core concepts](/ai-evals/core-concepts/targets.md) to understand the data model.
* Contact [Harness Support](mailto:support@harness.io) for platform questions or to request new integrations.
