> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/ai-evals/core-concepts/metrics.md).

# Metrics

Explore the 47+ built-in metrics across 9 categories and learn when to use each one.

**Metrics** define how AI Evals scores your outputs. Harness AI Evals ships with 47+ built-in metrics across 9 categories.

***

## What you will learn

* The 12 metric categories and when to use each
* Heuristic vs LLM-judge metrics
* How to configure thresholds and weights
* Which metrics require which dataset fields

***

## Metric categories

| Category          | Metrics                                                               | Use Case                                     |
| ----------------- | --------------------------------------------------------------------- | -------------------------------------------- |
| **Deterministic** | ExactMatch, Contains, Regex, NumericDiff                              | Fast, free scoring for exact matching        |
| **Structural**    | JsonDiff, SchemaValidation                                            | API response validation                      |
| **Similarity**    | BLEU, Levenshtein, EmbeddingSimilarity                                | Fuzzy matching and semantic similarity       |
| **LLM judge**     | GEval, RubricJudge, Pairwise, PromptAlignment                         | Qualitative scoring (tone, style, relevance) |
| **RAG**           | Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall        | Retrieval-augmented generation quality       |
| **Safety**        | PII, Toxicity, PromptInjection, Hallucination, Bias                   | Security and safety checks                   |
| **Agent / tool**  | TaskCompletion, ToolCorrectness, ToolSelectionAccuracy, PlanAdherence | Agentic system evaluation                    |
| **Conversation**  | GoalAccuracy, RoleAdherence, KnowledgeRetention, TopicAdherence       | Multi-turn dialogue quality                  |
| **Operational**   | LatencyMetric, TokenCostMetric, CostEfficiencyMetric                  | Performance and cost tracking                |

***

## Heuristic vs LLM-judge metrics

### Heuristic metrics

Deterministic, rule-based scoring. No LLM calls required.

**Advantages:**

* Free to run (no LLM judge cost)
* Fast (milliseconds per item)
* Reproducible (same input always produces same score)

**Examples:**

* `ExactMatchMetric`: Output exactly matches expected
* `ContainsMetric`: Output contains expected substring
* `RegexMetric`: Output matches a regular expression

**When to use:**

* Validating structured outputs (JSON, API responses)
* Checking for required keywords or patterns
* Cost-sensitive evaluations

### LLM-judge metrics

Use an LLM to score qualitative dimensions like tone, relevance, and helpfulness.

**Advantages:**

* Can evaluate subjective qualities (politeness, tone, clarity)
* Handles natural language variation
* No need to write brittle regex patterns

**Trade-offs:**

* Costs per judgment (token usage)
* Slower (seconds per item)
* Non-deterministic (slight score variation across runs)

**Examples:**

* `TaskCompletionMetric`: Did the agent complete the task?
* `AnswerRelevancyMetric`: Is the response on-topic?
* `PolitenessMetric`: Is the tone professional and courteous?

**When to use:**

* Evaluating tone, style, or subjective quality
* Complex reasoning that heuristics cannot capture
* When dataset lacks exact expected outputs

***

## Configuring metrics

### Threshold

The minimum score required for a row to pass.

**Example:**

```json
{
  "metric_id": "task-completion-uuid",
  "threshold": 0.8,
  "weight": 1.0
}
```

* Score >= 0.8: Pass
* Score < 0.8: Fail

**How to set thresholds:**

1. Run a baseline evaluation with default thresholds (0.5-0.7)
2. Review per-item scores to see where false positives/negatives occur
3. Adjust thresholds based on acceptable error rates
4. Use calibration to auto-derive thresholds from a baseline run

### Weight

Relative importance of this metric in the aggregate score.

**Example:**

```json
[
  {"metric_id": "faithfulness", "threshold": 0.9, "weight": 2.0},
  {"metric_id": "latency", "threshold": 0.5, "weight": 1.0}
]
```

* Faithfulness counts 2x as much as latency in the aggregate
* Use weights to prioritize critical metrics over nice-to-haves

***

## Common metric combinations

### Customer support agent

* **TaskCompletionMetric** (0.7): Did the agent answer the question?
* **PolitenessMetric** (0.8): Was the tone professional?
* **ContainsMetric** (1.0): Response mentions required keywords

### RAG pipeline

* **Faithfulness** (0.85): Answer is grounded in retrieved context
* **AnswerRelevancy** (0.8): Answer addresses the question
* **ContextRecall** (0.7): Retrieved context contains the answer

### API agent

* **ToolSelectionAccuracy** (0.8): Calls the correct tools
* **JsonDiffMetric** (0.9): API response matches expected structure
* **LatencyMetric** (0.5): Completes within acceptable time

***

## Next steps

* Go to [Get Started](/ai-evals/get-started/get-started.md) to configure your first metric set and run an evaluation.
