> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/ai-evals/core-concepts/targets.md).

# Targets

Understand the three target types in AI Evals and how to configure each one.

A **target** is the AI system you are evaluating. AI Evals supports three target types to accommodate different system architectures.

***

## What you will learn

* The three target types: Prompt, App, and Static
* When to use each target type
* Configuration requirements for each type
* How to test targets before running evaluations

***

## Target types overview

| Type                  | What It Evaluates                      | Example Use Case                                     |
| --------------------- | -------------------------------------- | ---------------------------------------------------- |
| **Prompt** (`prompt`) | Single LLM call with a prompt template | Testing prompt engineering changes before deployment |
| **App** (`agent`)     | Any HTTP endpoint (black box)          | Evaluating a deployed agent or RAG pipeline          |
| **Static** (`static`) | Precomputed outputs from a file        | Comparing model versions offline                     |

***

## Prompt targets

Prompt targets let you evaluate a prompt template with variable inputs.

**How it works:**

* Harness calls the LLM with your system message and user input
* Each dataset row's `input` is inserted into the prompt template
* The LLM response is scored against your metrics

**Configuration:**

* **System message**: Your prompt template (e.g., "You are a helpful assistant that...")
* **LLM connector**: Reference to a configured OpenAI, Anthropic, or other LLM connector
* **Temperature**: 0-2 (controls randomness)
* **Max tokens**: Response length limit
* **Model**: Specific model version (gpt-4, claude-3-opus, etc.)

**When to use:**

* Rapid iteration on prompt engineering
* Testing system message changes locally
* Comparing prompt variations before deployment
* Experimenting with temperature and model settings

**Example:**

```json
{
  "name": "Summarization Prompt",
  "type": "prompt",
  "config": {
    "llm_connector_ref": "connector_OpenAI_xyz",
    "system_message": "Summarize the following text in 2 sentences. Be concise and factual.",
    "temperature": 0.3,
    "max_tokens": 100,
    "model": "gpt-4"
  }
}
```

***

## App targets

App targets evaluate any HTTP endpoint as a black box.

**How it works:**

* Harness sends an HTTP request to your endpoint for each dataset row
* Your app processes the input and returns a response
* The response is scored against your metrics

**Configuration:**

* **Endpoint**: Full URL (e.g., `https://my-agent.example.com/api/eval`)
* **Method**: POST, GET, PUT, etc.
* **Request template**: How to map dataset `input` to your API's expected format
* **Response path**: JSONPath to extract the output from your API response
* **Timeout**: Maximum time to wait for a response (milliseconds)
* **Headers**: Optional custom headers (auth tokens, content-type)

**When to use:**

* Evaluating deployed systems in staging or production
* Testing complex multi-step agents or RAG pipelines
* Black box evaluation where you don't control the LLM directly
* Integration testing with external APIs

**Example:**

```json
{
  "name": "Flight Agent",
  "type": "agent",
  "config": {
    "endpoint": "https://my-agent.example.com/api/eval",
    "method": "POST",
    "request_template": {
      "input": "{{input.content}}",
      "session_id": "eval-session"
    },
    "response_path": "output.text",
    "timeout_ms": 30000,
    "headers": {
      "Authorization": "Bearer {{secrets.API_KEY}}"
    }
  }
}
```

{% hint style="warning" %}
**Request template is required.** Without `request_template`, the runner will double-wrap the input as `{input: {input: ...}}` and your API will likely reject it with HTTP 400. Always specify how to map dataset fields to your API's expected format.
{% endhint %}

***

## Static targets

Static targets evaluate precomputed outputs without making any API calls.

**How it works:**

* Upload a file with precomputed outputs (one per dataset row)
* Harness scores the precomputed outputs against your metrics
* No runtime invocation, evaluation is instantaneous

**Configuration:**

* **Output file**: JSONL file where each line matches a dataset row by ID
* **Output path**: Field name containing the output (e.g., `output`, `response`)

**When to use:**

* Comparing outputs from multiple models or systems
* Scoring outputs generated externally (e.g., batch jobs)
* Offline evaluation without re-running inference
* Cost-free metric tuning and threshold calibration

**Example:**

```json
{
  "name": "GPT-4 vs Claude Comparison",
  "type": "static",
  "config": {
    "output_file": "outputs/gpt4-baseline.jsonl",
    "output_path": "response"
  }
}
```

**Output file format:**

```jsonl
{"id": "case-1", "response": "The capital of France is Paris."}
{"id": "case-2", "response": "The result is 42."}
{"id": "case-3", "response": "I don't have enough information to answer that."}
```

***

## Testing targets

Before running a full evaluation, test your target to catch configuration issues early.

**Test workflow:**

1. Create or select a target in the UI
2. Use the **Test Target** panel to send a sample input
3. Verify the output looks correct
4. Fix any issues (endpoint unreachable, wrong response path, timeout too short)
5. Save and attach to an evaluation

**Common test failures:**

* **HTTP 400**: Missing or incorrect `request_template`
* **Timeout**: Increase `timeout_ms` or optimize your endpoint
* **No output extracted**: Verify `response_path` matches your API's response structure
* **Authentication error**: Check headers and API key secrets

***

## Next steps

Now that you understand targets, explore the other building blocks:

* Go to [Datasets](/ai-evals/core-concepts/datasets.md) to learn how to structure golden test cases.
* Go to [Metrics](/ai-evals/core-concepts/metrics.md) to review the 47+ built-in metrics.
* Go to [Get Started](/ai-evals/get-started/get-started.md) to create your first target and run an evaluation.
