> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/resilience-testing/chaos-testing/probes/probe-template-library/kubernetes/pod-resource-utilisation-check.md).

# Pod Resource Utilisation Check

Pod Resource Utilisation Check is a built-in Command Probe template that validates whether the CPU or memory usage of Kubernetes pods stays within a defined limit during a chaos experiment. Use it to confirm that a workload does not exceed its resource budget while a fault stresses the application. You select pods by label, by name, or by the owning workload kind and namespace.

The probe runs the `healthchecks` utility bundled in the chaos probe image, reads pod metrics from the Kubernetes Metrics API, and prints `[Pass]` when usage for the selected `METRIC_TYPE` is at or below the configured limit. The comparator marks the probe as passed when the output contains `[Pass]`.

{% hint style="info" %}
**BUILT-IN PROBE TEMPLATE**

This is a built-in Command Probe template that runs on Kubernetes chaos infrastructure. Add it to an experiment from the probe library and customize its inputs. Go to [Built-in probe templates](/resilience-testing/chaos-testing/probes/probe-templates.md) to browse the full library, or go to [Command probe](/resilience-testing/chaos-testing/probes/command-probe.md) to understand how command probes work.
{% endhint %}

***

### Use cases <a href="#use-cases" id="use-cases"></a>

Use this probe template to:

* Monitor resource usage during stress chaos experiments.
* Verify that resource limits are respected.
* Validate application performance under load.
* Confirm that pods do not exceed resource thresholds.

***

### How the probe works <a href="#how-the-probe-works" id="how-the-probe-works"></a>

The template configures a Command Probe that runs `healthchecks -name pod-resource-metrics-check`. The utility resolves the target pods from `TARGET_LABELS`, `TARGET_NAMES`, `TARGET_KIND`, and `TARGET_NAMESPACE`, reads CPU or memory usage for the container in `TARGET_CONTAINER` from the Metrics API, and prints `[Pass]` when usage for the selected `METRIC_TYPE` is at or below `CPU_LIMIT` or `MEMORY_LIMIT`. The comparator passes the probe when the output contains `[Pass]`, and fails it otherwise.

***

### Prerequisites <a href="#prerequisites" id="prerequisites"></a>

* **Chaos infrastructure:** A Kubernetes chaos infrastructure installed in the target cluster.
* **Metrics server:** The Kubernetes metrics server installed and running in the cluster.
* **Namespace access:** Access to the target namespace and pods.
* **RBAC permissions:** Permissions for the chaos service account to query pod metrics.

***

### Probe properties <a href="#probe-properties" id="probe-properties"></a>

#### Command <a href="#command" id="command"></a>

```bash
healthchecks -name pod-resource-metrics-check
```

#### Comparator <a href="#comparator" id="comparator"></a>

| Type   | Criteria | Value    |
| ------ | -------- | -------- |
| string | contains | `[Pass]` |

The probe passes when the command output contains `[Pass]`, which indicates that pod resource utilisation is at or below the acceptable limit.

#### Environment variables <a href="#environment-variables" id="environment-variables"></a>

| Variable               | Description                                                                          | Required | Default      |
| ---------------------- | ------------------------------------------------------------------------------------ | -------- | ------------ |
| `TARGET_LABELS`        | Comma-separated list of labels used to filter pods.                                  | No       | -            |
| `TARGET_NAMES`         | Comma-separated list of target pod names.                                            | No       | -            |
| `TARGET_NAMESPACE`     | Namespace of the target pods.                                                        | Yes      | -            |
| `TARGET_KIND`          | Kind of the owning workload (for example, `deployment`, `statefulset`, `daemonset`). | No       | `deployment` |
| `TARGET_CONTAINER`     | Name of the container to check resource metrics for.                                 | No       | -            |
| `METRIC_TYPE`          | Metric to check. Accepted values are `cpu` and `memory`.                             | No       | `cpu`        |
| `CPU_LIMIT`            | Maximum allowed CPU usage in millicores. The usage must be at or below this value.   | No       | `1000`       |
| `MEMORY_LIMIT`         | Maximum allowed memory usage in MB. The usage must be at or below this value.        | No       | `1024`       |
| `STATUS_CHECK_TIMEOUT` | Maximum time in seconds to wait for the status check.                                | No       | `180`        |
| `STATUS_CHECK_DELAY`   | Delay in seconds between status checks.                                              | No       | `2`          |

***

### Run properties <a href="#run-properties" id="run-properties"></a>

| Property          | Description                                                                      | Type    | Default |
| ----------------- | -------------------------------------------------------------------------------- | ------- | ------- |
| `timeout`         | Maximum time to wait for the probe to complete (for example, `30s`, `1m`, `5m`). | String  | `180s`  |
| `interval`        | Time between probe executions (for example, `1s`, `5s`, `10s`).                  | String  | `1s`    |
| `attempt`         | Number of retry attempts before the probe is marked as failed.                   | Integer | `1`     |
| `pollingInterval` | Time between retry attempts (for example, `1s`, `5s`, `10s`).                    | String  | -       |
| `initialDelay`    | Initial delay before the probe starts (for example, `0s`, `10s`, `30s`).         | String  | -       |
| `stopOnFailure`   | Stop the experiment if the probe fails.                                          | Boolean | `false` |
| `verbosity`       | Log verbosity level (`info`, `debug`, `trace`).                                  | String  | -       |

***

### Troubleshooting <a href="#troubleshooting" id="troubleshooting"></a>

<details>

<summary>Pod Resource Utilisation Check probe fails because usage exceeded the limit</summary>

CPU or memory usage rose above CPU\_LIMIT or MEMORY\_LIMIT, which can be the expected result of a stress fault. Confirm that METRIC\_TYPE matches the resource you want to bound, check live usage with kubectl top pod, and adjust the limit if the higher usage is acceptable for the workload.

</details>

<details>

<summary>Pod Resource Utilisation Check probe fails because metrics are unavailable</summary>

The probe could not read pod metrics, usually because the Kubernetes metrics server is not installed or not ready. Verify that kubectl top pod returns data, install or repair metrics-server, and confirm the chaos service account can read the metrics.k8s.io API.

</details>

<details>

<summary>Pod Resource Utilisation Check probe fails because no pods matched the target</summary>

The selectors did not resolve any pods, or TARGET\_CONTAINER did not match a container in the resolved pods. Confirm that TARGET\_LABELS, TARGET\_NAMES, TARGET\_NAMESPACE, and TARGET\_KIND identify the workload, and that TARGET\_CONTAINER matches a container name.

</details>

***

### Related probe templates <a href="#related-probe-templates" id="related-probe-templates"></a>

* [Pod Status Check](/resilience-testing/chaos-testing/probes/probe-template-library/kubernetes/pod-status-check.md): Validate that pods stay in the Running state.
* [Container Restart Check](/resilience-testing/chaos-testing/probes/probe-template-library/kubernetes/container-restart-check.md): Validate that container restart counts stay within a threshold.
* [Built-in probe templates](/resilience-testing/chaos-testing/probes/probe-templates.md): Browse the full probe template library.
