> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/resilience-testing/chaos-engineering/use-chaos-engineering/probes/probe-templates/kubernetes/pod-resource-utilisation-check.md).

# Pod Resource Utilisation Check

Pod resource utilisation check validates the current resource utilisation metrics of Kubernetes pods.

## Infrastructure type <a href="#infrastructure-type" id="infrastructure-type"></a>

* **Kubernetes**

## Use cases <a href="#use-cases" id="use-cases"></a>

Pod Resource Utilisation Check probe helps you:

* Monitor resource usage during stress chaos experiments
* Verify resource limits are respected
* Validate application performance under load
* Ensure pods don't exceed resource thresholds

***

## Overview <a href="#overview" id="overview"></a>

This probe validates that pod resource utilisation (CPU and memory) remains within acceptable limits during chaos experiments. It requires metrics-server to be installed in the cluster.

### Probe type <a href="#probe-type" id="probe-type"></a>

**Command Probe**

### Prerequisites <a href="#prerequisites" id="prerequisites"></a>

* Kubernetes cluster with chaos infrastructure installed
* Metrics server installed and running in the cluster
* Access to target namespace and pods
* Sufficient RBAC permissions to query pod metrics

***

## Probe properties <a href="#probe-properties" id="probe-properties"></a>

### Command <a href="#command" id="command"></a>

```
healthchecks -name pod-resource-metrics-check
```

### Comparator <a href="#comparator" id="comparator"></a>

| Type   | Criteria | Value   |
| ------ | -------- | ------- |
| string | contains | \[Pass] |

The probe passes when the command output contains `[Pass]`, indicating pod resource utilisation is within acceptable limits.

### Environment variables <a href="#environment-variables" id="environment-variables"></a>

| Variable               | Description                                                                   | Required | Default    |
| ---------------------- | ----------------------------------------------------------------------------- | -------- | ---------- |
| `TARGET_LABELS`        | Comma-separated list of target labels to filter pods.                         | No       | -          |
| `TARGET_NAMES`         | Comma-separated list of target pod names.                                     | No       | -          |
| `TARGET_NAMESPACE`     | Namespace of the target pods.                                                 | No       | -          |
| `TARGET_KIND`          | Kind of the target resource (e.g., `deployment`, `statefulset`, `daemonset`). | No       | deployment |
| `TARGET_CONTAINER`     | Name of the container to check resource metrics.                              | No       | -          |
| `METRIC_TYPE`          | Metric type to check: `cpu` or `memory`.                                      | No       | cpu        |
| `CPU_LIMIT`            | Pods should have CPU usage (in millicores) less than or equal to this value.  | No       | 1000       |
| `MEMORY_LIMIT`         | Pods should have memory usage (in MB) less than or equal to this value.       | No       | 1024       |
| `STATUS_CHECK_TIMEOUT` | Maximum time in seconds to wait for status check.                             | No       | 180        |
| `STATUS_CHECK_DELAY`   | Delay in seconds between status checks.                                       | No       | 2          |

***

## Run properties <a href="#run-properties" id="run-properties"></a>

| Property          | Description                                                              | Type    | Default |
| ----------------- | ------------------------------------------------------------------------ | ------- | ------- |
| `timeout`         | Maximum time to wait for the probe to complete (e.g., `30s`, `1m`, `5m`) | String  | 180s    |
| `interval`        | Time between probe executions (e.g., `1s`, `5s`, `10s`)                  | String  | 1s      |
| `attempt`         | Number of retry attempts before marking the probe as failed              | Integer | 1       |
| `pollingInterval` | Time between retry attempts (e.g., `1s`, `5s`, `10s`)                    | String  | -       |
| `initialDelay`    | Initial delay before starting the probe (e.g., `0s`, `10s`, `30s`)       | String  | -       |
| `stopOnFailure`   | Stop the experiment if the probe fails                                   | Boolean | false   |
| `verbosity`       | Log verbosity level (`info`, `debug`, `trace`)                           | String  | -       |

***

## Probe definition <a href="#probe-definition" id="probe-definition"></a>

You can define this probe in your chaos experiment as follows:

### CPU utilisation check <a href="#cpu-utilisation-check" id="cpu-utilisation-check"></a>

```yaml
probe:
  - name: "cpu-utilisation-check"
    type: "cmdProbe"
    mode: "Continuous"
    cmdProbe/inputs:
      command: "healthchecks -name pod-resource-metrics-check"
      comparator:
        type: "string"
        criteria: "contains"
        value: "[Pass]"
      env:
        - name: TARGET_LABELS
          value: "app=nginx"
        - name: TARGET_NAMESPACE
          value: "production"
        - name: METRIC_TYPE
          value: "cpu"
        - name: CPU_LIMIT
          value: "800"
    runProperties:
      timeout: 180s
      interval: 1s
      attempt: 1
      stopOnFailure: false
```

### Memory utilisation check <a href="#memory-utilisation-check" id="memory-utilisation-check"></a>

```yaml
probe:
  - name: "memory-utilisation-check"
    type: "cmdProbe"
    mode: "Continuous"
    cmdProbe/inputs:
      command: "healthchecks -name pod-resource-metrics-check"
      comparator:
        type: "string"
        criteria: "contains"
        value: "[Pass]"
      env:
        - name: TARGET_NAMES
          value: "my-app-pod"
        - name: TARGET_NAMESPACE
          value: "default"
        - name: TARGET_CONTAINER
          value: "app-container"
        - name: METRIC_TYPE
          value: "memory"
        - name: MEMORY_LIMIT
          value: "2048"
    runProperties:
      timeout: 180s
      interval: 2s
      attempt: 3
```
