> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/resilience-testing/chaos-testing/probes/probe-template-library/kubernetes/pod-replica-count-check.md).

# Pod Replica Count Check

Pod Replica Count Check is a built-in Command Probe template that validates whether a Kubernetes workload keeps at least a minimum number of healthy replicas during a chaos experiment. Use it to confirm that a Deployment, StatefulSet, or DaemonSet maintains availability while a fault removes or disrupts pods. You select the workload by label, by name, or by the owning workload kind and namespace.

The probe runs the `healthchecks` utility bundled in the chaos probe image, queries the Kubernetes API, and prints `[Pass]` when the workload has at least `MINIMUM_HEALTHY_REPLICA_COUNT` healthy replicas. The comparator marks the probe as passed when the output contains `[Pass]`.

{% hint style="info" %}
**BUILT-IN PROBE TEMPLATE**

This is a built-in Command Probe template that runs on Kubernetes chaos infrastructure. Add it to an experiment from the probe library and customize its inputs. Go to [Built-in probe templates](/resilience-testing/chaos-testing/probes/probe-templates.md) to browse the full library, or go to [Command probe](/resilience-testing/chaos-testing/probes/command-probe.md) to understand how command probes work.
{% endhint %}

***

### Use cases <a href="#use-cases" id="use-cases"></a>

Use this probe template to:

* Verify that deployments maintain the desired replica count.
* Validate auto-scaling behavior during load chaos.
* Monitor application availability during pod failures.
* Confirm high availability during chaos experiments.

***

### How the probe works <a href="#how-the-probe-works" id="how-the-probe-works"></a>

The template configures a Command Probe that runs `healthchecks -name validate-pod-replica`. The utility resolves the target workloads from `TARGET_LABELS`, `TARGET_NAMES`, `TARGET_KIND`, and `TARGET_NAMESPACE`, queries the Kubernetes API, and prints `[Pass]` when the number of healthy replicas is at or above `MINIMUM_HEALTHY_REPLICA_COUNT`. The comparator passes the probe when the output contains `[Pass]`, and fails it otherwise.

***

### Prerequisites <a href="#prerequisites" id="prerequisites"></a>

* **Chaos infrastructure:** A Kubernetes chaos infrastructure installed in the target cluster.
* **Namespace access:** Access to the target namespace and resources.
* **RBAC permissions:** Permissions for the chaos service account to query resource status.

***

### Probe properties <a href="#probe-properties" id="probe-properties"></a>

#### Command <a href="#command" id="command"></a>

```bash
healthchecks -name validate-pod-replica
```

#### Comparator <a href="#comparator" id="comparator"></a>

| Type   | Criteria | Value    |
| ------ | -------- | -------- |
| string | contains | `[Pass]` |

The probe passes when the command output contains `[Pass]`, which indicates that the workload has at least the minimum required healthy replicas.

#### Environment variables <a href="#environment-variables" id="environment-variables"></a>

| Variable                        | Description                                                                                       | Required | Default      |
| ------------------------------- | ------------------------------------------------------------------------------------------------- | -------- | ------------ |
| `TARGET_LABELS`                 | Comma-separated list of labels used to filter resources (for example, `app=nginx,tier=frontend`). | No       | -            |
| `TARGET_NAMES`                  | Comma-separated list of target resource names.                                                    | No       | -            |
| `TARGET_NAMESPACE`              | Namespace of the target resources.                                                                | Yes      | -            |
| `TARGET_KIND`                   | Kind of the target resource (for example, `deployment`, `statefulset`, `daemonset`).              | No       | `deployment` |
| `MINIMUM_HEALTHY_REPLICA_COUNT` | Minimum healthy replica count required for the target.                                            | No       | `1`          |
| `STATUS_CHECK_TIMEOUT`          | Maximum time in seconds to wait for the status check.                                             | No       | `180`        |
| `STATUS_CHECK_DELAY`            | Delay in seconds between status checks.                                                           | No       | `2`          |

***

### Run properties <a href="#run-properties" id="run-properties"></a>

| Property          | Description                                                                      | Type    | Default |
| ----------------- | -------------------------------------------------------------------------------- | ------- | ------- |
| `timeout`         | Maximum time to wait for the probe to complete (for example, `30s`, `1m`, `5m`). | String  | `180s`  |
| `interval`        | Time between probe executions (for example, `1s`, `5s`, `10s`).                  | String  | `1s`    |
| `attempt`         | Number of retry attempts before the probe is marked as failed.                   | Integer | `1`     |
| `pollingInterval` | Time between retry attempts (for example, `1s`, `5s`, `10s`).                    | String  | -       |
| `initialDelay`    | Initial delay before the probe starts (for example, `0s`, `10s`, `30s`).         | String  | -       |
| `stopOnFailure`   | Stop the experiment if the probe fails.                                          | Boolean | `false` |
| `verbosity`       | Log verbosity level (`info`, `debug`, `trace`).                                  | String  | -       |

***

### Troubleshooting <a href="#troubleshooting" id="troubleshooting"></a>

<details>

<summary>Pod Replica Count Check probe fails because the workload dropped below the minimum</summary>

The number of healthy replicas is below MINIMUM\_HEALTHY\_REPLICA\_COUNT, which usually means the injected fault removed more pods than the workload could replace in time. Inspect the workload with kubectl get deployment or kubectl describe to confirm replica counts, check whether the cluster has capacity to schedule replacements, and increase STATUS\_CHECK\_TIMEOUT so the probe waits for recovery.

</details>

<details>

<summary>Pod Replica Count Check probe fails because no resources matched the target</summary>

The selectors did not resolve any workloads. Confirm that TARGET\_LABELS, TARGET\_NAMES, and TARGET\_NAMESPACE match the workload, and that TARGET\_KIND matches the resource type (deployment, statefulset, or daemonset). An empty match is treated as a failure.

</details>

<details>

<summary>Pod Replica Count Check probe fails with a forbidden or RBAC error</summary>

The chaos service account does not have permission to read the target workload in the namespace. Grant get and list on the relevant resource type (deployments, statefulsets, or daemonsets) and on pods for the chaos service account, then rerun the experiment.

</details>

***

### Related probe templates <a href="#related-probe-templates" id="related-probe-templates"></a>

* [Pod Status Check](/resilience-testing/chaos-testing/probes/probe-template-library/kubernetes/pod-status-check.md): Validate that pods stay in the Running state.
* [Pod Startup Time Check](/resilience-testing/chaos-testing/probes/probe-template-library/kubernetes/pod-startup-time-check.md): Validate that pods start within a duration.
* [Built-in probe templates](/resilience-testing/chaos-testing/probes/probe-templates.md): Browse the full probe template library.
