> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/resilience-testing/chaos-engineering/faults/chaos-fault-categories/chaos-faults-reference.md).

# Chaos Faults

Chaos faults are the failures injected into the chaos infrastructure as part of a chaos experiment. Every fault is associated with a target resource, and you can customize the fault using the fault tunables, which you can define as part of the Chaos Experiment CR and Chaos Engine CR.

The fault execution is triggered when the chaos engine resource is created. Typically, the chaos engine is embedded within the **steps** of a chaos fault. However, you can also create the chaos engine manually, and the chaos operator reconciles this resource and triggers the fault execution.

You can customize a fault execution by changing the tunables (or parameters). Some tunables are common across all the faults (for example, **chaos duration**), and every fault has its own set of tunables: default and mandatory ones. You can update the default tunables when required and always provide values for mandatory tunables (as the name suggests).

## Fault Status <a href="#fault-status" id="fault-status"></a>

Fault status indicates the current status of the fault executed as a part of the chaos experiment. A fault can have 0, 1, or more associated [probes](/resilience-testing/chaos-engineering/use-chaos-engineering/probes/index.md). Other steps in a chaos experiment include resource creation and cleanup.

In a chaos experiment, a fault can be in one of six different states. It transitions from **running**, **stopped** or **skipped** to **completed**, **completed with error** or **error** state.

* **Running**: The fault is currently being executed.
* **Stopped**: The fault stopped after running for some time.
* **Skipped**: The fault skipped, that is, the fault is not executed.
* **Completed**: The fault completes execution without any **failed** or **N/A** probe statuses.
* **Completed with Error**: When the fault completes execution with at least one **failed** probe status but no **N/A** probe status, it is considered to be **completed with error**.
* **Error**: When the fault completes execution with at least one **N/A** probe status, it is considered to be **error** because you can't determine if the probe status was **passed** or **failed**. A fault is considered to be in an **error** state when it has 0 probes because there are no health checks to validate the sanity of the chaos experiment.

## Fault Categories <a href="#fault-categories" id="fault-categories"></a>

Harness Chaos Engineering provides a comprehensive library of pre-built chaos faults organized by target infrastructure and platform. Below are tables with links to individual fault documentation for easy navigation.

<table data-column-title-hidden data-view="cards"><thead><tr><th align="center"></th><th align="center"></th><th align="center"></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td align="center"><img src="/files/ZWYWuwYcdKpR2VAkpCx7" alt="" data-size="original"></td><td align="center"><strong>AWS</strong></td><td align="center">Chaos faults for AWS</td><td><a href="/pages/956TbSBJOAieGsCqVHzF">/pages/956TbSBJOAieGsCqVHzF</a></td></tr><tr><td align="center"><img src="/files/8SoIzTfyJ5nbRyXhKvmB" alt="" data-size="original"></td><td align="center"><strong>Azure</strong></td><td align="center">Chaos faults for Azurwe</td><td><a href="/pages/So6xND5WaMqPMDNykjXE">/pages/So6xND5WaMqPMDNykjXE</a></td></tr><tr><td align="center"><img src="/files/tt06yOm6HYOFqhzuyBu7" alt="" data-size="line"></td><td align="center"><strong>Cloud Foundry</strong></td><td align="center">Chaos faults for Cloud Foundry</td><td><a href="/pages/PDd9YaUpkTY1WsaqK5qC">/pages/PDd9YaUpkTY1WsaqK5qC</a></td></tr><tr><td align="center"><img src="/files/8cVPKEF7kjcBi6VdSPKW" alt="" data-size="original"></td><td align="center"><strong>GCP</strong></td><td align="center">Chaos faults for GCP</td><td><a href="/pages/LYdH2wPOfMr2x41jc3wL">/pages/LYdH2wPOfMr2x41jc3wL</a></td></tr><tr><td align="center"><img src="/files/GeaFZh3sduF9toMFz4VJ" alt="" data-size="original"></td><td align="center"><strong>Kube-Resilience</strong></td><td align="center">Chaos faults for Kube-resilience</td><td><a href="/pages/4frzya3VyfyVFrQAOEfk">/pages/4frzya3VyfyVFrQAOEfk</a></td></tr><tr><td align="center"><img src="/files/w5pnENxWzcrHHgWKK1pZ" alt="" data-size="original"></td><td align="center"><strong>Kubernetes</strong></td><td align="center">Chaos faults for Kubernetes</td><td><a href="/pages/0emEoFvIAXn0VDP04nst">/pages/0emEoFvIAXn0VDP04nst</a></td></tr><tr><td align="center"><img src="/files/Gb5wiqhA0FG4lZIZxMU2" alt="" data-size="original"></td><td align="center"><strong>Linux</strong></td><td align="center">Chaos faults for Linux</td><td><a href="/pages/YHyBT08KCETT8PtvSGqN">/pages/YHyBT08KCETT8PtvSGqN</a></td></tr><tr><td align="center"><img src="/files/9WKMHAh2TSjqFvMIRJuB" alt="" data-size="line"></td><td align="center"><strong>Load</strong></td><td align="center">Chaos faults for Load</td><td><a href="/pages/Pi3MI5ChngreJDWvLSOq">/pages/Pi3MI5ChngreJDWvLSOq</a></td></tr><tr><td align="center"><img src="/files/KQfgvNrW7WlwzCReiJvD" alt="" data-size="line"></td><td align="center"><strong>SSH</strong></td><td align="center">Chaos faults for SSH</td><td><a href="/pages/dfuJq0Iy2rvhraukLbs1">/pages/dfuJq0Iy2rvhraukLbs1</a></td></tr><tr><td align="center"><img src="/files/y2mt25RDZ07KliLtdiBx" alt="" data-size="original"></td><td align="center"><strong>VMware</strong></td><td align="center">Chaos faults for VMware</td><td><a href="/pages/D2F0QzoHrkTYuZzacLVn">/pages/D2F0QzoHrkTYuZzacLVn</a></td></tr><tr><td align="center"><img src="/files/pNPeguv4f4fGauocN1Jk" alt="" data-size="line"></td><td align="center"><strong>Windows</strong></td><td align="center">Chaos faults for Windows</td><td><a href="/pages/LCs988TQxh5WjzBOlLML">/pages/LCs988TQxh5WjzBOlLML</a></td></tr></tbody></table>

## Common Fault Tunables <a href="#common-fault-tunables" id="common-fault-tunables"></a>

All chaos faults share a set of common tunables that control basic execution parameters. These are provided at `.spec.experiment[*].spec.components.env` in the chaosengine.

### Duration of the Chaos <a href="#duration-of-the-chaos" id="duration-of-the-chaos"></a>

Total duration of the chaos injection (in seconds). Tune it by using the `TOTAL_CHAOS_DURATION` environment variable.

```yaml
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: engine-nginx
spec:
  engineState: "active"
  experiments:
  - name: pod-delete
    spec:
      components:
        env:
        - name: TOTAL_CHAOS_DURATION
          value: '60'
```

### Chaos Interval <a href="#chaos-interval" id="chaos-interval"></a>

The delay between each chaos iteration. Multiple iterations of chaos are tuned by setting the `CHAOS_INTERVAL` environment variable.

```yaml
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: engine-nginx
spec:
  engineState: "active"
  experiments:
  - name: pod-delete
    spec:
      components:
        env:
        - name: CHAOS_INTERVAL
          value: '15'
        - name: TOTAL_CHAOS_DURATION
          value: '60'
```

### Ramp Time <a href="#ramp-time" id="ramp-time"></a>

Period to wait before and after injecting chaos. It is in units of seconds. Tune it by using the `RAMP_TIME` environment variable.

```yaml
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: engine-nginx
spec:
  engineState: "active"
  experiments:
  - name: pod-delete
    spec:
      components:
        env:
        - name: RAMP_TIME
          value: '10'
```

### Sequence of Chaos Execution <a href="#sequence-of-chaos-execution" id="sequence-of-chaos-execution"></a>

The sequence of the chaos execution for multiple targets. Its default value is **parallel**. Tune it by using the `SEQUENCE` environment variable. It supports the following modes:

* `parallel`: The chaos is injected in all the targets at once.
* `serial`: The chaos is injected in all the targets one by one.

```yaml
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: engine-nginx
spec:
  engineState: "active"
  experiments:
  - name: pod-delete
    spec:
      components:
        env:
        - name: SEQUENCE
          value: 'parallel'
```

## Getting Started with Chaos Faults <a href="#getting-started-with-chaos-faults" id="getting-started-with-chaos-faults"></a>

### Using ChaosHub <a href="#using-chaoshub" id="using-chaoshub"></a>

1. **Browse Faults**: Explore the fault library in ChaosHub
2. **Select Fault**: Choose appropriate fault for your test scenario
3. **Configure Parameters**: Set fault-specific parameters and targets
4. **Add to Experiment**: Include fault in your chaos experiment

### Best Practices <a href="#best-practices" id="best-practices"></a>

* **Start Small**: Begin with low-impact faults in non-production environments
* **Gradual Progression**: Increase complexity and scope over time
* **Monitor Closely**: Use probes and monitoring to track system behavior
* **Document Results**: Record learnings and system improvements
* **Safety First**: Always have rollback procedures and safety mechanisms in place
