> For the complete documentation index, see [llms.txt](https://developer.harness.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://developer.harness.io/resilience-testing/chaos-engineering/faults/chaos-fault-categories/azure/azure-web-app-stop.md).

# Azure web app stop

Azure web app stop is an Azure chaos fault that stops one or more App Service web apps listed in `AZURE_WEB_APP_NAMES` (in `RESOURCE_GROUP`, subscription `AZURE_SUBSCRIPTION_ID`) for `TOTAL_CHAOS_DURATION` seconds, then starts them again. While stopped, the web app returns `403 Web App Stopped` to all incoming requests.

Use this fault to test how clients behave when an App Service web app is unavailable: whether dependent services degrade gracefully, whether Traffic Manager / Front Door fail traffic over to a healthy region, and whether monitoring detects the outage within the alerting SLA.

{% hint style="info" %}
**RUN YOUR FIRST EXPERIMENT**

If you have not configured the chaos infrastructure yet, go to [Quickstart](/resilience-testing/chaos-engineering/new-to-chaos-engineering/quickstart.md) to install the chaos infrastructure and run an experiment end to end.
{% endhint %}

***

### Use cases <a href="#use-cases" id="use-cases"></a>

Run this fault when you want to answer concrete questions like:

* **Web app unavailable:** When the web app stops, do dependent services degrade gracefully?
* **Traffic Manager / Front Door:** Does traffic shift to a healthy region inside the failover SLA?
* **Slot-aware deployments:** With deployment slots configured, does the staging slot take production traffic correctly?
* **Monitoring fidelity:** Do alerts on `Microsoft.Web/sites/Availability` and HTTP-error metrics fire within the alerting SLA?

***

### Prerequisites <a href="#prerequisites" id="prerequisites"></a>

* **Kubernetes version:** 1.21 or later for the chaos infrastructure cluster.
* **Target web apps reachable:** Each entry in `AZURE_WEB_APP_NAMES` exists in `RESOURCE_GROUP`.
* **Web app in `Running` state:** The fault refuses to stop a web app that is already `Stopped`.
* **Azure credentials available:** Service principal File Secret, workload identity, or managed identity.
* **RBAC granted:** The principal includes the role listed below.

***

### Supported environments <a href="#supported-environments" id="supported-environments"></a>

| Platform                             | Support status                               |
| ------------------------------------ | -------------------------------------------- |
| App Service (Windows or Linux plans) | Supported                                    |
| App Service slots (production only)  | Supported                                    |
| Function Apps                        | Supported (same Microsoft.Web actions apply) |
| App Service Environment (ASE) v3     | Supported                                    |

***

### Permissions required <a href="#permissions-required" id="permissions-required"></a>

The Azure principal used by the chaos pod needs the following role on the target resource group or subscription.

**Recommended built-in role:** `Website Contributor`

**Custom role (minimum actions):**

```json
{
  "Name": "Harness Chaos Web App Stop",
  "Actions": [
    "Microsoft.Web/sites/read",
    "Microsoft.Web/sites/start/action",
    "Microsoft.Web/sites/stop/action"
  ],
  "AssignableScopes": ["/subscriptions/<SUBSCRIPTION_ID>"]
}
```

Go to [Azure fault permissions](/resilience-testing/chaos-engineering/faults/chaos-fault-categories/azure/security-configurations/fault-permissions.md) to read the full permission catalog.

***

### Authentication <a href="#authentication" id="authentication"></a>

Go to [Azure authentication methods](/resilience-testing/chaos-engineering/faults/chaos-fault-categories/azure/security-configurations/azure-authentication-methods.md) to set up Service principal, Workload identity, or Managed identity.

***

### Fault tunables <a href="#fault-tunables" id="fault-tunables"></a>

Configure the following fault parameters when you add Azure web app stop to an experiment in Chaos Studio. Defaults are shown for reference.

**Required parameters**

| Tunable               | Description                                | Default    |
| --------------------- | ------------------------------------------ | ---------- |
| `AZURE_WEB_APP_NAMES` | Comma-separated list of web app names.     | (required) |
| `RESOURCE_GROUP`      | Resource group that contains the web apps. | (required) |

**Chaos parameters**

| Tunable                | Description                                                                                                                                                                                                      | Default    |
| ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- |
| `TOTAL_CHAOS_DURATION` | Total duration of the fault in seconds. The web apps stay stopped for this period.                                                                                                                               | `30`       |
| `CHAOS_INTERVAL`       | Delay in seconds between successive iterations when running for more than one cycle.                                                                                                                             | `30`       |
| `SEQUENCE`             | Order in which multiple web apps are stopped: `parallel` or `serial`.                                                                                                                                            | `parallel` |
| `RAMP_TIME`            | Wait period in seconds before and after the fault. Go to [ramp time](/resilience-testing/chaos-engineering/faults/chaos-fault-categories/common-tunables-for-all-faults.md#ramp-time) to read how it is applied. | `0`        |

**Authentication**

| Tunable                       | Description                                                                                           | Default |
| ----------------------------- | ----------------------------------------------------------------------------------------------------- | ------- |
| `AZURE_SUBSCRIPTION_ID`       | Target Azure subscription ID.                                                                         | `""`    |
| `AZURE_CLIENT_ID`             | Client ID of a user-assigned managed identity.                                                        | `""`    |
| `AZURE_AUTHENTICATION_SECRET` | Identifier of the **File Secret in Harness Secret Manager** that contains the service principal JSON. | `""`    |

Tunables that apply to every fault are documented in [common tunables for all faults](/resilience-testing/chaos-engineering/faults/chaos-fault-categories/common-tunables-for-all-faults.md).

***

### Fault execution in brief <a href="#fault-execution-in-brief" id="fault-execution-in-brief"></a>

Calls the Azure App Service `stop` API on each web app in `AZURE_WEB_APP_NAMES` (in `RESOURCE_GROUP`), waits for `TOTAL_CHAOS_DURATION` seconds, then calls the `start` API to bring them back.

***

### Expected behavior during fault execution <a href="#expected-behavior-during-fault-execution" id="expected-behavior-during-fault-execution"></a>

* Each affected web app returns `403 Web App Stopped` (or `404 Site Not Available`) for the duration.
* Inbound HTTP/HTTPS connections to the web app are immediately closed.
* Traffic Manager / Front Door health probes start failing on the affected backends and traffic shifts.
* Azure Monitor `Microsoft.Web/sites/Availability` drops to 0.
* After the duration ends, the web apps restart; cold-start time depends on the runtime stack.

{% hint style="info" %}
**WHEN THE FAULT ENDS**

The chaos pod calls `start` on every targeted web app. The web app is reachable as soon as Azure marks it `Running`, typically within 5-30 seconds.
{% endhint %}

#### Signals to watch <a href="#signals-to-watch" id="signals-to-watch"></a>

* **Web app availability:** Use an [HTTP probe](/resilience-testing/chaos-engineering/use-chaos-engineering/probes/http-probe.md) on the web app URL and assert the expected error status during the chaos window.
* **Failover:** If Traffic Manager / Front Door is in front, use an HTTP probe on the public endpoint and assert it stays available.

***

### Verify the fault execution effect <a href="#verify-the-fault-execution-effect" id="verify-the-fault-execution-effect"></a>

1. **Inspect web app state.**

   ```bash
   az webapp show --resource-group <rg> --name <webapp> --query state
   ```

   The state should be `Stopped` during the chaos window and `Running` afterwards.
2. **Hit the web app URL.**

   ```bash
   curl -i https://<webapp>.azurewebsites.net
   ```

   The response should be `403 Web App Stopped` during the chaos window.

***

### Recovery and cleanup <a href="#recovery-and-cleanup" id="recovery-and-cleanup"></a>

* **End of duration:** The chaos pod calls `start` on every web app.
* **Abort the experiment:** Stopping the experiment from Chaos Studio also calls `start`.
* **Manual recovery:** Run `az webapp start --resource-group <rg> --name <webapp>` for any web app that stayed stopped.

***

### Limitations <a href="#limitations" id="limitations"></a>

* **Resource group scope:** All entries in `AZURE_WEB_APP_NAMES` must be in `RESOURCE_GROUP`.
* **Deployment slots:** The fault stops the production slot. To target a non-production slot, manage the slot separately.
* **Cold start:** Restart time depends on the runtime stack and app size; budget at least 30 seconds for large apps.

***

### Troubleshooting <a href="#troubleshooting" id="troubleshooting"></a>

<details>

<summary>Azure web app stop fails with AuthorizationFailed in Harness Chaos Engineering</summary>

The Azure principal is missing Microsoft.Web/sites/stop/action or Microsoft.Web/sites/start/action. Assign Website Contributor (or a custom role with these actions) on the target resource group or subscription.

</details>

<details>

<summary>Web app did not return 403 during the chaos window</summary>

Verify the web app actually stopped with az webapp show -g -n --query state. If state is Running, the chaos pod's stop call may have failed (check fault pod logs). Some custom error pages or front-end caches may also mask the 403.

</details>

<details>

<summary>Web app stayed Stopped after the experiment ended</summary>

Run az webapp start --resource-group --name to start it manually. Inspect Azure Activity Log for failed start actions to root-cause why the chaos pod could not start it.

</details>

***

### Related faults <a href="#related-faults" id="related-faults"></a>

* [Azure web app access restrict](/resilience-testing/chaos-engineering/faults/chaos-fault-categories/azure/azure-web-app-access-restrict.md): Block traffic with an access restriction rule instead of stopping the web app.
* [Azure instance stop](/resilience-testing/chaos-engineering/faults/chaos-fault-categories/azure/azure-instance-stop.md): Stop a VM hosting the workload instead of an App Service web app.
