Pod application function error
Inject a configurable error into a specific function of an instrumented application running in a Kubernetes pod so you can test how callers and dependents handle the failure.
Pod application function error is a Kubernetes pod-level chaos fault that causes a specific function in an instrumented application to raise a configurable error for a configurable duration. Only the named function is affected; other application code paths run normally. When the fault ends, the function returns to its normal behavior immediately.
Use this fault to validate how callers and dependents behave when a specific business function starts failing: a third-party SDK that throws on a code path, a wrapper that emits a custom exception, or a side-effecting operation that suddenly returns an error.
Use cases
Run this fault when you want to answer concrete questions like:
Third-party SDK failure simulation: Inject errors into a wrapper around a payment or auth SDK and confirm callers fall back or queue retries.
Core-function failure modes: When a critical business function raises an unexpected error, does the request fail closed, or does partial work leak through?
Retry and fallback validation: Do retry budgets respect maximum attempts, or do retries amplify load on downstream services?
Circuit breaker behavior: Does a circuit breaker around the failing function open after the configured failure threshold and short-circuit subsequent calls?
Observability coverage: Does the failure surface in dashboards, traces, and alerts as expected, or does it slip past monitoring?
Prerequisites
Kubernetes version: 1.21 or later. Go to What's supported to confirm distribution support.
Target pod is Running: The application pod is in the
Runningstate.Application is instrumented: The application registers a name and exposes the target function to the chaos infrastructure. Without instrumentation, the chaos pod cannot reach the function.
Function is identifiable: The function to fail is reachable by the name set in
TARGET_APPLICATION_FUNCTION.Workload selector defined: The chaos experiment knows the target application by name.
Supported environments
Amazon EKS
Supported
Azure AKS
Supported
Google GKE
Supported
Red Hat OpenShift
Supported
Rancher
Supported
VMware Tanzu
Supported
Self-managed Kubernetes (CNCF-certified)
Supported
GKE Autopilot
Supported (no privileged access required)
EKS Fargate, ACI virtual nodes
Supported (no container runtime socket required)
Permissions required
The fault runs under the chaos infrastructure's service account.
Resource (apiGroup)
Verbs
Why it is needed
pods ("")
get, list, create, delete, deletecollection, patch, update
Discover the target pod and run the chaos pod
pods/log ("")
get, list, watch
Stream chaos pod logs for status and debugging
deployments, statefulsets, replicasets, daemonsets (apps)
get, list
Resolve the target workload to the pods it owns
events ("")
get, list, create, patch, update
Record fault progress as Kubernetes events
jobs (batch)
get, list, create, delete, deletecollection
Run the chaos job that drives the fault
The default Harness chaos infrastructure service account already includes these permissions.
Fault tunables
Configure the following fault parameters when you add Pod application function error to an experiment in Chaos Studio. Defaults are shown for reference.
Required parameters
TARGET_APPLICATION_NAME
Name of the target application as registered with the chaos infrastructure.
(required)
TARGET_APPLICATION_FUNCTION
Name of the function inside the target application to inject the error into.
(required)
Chaos parameters
MESSAGE
Error message attached to the injected failure. Empty uses a default message.
""
TOTAL_CHAOS_DURATION
Duration of the fault in seconds.
60
RAMP_TIME
Wait period in seconds before and after the fault. Go to ramp time to read how it is applied.
0
Tunables that apply to every fault are documented in common tunables for all faults.
Fault execution in brief
Signals the instrumented application to make the function named in TARGET_APPLICATION_FUNCTION raise an error containing MESSAGE for TOTAL_CHAOS_DURATION seconds.
Expected behavior during fault execution
Calls to the named function fail with the configured error message. Other functions in the same application run normally.
Direct callers of the function surface the failure as exceptions, error responses, or queued retries depending on the language and framework.
Downstream services may see reduced or absent traffic if the failing function fronted upstream calls.
Error dashboards and traces should show the injected error message alongside any cascading failures.
Signals to watch
Attach resilience probes to assert each layer:
Application error rate: Use an HTTP probe against endpoints that exercise the function to detect 4xx/5xx spikes.
Function-level metrics: Use a Prometheus probe on the function's error counter or success rate to confirm the failure injection.
Application logs: Use a command probe to grep container logs for the configured
MESSAGE.
Verify the fault execution effect
While the experiment is running, confirm the function is failing:
Exercise the function from a client.
The response should reflect the failure, either as an HTTP error or an error payload.
Confirm the failure surfaces in logs.
The configured
MESSAGEshould appear with each failed invocation.
Recovery and cleanup
End of duration: The function returns to its normal behavior automatically.
Abort the experiment: Stopping the experiment from Chaos Studio triggers the same cleanup path.
Stuck state: If the application caches the failure (for example by tripping a circuit breaker that does not reset on its own), restart the pod to clear the cached state.
Limitations
Instrumentation required: The fault only affects applications that have registered themselves and their functions with the chaos infrastructure. Uninstrumented applications cannot be targeted.
Function-name granularity: Only one function at a time is targeted. Use multiple experiments in sequence for multi-function scenarios.
No payload control: This fault raises an error from the function; it does not modify return values or arguments. Use Pod JVM modify return for JVM-specific return-value modification.
Troubleshooting
Related faults
Pod application function latency: Inject latency into a function instead of failing it.
Pod JVM method exception: JVM-specific method-level exception injection.
Pod JVM modify return: Modify the return value of a JVM method instead of raising an exception.
Common pod fault tunables: Shared environment variables for selecting target pods and workloads.
Last updated
Was this helpful?