Overview
Introduction to Harness Chaos Engineering
Chaos Engineering is the practice of proactively introducing controlled faults into applications and infrastructure to test the resilience of business services. Developers, QA teams, performance engineers, and Site Reliability Engineers (SREs) run chaos experiments to measure system resilience and discover weaknesses before they impact production.
Harness Chaos Engineering provides end-to-end tooling for resilience testing at enterprise scale through proven chaos engineering principles.
What you will learn from this topic
This topic introduces the capabilities, use cases, and deployment modes of Harness Chaos Engineering.
Core capabilities
Harness Chaos Engineering provides the following capabilities:
Chaos experiments: Run chaos experiments with more than 200 built-in faults, probes, and actions. These cover Kubernetes, cloud platforms, Linux, Windows, and application runtimes.
Resilience probes: Use probes to programmatically observe the expected behavior or steady state. Integrate with application performance monitoring (APM) tools and applications.
Actions: Perform custom tasks within a chaos experiment. Use actions for notifications, webhooks, and load-testing scripts.
Enterprise governance: Use ChaosGuard to control who runs experiments, where, and when.
Centralized chaos execution plane: Use scalable architecture with centralized execution and distributed agents through Harness Delegate.
Connectors: Integrate with CI/CD pipelines, monitoring tools, and cloud service providers.
AI-powered capabilities: Use the AI Reliability Agent for experiment-creation, optimization, and failure-resolution recommendations.
MCP tools: Use Harness MCP server tools from AI editors, such as Claude Desktop, Windsurf, and Cursor.
GameDay portal: Run controlled production experiments to validate incident-response procedures and system recovery.
The platform includes RBAC, single sign-on (SSO), comprehensive logging, and audit capabilities. It is available in SaaS and on-premises deployments. Go to Harness Platform key concepts to understand general Harness Platform concepts and features.
Use cases
Harness Chaos Engineering supports the following use cases:
Resilience testing in deployment pipelines: Add chaos experiments to deployment pipelines for continuous resilience validation.
Load testing with resilience testing: Run chaos experiments with load-testing tools under traffic stress.
GameDay exercises: Run controlled production tests to validate incident-response procedures and recovery.
Disaster recovery testing: Validate backups, failover mechanisms, and recovery procedures through fault injection.
Deployment modes
Choose one of the following deployment modes:
SaaS: Use a managed cloud service with automatic updates and scaling.
On-premises: Deploy in your infrastructure for complete control.
Chaos fault library
Browse more than 200 ready-to-use chaos faults across your infrastructure.
Go to Chaos Faults to browse the fault library.
New Chaos Studio
Last updated
Was this helpful?