
Chaos Engineering Experiment Designer
Design chaos experiments with controlled risks and validated hypotheses
What You Can Do
You'll create comprehensive chaos engineering experiment plans that include hypothesis statements, precisely defined blast radius limits, and systematic risk mitigation strategies. Claude helps you validate your assumptions about system resilience before executing potentially disruptive tests, ensuring experiments are safe, focused, and generate actionable insights.
Features
structured approach to defining hypotheses, scope, and success criteria
precise boundaries and scope limits to contain experiment impact
identify, quantify, and prioritize potential failure scenarios
define measurable criteria for verifying system behavior assumptions
staged execution plans from canary to production-scale tests
automatic rollback and mitigation steps with timing and dependencies
specify observation points, thresholds, and data collection methods
automated incident response playbooks with decision trees
Example Output
Experiment Design Document:
Hypothesis: A 30-minute outage in the cache layer
will not impact customer transactions
Blast Radius:
- Scope: Cache nodes in us-west-2 only
- Impact: ~2% of traffic, non-critical features
- Duration: 5 minutes (strict timeout)
Success Criteria:
✓ Transaction success rate stays >99%
✓ Error rate <0.5%
✓ Automatic failover triggers within 10s
Risk Assessment Matrix:
Scenario | Probability | Severity | Mitigation
Unplanned propagation | Medium | High | Kill switch, regional isolation
Data corruption | Low | Critical | Snapshot restore, validation
Customer-facing impact | Low | High | Real-time alerting, abort criteria
Recovery Procedures:
1. Monitor metrics (0-30s) → abort if any threshold breached
2. Trigger automated failover (30-60s)
3. Validate data consistency (60-120s)
4. Gradual traffic restoration (2-5min)
5. Post-incident review documentation
What's Included
- SKILL.md: Complete chaos engineering framework and methodology
- Experiment Design Template: Structured worksheet with hypothesis, scope, and success criteria
- Risk Assessment Matrix: Identify and quantify failure scenarios with impact/likelihood
- Recovery Procedures Checklist: Automated rollback and mitigation step sequences
- Rollout Strategy Guide: Canary → staging → production execution roadmap
- Hypothesis Validation Framework: Measurement criteria and validation methods
Who It's For
- Site Reliability Engineers (SREs) — validate system resilience and design safe experiments
- DevOps Engineers — plan infrastructure failure testing with confidence
- Cloud/Platform Architects — assess system design assumptions before incidents occur
- Infrastructure Engineering Teams — coordinate chaos testing across distributed systems
- Technical Program Managers — build organizational confidence in system reliability
Best For
- Designing controlled chaos engineering experiments
- Validating system resilience and failover assumptions
- Planning and executing failure mode testing
- Assessing blast radius and containment strategies
- Building incident response runbooks and recovery procedures







