
SRE Incident Response Orchestration
Orchestrate production incidents with structured response frameworks
What You Can Do
You'll coordinate rapid incident response with structured triage, automated severity assessment, and coordinated team communication. Claude generates incident decision trees, impact timelines, and post-mortem reports that align teams on what happened, why it happened, and how to prevent recurrence.
Features
Automatically categorize incidents by business impact, customer reach, and system scope
Structured troubleshooting frameworks to guide escalation, context-switching, and resolution steps
Pre-formatted status updates, customer notifications, and team alerts for rapid stakeholder alignment
Generate blameless post-mortem templates with timeline reconstruction, root cause analysis, and action item tracking
Quantify affected services, customer blast radius, data exposure, and financial or reputational risk
Clear decision criteria for when to page on-call, engage platform teams, or invoke major incident protocol
Structured event sequencing to establish incident duration and critical failure moments
Guided questions to drive blameless root cause analysis with the engineering team
Example Output
Incident Triage Output:
INCIDENT: Payment processing API returning 503
SEVERITY: SEV-1 (customer-facing, revenue impact)
AFFECTED: 45% of transactions, ~8,000 users, $12K/min revenue loss
ESTIMATED_DURATION: 8 minutes elapsed
ACTION: Page infrastructure team lead + payment platform owner
COMMUNICATION: Send customer status page update + internal #incident channel alert
Post-Mortem Template (Generated):
## What Happened
Payment API replica set lost quorum during DB failover window (planned maintenance window overlap).
## Timeline
14:32 — DB maintenance begins; replica promoted incorrectly
14:34 — Payment service detects loss of write capability → returns 503
14:42 — Incident declared; infrastructure team engaged
14:46 — Manual promotion to standby replica completed; traffic restored
## Root Cause
Maintenance procedure lacked synchronization with load balancer health checks.
## Action Items
☐ Add pre-maintenance checklist: Verify load balancer exclusion
☐ Implement automated replica health verification
☐ Update runbook with cross-team coordination step
What's Included
- SKILL.md: Full incident response orchestration framework
- Incident Triage Template: Severity classification checklist and decision matrix
- Communication Templates: Status page updates, Slack alerts, customer notifications
- Post-Mortem Checklist: Timeline reconstruction, root cause analysis guide, action item template
- Escalation Runbook: Decision tree for severity levels and on-call paging criteria
- Impact Assessment Worksheet: Customer blast radius, revenue impact, data scope quantification
Who It's For
- Site Reliability Engineers (SREs) — Triage and coordinate incident response across distributed systems
- On-Call Engineers — Rapidly assess severity, scope, and required escalations during incidents
- DevOps/Platform Engineers — Facilitate post-mortems and document learnings for future prevention
- Incident Commanders — Manage multi-team communication and decision-making under time pressure
- Engineering Managers — Track incident patterns and drive systemic reliability improvements
Best For
- Production incident triage — Classify severity, impact scope, and immediate escalation needs in under 2 minutes
- Cross-team coordination — Generate communication templates and decision frameworks for distributed incident response
- Incident post-mortems — Reconstruct timelines, identify root causes, and track prevention action items
- Severity assessment — Quantify customer impact, revenue loss, and system scope to align stakeholders on priority
- Runbook and playbook development — Encode incident response decision criteria for recurring incident types







