
Incident Response
Execute structured incident response from detection through blameless post-mortems
What You Can Do
You can execute a complete incident response workflow that minimizes time-to-recovery and captures learning. The skill guides you through detection and triage, coordinates the right responders based on severity (S1–S4), coordinates mitigation actions, and produces blameless post-mortems with follow-up items within 48 hours. This ensures production incidents are handled consistently and drive continuous improvement.
Features
quickly determine urgency and resource allocation
centralized documentation of all critical details and timeline events
assemble the right team based on severity and assign an Incident Commander
prioritize service restoration (rollback, scaling, feature flags) before investigation
structured post-mortem template with blameless language and timeline
capture preventive and corrective actions tied to each incident
templates for status updates during active incidents and post-resolution reporting
Example Output
Incident Record (INC-0847)
- Severity: S2 | Duration: 23 minutes | IC: Sarah Chen
- Impact: Payment processing failures affecting 12% of transactions
- Timeline: Alert at 14:05 UTC → IC assigned 14:08 → Rollback deployed 14:18 → Service restored 14:28
- Root Cause: Database connection pool exhaustion from recent query optimization change
- Follow-ups: [✓ Add connection pool monitoring] [✓ Implement query performance regression tests] [✓ Update deployment runbook]
Post-Mortem Summary
- What happened: Query optimization intended to improve performance instead created N+1 queries under load
- Why it happened: Load testing environment didn't replicate production transaction volume
- What we're doing: Automated performance regression testing in CI/CD + alerting on connection pool utilization above 80%
What's Included
- incident-response SKILL.md: full lifecycle framework with severity table and coordinator responsibilities
- Incident Record Template: structured markdown for capturing S1–S4 incident details, timeline, and impact
- Post-Mortem Template: blameless RCA format with timeline, root cause, contributing factors, and action items
- Severity Classification Guide: decision tree for assigning S1–S4 with response time SLAs
- Communication Templates: status update and resolution notice templates for stakeholders
Who It's For
- Site Reliability Engineers (SREs) — managing production incidents and improving system reliability
- DevOps Engineers — coordinating incident response and tracking infrastructure issues
- Engineering Managers — ensuring incidents are documented and actionable improvements are tracked
- On-Call Rotations — incident commanders and responders need structured guidance during active incidents
- Platform/Infrastructure Teams — responding to and learning from system-level outages
Best For
- Production outages and service degradation requiring rapid response
- Critical infrastructure failures needing responder coordination
- Post-incident root cause analysis and blameless post-mortems
- Building incident response culture and organizational learning
- Incident tracking and follow-up item management across teams







