SRE Incident Management Framework
Rapid incident response, severity triage, and blameless postmortems
What You Can Do
You use this skill to transform chaotic incident response into a structured, repeatable process. It guides you through severity classification, helps you document incident timelines accurately, and generates comprehensive post-mortem reports that identify root causes without blame. This framework reduces mean time to recovery (MTTR), improves team communication during crises, and ensures every incident becomes a learning opportunity.
Features
Automatically triage incidents by impact (P1 critical through P4 low) using objective criteria
Structured prompts to capture events chronologically with accurate timestamps and impact assessment
Clear logic for when to wake up executives, declare SEV-2, or involve external teams
Standardized format for blameless postmortems with root cause analysis and action items
Pre-written templates for status updates, stakeholder notifications, and all-hands reports
5-Why technique, failure mode analysis, and system resilience assessment
Capture MTTR, detection time, and recovery details for trend analysis
Ensure no stakeholder is left out during communication phases
Example Output
Example 1: Severity Classification
Severity: P2 (High)
Impact: Database connection pool exhaustion affecting 15% of user traffic
Detection Time: 3 minutes (automated alerts)
Escalation: Page on-call engineer + notify platform lead
Expected Recovery: 30-45 minutes
Notify: VP Eng, customer success team
Example 2: Incident Timeline
14:23 UTC — Alert fired: DB connection pool at 95%
14:25 UTC — On-call investigates, identifies runaway query
14:27 UTC — Query killed, pool normalizes to 40%
14:30 UTC — Incident declared resolved
14:32 UTC — Status page updated, customers notified
Example 3: Post-Mortem Summary
## Root Cause
Unoptimized query in reporting job after schema migration.
## Action Items
- [x] Add query performance monitoring to CI/CD
- [ ] Re-optimize 3 similar queries by Aug 15
- [ ] Implement automatic connection pool scaling
What's Included
- SKILL.md: Full incident response framework and decision flows
- Severity Classification Guide: Criteria and examples for each severity level
- Incident Timeline Template: Structured format for capturing event sequence
- Post-Mortem Template: Blameless review format with root cause and action items
- Communication Templates: Status update, stakeholder notification, all-hands scripts
- Escalation Decision Tree: When to page, declare SEV-2, involve executives
- Root Cause Analysis Toolkit: 5-Why, fishbone, and resilience assessment prompts
Who It's For
- SRE Engineers — Structured framework for all incident types
- On-Call Engineers — Quick-reference severity guide and communication templates
- DevOps Leads — Post-mortem ownership and trend analysis
- Infrastructure Teams — Root cause analysis and system resilience improvement
- Tech Leads — Cross-team incident coordination and executive communication
Best For
- Production incident response — Triage and respond to live incidents
- Post-incident reviews — Blameless postmortems with actionable improvements
- Severity classification — Objective triage without guesswork
- Cross-team communication — Consistent updates for stakeholders at all levels
- Root cause analysis — Structured techniques to find systemic issues, not just symptoms






