
Production Incident Post-Mortem Framework
Systematize incident investigation and extract actionable learnings from production failures
What You Can Do
You can systematically investigate production incidents, document root causes, and generate comprehensive post-mortems that capture lessons learned and preventive measures. The skill guides you through incident timeline reconstruction, stakeholder interviews, impact analysis, and action item prioritization. Your team gains structured insights that reduce mean time to resolution (MTTR) and prevent similar incidents from recurring.
Features
Chronologically order events with precise timestamps, identify key decision points, and distinguish facts from assumptions to build an accurate sequence of what happened.
Apply structured investigation techniques (5 Whys, fault trees, causal maps) to uncover underlying systemic issues rather than surface symptoms.
Measure incident severity: customers affected, revenue impact, data exposure, service downtime, and organizational reputation risk.
Generate role-specific interview prompts for engineers, ops teams, managers, and customer-facing staff to gather complete incident perspectives.
Transform findings into prioritized remediation steps with clear ownership, deadlines, and success criteria to drive organizational change.
Identify recurring failure modes and systemic vulnerabilities across multiple incidents to address root systemic issues.
Design safeguards, monitoring improvements, alerting thresholds, and architectural changes that reduce recurrence risk.
Generate professionally formatted post-mortem reports suitable for internal audits, executive review, and regulatory compliance.
Example Output
Incident Post-Mortem: Database Connection Pool Exhaustion (2026-07-28)
Timeline:
- 14:23 UTC: Surge in API requests from marketing campaign
- 14:25 UTC: Database connection pool reaches 95% capacity
- 14:27 UTC: Application begins returning 503 Service Unavailable
- 14:31 UTC: On-call engineer paged
- 14:45 UTC: Connection pool manually reset; service restored
- 15:00 UTC: Load returned to baseline
Root Cause: Application code was not properly closing idle database connections; under sustained traffic, the pool exhausted. Load balancer sent traffic to all backend instances simultaneously without graceful degradation.
Impact: 18 minutes downtime, 45K failed API requests, ~2% of daily revenue lost, zero data loss.
Action Items:
- Implement connection pool monitoring with alerts at 75% capacity (P0, due 2026-08-04)
- Add connection timeout validation to CI/CD tests (P1, due 2026-08-11)
- Design circuit breaker for graceful service degradation (P1, due 2026-08-18)
What's Included
- Incident Classification System: Severity levels (P1–P4), incident categories (infrastructure, application, data, vendor), and classification framework for consistent triage.
- Root Cause Analysis Checklist: Multi-layer investigation templates with decision trees, fault tree diagrams, and guidance for distinguishing direct causes from systemic contributors.
- Post-Mortem Report Template: Professional structure including executive summary, incident timeline, root cause findings, impact analysis, action items, and lessons learned sections.
- Interview Question Bank: Role-specific prompts for engineers, on-call responders, managers, and customer-facing teams to gather diverse perspectives and context.
- Action Item Tracker: Template for converting findings into actionable improvements with priority levels, ownership, deadlines, and success metrics.
- Lessons Learned Framework: Structured approach to capturing preventive insights, design improvements, and cultural learnings that reduce future incident likelihood.
Who It's For
- Engineering Managers
- Site Reliability Engineers (SRE)
- DevOps Engineers
- Platform Engineers
- Tech Leads
Best For
- Production outage post-analysis
- Data incident investigations
- Customer-impacting service failures
- System reliability improvements
- Incident trend analysis and prevention







