
Infrastructure Incident Postmortem & Root Cause Analysis
Transform infrastructure incidents into actionable root causes
What You Can Do
Claude systematically analyzes production incidents to identify root causes, contributing factors, and systemic vulnerabilities in your infrastructure. You'll receive structured postmortems that guide your team through incident timelines, pinpoint what actually failed, and prioritize infrastructure improvements that prevent recurrence. The skill also detects patterns across your incident history—revealing repeated failures that point to deeper architectural weaknesses.
Features
Convert raw logs, alerts, and team observations into a precise, timestamped sequence showing exactly when the incident began, escalated, and resolved.
Apply structured RCA frameworks (5 Whys, fishbone diagram, fault tree analysis) to trace surface symptoms back to underlying infrastructure or process failures.
Distinguish between direct causes and contributing conditions—missing monitoring, outdated runbooks, alert gaps, configuration drift—that enabled the failure.
Produce comprehensive postmortem documents with executive summary, timeline, RCA findings, action items, and lessons learned—ready to share across teams.
Automatically rank follow-up actions by impact and urgency: quick fixes, infrastructure hardening, process improvements, and monitoring enhancements.
Generate tailored communication for different audiences—executive summary for leadership, technical deep-dive for engineers, and transparent customer-facing explanation.
Identify recurring failure patterns across your incident backlog, revealing systemic infrastructure gaps (e.g., connection pooling, memory leaks, cascading failures).
Analyze MTTR trends and recommend monitoring, runbook, or architectural changes to reduce recovery windows for similar future incidents.
Example Output
Incident Postmortem: Database Connection Pool Exhaustion
Incident ID: INC-2026-0715-001
Duration: 14:23–15:47 UTC (84 minutes)
Impact: API timeout errors affecting 1–3% of requests; ~15,000 failed transactions
Timeline
| Time (UTC) | Event |
|---|---|
| 14:23 | Traffic spike detected (+280% RPS above baseline) |
| 14:25 | Database connection pool utilization reaches 95% |
| 14:27 | First API timeout errors appear in logs |
| 14:31 | On-call SRE paged; 4-minute response delay |
| 14:45 | Root cause identified: connection pool size misconfigured |
| 15:12 | Temporary fix deployed (increased pool from 20 to 80 connections) |
| 15:47 | Permanent fix + monitoring deployed; traffic normalized |
Root Cause
Primary: Postgres connection pool size (20 concurrent connections) was incompatible with post-microservice architecture. Migration doubled concurrent client count without revisiting pool configuration.
Contributing Factors
- Alert gap: No monitoring on connection pool utilization; team discovered issue only after timeout cascade
- Load testing mismatch: Staging tests simulated throughput (RPS) but not concurrent connection patterns
- Runbook drift: On-call runbook referenced pre-migration database architecture; initial troubleshooting wasted 8 minutes
- Configuration blind spot: No automated validation that pool sizing matched concurrent client estimates
Action Items
| Priority | Action | Owner | Target Date |
|---|---|---|---|
| P0 | Deploy CloudWatch alarm: connection pool utilization > 80% | Platform | 2026-08-01 |
| P0 | Update connection pool sizing formula based on client concurrency model | Database SRE | 2026-08-07 |
| P1 | Add connection pool stress test to CI/CD pipeline | QA Engineer | 2026-08-14 |
| P1 | Audit similar services for identical misconfiguration | Platform Lead | 2026-08-05 |
| P2 | Revise on-call runbook with current database topology | SRE | 2026-08-10 |
Lessons Learned
- Load testing must simulate concurrent connection patterns, not just throughput. A service under high RPS is not the same as a service with high concurrent clients.
- Configuration drift happens silently. Add health-check validations that alert if actual connection pool size doesn't match predicted concurrency.
- MTTR depends on alert quality. Implementing pool utilization monitoring would have surfaced this in 30 seconds instead of 4 minutes.
What's Included
- RCA Framework Templates: Structured prompts for 5 Whys, fishbone diagrams, and fault trees—Claude selects the best framework for your incident type.
- Postmortem Document Template: Pre-formatted sections (timeline, RCA, contributing factors, action items, lessons learned) that guide your team through structured incident analysis.
- Stakeholder Communication Guides: Ready-to-customize templates for executive summaries, customer-facing incident explanations, and technical deep-dives for different audiences.
- Contributing Factors Checklist: Systematic checklist covering monitoring gaps, alert blindness, runbook accuracy, load testing coverage, and configuration management—ensures nothing is overlooked.
- Pattern Detection Prompts: Guided analysis to spot recurring failure modes across your incident history and identify systemic infrastructure vulnerabilities.
Who It's For
- Site Reliability Engineers (SREs)
- Infrastructure and DevOps Engineers
- Platform and Systems Engineers
- Engineering Managers and Tech Leads
- On-Call Incident Commanders and Responders
Best For
- Writing production incident postmortems after outages
- Identifying root causes of system failures and cascading incidents
- Analyzing incident timelines and recovery processes
- Prioritizing infrastructure hardening and preventive actions
- Communicating incident impact to executives and customers







