
Incident Response & Postmortem Facilitator
Structure incident response and postmortems with guided facilitation and RCA automation
What You Can Do
Turn chaotic incident response into a structured, documented process. Claude guides your team through real-time triage decisions, collaboratively builds accurate root cause analysis, and generates complete postmortem documentation with tracked action items. From alert to closed ticket, you'll have auditable incident records and accountability.
Features
Guides teams through severity assessment, impact estimation, escalation decisions, and immediate response steps with structured decision trees.
Structures root cause analysis using Ishikawa diagrams, timeline reconstruction, failure mode analysis, and targeted questioning to uncover systemic issues.
Produces polished incident reports with executive summary, detailed timeline, impact metrics, findings, and compliance-ready formatting.
Creates prioritized action items with assigned owners, due dates, dependencies, and verification criteria for follow-up and closure.
Categorizes incidents by type, severity, and affected systems to enable trend analysis and pattern recognition across your infrastructure.
Maintains non-punitive language throughout analysis, focusing on system improvements rather than individual accountability.
Helps teams build accurate chronological narratives of what happened, when it happened, and who was involved at each stage.
Ensures documentation meets regulatory requirements, internal policies, and creates searchable incident records for future reference.
Example Output
Example 1: Triage Output
- 🚨 SEVERITY: Critical (SEV-1)
- ⏱️ DECISION: Activate full incident response team
- 📢 ESCALATE TO: VP of Engineering + Database on-call
- 🎯 IMMEDIATE ACTIONS:
1. Page database team immediately
2. Begin customer communications
3. Start incident timeline log
4. Establish war room on Slack #incident-202407
- 🔍 IMPACT ESTIMATE:
- Affected customers: ~2,300
- Revenue impact: ~$4,200/minute
- Estimated duration: 10-30 minutes
Example 2: RCA Session Output
## Root Cause Analysis
**Primary Root Cause:** Connection pool exhaustion in primary database due to cascading query timeouts in reporting service.
**Failure Timeline:**
1. 14:32 - Reporting service deployed with unoptimized query
2. 14:33 - Query timeout begins blocking connection pool
3. 14:35 - API service depletes backup connections
4. 14:36 - User traffic starts failing; alerts fire
5. 14:52 - Database team restarts connection pool
6. 14:55 - Services recover; traffic normalizes
**Contributing Factors:**
- Reporting query not tested against production scale
- Connection pool monitoring wasn't granular enough
- No circuit breaker on reporting service queries
Example 3: Postmortem Action Items
## Action Items (Prioritized)
1. **[P0] Query optimization review** — Reporting team — Due: Aug 3
- Verify all queries use indexes correctly
- Add query timeout limits
2. **[P1] Implement circuit breaker** — Platform team — Due: Aug 10
- Prevent timeouts in one service from cascading
- Configure graceful degradation
3. **[P1] Improve connection pool monitoring** — DBA — Due: Aug 5
- Add per-service connection tracking
- Create alerts at 70% and 90% utilization
What's Included
- Real-time triage facilitator: Interactive workflow for incident severity assessment, escalation decisions, and response planning during active incidents.
- RCA workshop moderator: Structured questioning framework and analysis templates to guide teams through collaborative root cause investigation.
- Postmortem template generator: Creates polished, compliance-ready incident reports with automatic formatting and required section structure.
- Action item tracker: Manages follow-up tasks with ownership, prioritization, due dates, and verification checkpoints for continuous improvement.
- Incident classification system: Categorizes and tags incidents for trend analysis, pattern recognition, and historical incident correlation.
- Compliance and audit checklist: Ensures all incident documentation meets regulatory and internal requirements with complete audit trails and traceability.
Who It's For
- SRE and Infrastructure Engineers
- On-call Incident Commanders
- DevOps and Platform Engineering Teams
- Infrastructure and Systems Engineering Leaders
- Site Reliability Engineering Managers
Best For
- Post-incident postmortem analysis and documentation
- Real-time incident triage and severity assessment
- Facilitating cross-functional root cause analysis sessions
- Standardizing incident response processes across teams
- Building incident trends and identifying preventive measures







