
Postmortem Report Writer
Generate blameless postmortem reports following Google SRE culture
What You Can Do
You can convert raw incident information—timelines, logs, metrics, and witness accounts—into professional postmortem reports that follow Google SRE blameless culture principles. The skill structures complex incidents into clear narratives covering what happened, why it happened, impact analysis, and actionable follow-up items, enabling your team to learn from failures and prevent recurrence.
Features
frames incidents as systemic failures, not individual mistakes
organizes events chronologically with exact timestamps and decision points
documents customer-facing and internal effects with metrics
identifies underlying causes using 5-why methodology or equivalent
surfaces process gaps, monitoring blind spots, and design limitations
generates specific, trackable remediation items with owners and deadlines
follows industry-standard postmortem structure and cultural guidelines
produces polished documents suitable for leadership review and team discussion
Example Output
Example 1: Database Failover Incident
Incident Title: Primary Database Failover Delay (45-minute customer impact)
Timeline:
- 14:22 UTC: Disk IO saturation on primary database detected
- 14:24 UTC: Automated health checks fail; failover initiated
- 14:28 UTC: Secondary replica promoted; DNS update propagation begins
- 14:67 UTC: 95% of traffic restored; final clients reconnect
Root Cause: Monitoring thresholds set too high to catch IO saturation early. Failover runbook was undocumented, causing 3-minute delay in decision-making.
Contributing Factors: No load testing in 6 months; cache invalidation strategy created unexpected query surge; secondary replica 45 seconds behind primary.
Follow-ups:
- Implement predictive IO monitoring with earlier alerts (Owner: Platform team, Due: 2 weeks)
- Automate failover decision at 80% IO utilization (Owner: SRE, Due: 3 weeks)
Example 2: API Rate Limiting Bug
Incident Title: Unintended 429 errors for legitimate traffic (72-minute outage)
Root Cause: Rate limiter configuration deployed with typo: per-user limits applied per-IP instead. Legitimate traffic from shared NAT hit limits immediately.
Contributing Factors: Rate limiter change deployed without load test; no staging environment matches production traffic patterns; missing validation in deployment pipeline.
Follow-ups:
- Add integration tests for rate limiter edge cases (Owner: Backend team)
- Require load test sign-off for traffic-facing changes (Owner: Release mgmt)
What's Included
- SKILL.md instruction file with blameless culture principles and incident severity guidelines:
- Postmortem template covering timeline, impact, root cause, contributing factors, and follow-ups:
- Blameless writing checklist to catch blame language and reframe incidents systemically:
- 5-Why analysis framework to dig deeper into root causes:
- Follow-up action tracker with ownership, severity, and deadline fields:
Who It's For
- SRE/DevOps Engineers — leading incident response and postmortem writing
- Engineering Managers — reviewing postmortems and driving follow-up actions
- Technical Leads — documenting incidents and building institutional knowledge
- On-call Responders — capturing incident details immediately after resolution
- Quality/Process Leads — tracking systemic improvements across incidents
Best For
- Production incidents causing customer-facing downtime or data impact (>15 min)
- Near-misses that could have caused major incidents without detection
- Process failures requiring organizational learning and systemic change
- Recurring incidents with similar root causes across different teams
- High-complexity incidents involving multiple systems, teams, or contributing factors







