SkillsLib.ai

Postmortem Report Writer

Generate blameless postmortem reports following Google SRE culture

4.2(46 reviews)
100+ downloads
Updated Oct 2026
Verified SafeSecurity VerifiedThis skill was analyzed by our AI security scanner for harmful content including data exfiltration, system manipulation, credential theft, and prompt injection. No threats were detected.

What You Can Do

You can convert raw incident information—timelines, logs, metrics, and witness accounts—into professional postmortem reports that follow Google SRE blameless culture principles. The skill structures complex incidents into clear narratives covering what happened, why it happened, impact analysis, and actionable follow-up items, enabling your team to learn from failures and prevent recurrence.

Features

Blameless narrative writing

frames incidents as systemic failures, not individual mistakes

Timeline reconstruction

organizes events chronologically with exact timestamps and decision points

Impact quantification

documents customer-facing and internal effects with metrics

Root cause analysis

identifies underlying causes using 5-why methodology or equivalent

Contributing factors

surfaces process gaps, monitoring blind spots, and design limitations

Actionable follow-ups

generates specific, trackable remediation items with owners and deadlines

Google SRE compliance

follows industry-standard postmortem structure and cultural guidelines

Stakeholder-ready format

produces polished documents suitable for leadership review and team discussion

Example Output

Example 1: Database Failover Incident

Incident Title: Primary Database Failover Delay (45-minute customer impact)

Timeline:

  • 14:22 UTC: Disk IO saturation on primary database detected
  • 14:24 UTC: Automated health checks fail; failover initiated
  • 14:28 UTC: Secondary replica promoted; DNS update propagation begins
  • 14:67 UTC: 95% of traffic restored; final clients reconnect

Root Cause: Monitoring thresholds set too high to catch IO saturation early. Failover runbook was undocumented, causing 3-minute delay in decision-making.

Contributing Factors: No load testing in 6 months; cache invalidation strategy created unexpected query surge; secondary replica 45 seconds behind primary.

Follow-ups:

  • Implement predictive IO monitoring with earlier alerts (Owner: Platform team, Due: 2 weeks)
  • Automate failover decision at 80% IO utilization (Owner: SRE, Due: 3 weeks)

Example 2: API Rate Limiting Bug

Incident Title: Unintended 429 errors for legitimate traffic (72-minute outage)

Root Cause: Rate limiter configuration deployed with typo: per-user limits applied per-IP instead. Legitimate traffic from shared NAT hit limits immediately.

Contributing Factors: Rate limiter change deployed without load test; no staging environment matches production traffic patterns; missing validation in deployment pipeline.

Follow-ups:

  • Add integration tests for rate limiter edge cases (Owner: Backend team)
  • Require load test sign-off for traffic-facing changes (Owner: Release mgmt)

What's Included

  • SKILL.md instruction file with blameless culture principles and incident severity guidelines:
  • Postmortem template covering timeline, impact, root cause, contributing factors, and follow-ups:
  • Blameless writing checklist to catch blame language and reframe incidents systemically:
  • 5-Why analysis framework to dig deeper into root causes:
  • Follow-up action tracker with ownership, severity, and deadline fields:

Who It's For

  • SRE/DevOps Engineers — leading incident response and postmortem writing
  • Engineering Managers — reviewing postmortems and driving follow-up actions
  • Technical Leads — documenting incidents and building institutional knowledge
  • On-call Responders — capturing incident details immediately after resolution
  • Quality/Process Leads — tracking systemic improvements across incidents

Best For

  • Production incidents causing customer-facing downtime or data impact (>15 min)
  • Near-misses that could have caused major incidents without detection
  • Process failures requiring organizational learning and systemic change
  • Recurring incidents with similar root causes across different teams
  • High-complexity incidents involving multiple systems, teams, or contributing factors

You might also like

Database Performance Tuning Analyzer
$45
Database Performance Tuning Analyzer

You can systematically diagnose database performance bottlenecks by sharing your schema, slow query logs, and execution plans with Claude. It identifies root causes—missing indexes, inefficient joins, lock contention—and provides prioritized recommendations with ready-to-implement SQL. Skip the manual log analysis and get tuning strategies tailored to your workload.

Database Performance Tuning Analyst
$30
Database Performance Tuning Analyst

Use Claude to systematically analyze your database queries, execution plans, and schema to identify performance bottlenecks. The skill generates actionable optimization recommendations with SQL rewrites, index strategies, and configuration tuning. You'll receive detailed before-and-after performance analysis to validate improvements and prioritize work by impact.

IoT Firmware Analysis & Device Debugger
$40
IoT Firmware Analysis & Device Debugger

Rapidly analyze firmware logs and diagnose hardware issues that cause device failures, connectivity problems, and performance degradation. You'll identify root causes from stack traces, crash dumps, and sensor data, then generate specific optimization recommendations. This skill transforms raw device logs into actionable debugging plans that reduce time-to-resolution from hours to minutes.

Injectable Formulation Development & Troubleshooting
$40
Injectable Formulation Development & Troubleshooting

You'll develop systematic approaches to injectable formulation design, from API selection through sterilization strategy. Claude helps you troubleshoot failed batches by analyzing root causes, recommends regulatory pathways (505(b)(2), ANDA, NDA), and provides science-backed solutions for stability, compatibility, and manufacturability challenges.

Mobile Feature Architecture & Implementation
$40
Mobile Feature Architecture & Implementation

You'll design and implement mobile features with architectural rigor, cross-platform considerations, and edge-case handling built-in. This skill generates complete system designs, platform-specific implementation strategies, performance optimization approaches, and testing frameworks. The output is production-ready guidance spanning iOS and Android with security, offline resilience, and deployment strategies included.

Git Commit Message Writer
$45
CI/CD4.3(47)
Git Commit Message Writer

Claude analyzes your code diffs and generates standardized commit messages that follow the Conventional Commits specification. The skill automatically determines the correct commit type, scope, and description based on the changes you've made, ensuring your messages are parseable by automation tools while remaining human-readable for code reviewers.

Structured NLP Analysis and Annotation with Claude
$35
NLP3.3(6)
Structured NLP Analysis and Annotation with Claude

You can transform raw text into structured, labeled datasets for machine learning, analysis, and research. This skill performs named entity recognition, sentiment classification, part-of-speech tagging, and dependency parsing—generating consistent, validated annotations at scale. Use it to prepare corpora, extract entities, classify documents, or perform linguistic analysis without manual annotation.

ROS Control Architecture & Debugging
$30
ROS Control Architecture & Debugging

You can architect multi-node ROS control systems from scratch, including node design patterns, communication flows, and real-time constraints. You'll debug complex node interactions using publisher/subscriber analysis, service call tracing, and action server diagnostics. You can optimize motion controllers through PID tuning, trajectory planning validation, and performance profiling to achieve precise, responsive robotic behavior.

$45.00