SkillsLib.ai

Postmortem Report Writer

Generate blameless postmortem reports following Google SRE culture

4.2(46 reviews)
100+ downloads
Updated Sep 2026
Verified SafeSecurity VerifiedThis skill was analyzed by our AI security scanner for harmful content including data exfiltration, system manipulation, credential theft, and prompt injection. No threats were detected.

What You Can Do

You can convert raw incident information—timelines, logs, metrics, and witness accounts—into professional postmortem reports that follow Google SRE blameless culture principles. The skill structures complex incidents into clear narratives covering what happened, why it happened, impact analysis, and actionable follow-up items, enabling your team to learn from failures and prevent recurrence.

Features

Blameless narrative writing

frames incidents as systemic failures, not individual mistakes

Timeline reconstruction

organizes events chronologically with exact timestamps and decision points

Impact quantification

documents customer-facing and internal effects with metrics

Root cause analysis

identifies underlying causes using 5-why methodology or equivalent

Contributing factors

surfaces process gaps, monitoring blind spots, and design limitations

Actionable follow-ups

generates specific, trackable remediation items with owners and deadlines

Google SRE compliance

follows industry-standard postmortem structure and cultural guidelines

Stakeholder-ready format

produces polished documents suitable for leadership review and team discussion

Example Output

Example 1: Database Failover Incident

Incident Title: Primary Database Failover Delay (45-minute customer impact)

Timeline:

  • 14:22 UTC: Disk IO saturation on primary database detected
  • 14:24 UTC: Automated health checks fail; failover initiated
  • 14:28 UTC: Secondary replica promoted; DNS update propagation begins
  • 14:67 UTC: 95% of traffic restored; final clients reconnect

Root Cause: Monitoring thresholds set too high to catch IO saturation early. Failover runbook was undocumented, causing 3-minute delay in decision-making.

Contributing Factors: No load testing in 6 months; cache invalidation strategy created unexpected query surge; secondary replica 45 seconds behind primary.

Follow-ups:

  • Implement predictive IO monitoring with earlier alerts (Owner: Platform team, Due: 2 weeks)
  • Automate failover decision at 80% IO utilization (Owner: SRE, Due: 3 weeks)

Example 2: API Rate Limiting Bug

Incident Title: Unintended 429 errors for legitimate traffic (72-minute outage)

Root Cause: Rate limiter configuration deployed with typo: per-user limits applied per-IP instead. Legitimate traffic from shared NAT hit limits immediately.

Contributing Factors: Rate limiter change deployed without load test; no staging environment matches production traffic patterns; missing validation in deployment pipeline.

Follow-ups:

  • Add integration tests for rate limiter edge cases (Owner: Backend team)
  • Require load test sign-off for traffic-facing changes (Owner: Release mgmt)

What's Included

  • SKILL.md instruction file with blameless culture principles and incident severity guidelines:
  • Postmortem template covering timeline, impact, root cause, contributing factors, and follow-ups:
  • Blameless writing checklist to catch blame language and reframe incidents systemically:
  • 5-Why analysis framework to dig deeper into root causes:
  • Follow-up action tracker with ownership, severity, and deadline fields:

Who It's For

  • SRE/DevOps Engineers — leading incident response and postmortem writing
  • Engineering Managers — reviewing postmortems and driving follow-up actions
  • Technical Leads — documenting incidents and building institutional knowledge
  • On-call Responders — capturing incident details immediately after resolution
  • Quality/Process Leads — tracking systemic improvements across incidents

Best For

  • Production incidents causing customer-facing downtime or data impact (>15 min)
  • Near-misses that could have caused major incidents without detection
  • Process failures requiring organizational learning and systemic change
  • Recurring incidents with similar root causes across different teams
  • High-complexity incidents involving multiple systems, teams, or contributing factors

You might also like

Injectable Formulation Development & Troubleshooting
$40
Injectable Formulation Development & Troubleshooting

You'll develop systematic approaches to injectable formulation design, from API selection through sterilization strategy. Claude helps you troubleshoot failed batches by analyzing root causes, recommends regulatory pathways (505(b)(2), ANDA, NDA), and provides science-backed solutions for stability, compatibility, and manufacturability challenges.

$50
Integration Architecture Assessor

This skill helps you systematically assess integration needs across your systems, design architecture patterns that scale with your organization, and identify technical and operational risks before implementation. You'll receive architecture recommendations aligned to your business constraints, clear integration roadmaps, and risk mitigation strategies that reduce deployment surprises. Get structured decision records suitable for architecture review boards and engineering teams.

Git Commit Message Writer
$45
CI/CD4.3(47)
Git Commit Message Writer

Claude analyzes your code diffs and generates standardized commit messages that follow the Conventional Commits specification. The skill automatically determines the correct commit type, scope, and description based on the changes you've made, ensuring your messages are parseable by automation tools while remaining human-readable for code reviewers.

Structured NLP Analysis and Annotation with Claude
$35
NLP3.3(6)
Structured NLP Analysis and Annotation with Claude

You can transform raw text into structured, labeled datasets for machine learning, analysis, and research. This skill performs named entity recognition, sentiment classification, part-of-speech tagging, and dependency parsing—generating consistent, validated annotations at scale. Use it to prepare corpora, extract entities, classify documents, or perform linguistic analysis without manual annotation.

Service Mesh Architecture & Troubleshooting
$25
Service Mesh Architecture & Troubleshooting

This skill helps you systematically analyze service mesh architectures, identify inter-service communication failures, and design optimal routing and security policies. You'll receive step-by-step troubleshooting guidance tailored to your mesh platform (Istio, Linkerd, Consul), configuration validation reports, and architectural recommendations that reduce latency, improve observability, and tighten security posture.

Production ML Deployment Validation & Runbook Automation
$25
Production ML Deployment Validation & Runbook Automation

This skill automates the creation of production-ready deployment validation checklists, infrastructure-as-code templates, and incident response runbooks tailored to your ML stack. You get comprehensive pre-deployment checks covering model validation, data pipeline integrity, infrastructure readiness, and monitoring setup—all customized for your specific models and cloud provider. The skill generates executable runbooks that teams can follow during incidents, including rollback procedures, failover strategies, and diagnostic commands.

Internal Developer Platform Architecture & Golden Paths
$15
Internal Developer Platform Architecture & Golden Paths

Build scalable internal developer platform architectures that reduce cognitive load and standardize workflows for your development organization. Document golden paths that guide developers through common tasks like onboarding, deployment, and troubleshooting. Map platform capabilities, integrations, and service topology to align with your engineering scale and technical strategy.

IoT Firmware Analysis & Device Debugger
$40
IoT Firmware Analysis & Device Debugger

Rapidly analyze firmware logs and diagnose hardware issues that cause device failures, connectivity problems, and performance degradation. You'll identify root causes from stack traces, crash dumps, and sensor data, then generate specific optimization recommendations. This skill transforms raw device logs into actionable debugging plans that reduce time-to-resolution from hours to minutes.

$45.00