SkillsLib.ai

Production Incident Post-Mortem Framework

Systematize incident investigation and extract actionable learnings from production failures

3.6(5 reviews)
100+ downloads
Updated Sep 2026

What You Can Do

You can systematically investigate production incidents, document root causes, and generate comprehensive post-mortems that capture lessons learned and preventive measures. The skill guides you through incident timeline reconstruction, stakeholder interviews, impact analysis, and action item prioritization. Your team gains structured insights that reduce mean time to resolution (MTTR) and prevent similar incidents from recurring.

Features

Incident Timeline Reconstruction

Chronologically order events with precise timestamps, identify key decision points, and distinguish facts from assumptions to build an accurate sequence of what happened.

Root Cause Analysis

Apply structured investigation techniques (5 Whys, fault trees, causal maps) to uncover underlying systemic issues rather than surface symptoms.

Impact Quantification

Measure incident severity: customers affected, revenue impact, data exposure, service downtime, and organizational reputation risk.

Stakeholder Interview Guide

Generate role-specific interview prompts for engineers, ops teams, managers, and customer-facing staff to gather complete incident perspectives.

Action Item Extraction

Transform findings into prioritized remediation steps with clear ownership, deadlines, and success criteria to drive organizational change.

Pattern Recognition

Identify recurring failure modes and systemic vulnerabilities across multiple incidents to address root systemic issues.

Prevention Roadmap

Design safeguards, monitoring improvements, alerting thresholds, and architectural changes that reduce recurrence risk.

Compliance Documentation

Generate professionally formatted post-mortem reports suitable for internal audits, executive review, and regulatory compliance.

Example Output

Incident Post-Mortem: Database Connection Pool Exhaustion (2026-07-28)

Timeline:

  • 14:23 UTC: Surge in API requests from marketing campaign
  • 14:25 UTC: Database connection pool reaches 95% capacity
  • 14:27 UTC: Application begins returning 503 Service Unavailable
  • 14:31 UTC: On-call engineer paged
  • 14:45 UTC: Connection pool manually reset; service restored
  • 15:00 UTC: Load returned to baseline

Root Cause: Application code was not properly closing idle database connections; under sustained traffic, the pool exhausted. Load balancer sent traffic to all backend instances simultaneously without graceful degradation.

Impact: 18 minutes downtime, 45K failed API requests, ~2% of daily revenue lost, zero data loss.

Action Items:

  • Implement connection pool monitoring with alerts at 75% capacity (P0, due 2026-08-04)
  • Add connection timeout validation to CI/CD tests (P1, due 2026-08-11)
  • Design circuit breaker for graceful service degradation (P1, due 2026-08-18)

What's Included

  • Incident Classification System: Severity levels (P1–P4), incident categories (infrastructure, application, data, vendor), and classification framework for consistent triage.
  • Root Cause Analysis Checklist: Multi-layer investigation templates with decision trees, fault tree diagrams, and guidance for distinguishing direct causes from systemic contributors.
  • Post-Mortem Report Template: Professional structure including executive summary, incident timeline, root cause findings, impact analysis, action items, and lessons learned sections.
  • Interview Question Bank: Role-specific prompts for engineers, on-call responders, managers, and customer-facing teams to gather diverse perspectives and context.
  • Action Item Tracker: Template for converting findings into actionable improvements with priority levels, ownership, deadlines, and success metrics.
  • Lessons Learned Framework: Structured approach to capturing preventive insights, design improvements, and cultural learnings that reduce future incident likelihood.

Who It's For

  • Engineering Managers
  • Site Reliability Engineers (SRE)
  • DevOps Engineers
  • Platform Engineers
  • Tech Leads

Best For

  • Production outage post-analysis
  • Data incident investigations
  • Customer-impacting service failures
  • System reliability improvements
  • Incident trend analysis and prevention

You might also like

Industrial Deal Analyzer & Market Evaluator
$35
Industrial Deal Analyzer & Market Evaluator

This skill evaluates industrial real estate deals by analyzing property fundamentals, comparable market data, risk factors, and investment metrics. You get detailed assessments of deal viability, market positioning, and potential returns—all structured to support your investment thesis with hard data. Whether you're analyzing a single property or comparing multiple deals, you'll receive an executive summary with a clear buy/hold/pass recommendation.

Industrial Lease Intelligence for Brokers
$30
Industrial Lease Intelligence for Brokers

Extract critical lease terms and financial obligations in minutes, identify risks and red flags that could impact deal profitability, and generate professional broker memos with negotiation strategy recommendations. You'll transform lease documents into actionable intelligence for faster, more confident decision-making.

Digital Transformation Strategy Architect
$30
Digital Transformation Strategy Architect

This skill guides you through creating executive-aligned digital transformation roadmaps that balance innovation with operational stability. You'll conduct technology maturity assessments, identify modernization priorities, and build phased implementation strategies with risk mitigation. The result is a comprehensive strategy document that stakeholders and technical teams can execute against immediately.

Frontend Engineering Manager: Code Review & Quality Framework
$15
Frontend3.3(3)
Frontend Engineering Manager: Code Review & Quality Framework

You can leverage Claude to conduct thorough code reviews, evaluate code quality, identify architectural issues, and provide actionable feedback to your frontend engineers. This skill helps you standardize code review practices across your team, catch bugs before production, and maintain consistent quality standards while scaling your engineering capacity.

Strategic Technology Assessment Framework
$35
Strategic Technology Assessment Framework

This skill guides you through systematic technology assessments using proven evaluation frameworks. You provide your current environment, business goals, and constraints—Claude produces detailed assessment reports with technology recommendations, risk analyses, implementation roadmaps, and ROI projections. Get actionable strategic advice grounded in your specific business context, not generic tech trends.

Lease Term Analysis & Negotiation Assistant
$40
Office3.7(3)
Lease Term Analysis & Negotiation Assistant

This skill transforms complex lease agreements into clear, actionable analysis. It evaluates terms against market benchmarks, identifies negotiation leverage points, and generates professional recommendation letters that strengthen your position in lease discussions.

Engineering Leadership Strategy for Scale-Ups
$45
Scale-Up4.0(5)
Engineering Leadership Strategy for Scale-Ups

You get a structured decision-making framework tailored to scale-ups navigating rapid growth. Claude helps you evaluate technical architecture choices, plan team scaling with hiring rubrics, prioritize technical debt against velocity demands, and craft engineering strategy communications for leadership. The framework balances short-term velocity with long-term maintainability.

Foreclosure Auction Specialist
$35
Foreclosure Auction Specialist

Generate legally compliant foreclosure auction packages with complete property analysis, title review, and risk assessment in minutes. Analyze distressed property values and market conditions to develop data-driven bidding strategies that maximize your ROI. Create professional auction catalogs with embedded compliance checklists and competitive bid analysis.

$35.00