SkillsLib.ai

Infrastructure Incident Postmortem & Root Cause Analysis

Transform infrastructure incidents into actionable root causes

3.5(6 reviews)
100+ downloads
Updated Sep 2026

What You Can Do

Claude systematically analyzes production incidents to identify root causes, contributing factors, and systemic vulnerabilities in your infrastructure. You'll receive structured postmortems that guide your team through incident timelines, pinpoint what actually failed, and prioritize infrastructure improvements that prevent recurrence. The skill also detects patterns across your incident history—revealing repeated failures that point to deeper architectural weaknesses.

Features

Incident Timeline Reconstruction

Convert raw logs, alerts, and team observations into a precise, timestamped sequence showing exactly when the incident began, escalated, and resolved.

Root Cause Identification

Apply structured RCA frameworks (5 Whys, fishbone diagram, fault tree analysis) to trace surface symptoms back to underlying infrastructure or process failures.

Contributing Factors Analysis

Distinguish between direct causes and contributing conditions—missing monitoring, outdated runbooks, alert gaps, configuration drift—that enabled the failure.

Postmortem Document Generation

Produce comprehensive postmortem documents with executive summary, timeline, RCA findings, action items, and lessons learned—ready to share across teams.

Action Item Prioritization

Automatically rank follow-up actions by impact and urgency: quick fixes, infrastructure hardening, process improvements, and monitoring enhancements.

Stakeholder Communication Templates

Generate tailored communication for different audiences—executive summary for leadership, technical deep-dive for engineers, and transparent customer-facing explanation.

Incident Pattern Detection

Identify recurring failure patterns across your incident backlog, revealing systemic infrastructure gaps (e.g., connection pooling, memory leaks, cascading failures).

Recovery Time Improvement Recommendations

Analyze MTTR trends and recommend monitoring, runbook, or architectural changes to reduce recovery windows for similar future incidents.

Example Output

Incident Postmortem: Database Connection Pool Exhaustion

Incident ID: INC-2026-0715-001
Duration: 14:23–15:47 UTC (84 minutes)
Impact: API timeout errors affecting 1–3% of requests; ~15,000 failed transactions

Timeline

Time (UTC)Event
14:23Traffic spike detected (+280% RPS above baseline)
14:25Database connection pool utilization reaches 95%
14:27First API timeout errors appear in logs
14:31On-call SRE paged; 4-minute response delay
14:45Root cause identified: connection pool size misconfigured
15:12Temporary fix deployed (increased pool from 20 to 80 connections)
15:47Permanent fix + monitoring deployed; traffic normalized

Root Cause

Primary: Postgres connection pool size (20 concurrent connections) was incompatible with post-microservice architecture. Migration doubled concurrent client count without revisiting pool configuration.

Contributing Factors

  1. Alert gap: No monitoring on connection pool utilization; team discovered issue only after timeout cascade
  2. Load testing mismatch: Staging tests simulated throughput (RPS) but not concurrent connection patterns
  3. Runbook drift: On-call runbook referenced pre-migration database architecture; initial troubleshooting wasted 8 minutes
  4. Configuration blind spot: No automated validation that pool sizing matched concurrent client estimates

Action Items

PriorityActionOwnerTarget Date
P0Deploy CloudWatch alarm: connection pool utilization > 80%Platform2026-08-01
P0Update connection pool sizing formula based on client concurrency modelDatabase SRE2026-08-07
P1Add connection pool stress test to CI/CD pipelineQA Engineer2026-08-14
P1Audit similar services for identical misconfigurationPlatform Lead2026-08-05
P2Revise on-call runbook with current database topologySRE2026-08-10

Lessons Learned

  • Load testing must simulate concurrent connection patterns, not just throughput. A service under high RPS is not the same as a service with high concurrent clients.
  • Configuration drift happens silently. Add health-check validations that alert if actual connection pool size doesn't match predicted concurrency.
  • MTTR depends on alert quality. Implementing pool utilization monitoring would have surfaced this in 30 seconds instead of 4 minutes.

What's Included

  • RCA Framework Templates: Structured prompts for 5 Whys, fishbone diagrams, and fault trees—Claude selects the best framework for your incident type.
  • Postmortem Document Template: Pre-formatted sections (timeline, RCA, contributing factors, action items, lessons learned) that guide your team through structured incident analysis.
  • Stakeholder Communication Guides: Ready-to-customize templates for executive summaries, customer-facing incident explanations, and technical deep-dives for different audiences.
  • Contributing Factors Checklist: Systematic checklist covering monitoring gaps, alert blindness, runbook accuracy, load testing coverage, and configuration management—ensures nothing is overlooked.
  • Pattern Detection Prompts: Guided analysis to spot recurring failure modes across your incident history and identify systemic infrastructure vulnerabilities.

Who It's For

  • Site Reliability Engineers (SREs)
  • Infrastructure and DevOps Engineers
  • Platform and Systems Engineers
  • Engineering Managers and Tech Leads
  • On-Call Incident Commanders and Responders

Best For

  • Writing production incident postmortems after outages
  • Identifying root causes of system failures and cascading incidents
  • Analyzing incident timelines and recovery processes
  • Prioritizing infrastructure hardening and preventive actions
  • Communicating incident impact to executives and customers

You might also like

Industrial Lease Intelligence for Brokers
$30
Industrial Lease Intelligence for Brokers

Extract critical lease terms and financial obligations in minutes, identify risks and red flags that could impact deal profitability, and generate professional broker memos with negotiation strategy recommendations. You'll transform lease documents into actionable intelligence for faster, more confident decision-making.

Enterprise Architecture Review & Decision Framework
$30
Enterprise Architecture Review & Decision Framework

You systematically analyze multi-layered enterprise architectures, identify critical trade-offs across scalability, cost, and risk, and generate documented decisions that align technical choices with business strategy. This skill produces comprehensive architecture reviews that expose hidden dependencies, quantify implementation trade-offs, and deliver prioritized, actionable recommendations backed by formal Architecture Decision Records.

Digital Transformation Strategy Architect
$30
Digital Transformation Strategy Architect

This skill guides you through creating executive-aligned digital transformation roadmaps that balance innovation with operational stability. You'll conduct technology maturity assessments, identify modernization priorities, and build phased implementation strategies with risk mitigation. The result is a comprehensive strategy document that stakeholders and technical teams can execute against immediately.

Frontend Engineering Manager: Code Review & Quality Framework
$15
Frontend3.3(3)
Frontend Engineering Manager: Code Review & Quality Framework

You can leverage Claude to conduct thorough code reviews, evaluate code quality, identify architectural issues, and provide actionable feedback to your frontend engineers. This skill helps you standardize code review practices across your team, catch bugs before production, and maintain consistent quality standards while scaling your engineering capacity.

Strategic Technology Assessment Framework
$35
Strategic Technology Assessment Framework

This skill guides you through systematic technology assessments using proven evaluation frameworks. You provide your current environment, business goals, and constraints—Claude produces detailed assessment reports with technology recommendations, risk analyses, implementation roadmaps, and ROI projections. Get actionable strategic advice grounded in your specific business context, not generic tech trends.

Engineering Leadership Strategy for Scale-Ups
$45
Scale-Up4.0(5)
Engineering Leadership Strategy for Scale-Ups

You get a structured decision-making framework tailored to scale-ups navigating rapid growth. Claude helps you evaluate technical architecture choices, plan team scaling with hiring rubrics, prioritize technical debt against velocity demands, and craft engineering strategy communications for leadership. The framework balances short-term velocity with long-term maintainability.

Lease Term Analysis & Negotiation Assistant
$40
Office3.7(3)
Lease Term Analysis & Negotiation Assistant

This skill transforms complex lease agreements into clear, actionable analysis. It evaluates terms against market benchmarks, identifies negotiation leverage points, and generates professional recommendation letters that strengthen your position in lease discussions.

Industrial Deal Analyzer & Market Evaluator
$35
Industrial Deal Analyzer & Market Evaluator

This skill evaluates industrial real estate deals by analyzing property fundamentals, comparable market data, risk factors, and investment metrics. You get detailed assessments of deal viability, market positioning, and potential returns—all structured to support your investment thesis with hard data. Whether you're analyzing a single property or comparing multiple deals, you'll receive an executive summary with a clear buy/hold/pass recommendation.

$30.00