
Incident Response & Root Cause Analysis
Correlate logs, traces, and metrics to surface root causes
What You Can Do
Rapidly analyze production incidents by correlating observability signals—logs, metrics, traces, and events—to identify the true root cause instead of symptoms. You'll reconstruct incident timelines, quantify impact scope, and generate both immediate remediation steps and long-term prevention measures.
Features
Synthesizes logs, metrics, traces, and events to uncover patterns invisible in any single signal
Builds precise chronological sequences from trigger event through symptoms to recovery
Identifies deviations from baseline across all observability dimensions simultaneously
Measures scope (affected users, services, SLA breach count, revenue impact)
Categorizes causes (infrastructure, application, config, dependency, external)
Recommends immediate fixes, short-term patches, and long-term architectural changes
Generates executive summaries and technical deep-dives for stakeholders
Example Output
Input: Incident data (logs, metrics, traces spanning 45-minute window of payment-api 500 errors)
Output:
## Root Cause: Database Connection Pool Exhaustion
Evidence:
- Logs: [CRITICAL] connection pool max=100, current=100, queue_depth=500
- Metrics: DB latency ↑450%, connection wait time avg 3.2s
- Traces: End-to-end P99 spans show 2.8s median time in pool queue
Timeline:
14:20 UTC - Deploy v2.3.1 (connection retry logic added)
14:23 UTC - Retry storm hits DB during transient network blip
14:25 UTC - Pool exhaustion, requests timeout after 60s
14:28 UTC - Circuit breaker trips, 500 errors cascade
14:35 UTC - Alert fires (too late)
14:40 UTC - Rollback + pool-size increase deployed
14:42 UTC - Full recovery
Impact: 15,284 users affected, 2,100 SLA breaches, ~$12K lost revenue
Remediation:
- Immediate: Increase pool max from 100→200, deploy circuit breaker tuning
- Short-term: Add connection pool health metric + automated scaling policy
- Long-term: Implement connection pooling sidecar, replace retry logic with exponential backoff
What's Included
- SKILL.md: Multi-signal incident analysis workflow with decision trees for different incident types
- Incident Analysis Checklist: Step-by-step guide for gathering and correlating signals
- Signal Correlation Template: Worksheet to map symptoms → signals → root cause
- Impact Assessment Template: Structured format for quantifying blast radius and customer impact
- Root Cause Classification Guide: Decision tree for categorizing cause types (infrastructure, app, config, dependency, external)
- Post-Incident Report Template: Executive summary + technical deep-dive + prevention recommendations
Who It's For
- On-call engineers and incident responders — Teams handling production incidents under time pressure need structured analysis
- Site reliability engineers (SREs) — SREs own observability strategy and need deep root cause analysis for incident prevention
- DevOps and platform engineers — Operators managing infrastructure need to quickly isolate whether issues are infra, config, or application
- Engineering managers and tech leads — Leadership conducting post-mortems needs clear evidence and remediation plans
- Support engineers (L2/L3 technical) — Technical support teams escalating production issues need structured hand-offs to engineering
Best For
- Production outage investigation — When a service is down, rapidly identify whether it's infrastructure, deployment, config, or dependency
- Performance degradation troubleshooting — When latency or throughput degrades, correlate metrics and traces to pinpoint the bottleneck
- Failed deployment analysis — When a recent deploy broke production, find which code change or config caused the issue
- Cascading failure diagnosis — When multiple services fail, trace the dependency chain to find the initial trigger
- Post-incident root cause analysis — Structured analysis for incident reviews and blameless postmortems






