
Incident Diagnosis & Root Cause Analysis via Observability Signals
Diagnose production incidents by correlating observability signals
What You Can Do
Systematically investigate production incidents by correlating metrics, logs, and distributed traces to pinpoint root causes. You provide Claude with observability data from your infrastructure, and it builds a timeline, identifies anomalies, tests hypotheses, and delivers a structured incident report with actionable remediation steps. This accelerates MTTR and prevents recurrence through data-driven diagnostics.
Features
Maps events across metrics, logs, and traces to establish precise temporal relationships
Identifies deviations in CPU, memory, latency, error rates correlated with incident onset
Maps affected service dependencies and determines blast radius from root cause
Systematically eliminates suspects by cross-referencing observability evidence
Prioritizes high-impact signals (resource utilization, error spikes, latency) by relevance
Groups firing alerts to identify common triggers and suppress noise
Suggests immediate mitigation actions and long-term fixes based on findings
Produces structured analysis with timeline, impact quantification, and prevention steps
Example Output
Timeline
- 14:23:47 UTC — Alert: API p95 latency jumped 3.2s → 18s
- 14:23:51 UTC — Database connection pool exhausted (150/150 active connections)
- 14:24:02 UTC — CPU on database host spiked 45% → 92%
- 14:24:15 UTC — Slow query log: full table scan on
userstable (missing index on email column)
Root Cause
New feature deploy introduced bulk email validation without indexing the email column, causing full table scans. This exhausted the database connection pool, cascading to the API layer and degrading all downstream services.
Remediation
- Immediate: Execute
CREATE INDEX idx_users_email ON users(email)— resolves blockage in <2 minutes - Long-term: Add database query review to CI/CD; add observability alert on slow query duration exceeding 500ms
Impact Customer-facing outage: 12 minutes affecting 8% of user logins. MTTR: 3 minutes (time to add index). Estimated revenue impact: $12K.
What's Included
- SKILL.md: Complete incident diagnosis workflow with multi-signal correlation logic
- Investigation checklist: Step-by-step process for systematically isolating root cause
- Timeline builder template: Structured format for mapping events across observability sources
- Hypothesis testing worksheet: Framework for validating and eliminating suspect components
- Remediation decision tree: Guide for prioritizing immediate mitigation vs. permanent fixes
- Query pattern library: Pre-built observability queries for Datadog, Prometheus, Splunk, and ELK
Who It's For
- Site Reliability Engineers (SREs) managing incident response and MTTR
- DevOps and platform engineers diagnosing infrastructure failures
- On-call engineers triaging P1/P2 incidents under time pressure
- Engineering managers analyzing recurring failure patterns
- Cloud architects investigating multi-service failure cascades
Best For
- P1/P2 incident diagnosis and MTTR reduction
- Production performance degradation investigation
- Error rate or latency spike root cause analysis
- Service failure diagnosis in multi-tier architectures
- Post-incident analysis and recurrence prevention planning







