
Cloud Incident Root Cause Analysis
Analyze cloud logs and metrics to identify root causes and remediation steps
What You Can Do
You can upload logs and metrics from distributed cloud services—compute, networking, storage, databases—and Claude will parse them to identify failure patterns, map service dependencies, and trace how failures cascade across your infrastructure. The skill generates prioritized remediation steps backed by forensic evidence, plus customer-ready explanations that translate technical findings into business impact.
Features
Parse and cross-reference logs from compute, networking, storage, and database services to surface hidden failure chains
Identify specific root causes (throttling, resource exhaustion, authentication failures, configuration drift) from symptom clusters
Visualize how services depend on each other and trace failure propagation through your architecture
Generate actionable steps ranked by impact and urgency with success criteria and validation steps
Create detailed incident reports with timeline reconstruction, evidence summaries, and customer-ready explanations
Compare current configs against baselines to identify unintended changes that triggered incidents
Correlate metrics spikes with log events to pinpoint root causes of latency and throughput issues
Example Output
Example 1: Cascading Database Failure
Input: CloudWatch logs, RDS metrics, application error logs
Output:
ROOT CAUSE: RDS connection pool exhaustion → API request timeout → Client errors
Evidence Timeline:
- 14:23 UTC: RDS CPU spike to 87% (slow query detected)
- 14:24 UTC: Connection count exceeds pool limit (150/150)
- 14:25 UTC: API latency increases from 200ms to 8s
- 14:26 UTC: Client receives 503 Service Unavailable
Remediation (Priority Order):
1. Immediate: Terminate long-running queries blocking pool
2. Short-term: Increase RDS connection pool size to 250
3. Medium-term: Add query timeout enforcement (30s max)
4. Long-term: Implement connection pooling proxy layer
Example 2: Configuration Drift Post-Deployment
Input: Git deployment logs, security group rules, VPC flow logs
Output:
ROOT CAUSE: Security group rule misconfiguration post-deployment blocks internal service-to-service communication
Drift Detected:
- Expected: sg-internal allows port 5432 from app tier
- Actual: sg-internal denies all inbound (reverted during rollback)
Impact: Database unreachable → 45 failed requests/second
Fix: Restore security group rule, validate with: `nc -zv database.internal 5432`
What's Included
- SKILL.md: Complete skill instruction file with prompting templates
- Multi-Service Log Template: Pre-formatted checklist for gathering logs from AWS/GCP/Azure services
- RCA Report Framework: Structured markdown template for documenting findings, timeline, and remediation
- Service Dependency Map Canvas: Visual template for mapping service interactions and failure propagation
- Evidence Checklist: Validation list ensuring all relevant logs, metrics, and configs are analyzed before conclusions
Who It's For
- Cloud Technical Support Engineers — Diagnose multi-service incidents with structured methodology backed by forensic evidence
- DevOps/SRE Teams — Investigate production incidents and document RCAs for post-incident reviews and knowledge base updates
- Cloud Architects — Validate configurations before deployments and analyze performance degradation root causes
- Incident Commanders — Generate prioritized remediation steps and customer-ready impact summaries during active incidents
- AWS/GCP/Azure Support Escalation Teams — Provide detailed technical analysis for customer escalations demanding evidence-backed explanations
Best For
- Multi-service outages with unclear failure origin (database cascading to API timeouts)
- Intermittent issues requiring pattern analysis across time periods and service boundaries
- Customer escalations requiring detailed RCA documentation with forensic evidence and timeline reconstruction
- Post-incident reviews and root cause documentation for knowledge base and prevention strategies
- Configuration drift detection and validation before major deployments or after incidents







