
Cloud Incident Diagnosis and Resolution
Diagnose cloud infrastructure incidents with structured troubleshooting and log analysis
What You Can Do
You can rapidly diagnose cloud infrastructure incidents across AWS, Azure, or GCP by following structured troubleshooting frameworks that work from application layer down through platform and infrastructure components. The skill guides you through incident classification, multi-layer dependency mapping, log pattern analysis, and escalation criteria—enabling you to isolate root causes efficiently and deliver customers both immediate workarounds and permanent fixes.
Features
categorize issues by severity, scope, and affected service layers
systematically test application, platform, and infrastructure components to eliminate possibilities
parse cloud service outputs (CloudTrail, Activity Logs, Stackdriver) to extract diagnostic signals
identify cascading failures across compute, storage, networking, and database services
determine when to escalate to cloud provider support with complete evidence packages
deliver both immediate workarounds and permanent resolution steps
capture logs, timestamps, and configuration details needed for provider escalation
Example Output
Example 1: EC2 Instance Unavailability
- Classification: P2 (service degradation, single customer impact)
- Layer 1 findings: Application logs show successful startup, no errors
- Layer 2 findings: Security group ingress rule blocking port 443
- Layer 3 findings: CloudTrail shows rule modified 2 hours ago
- Root cause: Overly permissive security group rule reverted by auto-remediation policy
- Immediate fix: Manually restore correct ingress rules (5 min)
- Permanent fix: Update auto-remediation policy logic to exclude production security groups
Example 2: Database Connection Pool Exhaustion
- Classification: P1 (application unavailable, revenue-impacting)
- Layer 1 findings: RDS logs show max connections (300) exceeded
- Layer 2 findings: CloudWatch metrics show 3x normal connection spike 15 minutes ago
- Layer 3 findings: Check deployment history—new app version deployed 20 min ago
- Root cause: Application connection pooling configuration incompatible with new ORM version
- Immediate fix: Temporarily scale RDS instance (10 min), notify customer
- Escalation: Not needed; issue is application-side
- Permanent fix: Revert ORM version, run connection pool load testing
What's Included
- SKILL.md: Complete troubleshooting framework with layered diagnostic sequences
- Incident Classification Matrix: Severity and scope assessment template for quick triage
- Cloud Service Log Patterns Checklist: Keywords and error signatures for AWS CloudTrail, Azure Activity Logs, and GCP Cloud Logging
- Multi-Layer Troubleshooting Checklist: Step-by-step tests for application, platform, and infrastructure layers
- Escalation Decision Tree: Criteria and evidence requirements for cloud provider support handoff
- Cross-Service Dependency Map Template: Worksheet for mapping service interactions and failure cascades
Who It's For
- Cloud Support Engineers — Rapidly diagnose and resolve customer infrastructure incidents
- Site Reliability Engineers (SREs) — Systematic root cause analysis for production outages
- DevOps Engineers — Troubleshoot infrastructure issues in distributed cloud environments
- Solutions Architects — Understand common failure modes when designing customer cloud solutions
- Technical Account Managers — Escalate to cloud providers with complete diagnostic evidence
Best For
- Production outages and degraded performance in cloud services (compute, storage, database, networking)
- Multi-service failures where root cause is unclear and spans multiple layers
- Escalations to cloud provider support requiring comprehensive evidence packages
- Post-incident diagnosis to identify systemic configuration or policy issues
- Rapid customer communication with triage status, workarounds, and resolution timelines
