SkillsLib.ai

Cloud Incident Diagnosis and Resolution

Diagnose cloud infrastructure incidents with structured troubleshooting and log analysis

4.0(32 reviews)
500+ downloads
Updated Oct 2026
Verified SafeSecurity VerifiedThis skill was analyzed by our AI security scanner for harmful content including data exfiltration, system manipulation, credential theft, and prompt injection. No threats were detected.

What You Can Do

You can rapidly diagnose cloud infrastructure incidents across AWS, Azure, or GCP by following structured troubleshooting frameworks that work from application layer down through platform and infrastructure components. The skill guides you through incident classification, multi-layer dependency mapping, log pattern analysis, and escalation criteria—enabling you to isolate root causes efficiently and deliver customers both immediate workarounds and permanent fixes.

Features

Incident classification matrices

categorize issues by severity, scope, and affected service layers

Multi-layer troubleshooting sequences

systematically test application, platform, and infrastructure components to eliminate possibilities

Structured log analysis patterns

parse cloud service outputs (CloudTrail, Activity Logs, Stackdriver) to extract diagnostic signals

Cross-service dependency mapping

identify cascading failures across compute, storage, networking, and database services

Escalation decision trees

determine when to escalate to cloud provider support with complete evidence packages

Remediation frameworks

deliver both immediate workarounds and permanent resolution steps

Evidence documentation requirements

capture logs, timestamps, and configuration details needed for provider escalation

Example Output

Example 1: EC2 Instance Unavailability

  • Classification: P2 (service degradation, single customer impact)
  • Layer 1 findings: Application logs show successful startup, no errors
  • Layer 2 findings: Security group ingress rule blocking port 443
  • Layer 3 findings: CloudTrail shows rule modified 2 hours ago
  • Root cause: Overly permissive security group rule reverted by auto-remediation policy
  • Immediate fix: Manually restore correct ingress rules (5 min)
  • Permanent fix: Update auto-remediation policy logic to exclude production security groups

Example 2: Database Connection Pool Exhaustion

  • Classification: P1 (application unavailable, revenue-impacting)
  • Layer 1 findings: RDS logs show max connections (300) exceeded
  • Layer 2 findings: CloudWatch metrics show 3x normal connection spike 15 minutes ago
  • Layer 3 findings: Check deployment history—new app version deployed 20 min ago
  • Root cause: Application connection pooling configuration incompatible with new ORM version
  • Immediate fix: Temporarily scale RDS instance (10 min), notify customer
  • Escalation: Not needed; issue is application-side
  • Permanent fix: Revert ORM version, run connection pool load testing

What's Included

  • SKILL.md: Complete troubleshooting framework with layered diagnostic sequences
  • Incident Classification Matrix: Severity and scope assessment template for quick triage
  • Cloud Service Log Patterns Checklist: Keywords and error signatures for AWS CloudTrail, Azure Activity Logs, and GCP Cloud Logging
  • Multi-Layer Troubleshooting Checklist: Step-by-step tests for application, platform, and infrastructure layers
  • Escalation Decision Tree: Criteria and evidence requirements for cloud provider support handoff
  • Cross-Service Dependency Map Template: Worksheet for mapping service interactions and failure cascades

Who It's For

  • Cloud Support Engineers — Rapidly diagnose and resolve customer infrastructure incidents
  • Site Reliability Engineers (SREs) — Systematic root cause analysis for production outages
  • DevOps Engineers — Troubleshoot infrastructure issues in distributed cloud environments
  • Solutions Architects — Understand common failure modes when designing customer cloud solutions
  • Technical Account Managers — Escalate to cloud providers with complete diagnostic evidence

Best For

  • Production outages and degraded performance in cloud services (compute, storage, database, networking)
  • Multi-service failures where root cause is unclear and spans multiple layers
  • Escalations to cloud provider support requiring comprehensive evidence packages
  • Post-incident diagnosis to identify systemic configuration or policy issues
  • Rapid customer communication with triage status, workarounds, and resolution timelines

You might also like

Cloud Incident Root Cause Analysis
$45
Cloud4.0(26)
Cloud Incident Root Cause Analysis

You can upload logs and metrics from distributed cloud services—compute, networking, storage, databases—and Claude will parse them to identify failure patterns, map service dependencies, and trace how failures cascade across your infrastructure. The skill generates prioritized remediation steps backed by forensic evidence, plus customer-ready explanations that translate technical findings into business impact.

$40.00