
Kubernetes Production Incident Response & Root Cause Analysis
Diagnose Kubernetes incidents and pinpoint root causes systematically
What You Can Do
You'll use structured workflows to triage production incidents in your Kubernetes cluster, gather forensic evidence from logs and metrics, and systematically identify root causes. The skill provides command templates for common failure modes—pod crashes, node failures, network issues—and generates incident timelines and remediation recommendations that you can act on immediately.
Features
Classify failures by symptom (pod, node, network, storage) and severity (critical, high, medium) with structured decision trees
Pre-built commands for gathering logs, metrics, events, and resource utilization from kubectl, container runtimes, and cloud APIs
Systematic hypothesis testing with validation steps to distinguish between application bugs, configuration errors, resource exhaustion, and infrastructure failures
Copy-paste kubectl commands, grep patterns, and jq filters for extracting relevant signals from cluster state
Parse logs and events to build a chronological view of what happened before and after the symptom appeared
Suggests fixes (pod restart, config correction, resource scaling, node drain) based on identified root cause
Structured questions for blameless postmortems, timeline verification, and prevention strategies
Covers application, runtime, Kubernetes API, node OS, and cloud infrastructure layers
Example Output
Scenario 1: Pod CrashLoopBackOff
✓ Symptom: nginx-deploymnet-abc123 restarting every 10s
✓ Evidence:
- Exit code 1 (from 'kubectl logs pod-abc123 --previous')
- /var/log/nginx/error.log shows: port 80 in use
- No resource limits exceeded (memory 512Mi, using 45Mi)
✓ Root Cause: Sidecar container binding to port 80 before nginx
✓ Fix: Update container startup order in Pod spec
Scenario 2: Node Memory Exhaustion
✓ Symptom: Pods on node-prod-07 evicted with OOMKilled
✓ Evidence:
- Node allocatable: 32Gi, allocated: 28Gi (88%)
- Prometheus: memory pressure detected at 2026-07-30T14:32:00Z
- 'kubectl describe node' shows: MemoryPressure=True
- kubelet log: evicting pod default/worker-456 due to memory
✓ Root Cause: Memory leak in workload; requests underprovisioned
✓ Fix: Increase memory limit, enable resource quota enforcement
Scenario 3: Network Connectivity Issue
✓ Symptom: Service-to-service calls timeout; 10% error rate
✓ Evidence:
- Network Policy allows traffic (verified with 'kubectl get networkpolicies')
- DNS resolution working (nslookup in debug pod successful)
- MTU mismatch detected: node=1500, overlay=1450
✓ Root Cause: Fragmented packets dropped by overlay network
✓ Fix: Update MTU on node network interface or adjust application socket buffer
What's Included
- SKILL.md: Complete incident response framework with workflows, decision trees, and configuration
- Incident Triage Checklist: Symptoms-to-action mapping for pod, node, network, and storage failures
- Command Reference Library: 50+ kubectl, container, and observability commands organized by layer
- Root Cause Analysis Templates: Hypothesis-driven investigation steps for each failure class
- Evidence Collection Script: Pre-written bash commands to extract logs, metrics, and events
- Incident Timeline Reconstruction Guide: Methods for parsing timestamps and correlating events
- Remediation Decision Tree: Maps root causes to fixes (restart, config, scale, drain, upgrade)
- Post-Incident Review Template: Blameless postmortem format with prevention checklist
- Troubleshooting Flowchart: Visual decision path from symptom to root cause to fix
Who It's For
- Site Reliability Engineers (SREs) on-call for cluster incidents
- Kubernetes cluster operators and platform engineers managing production clusters
- DevOps engineers troubleshooting application deployment and runtime failures
- Infrastructure teams responsible for Kubernetes uptime and incident response SLOs
- Development teams responding to production issues in their services
Best For
- Pod restarts and CrashLoopBackOff failures with unknown causes
- Node failures, resource exhaustion (OOMKilled), and evictions
- Service-to-service connectivity issues and network policy violations
- Persistent volume attach/mount failures and storage timeouts
- Performance degradation: latency spikes, request errors, throughput drops with unclear root cause







