SkillsLib.ai

Kubernetes Production Incident Response & Root Cause Analysis

Diagnose Kubernetes incidents and pinpoint root causes systematically

0.0(0 reviews)
100+ downloads
Updated Oct 2026

What You Can Do

You'll use structured workflows to triage production incidents in your Kubernetes cluster, gather forensic evidence from logs and metrics, and systematically identify root causes. The skill provides command templates for common failure modes—pod crashes, node failures, network issues—and generates incident timelines and remediation recommendations that you can act on immediately.

Features

Incident triage workflow

Classify failures by symptom (pod, node, network, storage) and severity (critical, high, medium) with structured decision trees

Evidence collection templates

Pre-built commands for gathering logs, metrics, events, and resource utilization from kubectl, container runtimes, and cloud APIs

Root cause analysis framework

Systematic hypothesis testing with validation steps to distinguish between application bugs, configuration errors, resource exhaustion, and infrastructure failures

Command reference library

Copy-paste kubectl commands, grep patterns, and jq filters for extracting relevant signals from cluster state

Incident timeline reconstruction

Parse logs and events to build a chronological view of what happened before and after the symptom appeared

Remediation recommendation engine

Suggests fixes (pod restart, config correction, resource scaling, node drain) based on identified root cause

Post-incident review template

Structured questions for blameless postmortems, timeline verification, and prevention strategies

Multi-layer diagnostic checklist

Covers application, runtime, Kubernetes API, node OS, and cloud infrastructure layers

Example Output

Scenario 1: Pod CrashLoopBackOff

code
✓ Symptom: nginx-deploymnet-abc123 restarting every 10s
✓ Evidence:
  - Exit code 1 (from 'kubectl logs pod-abc123 --previous')
  - /var/log/nginx/error.log shows: port 80 in use
  - No resource limits exceeded (memory 512Mi, using 45Mi)
✓ Root Cause: Sidecar container binding to port 80 before nginx
✓ Fix: Update container startup order in Pod spec

Scenario 2: Node Memory Exhaustion

code
✓ Symptom: Pods on node-prod-07 evicted with OOMKilled
✓ Evidence:
  - Node allocatable: 32Gi, allocated: 28Gi (88%)
  - Prometheus: memory pressure detected at 2026-07-30T14:32:00Z
  - 'kubectl describe node' shows: MemoryPressure=True
  - kubelet log: evicting pod default/worker-456 due to memory
✓ Root Cause: Memory leak in workload; requests underprovisioned
✓ Fix: Increase memory limit, enable resource quota enforcement

Scenario 3: Network Connectivity Issue

code
✓ Symptom: Service-to-service calls timeout; 10% error rate
✓ Evidence:
  - Network Policy allows traffic (verified with 'kubectl get networkpolicies')
  - DNS resolution working (nslookup in debug pod successful)
  - MTU mismatch detected: node=1500, overlay=1450
✓ Root Cause: Fragmented packets dropped by overlay network
✓ Fix: Update MTU on node network interface or adjust application socket buffer

What's Included

  • SKILL.md: Complete incident response framework with workflows, decision trees, and configuration
  • Incident Triage Checklist: Symptoms-to-action mapping for pod, node, network, and storage failures
  • Command Reference Library: 50+ kubectl, container, and observability commands organized by layer
  • Root Cause Analysis Templates: Hypothesis-driven investigation steps for each failure class
  • Evidence Collection Script: Pre-written bash commands to extract logs, metrics, and events
  • Incident Timeline Reconstruction Guide: Methods for parsing timestamps and correlating events
  • Remediation Decision Tree: Maps root causes to fixes (restart, config, scale, drain, upgrade)
  • Post-Incident Review Template: Blameless postmortem format with prevention checklist
  • Troubleshooting Flowchart: Visual decision path from symptom to root cause to fix

Who It's For

  • Site Reliability Engineers (SREs) on-call for cluster incidents
  • Kubernetes cluster operators and platform engineers managing production clusters
  • DevOps engineers troubleshooting application deployment and runtime failures
  • Infrastructure teams responsible for Kubernetes uptime and incident response SLOs
  • Development teams responding to production issues in their services

Best For

  • Pod restarts and CrashLoopBackOff failures with unknown causes
  • Node failures, resource exhaustion (OOMKilled), and evictions
  • Service-to-service connectivity issues and network policy violations
  • Persistent volume attach/mount failures and storage timeouts
  • Performance degradation: latency spikes, request errors, throughput drops with unclear root cause

You might also like

Heritage Tourism Narrative Development
$35
Heritage Tourism Narrative Development

Transform archival research and complex historical information into engaging narratives that resonate with diverse visitor segments while maintaining scholarly rigor. You can create consistent interpretive strategies across multiple touchpoints—from wayfinding copy and museum panels to audio guides and digital experiences—that drive tourism engagement without compromising cultural sensitivity or historical accuracy. This skill systematically identifies interpretation opportunities within existing spaces and integrates primary sources into compelling storytelling frameworks.

Grafana Dashboard & Alert Optimization
$30
Grafana Dashboard & Alert Optimization

This skill positions Claude as your Grafana technical advisor. You can optimize PromQL queries for performance, design clear dashboards for incident response, create effective alert rules that minimize false positives, and troubleshoot visualization issues—all grounded in monitoring best practices and your specific metrics context.

Stratigraphic Analysis & Deposit Interpretation
$45
Stratigraphic Analysis & Deposit Interpretation

You can leverage Claude as a structured thinking partner to organize complex multi-layered deposit sequences, identify stratigraphic anomalies, build robust site phasing models, and generate defensible interpretive narratives. Claude processes your detailed context descriptions, deposit documentation, and relationship matrices to identify patterns, flag inconsistencies, and suggest alternative interpretations you might have overlooked during field analysis.

NHPA Section 106 Compliance Reviewer
$40
NHPA Section 106 Compliance Reviewer

You can evaluate whether a project requires Section 106 review, determine the appropriate federal agency lead, define the Area of Potential Effects (APE), identify consulting parties, assess adverse effects to historic properties, and generate compliance documentation including MOAs and Programmatic Agreements. This skill helps you scope compliance packages, review existing Section 106 documents for completeness, and prepare for consultation with SHPOs, THPOs, and Indian tribes.

Historic Material Condition Assessment for Conservation Planning
$40
Historic Material Condition Assessment for Conservation Planning

You can conduct comprehensive material condition assessments that capture degradation patterns, environmental factors, and damage mechanisms in historic buildings. This skill guides you through structured evaluation protocols that transform subjective observations into defensible conservation decisions, enabling you to prioritize repairs across multiple elements, justify preservation budgets to stakeholders, and specify appropriate materials and methods for contractors.

Historic Structure Documentation Compiler
$50
Historic Structure Documentation Compiler

You can generate complete technical documentation packages for historic structures that meet professional preservation standards and regulatory requirements. This skill produces defensible, archival-quality records including condition assessments, measured drawing descriptions, photographic documentation protocols, material analysis frameworks, and standardized compliance reports. Your documentation serves multiple purposes—regulatory filings, restoration guidance, insurance claims, grant applications, and long-term institutional stewardship.

Source Prep
$30
Source Prep

You can transform a scheduled expert interview into a strategic session by mapping your source's public record, incentives, and known positions, then building a tiered question architecture that moves from foundational to provocative with built-in pivots. Claude helps you anticipate likely evasions, scripts pre-emptive follow-ups, flags attribution constraints, and designs your interview flow for narrative payoff—not just information collection.

Historic Documentation Assessment for Cultural Resources
$45
Historic Documentation Assessment for Cultural Resources

You can conduct comprehensive documentation and evaluation of historic properties aligned with NRHP Bulletin 16A criteria, Section 106 processes, and state preservation office expectations. The skill produces structured evaluation reports that support preservation planning, funding applications, and regulatory compliance workflows while organizing historical research, property analysis, and significance assessment into actionable documentation.

$40.00