
Docker Container Troubleshooting & Performance Optimization
Debug containers faster—diagnose failures and optimize performance in minutes
What You Can Do
Analyze Docker container logs instantly to identify root causes of failures, crashes, and resource bottlenecks. Get actionable recommendations for resource allocation, network configuration, and architectural improvements that eliminate production issues before they cascade.
Features
Claude identifies stack traces, error patterns, and failure root causes from container logs
Analyzes CPU/memory/disk usage patterns to right-size container limits and requests
Generate optimal health check configurations (startup, liveness, readiness probes)
Pre-built configurations for security, restart policies, and orchestration best practices
Identify bottlenecks in application and container-level metrics
Diagnose port conflicts, DNS issues, overlay network problems, and connectivity failures
Reduce image size, layer analysis, and multi-stage build recommendations
Example Output
Example 1: Container OOM Kill Analysis
Input: "My PostgreSQL container keeps getting OOMKilled. Here are the logs..."
Output:
- Root cause: Process memory limit (256Mi) insufficient for buffer cache
- Current usage: 312Mi (122% of limit)
- Recommendation: Increase memory request to 512Mi, limit to 1Gi
- Config snippet:
resources: {requests: {memory: "512Mi"}, limits: {memory: "1Gi"}} - Timeline: Issue first appeared after 5.2GB dataset import
Example 2: Dockerfile Performance Optimization
Input: "My build takes 15 minutes. Here's my Dockerfile..."
Output:
- Layer analysis: Base image (1.2GB) + dependencies (800MB) + app code (150MB)
- Optimization: Reorder RUN commands, use BuildKit cache mount
- Size reduction: 2.2GB → 380MB (82% reduction)
- Build time: 15min → 2.3min (85% faster)
- Refactored Dockerfile with inline comments and best practices
What's Included
- SKILL.md file with diagnostic workflows and decision trees:
- Container failure troubleshooting checklist (OOM, CrashLoop, ImagePullBackOff, Pending, etc.):
- Resource optimization templates (requests/limits for common services: PostgreSQL, Redis, Node.js, Python, Go):
- Dockerfile analysis and multi-stage build templates:
- Health check configuration templates (Kubernetes probes + Docker HEALTHCHECK):
- Network troubleshooting flowchart and common connectivity patterns:
- Performance baseline metrics collection guide:
- Production hardening security checklist:
Who It's For
- DevOps engineers — Rapidly diagnose container failures in CI/CD pipelines and production
- Platform engineers — Design resource policies and health checks for microservices
- SREs — Troubleshoot multi-container production incidents and optimize infrastructure
- Backend developers — Debug containerized applications locally and optimize Dockerfiles
- Cloud architects — Implement container best practices and production hardening
Best For
- Debugging container crashes and restarts (CrashLoopBackOff, OOMKilled, ImagePullBackOff, Pending)
- Right-sizing CPU/memory requests and limits for containerized workloads
- Optimizing Docker image size and build performance
- Designing health checks and startup probes for reliable container orchestration
- Analyzing container logs for security, performance, and reliability issues







