
ML Infrastructure Optimization & Troubleshooting
Debug ML infrastructure bottlenecks, fix distributed training failures, optimize performance
What You Can Do
You can systematically diagnose performance issues across your ML infrastructure, identify root causes of distributed training failures, and receive actionable optimization recommendations. This skill analyzes metrics, logs, and configuration data to pinpoint resource constraints, communication overhead, and configuration misalignments that impact training speed and reliability.
Features
Analyzes CPU, GPU, memory, and I/O utilization patterns to identify where your pipeline loses time
Traces collective communication failures, gradient synchronization issues, and node-to-node connectivity problems
Maps compute, data loading, and network stages to spot inefficiencies from data ingestion through model checkpointing
Reviews hyperparameters, batch sizes, gradient accumulation, and communication backends for alignment with your hardware
Recommends GPU/CPU/memory configurations based on model size, data throughput, and fault tolerance requirements
Extracts and correlates anomalies across training logs, system metrics, and error traces to surface hidden dependencies
Diagnoses clock skew, network latency, and synchronization issues across distributed clusters
Provides priority-ordered fixes with implementation guidance and expected performance gains
Example Output
Example 1: Bottleneck Analysis
- π΄ CRITICAL: GPU utilization 32% (target 80%+)
ββ Root cause: Data loading blocking GPU pipeline
ββ Symptom: 4.2s train step (2.8s data load, 1.4s forward/backward)
ββ Fix: Enable async prefetching, increase DataLoader workers from 4 β 8
Expected improvement: 65% reduction in step time
- β οΈ WARNING: AllGather communication overhead 18% of step time
ββ Symptom: Gradient synchronization across 8 GPUs takes 1.1s
ββ Recommendation: Switch NCCL backend to use NVLink + enable gradient compression
Example 2: Distributed Training Failure
- π¨ Failure pattern detected across nodes:
ββ Node 3 shows hanging AllReduce at iteration 45
ββ Network trace: Packet loss 0.2% between Node 1 β Node 3
ββ Root cause: Network congestion during weight synchronization
ββ Temporary fix: Increase NCCL timeout from 30s β 120s (masks symptom)
ββ Permanent fix: Route collective ops through dedicated network interface
What's Included
- SKILL.md: Full infrastructure analysis workflow with decision trees for diagnosis
- Performance profiling checklist: Metrics to collect, interpret, and correlate
- Configuration templates: DDP, FSDP, Horovod, Ray Train setup examples
- Log parsing patterns: Regular expressions and extraction rules for common ML frameworks
- Troubleshooting flowchart: Decision tree from symptom β root cause β fix
- Benchmark comparison tool: Script to baseline performance against known good configurations
Who It's For
- ML infrastructure engineers β Debug production training pipelines and scale distributed models
- Research scientists β Optimize hyperparameter sweeps and reduce iteration time on clusters
- DevOps/platform engineers β Right-size hardware allocation and improve cluster utilization
- MLOps teams β Implement monitoring and automated diagnosis for training pipeline reliability
- Data center operators β Identify network and resource constraints affecting ML workloads
Best For
- Diagnosing why distributed training hangs or crashes across multiple nodes
- Identifying data loading bottlenecks blocking GPU pipeline efficiency
- Analyzing collective communication overhead (AllGather, AllReduce) in multi-GPU setups
- Right-sizing batch size, gradient accumulation, and communication backends for your hardware
- Troubleshooting out-of-memory errors and resource allocation failures
- Profiling end-to-end pipeline performance from data ingestion through checkpoint save







