
MLOps System Diagnostic & Troubleshooting
Diagnose ML pipeline failures in minutes with structured log analysis
What You Can Do
You can rapidly identify root causes of training failures, data pipeline breakdowns, and model serving issues by analyzing logs, metrics, and infrastructure state. Claude examines your system events, resource utilization patterns, and error traces to pinpoint performance bottlenecks, data quality problems, and deployment misconfigurations—then recommends targeted remediation steps.
Features
Convert unstructured logs into searchable diagnostic timelines with event correlation
Trace failures back to their origin (data, compute, configuration, dependency)
Spot training slowdowns, GPU utilization issues, and I/O contention
Detect schema drift, missing values, distribution anomalies in training data
Analyze latency, throughput, and error rates in inference pipelines
Correlate errors across microservices, workers, and cloud infrastructure
Identify memory leaks, CPU saturation, disk space issues, and network congestion
Create step-by-step remediation guides from diagnostic findings
Example Output
Example 1: Training Failure Diagnosis
Symptom: Training job OOM-killed after 2 hours
Claude Analysis:
✓ Identified: Batch size 256 + gradient accumulation 4 = effective batch 1024
✓ Bottleneck: GPU memory 80GB insufficient for model + optimizer state
✓ Timeline: Memory usage grew 15% per epoch (no weight decay regularization)
✓ Recommendation: Reduce batch to 128, enable gradient checkpointing, add L2 regularization
Example 2: Data Pipeline Stall
Symptom: Daily ETL hangs, training waits 6+ hours for data
Claude Analysis:
✓ Root cause: CSV read operation on 500GB file via single process
✓ Bottleneck: S3 download bandwidth 50MB/s (3 hours), parsing 200MB/s (50 mins)
✓ Contributing factor: No partitioning strategy; full scan on every run
✓ Recommendation: Enable columnar format (Parquet), add time-range partitioning, parallelize reads (8 workers)
✓ Expected speedup: 6h → 45 min
Example 3: Model Serving Latency
Symptom: p95 inference latency 500ms (target 50ms)
Claude Analysis:
✓ Profiled request breakdown: Model inference 15ms + preprocessing 350ms + postprocessing 100ms
✓ Bottleneck: Preprocessing serializes NLP tokenization on CPU
✓ Issue: Model deployed on single GPU; 100 concurrent requests queue
✓ Recommendation: Batch tokenization, add GPU queue, deploy 4 replicas behind load balancer
✓ Expected p95 latency: 500ms → 65ms
What's Included
- SKILL.md: Full diagnostic workflows with decision trees for training, data, and serving failures
- Log Analysis Template: Structured checklist for parsing and correlating events
- Performance Profiling Guide: Step-by-step GPU/CPU/memory/I/O analysis framework
- Data Quality Verification Checklist: Schema, completeness, distribution validation
- Incident Response Runbook: Diagnostic questions and remediation workflow
- Metrics Collection Checklist: Essential metrics to capture for future debugging
- Troubleshooting Decision Tree: Flowchart to narrow failure category quickly
- Post-Incident Review Template: Lessons learned and prevention measures
Who It's For
- ML Engineers — Diagnosing failed training runs and production model issues
- MLOps Engineers — Troubleshooting infrastructure, pipeline orchestration, and resource constraints
- Data Engineers — Identifying data quality issues, ETL bottlenecks, and schema problems
- Platform/DevOps Teams — Debugging distributed systems, Kubernetes deployments, cloud service issues
- Incident Response Teams — Rapidly responding to model serving outages and training failures
Best For
- Training job failures — OOM errors, convergence issues, data loading timeouts, GPU/TPU utilization problems
- Data pipeline breakdowns — ETL stalls, schema drift, missing data, slow ingestion, partitioning issues
- Model serving incidents — High latency, throughput degradation, error spikes, cold start delays
- Performance optimization — Identifying and resolving bottlenecks in training, inference, and data pipelines
- Infrastructure debugging — Resource contention, Kubernetes pod evictions, cloud quota issues, network congestion







