
Process Troubleshooting & Root Cause Analysis Framework
Diagnose process failures with systematic root cause analysis
What You Can Do
You'll work through a structured 5-phase troubleshooting framework to identify the true root causes of process anomalies—not just symptoms. The framework guides you to gather relevant data, generate and test hypotheses, rank probable causes by evidence, and design corrective actions with prevention measures. You'll produce clear documentation suitable for incident reviews and process improvements.
Features
Problem definition → Data collection → Hypothesis generation → Evidence evaluation → Action planning
Checklist of metrics, logs, timestamps, and context to capture before analysis begins
Brainstorm likely causes using failure mode thinking, recent changes, and operating conditions
Score each hypothesis against available data (correlation, timeline, plausibility) to rank by likelihood
Generate immediate fixes, preventive measures, and monitoring strategies
Structure your findings as post-mortem reports with clear cause statements and lessons learned
Identify who to involve (ops, engineering, product, vendor) and what questions to ask them
Example Output
Example 1: Production Database Timeout
Problem: Customer-facing API timing out under normal load (noon UTC, 50% of baseline requests failed)
Root Cause Hypotheses (ranked by evidence):
- Recent query optimization rollback — Timeline matches last deploy 3h ago; DBA confirms revert happened
- Disk I/O saturation — Monitoring shows 85% utilization but stable; doesn't explain sudden failure
- Memory leak in connection pool — Memory flat; pool exhaustion would show queue delays (absent)
Immediate Action: Revert optimization change → Service restored in 8 minutes
Prevention: Add pre-deploy load test; require performance sign-off on all query changes
Example 2: Manufacturing Line Defect Spike
Problem: Widget defect rate jumped from 2% to 8% Thursday 2pm; no line shutdown events recorded
Root Cause Hypotheses:
- New batch of raw material (HIGH) — Supplier batch #47291 received Wed; QC passes but material properties drift slightly
- Calibration drift on press #3 — Last calibrated Monday; runs +0.3mm but within spec historically
- Humidity spike in production floor — Actually dropped Thurs; rules out moisture sensitivity
Immediate Action: Quarantine material batch; retest supplier lot → Defect rate returns to 2%
Prevention: Add incoming material batch traceability; tighten supplier specifications
What's Included
- SKILL.md: Complete troubleshooting framework with decision trees and phase workflows
- Data Collection Checklist: Metrics, logs, timeline, environmental context, recent changes
- Root Cause Hypothesis Template: Structure for brainstorming and documenting suspected causes
- Evidence Evaluation Matrix: Scoring rubric to rank hypotheses (correlation, timeline, plausibility, testability)
- Corrective Action Plan Template: Immediate fix, preventive measures, monitoring strategy, owner assignments
- Incident Post-Mortem Guide: Writing structure for incident reports and lessons learned
Who It's For
- Operations and DevOps engineers — Debugging production outages and service failures
- Process improvement specialists — Systematic analysis for Six Sigma and Lean initiatives
- Manufacturing/production leads — Root cause analysis for line stoppages and defect spikes
- Incident response teams — Structured post-mortem analysis and action planning
- Quality assurance managers — Problem trending and corrective action tracking
Best For
- Production outages and unexpected service failures requiring rapid diagnosis
- Process degradation — gradual performance drops, increasing defect rates, or rising error counts
- Recurring incidents — identifying systemic causes after multiple similar failures
- Equipment malfunctions and anomalies in manufacturing or infrastructure
- Post-mortem incident reviews and prevention planning






