
SCADA Troubleshooting & Root Cause Analysis
Diagnose SCADA faults and correlate alarms into actionable root cause analysis
What You Can Do
You can systematically analyze SCADA system failures by correlating alarms, logs, and sensor data to identify root causes. The skill generates detailed root cause analysis (RCA) reports that pinpoint failure modes, quantify impact, and recommend preventive actions—eliminating guesswork from troubleshooting and accelerating resolution times.
Features
automatically detect patterns in multi-source alarms and reconstruct failure sequences
extract and correlate timestamps across PLC logs, historian data, and sensor records
map observed symptoms to known failure modes and equipment degradation patterns
produce structured reports with evidence, contributing factors, and preventive measures
follow equipment state changes throughout failure window to identify cascading failures
quantify downtime, production loss, and risk exposure from each root cause
flag sensor drift, unexpected state transitions, and out-of-specification readings
recommend maintenance intervals, sensor calibration, or control logic changes
Example Output
Example 1: Pump Failure RCA
ROOT CAUSE: Cavitation due to suction pressure below minimum threshold
TIMELINE:
14:32:15 — Suction pressure drops to 0.8 bar (limit: 1.2 bar)
14:32:45 — Pump vibration alarm triggered
14:33:22 — Discharge pressure unstable; flow drops 40%
14:34:00 — Operator shuts down pump
CONTRIBUTING FACTORS:
• Intake valve partially blocked (visual inspection needed)
• Coolant reservoir below minimum level
• System running outside design specifications
PREVENTIVE ACTION: Schedule intake valve maintenance; lower pump intake pressure setpoint by 0.2 bar
Example 2: Alarm Storm Correlation
- 47 alarms in 3 minutes → 5 root causes identified
- Primary: Loss of communication to PLC-2 (caused 31 downstream alarms)
- Secondary: Temperature sensor drift in Zone C (caused 12 secondary alarms)
- False positives: 3 alarms from cascading sensor failures (not independent issues)
What's Included
- SKILL.md: Complete troubleshooting workflow with decision trees for pump, compressor, valve, and sensor failures
- RCA Report Template: Structured markdown template with root cause, timeline, contributing factors, and preventive actions
- Alarm Correlation Checklist: Step-by-step guide to identify true root causes vs. secondary/cascading alarms
- Failure Mode Decision Tree: Logic for mapping symptoms (pressure drop, vibration, leaks, temperature) to likely equipment issues
- Log Parsing Guidelines: Rules for extracting timestamps, severity levels, and state changes from common SCADA historian formats
- Impact Assessment Worksheet: Quantify downtime cost, production loss, and risk from each root cause
Who It's For
- SCADA & control systems engineers troubleshooting plant failures in real-time
- Industrial plant operators investigating alarm storms and unexpected shutdowns
- Maintenance technicians documenting post-incident root cause analysis for compliance
- Systems reliability engineers analyzing failure trends and improving system resilience
- Operations managers quantifying downtime impact and prioritizing equipment maintenance
Best For
- Emergency diagnostics during active SCADA failures with limited time windows
- Post-incident RCA documentation for regulatory compliance and continuous improvement
- Alarm storm triage to distinguish primary failures from cascading secondary alarms
- Sensor & actuator failure diagnosis with historical trend validation
- Preventive maintenance planning based on failure pattern analysis







