
Network Troubleshooting & Diagnostic Analysis
Diagnose network issues systematically with structured workflows and decision trees
What You Can Do
You'll work through complex network problems using structured diagnostic workflows and decision trees that guide you from symptoms to root cause. This skill accelerates your MTTR by eliminating guesswork, helping you quickly identify whether issues stem from connectivity, DNS, firewall, routing, or performance problems. You'll also capture findings into runbooks to prevent recurrence and build team knowledge.
Features
Navigate from symptoms to root cause using systematic flowcharts that narrow down the problem space
Extract actionable signals from syslog, tcpdump, netstat, and monitoring data to spot anomalies
Skip the guesswork and pinpoint issues in minutes instead of hours through structured triage
Collect, timestamp, and correlate relevant logs, metrics, and system state into structured reports
Generate step-by-step fix procedures tailored to your specific issue and environment
Turn incidents into reusable playbooks for on-call teams to follow during future recurrence
Document root causes, prevention measures, and process improvements after outages
Know when to involve database, application, or infrastructure teams based on diagnostic findings
Example Output
Example 1: Diagnostic Decision Tree
- π SYMPTOM: High packet loss on datacenter link
ββ Check 1: Is link physically up? (ethtool, ip link)
β ββ YES β Check 2: Any FCS/CRC errors? (ethtool -S)
β ββ YES β Action: Suspect cable/NIC hardware issue
β ββ NO β Check 3: CPU dropping packets? (netstat -s)
ββ NO β Action: Check optics/SFP, escalate to NOC
Example 2: Root Cause Analysis Output
| Evidence | Finding | Interpretation |
|---|---|---|
| TCP retransmit spike at 14:32 | Packet loss began after config deploy | Likely routing or ACL change |
| BGP flap log entries | Route convergence took 45 seconds | Temporary asymmetric routing |
| Latency spike correlates with retransmits | Client timeouts after 30s | Application impact: ~2M requests failed |
| Root Cause: BGP configuration deploy caused 45-second asymmetric routing window |
Example 3: Escalation Playbook
- If packet loss > 5% lasting > 5 min β Page network ops immediately (SEV-2)
- If DNS query RTT > 500ms β Contact DNS infrastructure team
- If inter-datacenter latency > 2x baseline β Check WAN provider status, then escalate
What's Included
- SKILL.md: Complete diagnostic skill with workflows and templates
- Diagnostic decision trees: Flowcharts for connectivity, DNS, firewall, routing, and performance issues
- Log analysis templates: Patterns for syslog, tcpdump, netstat, and route outputs
- Evidence gathering checklist: Standardized list of commands and data to collect during incidents
- Root cause analysis worksheet: Structured template to link symptoms to underlying causes
- Runbook template: Format for documenting fixes and creating preventive playbooks
- Escalation matrix: Severity thresholds and team assignment criteria
- Post-mortem template: Timeline, findings, and prevention measures format
Who It's For
- Network Engineers β Rapidly diagnose complex routing, BGP, and connectivity issues
- DevOps Engineers & SREs β Troubleshoot infrastructure connectivity and performance during incidents
- Systems Administrators β Debug firewall, DNS, and network stack problems on production systems
- Network Operations Center (NOC) Technicians β Follow structured workflows to escalate issues correctly
- On-call responders β Use generated runbooks to respond consistently to recurring issues
Best For
- Diagnosing connectivity outages β Pinpoint whether issue is link down, routing, or ACLs
- Analyzing packet loss and latency β Correlate symptoms with hardware, buffer, or congestion issues
- Troubleshooting DNS failures β Debug resolution timeouts, NXDOMAIN, and response delays
- Debugging firewall and routing problems β Trace packet path and identify policy misconfigurations
- Post-mortem analysis and runbook creation β Document incident findings and build team playbooks






