
Linux System Troubleshooting & Diagnostics
Diagnose Linux issues fast with structured troubleshooting
What You Can Do
Rapidly diagnose Linux system problems using a structured troubleshooting methodology that narrows down root causes systematically. You can analyze system logs, identify performance bottlenecks, and resolve service failures with actionable steps. This skill transforms scattered symptoms into focused solutions using proven diagnostic decision trees and Linux best practices.
Features
Parse system logs, application logs, and kernel messages to identify error patterns and correlations
Use decision trees and logical deduction to narrow issues from symptoms to underlying problems
Analyze CPU, memory, disk I/O, and network metrics to pinpoint bottlenecks and resource constraints
Diagnose failed services, dependency issues, and systemd/init problems with remediation steps
Check system settings, permissions, and configurations against expected states and best practices
Follow step-by-step diagnostic checklists for common Linux issues and error scenarios
Get specific commands for deeper investigation tailored to your symptom
Receive prioritized, step-by-step solutions with expected outcomes at each stage
Example Output
Example 1: Disk Space Issue
[DIAGNOSIS]
Symptom: "du -sh / shows 95% used, but find shows no large files"
→ Root Cause: Orphaned or hidden files in /var/log and journal
[REMEDIATION]
✓ Step 1: Check /var/log — du -sh /var/log
✓ Step 2: Compress old logs — gzip /var/log/*.old
✓ Step 3: Clear journal — journalctl --vacuum=30d
✓ Step 4: Verify — df -h (should show improvement)
Example 2: Service Failure
[DIAGNOSIS]
Symptom: "nginx won't start after config change"
Decision Tree:
1. Is systemd running? → systemctl status nginx
→ stderr: "nginx: configuration syntax error"
2. What changed? → nginx -t (test config)
→ Line 45: invalid directive in server block
[REMEDIATION]
✓ Step 1: Restore previous config — cp /etc/nginx/nginx.conf.bak /etc/nginx/nginx.conf
✓ Step 2: Verify syntax — nginx -t
✓ Step 3: Start service — systemctl start nginx
✓ Step 4: Confirm — systemctl status nginx
Example 3: High CPU Usage
[DIAGNOSIS]
Symptom: "System shows 85% CPU but no obvious process"
→ Investigation: top shows kswapd consuming 60% CPU
→ Root Cause: Memory pressure forcing excessive swap
[REMEDIATION]
✓ Step 1: Check memory — free -h (shows 99% used)
✓ Step 2: Identify memory hog — ps aux --sort=-%mem | head -5
✓ Step 3: Action — Terminate process or increase RAM
✓ Step 4: Verify — top (CPU should drop to <20%)
What's Included
- SKILL.md: Structured diagnostic framework with decision trees for common Linux issues
- Diagnostic Checklists: Step-by-step checklists for disk space, service failures, performance, memory, and network problems
- Linux Command Reference: Essential commands for logs (journalctl, tail), metrics (free, top, iostat), and processes (ps, lsof)
- Log Analysis Templates: Structured templates for parsing and correlating system logs, application logs, and error messages
- Decision Tree Workflows: Visual flowcharts for narrowing issues: symptom → investigation → root cause → remediation
Who It's For
- Linux system administrators managing production servers and infrastructure
- DevOps engineers diagnosing infrastructure and deployment failures
- Site reliability engineers (SREs) responding to incidents and service degradation
- Support engineers triaging customer-reported system problems
- Cloud infrastructure engineers troubleshooting container and VM issues
Best For
- Diagnosing unresponsive or slow servers when nothing obvious jumps out
- Analyzing application crashes and service failures across the Linux stack
- Resolving disk space exhaustion and I/O performance problems
- Investigating memory leaks, swapping, and resource exhaustion issues
- Debugging configuration errors, permission issues, and startup failures






