
Log Analyzer
Diagnose production incidents by correlating logs across services
What You Can Do
You can rapidly diagnose production incidents by uploading logs from multiple sources—application servers, databases, network infrastructure—and Claude will identify error patterns, correlate events across services to establish causality, extract key metrics and anomalies, and produce a structured timeline of what went wrong. This transforms raw log data into actionable incident reports with clear root cause identification and remediation priorities.
Features
automatically identifies error signatures and clusters related events from application, system, database, and network logs
establishes causality by correlating events across services with timestamps to show which failure triggered downstream impact
groups similar errors and anomalies to reveal the true scope and frequency of issues
visualizes how failures cascade through microservices and dependent systems
identifies threshold violations, performance degradation, and resource utilization spikes
produces structured chronological narrative with evidence links showing exactly what happened and when
synthesizes log patterns to pinpoint the triggering event and contributing factors
recommends fixes ranked by impact and urgency
Example Output
Example 1: API Service Outage
INCIDENT TIMELINE
09:15:23 — Database connection pool exhausted (PostgreSQL logs)
09:15:45 — API timeout errors spike to 87% (application logs)
09:16:12 — Load balancer marks all API instances unhealthy (infrastructure logs)
09:18:00 — Cache layer shows 0 hits, all requests hitting database (cache logs)
ROOT CAUSE: Deployment at 09:14:58 introduced N+1 query bug. Combined with cache flush during deployment, DB connection pool saturated within 90 seconds.
MITIGATION: Revert deployment, repopulate cache, monitor connection pool.
Example 2: Intermittent Payment Processing Failures
ERROR PATTERN ANALYSIS
Error Type: "timeout connecting to payment gateway"
Frequency: 8.3% of transactions between 09:00–11:30 UTC
Correlated Events: Network latency spike to payment provider (50ms → 2.5s) + circuit breaker opening after 3 consecutive timeouts
SERVICE IMPACT MAP
Payment Service → Retry Queue → Notification Service → Customer Alerts
ACTION ITEMS
1. Increase circuit breaker timeout threshold (1s → 3s)
2. Investigate network path to payment provider
3. Implement exponential backoff on retries
What's Included
- log-analyzer SKILL.md: Core instruction file with pattern recognition algorithms and analysis frameworks
- Multi-source log parsing template: Structure for ingesting logs from application servers, databases, containers, and infrastructure
- Incident timeline reconstruction checklist: Step-by-step format for establishing event causality and impact mapping
- Error pattern clustering worksheet: Framework for grouping similar failures and extracting root causes
- Service dependency impact worksheet: Template for visualizing how failures cascade across microservices
Who It's For
- DevOps Engineers — diagnosing production incidents and post-incident root cause analysis
- SRE/Platform Engineers — systematic troubleshooting and incident timeline reconstruction
- Backend Engineers — understanding application failure modes during active incidents
- On-call responders — rapid triage and diagnosis during incident response
- Infrastructure/Systems Engineers — correlating infrastructure events with application failures
Best For
- Active incident diagnosis — application errors, service failures, or performance degradation requiring rapid root cause identification
- Post-incident analysis — understanding what happened, when, and why after a resolved outage
- Cross-service failures — issues spanning multiple microservices requiring event correlation
- Error spike investigation — sudden increases in error rate with unknown cause
- Performance degradation troubleshooting — response time increases, throughput drops, or resource exhaustion without obvious triggers






