
GCP Troubleshooting & Root Cause Analysis
Debug GCP issues fast with systematic root cause analysis
What You Can Do
Systematically diagnose and resolve GCP performance, reliability, and cost issues by analyzing logs, metrics, and traces. You get structured troubleshooting workflows that pinpoint root causes—whether it's database latency, resource exhaustion, misconfigurations, or budget overruns—and deliver actionable remediation steps.
Features
Identify bottlenecks in compute, networking, and database layers using Cloud Monitoring metrics and trace data
Parse Cloud Logging to correlate errors, failures, and exceptions with specific GCP services and user impact
Analyze billing data and resource utilization to spot overspending, idle resources, and rightsizing opportunities
Walk through hypothesis testing, evidence collection, and elimination logic to find the actual cause—not symptoms
Review IAM, networking, quotas, and service settings for misconfigurations that trigger cascading failures
Cross-reference multiple signals (CPU, memory, latency, error rate) to isolate where problems originate
Get step-by-step fix recommendations ranked by impact and effort, with rollback strategies
Example Output
Example 1: Performance Issue
- Symptom: Cloud Run service responding slowly (p99 latency 8s, expected <500ms)
- Root Cause: Database connection pool exhausted; queries backing up
- Evidence: Cloud SQL CPU at 95%, active connections at max, query logs show 10+ second waits
- Fix: Increase connection pool size, optimize N+1 query patterns, add query caching
- Validation: Monitor metrics; confirm p99 drops to <600ms within 5 min
Example 2: Cost Spike
- Symptom: Monthly bill jumped 40% unexpectedly
- Root Cause: Unoptimized batch job running 24/7 on overpowered VMs
- Evidence: Compute Engine sustained use is 730 hours/month at e2-highmem-8
- Fix: Schedule job for peak hours only (6 hour window), downsize to e2-medium
- Savings: Est. $850/month
Example 3: Reliability Issue
- Symptom: Cloud Run health checks failing intermittently
- Root Cause: Insufficient memory; OOM kills during traffic spikes
- Evidence: Error Reporting shows OOM exceptions correlated with traffic bursts
- Fix: Increase memory allocation from 512MB to 1GB, set concurrency limit
- Validation: Health checks passing 100% for 24+ hours
What's Included
- SKILL.md: Complete troubleshooting framework with diagnostic decision trees, evidence collection templates, and remediation playbooks
- Diagnostic Checklist: Systematic workflow to eliminate causes: performance, reliability, cost, or configuration
- Evidence Collection Templates: Formatted queries for Cloud Logging, Cloud Monitoring dashboards, and metric exports
- Root Cause Analysis Matrix: Cross-reference symptoms to probable causes and validation steps
- Remediation Workflows: Ranked fixes by impact/effort for common GCP issues
- Validation Templates: Before/after metric comparisons to confirm fixes worked
- Incident Playbooks: Real-world runbooks for latency, errors, resource exhaustion, and budget overruns
Who It's For
- DevOps engineers — Respond to production incidents and resolve misconfigurations quickly
- Site reliability engineers (SREs) — Reduce MTTR by systematically isolating root causes
- Cloud architects — Audit existing deployments for performance and cost inefficiencies
- Backend developers — Debug application issues in production environments
- Platform engineers — Build runbooks and incident playbooks for your team
Best For
- Production incident response — Quickly isolate root causes and execute fixes during outages
- Performance optimization — Identify latency bottlenecks and improve user experience
- Cost reduction — Spot overspending patterns and rightsize resources
- Reliability improvements — Prevent recurring failures by fixing underlying misconfigurations
- Post-incident review — Analyze logs and metrics to understand what happened and why





