
GCP Infrastructure Troubleshooting & Diagnostics Guide
Diagnose and fix GCP incidents with structured diagnostics and root cause analysis
What You Can Do
Rapidly triage infrastructure incidents across all GCP services using guided diagnostic workflows, root cause decision trees, and remediation procedures. You get structured steps to isolate failures in Compute Engine, Cloud Run, Cloud SQL, networking, and storage, complete with gcloud commands, health checks, and rollback procedures.
Features
systematically rule out causes
identify likely failure mode (70% OOM, 20% config, 10% dependency)
copy-paste gcloud commands, SQL queries, and kubectl diagnostics
parse errors, trace multi-service failures, correlate timestamps across services
numbered fix steps with safety checks and prerequisites
map events and changes leading up to the failure
visualize service interactions to find cascade failures
safe revert procedures if fixes fail
Example Output
Issue: Cloud Run service returning 502 errors
Root Cause Analysis:
- Container OOM/startup timeout (72%)
- Service configuration mismatch (18%)
- Upstream API failure (10%)
Immediate Diagnostics:
gcloud run services describe my-api --region us-central1
gcloud run services logs read my-api --region us-central1 --limit=50
Remediation (OOM Case):
- Increase memory:
gcloud run services update my-api --memory 512Mi - Redeploy and verify
- Monitor:
gcloud monitoring timeseries list --filter='metric.type=run.googleapis.com/request_count'
Rollback Plan: gcloud run services update-traffic my-api --to-revisions=[PREVIOUS-ID]=100
Issue: Compute Engine VMs can't communicate
Diagnosis Tree:
- Network interface up? →
gcloud compute instances describe [VM] --zone [ZONE] - Firewall rule allows traffic? →
gcloud compute firewall-rules list --filter='sourceRanges:10.0.0.0/8' - Routes configured? →
gcloud compute routes list --filter='destRange:10.0.0.0/8'
Fix: Apply allow-internal firewall rule → Test ping → Verify connectivity → Monitor logs
What's Included
- SKILL.md: Complete diagnostic framework with multi-service workflows
- GCP Service Health Checks: Ready-to-run gcloud and gsutil commands for Compute, Cloud SQL, Cloud Storage, Networking
- Incident Response Runbooks: Step-by-step fixes for common scenarios (OOM, timeouts, permission errors, quota limits)
- Root Cause Decision Trees: Structured workflows to isolate failures
- Log Analysis Checklists: What to search for in Cloud Logging, Application Insights, and service-specific logs
- Dependency Mapping Templates: Visualize multi-tier failures
- Recovery Procedures: Rollback, data recovery, and safe remediation steps
Who It's For
- GCP Infrastructure Engineers — Design and maintain cloud infrastructure
- DevOps & SRE Teams — On-call incident response and system reliability
- Platform Engineers — Support internal developer platforms and service reliability
- Cloud Architects — Design resilient, observable systems and validate operational readiness
- On-Call Responders — Rapid triage and remediation during production incidents
Best For
- Production incident triage and diagnosis — Quickly identify root cause and severity
- GCP service health assessment — Validate compute, database, network, and storage layer health
- Performance degradation investigation — Track latency, CPU, memory, and I/O bottlenecks
- Cross-service dependency troubleshooting — Map and debug failures across multiple GCP services
- Post-incident root cause analysis — Reconstruct incident timeline and document prevention steps







