
GCP Infrastructure Troubleshooting & Optimization Analyzer
Diagnose, analyze, and optimize your GCP infrastructure automatically
What You Can Do
You can rapidly diagnose infrastructure problems across GCP services by feeding Claude your metrics, logs, and configuration data. Claude analyzes the data against best practices, identifies bottlenecks and misconfigurations, and generates actionable optimization recommendations organized as executable runbooks. You'll get a prioritized list of issues with step-by-step remediation steps, cost-saving opportunities, and performance tuning guidance.
Features
Parse CloudMonitoring data to identify latency, throughput, and resource utilization issues
Analyze GCP Cloud Logging entries to pinpoint error patterns and performance degradation
Evaluate GCP resource configuration (Compute Engine, Cloud SQL, GKE) against security and performance baselines
Detect overprovisioned resources, underutilized instances, and savings opportunities
Produce structured remediation steps with estimated impact and rollback procedures
Link issues across Compute, Networking, Database, and Kubernetes services
Flag deviations from GCP Well-Architected Framework and industry standards
Quantify expected improvements in latency, cost, or availability from each recommendation
Example Output
Example 1: Cost Optimization Report
Issues Found
- 7 n1-standard-4 instances running at <5% CPU utilization
- 2 idle e2-standard-2 VMs in us-central1 with 90 days no traffic
- Expired committed use discounts with no replacement
Top Recommendation: Right-size underutilized instances → Est. savings: $1,240/month
Example 2: GKE Performance Analysis
Symptoms: Pod startup latency 45s (baseline: 8s), Node CPU 87%, Memory 92%
Root Causes
- Insufficient node pool capacity — 3 nodes for 120 pods
- Missing resource requests/limits — 8 OOMKills in 24h
- Single-zone deployment — No redundancy
Fix Priority
- Increase node pool to 5 nodes (Est. impact: -65% eviction rate)
- Add resource requests/limits to all deployments
- Enable multi-zone with pod disruption budgets
Example 3: VPC Audit Report
Security Findings: 5 overly-permissive ingress rules, NAT gateway not using reserved IPs, Cloud SQL instance publicly accessible
Compliance Issues: VPC Flow Logs disabled, no encryption in transit on internal traffic
Runbook: 6 actionable steps with gcloud commands to close gaps
What's Included
- SKILL.md: Full troubleshooting workflows, decision trees, metric interpretation guides
- Cost optimization audit checklist: Compute, storage, networking, and managed services review
- GCP performance baseline reference: Expected metrics by instance type, workload, and region
- Log pattern detection examples: Common error signatures and remediation steps
- Runbook generation templates: Structured format for remediation steps with validation
- Security & compliance audit: IAM, encryption, VPC, logging, and industry standard checks
- Impact estimation worksheet: Quantify expected improvements before applying changes
- Rollback & risk mitigation procedures: Safe deployment patterns for high-risk changes
Who It's For
- Cloud Infrastructure Engineers — Manage and troubleshoot GCP deployments at scale
- DevOps/SRE teams — Resolve production incidents and optimize infrastructure
- FinOps specialists — Reduce cloud spend with data-driven recommendations
- Solutions Architects — Conduct GCP health checks and design governance policies
- Platform engineers — Establish GCP best practices across teams
Best For
- Investigating sudden spikes in latency, error rates, or resource utilization
- Reducing cloud costs across Compute, Networking, Database, and managed services
- Troubleshooting GKE cluster issues — pod evictions, node pressure, networking
- Conducting periodic audits of infrastructure health, security, and compliance
- Capacity planning with data-driven rightsizing recommendations






