SkillsLib.ai

GCP Infrastructure Troubleshooting & Diagnostics Guide

Diagnose and fix GCP incidents with structured diagnostics and root cause analysis

0.0(0 reviews)
100+ downloads
Updated Oct 2026

What You Can Do

Rapidly triage infrastructure incidents across all GCP services using guided diagnostic workflows, root cause decision trees, and remediation procedures. You get structured steps to isolate failures in Compute Engine, Cloud Run, Cloud SQL, networking, and storage, complete with gcloud commands, health checks, and rollback procedures.

Features

Structured diagnostic workflows for Compute, Networking, Databases, Storage, and Managed Services

systematically rule out causes

Root cause decision trees with probability scoring

identify likely failure mode (70% OOM, 20% config, 10% dependency)

Pre-built health check templates

copy-paste gcloud commands, SQL queries, and kubectl diagnostics

Log pattern analysis

parse errors, trace multi-service failures, correlate timestamps across services

Step-by-step remediation procedures

numbered fix steps with safety checks and prerequisites

Incident timeline reconstruction

map events and changes leading up to the failure

Dependency mapping

visualize service interactions to find cascade failures

Rollback and recovery planning

safe revert procedures if fixes fail

Example Output

Issue: Cloud Run service returning 502 errors

Root Cause Analysis:

  • Container OOM/startup timeout (72%)
  • Service configuration mismatch (18%)
  • Upstream API failure (10%)

Immediate Diagnostics:

code
gcloud run services describe my-api --region us-central1
gcloud run services logs read my-api --region us-central1 --limit=50

Remediation (OOM Case):

  1. Increase memory: gcloud run services update my-api --memory 512Mi
  2. Redeploy and verify
  3. Monitor: gcloud monitoring timeseries list --filter='metric.type=run.googleapis.com/request_count'

Rollback Plan: gcloud run services update-traffic my-api --to-revisions=[PREVIOUS-ID]=100


Issue: Compute Engine VMs can't communicate

Diagnosis Tree:

  1. Network interface up? → gcloud compute instances describe [VM] --zone [ZONE]
  2. Firewall rule allows traffic? → gcloud compute firewall-rules list --filter='sourceRanges:10.0.0.0/8'
  3. Routes configured? → gcloud compute routes list --filter='destRange:10.0.0.0/8'

Fix: Apply allow-internal firewall rule → Test ping → Verify connectivity → Monitor logs

What's Included

  • SKILL.md: Complete diagnostic framework with multi-service workflows
  • GCP Service Health Checks: Ready-to-run gcloud and gsutil commands for Compute, Cloud SQL, Cloud Storage, Networking
  • Incident Response Runbooks: Step-by-step fixes for common scenarios (OOM, timeouts, permission errors, quota limits)
  • Root Cause Decision Trees: Structured workflows to isolate failures
  • Log Analysis Checklists: What to search for in Cloud Logging, Application Insights, and service-specific logs
  • Dependency Mapping Templates: Visualize multi-tier failures
  • Recovery Procedures: Rollback, data recovery, and safe remediation steps

Who It's For

  • GCP Infrastructure Engineers — Design and maintain cloud infrastructure
  • DevOps & SRE Teams — On-call incident response and system reliability
  • Platform Engineers — Support internal developer platforms and service reliability
  • Cloud Architects — Design resilient, observable systems and validate operational readiness
  • On-Call Responders — Rapid triage and remediation during production incidents

Best For

  • Production incident triage and diagnosis — Quickly identify root cause and severity
  • GCP service health assessment — Validate compute, database, network, and storage layer health
  • Performance degradation investigation — Track latency, CPU, memory, and I/O bottlenecks
  • Cross-service dependency troubleshooting — Map and debug failures across multiple GCP services
  • Post-incident root cause analysis — Reconstruct incident timeline and document prevention steps

You might also like

Open Data Story Discovery for Journalists
$30
Open Data3.4(5)
Open Data Story Discovery for Journalists

You can rapidly evaluate public datasets to uncover newsworthy patterns and develop story angles backed by reproducible analysis. Claude helps you assess data quality, identify anomalies, generate multiple narrative angles, and document your methodology so editors and fact-checkers can verify your findings. Turn raw data into compelling stories faster than traditional research.

Docker Container Troubleshooting & Performance Optimization
$35
Docker Container Troubleshooting & Performance Optimization

Analyze Docker container logs instantly to identify root causes of failures, crashes, and resource bottlenecks. Get actionable recommendations for resource allocation, network configuration, and architectural improvements that eliminate production issues before they cascade.

Terraform Rapid Module Design & Review
$45
Terraform Rapid Module Design & Review

Quickly architect scalable, reusable Terraform modules that follow HashiCorp best practices and organizational standards. Get automated reviews that catch common pitfalls—variable naming, provider configuration, resource dependencies—before they reach production. Standardize your infrastructure-as-code across teams with instant feedback on module quality, security posture, and cost optimization opportunities.

Story Development & Editorial Workflow
$40
News3.3(6)
Story Development & Editorial Workflow

Manage your entire story development pipeline—from assignment briefs and source research guidance to fact-checking verification and copy editing—all within Claude's context. You'll generate assignment templates, receive real-time editorial feedback, identify verification gaps, and receive copyediting suggestions with tracked changes. The skill handles complex multi-source stories, deadline pressure, and maintains editorial standards across your publication.

Interactive Data Narrative Builder
$35
Interactive Data Narrative Builder

You can transform complex datasets into compelling interactive narratives that maintain coherence while inviting exploration. This skill helps you architect data-driven stories where insights unfold naturally through guided discovery points, keeping readers engaged without limiting their agency. Readers navigate your narrative with purpose, uncovering patterns and relationships that would remain hidden in static presentations.

Grafana Dashboard Design & Query Optimization
$30
Grafana Dashboard Design & Query Optimization

Create production-ready Grafana dashboards that visualize your metrics effectively and perform efficiently. Claude helps you write optimized PromQL and other database queries, configure smart alerts with custom thresholds, and design intuitive layouts that surface the metrics that matter most. You get dashboard JSON, query recommendations, and alert configurations ready to deploy.

Investigative Evidence Architecture
$30
Investigative Evidence Architecture

Map complex investigations into structured evidence chains that verify source credibility, identify logical gaps, and build defensible narratives before publication. You can validate citations, assess counter-arguments, and document your research methodology in publication-ready format. This ensures your findings withstand scrutiny and legal challenges.

Visual Story Angles & Assignment Analysis for News
$40
News3.8(5)
Visual Story Angles & Assignment Analysis for News

Transform breaking news briefs into compelling visual story frameworks that guide photographers toward impactful coverage. You'll receive multiple narrative angles, detailed shot lists organized by scene and purpose, and complete assignment briefs with sourcing guidance. Each output ensures comprehensive emotional and contextual storytelling from first frame to final edit.

$35.00