
Monitoring Setup Guide
Design observability stacks with Prometheus, Grafana, and SLO-aligned alerts
What You Can Do
You can design end-to-end observability architectures for cloud-native applications by defining golden signals (latency, traffic, errors, saturation) per service, generating production-ready Prometheus configurations and alert rules aligned to SLOs, and creating meaningful Grafana dashboards. This skill translates business requirements and SLOs into measurable observability coverage with runbook linkage and thresholds that reduce alert fatigue.
Features
Identifies and quantifies latency, traffic, error rate, and saturation metrics specific to each service
Converts Service Level Objectives into Prometheus alert rules with appropriate thresholds and severity levels
Generates scrape configs, recording rules, and metric relabeling strategies for multi-service environments
Structures dashboards with layered views (overview, service detail, dependency tracing) for different audience roles
Creates escalation policies and runbook linkage so incidents trigger appropriate responses
Audits existing monitoring against application topology to identify blind spots and redundancy
Maps distributed tracing requirements and log aggregation strategies across service boundaries
Designs isolation and billing-aware metrics for shared infrastructure and multi-team environments
Example Output
Example 1: Payment Service Golden Signals
Latency: p99 < 500ms (alert > 750ms)
Traffic: 100 req/sec nominal (alert < 20 req/sec)
Error Rate: < 0.1% (alert > 0.5%)
Saturation: CPU < 70% (alert > 85%)
Example 2: Prometheus Alert Rule
alert: PaymentServiceHighLatency
expr: histogram_quantile(0.99, payment_request_duration_seconds_bucket) > 0.75
for: 5m
annotations:
summary: "Payment service p99 latency > 750ms"
runbook: "https://wiki.internal/runbooks/payment-latency"
Example 3: Grafana Dashboard Section
- Service health status card (traffic, error rate, availability %)
- Latency distribution (p50, p95, p99)
- Resource usage gauges (CPU, memory, disk)
- Dependency health heatmap (downstream services)
What's Included
- SKILL.md: Full observability design framework and prompt patterns
- Golden Signals Template: Service-specific metric definitions and SLO thresholds
- Prometheus Config Samples: Scrape configs, recording rules, and alert rule examples
- Grafana Dashboard Blueprint: JSON structure and panel configurations for multi-layer views
- SLO-to-Alert Mapping Worksheet: Worksheet to translate business SLOs into actionable alert policies
Who It's For
- DevOps Engineers — Setting up or redesigning monitoring infrastructure for microservices
- SRE Teams — Translating SLOs into alert rules and observability strategies
- Platform Engineers — Building shared observability platforms for multiple teams
- Cloud Architects — Designing observability requirements for cloud migrations
- On-Call Engineers — Developing runbooks and alert policies to reduce incident response time
Best For
- Greenfield observability deployments for new applications or platforms
- SLO definition and alert rule design tied to business outcomes
- Auditing and filling observability gaps in existing monitoring stacks
- Multi-service architecture design requiring coordinated metrics and dashboards
- Alert policy redesign to reduce noise and improve signal-to-noise ratios






