SkillsLib.ai

Monitoring Setup Guide

Design observability stacks with Prometheus, Grafana, and SLO-aligned alerts

4.1(49 reviews)
100+ downloads
Updated Sep 2026
Verified SafeSecurity VerifiedThis skill was analyzed by our AI security scanner for harmful content including data exfiltration, system manipulation, credential theft, and prompt injection. No threats were detected.

What You Can Do

You can design end-to-end observability architectures for cloud-native applications by defining golden signals (latency, traffic, errors, saturation) per service, generating production-ready Prometheus configurations and alert rules aligned to SLOs, and creating meaningful Grafana dashboards. This skill translates business requirements and SLOs into measurable observability coverage with runbook linkage and thresholds that reduce alert fatigue.

Features

Golden Signal Definition

Identifies and quantifies latency, traffic, error rate, and saturation metrics specific to each service

SLO-to-Alerts Translation

Converts Service Level Objectives into Prometheus alert rules with appropriate thresholds and severity levels

Prometheus Configuration

Generates scrape configs, recording rules, and metric relabeling strategies for multi-service environments

Grafana Dashboard Design

Structures dashboards with layered views (overview, service detail, dependency tracing) for different audience roles

Alerting Policy Framework

Creates escalation policies and runbook linkage so incidents trigger appropriate responses

Observability Gap Analysis

Audits existing monitoring against application topology to identify blind spots and redundancy

Trace Context & Log Correlation

Maps distributed tracing requirements and log aggregation strategies across service boundaries

Multi-Tenant Observability

Designs isolation and billing-aware metrics for shared infrastructure and multi-team environments

Example Output

Example 1: Payment Service Golden Signals

code
Latency: p99 < 500ms (alert > 750ms)
Traffic: 100 req/sec nominal (alert < 20 req/sec)
Error Rate: < 0.1% (alert > 0.5%)
Saturation: CPU < 70% (alert > 85%)

Example 2: Prometheus Alert Rule

code
alert: PaymentServiceHighLatency
expr: histogram_quantile(0.99, payment_request_duration_seconds_bucket) > 0.75
for: 5m
annotations:
  summary: "Payment service p99 latency > 750ms"
  runbook: "https://wiki.internal/runbooks/payment-latency"

Example 3: Grafana Dashboard Section

  • Service health status card (traffic, error rate, availability %)
  • Latency distribution (p50, p95, p99)
  • Resource usage gauges (CPU, memory, disk)
  • Dependency health heatmap (downstream services)

What's Included

  • SKILL.md: Full observability design framework and prompt patterns
  • Golden Signals Template: Service-specific metric definitions and SLO thresholds
  • Prometheus Config Samples: Scrape configs, recording rules, and alert rule examples
  • Grafana Dashboard Blueprint: JSON structure and panel configurations for multi-layer views
  • SLO-to-Alert Mapping Worksheet: Worksheet to translate business SLOs into actionable alert policies

Who It's For

  • DevOps Engineers — Setting up or redesigning monitoring infrastructure for microservices
  • SRE Teams — Translating SLOs into alert rules and observability strategies
  • Platform Engineers — Building shared observability platforms for multiple teams
  • Cloud Architects — Designing observability requirements for cloud migrations
  • On-Call Engineers — Developing runbooks and alert policies to reduce incident response time

Best For

  • Greenfield observability deployments for new applications or platforms
  • SLO definition and alert rule design tied to business outcomes
  • Auditing and filling observability gaps in existing monitoring stacks
  • Multi-service architecture design requiring coordinated metrics and dashboards
  • Alert policy redesign to reduce noise and improve signal-to-noise ratios

You might also like

Chemical Process Safety Analyzer
$40
Chemical Process Safety Analyzer

Analyze chemical processes systematically to identify hazards, assess risks, and generate safety recommendations. You can perform HAZOP analyses, evaluate compliance with industry standards, conduct root-cause analysis of incidents, and develop emergency response procedures. The skill guides you through structured safety reviews that reduce the likelihood of accidents and regulatory violations.

Injectable Formulation Development & Troubleshooting
$40
Injectable Formulation Development & Troubleshooting

You'll develop systematic approaches to injectable formulation design, from API selection through sterilization strategy. Claude helps you troubleshoot failed batches by analyzing root causes, recommends regulatory pathways (505(b)(2), ANDA, NDA), and provides science-backed solutions for stability, compatibility, and manufacturability challenges.

$50
Integration Architecture Assessor

This skill helps you systematically assess integration needs across your systems, design architecture patterns that scale with your organization, and identify technical and operational risks before implementation. You'll receive architecture recommendations aligned to your business constraints, clear integration roadmaps, and risk mitigation strategies that reduce deployment surprises. Get structured decision records suitable for architecture review boards and engineering teams.

Chemical Process Hazard Analysis and Control Design
$30
Chemical Process Hazard Analysis and Control Design

You can conduct comprehensive hazard analyses for chemical processes using industry-standard methodologies like HAZOP and LOPA. The skill helps you assess risks quantitatively, identify control gaps, design engineered safeguards, and generate formal documentation for regulatory compliance and process safety management.

Git Commit Message Writer
$45
CI/CD4.3(47)
Git Commit Message Writer

Claude analyzes your code diffs and generates standardized commit messages that follow the Conventional Commits specification. The skill automatically determines the correct commit type, scope, and description based on the changes you've made, ensuring your messages are parseable by automation tools while remaining human-readable for code reviewers.

Structured NLP Analysis and Annotation with Claude
$35
NLP3.3(6)
Structured NLP Analysis and Annotation with Claude

You can transform raw text into structured, labeled datasets for machine learning, analysis, and research. This skill performs named entity recognition, sentiment classification, part-of-speech tagging, and dependency parsing—generating consistent, validated annotations at scale. Use it to prepare corpora, extract entities, classify documents, or perform linguistic analysis without manual annotation.

Service Mesh Architecture & Troubleshooting
$25
Service Mesh Architecture & Troubleshooting

This skill helps you systematically analyze service mesh architectures, identify inter-service communication failures, and design optimal routing and security policies. You'll receive step-by-step troubleshooting guidance tailored to your mesh platform (Istio, Linkerd, Consul), configuration validation reports, and architectural recommendations that reduce latency, improve observability, and tighten security posture.

IoT Firmware Analysis & Device Debugger
$40
IoT Firmware Analysis & Device Debugger

Rapidly analyze firmware logs and diagnose hardware issues that cause device failures, connectivity problems, and performance degradation. You'll identify root causes from stack traces, crash dumps, and sensor data, then generate specific optimization recommendations. This skill transforms raw device logs into actionable debugging plans that reduce time-to-resolution from hours to minutes.

$20.00