Capacity Planning & Scaling Decision Framework for SREs
Turn infrastructure metrics into scaling decisions with forecasting analysis
What You Can Do
Analyze historical infrastructure metrics to forecast resource needs 3-12 months ahead and build data-driven capacity plans. You provide CPU, memory, disk, and network utilization data, and Claude generates bottleneck assessments, scaling timelines, cost projections, and risk quantifications. This helps you stay ahead of performance constraints while avoiding expensive over-provisioning.
Features
Parse Prometheus, CloudWatch, Grafana, or CSV exports into standardized time series
Project demand using trend analysis, seasonality detection, and organic growth modeling
Identify which resource (CPU, memory, disk I/O, network) will fail first and when
Generate step-by-step upgrade plans with timing, instance sizing, and cost deltas
Compare vertical vs. horizontal scaling with ROI and payback period calculations
Quantify over-provisioning waste and under-provisioning outage risk in dollar terms
Export side-by-side comparisons of upgrade paths (upgrade vs. add nodes vs. re-architect)
Plan scaling for compute, databases, caches, and load balancers independently
Example Output
Example 1: Utilization & Forecast
- Current CPU: 74% peak, 52% average
- Growth trend: +7% quarterly
- Forecast: CPU exhaustion by March 2027 (7 months)
- Recommended action: Provision 2x vCPU upgrade in Q1 2027 (~$180/month delta)
Example 2: Risk vs. Cost
- Over-provisioning cost: $48K/year (currently 40% excess capacity)
- Under-provisioning risk: 22% probability of peak-load outage (revenue impact: $1.8M)
- Optimal target: 80% utilization (save $12K annually, accept 8% outage risk)
Example 3: Scaling Timeline
| Quarter | Action | Cost | Risk Level |
|---|---|---|---|
| Q3 2026 | Monitor baseline | $0 | Safe |
| Q1 2027 | Database replication | $18K | Reduced |
| Q3 2027 | Compute tier upgrade | $24K | Minimal |
What's Included
- SKILL.md: Complete capacity planning workflows with decision trees and verification checklists
- Metric parsing templates: Pre-built prompts for Prometheus, CloudWatch, Grafana, and CSV formats
- Forecasting models: Linear trend, exponential growth, and seasonal decomposition approaches
- Scaling recommendation checklist: Resource configs, migration steps, and rollback procedures
- Cost-benefit calculator: Template for comparing upgrade paths with ROI and payback periods
- Risk quantification sheet: Formula for outage probability + revenue impact at different utilization levels
- Sample reports: Example capacity plans showing charts, decision matrices, and executive summaries
Who It's For
- Site reliability engineers (SREs) planning quarterly infrastructure growth and budget forecasting
- DevOps engineers forecasting database, cache, and compute tier upgrades
- Cloud architects assessing multi-region or multi-cloud scaling strategies
- Engineering leaders building data-driven business cases for infrastructure investments
- Platform teams optimizing cost vs. availability tradeoffs and reducing cloud waste
Best For
- Quarterly capacity reviews — Forecast next 12 months, identify bottlenecks, plan upgrades
- Preventing outages — Detect capacity exhaustion 3-6 months ahead before it impacts users
- Right-sizing infrastructure — Eliminate over-provisioning and reduce cloud bills by 15-25%
- Investment justification — Present data-driven cases for infrastructure spending to leadership
- Comparing upgrade strategies — Vertical vs. horizontal vs. migrate: see costs, timelines, and risk tradeoffs side-by-side






