
SRE Capacity Planner
Forecast infrastructure needs and model scaling decisions with data analysis
What You Can Do
You can upload infrastructure metrics, historical usage data, and current resource allocation to get detailed capacity forecasts, growth projections, and scaling recommendations. The skill analyzes your data against industry best practices and models multiple scenarios to identify optimal scaling timing and strategy. It provides actionable recommendations for rightsizing, autoscaling configuration, cost optimization, and risk mitigation.
Features
Project CPU, memory, storage, and network demand using exponential smoothing and trend analysis
Automatically compute safe headroom thresholds (typically 30-50%) and identify when capacity will be breached
Compare vertical scaling, horizontal autoscaling, and multi-region expansion with cost and performance tradeoffs
Calculate capex vs opex impact of different strategies and estimate monthly savings from rightsizing
Identify capacity bottlenecks, single points of failure, and infrastructure dependencies affecting availability
Factor seasonal spikes, traffic events, and business growth projections into forecasts
Get prioritized actions with implementation effort, cost impact, and ROI for each option
Detect data gaps, anomalies, and outliers to ensure forecast reliability
Example Output
Example 1: Cloud Database Scaling
Input: PostgreSQL CPU averaging 62%, disk at 78% capacity, 12% monthly growth
Output:
- Current headroom: 4 weeks at current growth rate
- Recommended action: Upgrade to db.r6i.2xlarge in 3 weeks (cost increase: $580/month)
- Alternative: Implement read replicas for 40% cost savings
- Risk: Single-AZ deployment; recommend multi-AZ for HA (+$420/month)
Example 2: Kubernetes Infrastructure Planning
Input: 3-node cluster, 85% CPU during peak, 2 pending pods, 18% monthly growth
Output:
- HPA recommendation: Scale to 4 nodes when CPU reaches 75% sustained
- Forecast: Need 6 nodes by Q4 2026 based on traffic trends
- Cost comparison: Reserved instances save 38% vs on-demand
- Action: Enable Karpenter for spot optimization (potential 60% savings)
What's Included
- SKILL.md: Complete workflow for data intake, metrics analysis, forecasting, scenario modeling, and recommendation generation
- Metric templates: CSV templates for CPU, memory, storage, and network utilization data collection
- Scenario modeling worksheet: Structured prompts for comparing vertical, horizontal, and multi-region scaling options
- Cost calculator: Built-in analysis for capex vs opex decisions and ROI calculations
- Risk assessment checklist: Identify infrastructure bottlenecks, dependencies, and HA requirements
- How-to guide: Step-by-step walkthrough with real-world infrastructure examples and interpretation guide
Who It's For
- Site reliability engineers (SREs) — Planning infrastructure scaling and long-term capacity strategies
- DevOps engineers — Right-sizing cloud resources and optimizing infrastructure costs
- Infrastructure architects — Designing scalable systems and modeling expansion scenarios
- Engineering managers — Making data-driven decisions on infrastructure investment and timing
- Cloud platform teams — Forecasting resource needs across shared infrastructure environments
Best For
- Capacity forecasting and headroom planning for production systems
- Right-sizing cloud instances and databases to optimize cost without sacrificing performance
- Modeling and comparing scaling strategies with cost and complexity tradeoffs
- Planning infrastructure upgrades and budgeting capex and opex for growth
- Identifying capacity bottlenecks and designing risk mitigation for high-availability systems







