
Azure Deployment Diagnostics & Optimization
Diagnose and optimize Azure deployments in minutes
What You Can Do
This skill analyzes your Azure deployment logs, resource configurations, and performance metrics to identify bottlenecks, cost inefficiencies, and potential failures. You get actionable optimization recommendations, a prioritized remediation plan, and architectural improvements that reduce downtime and cut infrastructure costs by up to 30%.
Features
Parse deployment logs, error traces, and event streams to pinpoint root causes of failures
Identify CPU throttling, memory leaks, network latency, and database query bottlenecks
Detect oversized VMs, unused resources, inefficient storage tiers, and unnecessary replication
Review network security groups, RBAC assignments, encryption configuration, and compliance gaps
Visualize relationships between App Services, databases, caches, and networking to spot single points of failure
Forecast scaling needs based on growth patterns and suggest autoscaling configurations
Analyze IaC (Bicep/ARM templates) for anti-patterns and best practice violations
Verify staging slot configurations, traffic routing rules, and disaster recovery setup
Example Output
Example 1: Performance Diagnosis
- 🔴 CRITICAL: App Service CPU 95%+ sustained for 6+ hours
→ Root cause: N+1 database queries in product search endpoint
→ Fix: Implement query batching + Redis caching (est. 60% CPU reduction)
→ Cost impact: -$180/month on current plan
Example 2: Cost Optimization Report
- 💰 Potential savings: $4,200/month (42% reduction)
• Downsize 3x D4 VMs → D2s (traffic averages 30% utilization)
• Move cold archive data from hot tier → archive tier
• Replace Standard SSD → Premium SSD only for databases
• Auto-shutdown dev/test environments after business hours
Example 3: Remediation Priority Matrix
┌─ Priority 1 (Do Now)
│ • Enable geo-redundancy on SQL Database (RTO: 4 hours → 0)
│ • Restrict App Service public IP, use Application Gateway
│
└─ Priority 2 (This Sprint)
• Implement distributed caching (estimated 40% latency reduction)
• Configure alert rules for database connection timeouts
What's Included
- SKILL.md: Complete diagnostic workflow and optimization decision trees
- Azure Log Parser Template: Extract and categorize error patterns from Application Insights and resource logs
- Cost Analysis Checklist: Service-by-service breakdown of optimization opportunities
- Security Compliance Scanner: RBAC, NSG, encryption, and compliance requirement validation
- Capacity Planning Worksheet: Growth forecasting and autoscaling rule recommendations
- Deployment Validation Checklist: Multi-region failover, backup, and disaster recovery verification
Who It's For
- DevOps Engineers — Troubleshoot deployment failures and optimize CI/CD pipelines
- Cloud Architects — Plan cost-efficient, highly available Azure infrastructure
- Site Reliability Engineers — Reduce MTTR and improve system resilience
- FinOps / Cost Engineers — Identify cloud waste and right-size resource spending
- Platform Engineering Teams — Audit internal Azure platforms for performance and compliance
Best For
- Production troubleshooting — Diagnose live incidents and latency spikes in minutes
- Cost reduction initiatives — Find quick wins and multi-quarter optimization strategies
- Compliance audits — Verify security, encryption, and regulatory requirements
- Capacity planning — Forecast scaling needs and prevent performance degradation
- Post-mortem analysis — Understand failure chains and design preventive mitigations







