
Production ML Deployment Validation & Runbook Automation
Generate ML deployment checklists, infrastructure code, and incident runbooks
What You Can Do
This skill automates the creation of production-ready deployment validation checklists, infrastructure-as-code templates, and incident response runbooks tailored to your ML stack. You get comprehensive pre-deployment checks covering model validation, data pipeline integrity, infrastructure readiness, and monitoring setup—all customized for your specific models and cloud provider. The skill generates executable runbooks that teams can follow during incidents, including rollback procedures, failover strategies, and diagnostic commands.
Features
Creates comprehensive checklists for model testing, data validation, and infrastructure readiness
Generates Terraform, CloudFormation, or Kubernetes manifests tailored to your ML workload
Produces step-by-step incident procedures including detection, diagnosis, mitigation, and rollback
Generates schemas for model performance testing, data drift detection, and A/B test validation
Creates alert thresholds, log patterns, and metrics definitions specific to your models
Produces deployment security review items aligned with industry standards (SOC2, HIPAA, etc.)
Generates shell and Python scripts for health checks, model inference tests, and data quality validation
Creates lock files and versioning strategies for reproducible deployments
Example Output
Pre-Deployment Checklist (Real-Time Model Serving)
- Model artifacts signed and integrity verified with SHA256
- Performance benchmarks pass: p99 latency < 100ms, throughput > 10K req/sec
- Data pipeline validation: NaN/null rates < 0.1%, feature distributions match training
- A/B test plan approved with 95% confidence interval calculator
- Rollback procedure documented and tested in staging
- CloudWatch dashboards created with 1-min granularity metrics
- Canary deployment configured: 5% traffic for 2 hours, then 50%, then 100%
Incident Runbook: Model Inference Latency Spike
1. Alert triggered: p99 latency > 500ms for 5 consecutive minutes
2. Gather data: kubectl logs -l app=model --tail=100 | grep ERROR
3. Diagnose: Check memory (>85%?), GPU utilization (100%?), batch queue depth
4. Mitigate: Scale replicas → kubectl scale deployment model --replicas=5
5. If unresolved after 10m: kubectl rollout undo deployment/model
Terraform Template (SageMaker Endpoint)
resource "aws_sagemaker_endpoint" "prod_model" {
endpoint_name = "fraud-detector-v2"
endpoint_config_name = aws_sagemaker_endpoint_config.prod.name
data_capture_config {
enabled = true
s3_output_path = "s3://model-monitoring/data-capture"
}
tags = { Environment = "production", Model = "fraud-v2" }
}
What's Included
- SKILL.md file with production ML deployment expertise:
- Pre-deployment validation checklist templates (model, data, infrastructure):
- Infrastructure-as-code templates (Terraform, CloudFormation, Kubernetes):
- Incident response runbook template with decision trees:
- Model validation protocols and test schemas:
- Monitoring and alerting configuration templates:
- Security and compliance review checklists:
- Post-deployment verification scripts (Bash, Python):
- Canary and blue-green deployment strategies:
Who It's For
- MLOps engineers and DevOps teams managing production ML systems
- ML engineering managers coordinating safe deployments and incident response
- Data scientists transitioning models from notebooks to production
- Platform engineers building ML infrastructure and deployment frameworks
- Compliance and security teams auditing ML production deployments
Best For
- Preparing comprehensive checklists before deploying new ML models to production
- Creating incident response procedures for ML service failures and model degradation
- Defining infrastructure automation and deployment strategies for ML workloads
- Establishing monitoring and alerting strategies for model performance and data drift
- Conducting security and compliance audits of ML infrastructure and data pipelines







