
Systematic Feature Engineering & Selection for Production ML
Engineer and select high-impact features for production ML models
What You Can Do
This skill guides you through systematic feature engineering to identify and validate features that drive real model performance. You'll leverage Claude to analyze your data, generate feature hypotheses, validate statistical significance, and quantify business impact—all with workflows designed to catch pitfalls before production deployment.
Features
Analyze distributions, correlations, and multicollinearity to identify candidate features from raw data
Brainstorm domain-specific engineered features using Claude's domain knowledge
Test feature significance, stability, and collinearity across data splits
Quantify expected model lift, ROI, and operational cost-benefit for each feature
Verify data quality, monitoring, and deployment stability before going live
Rank candidates by performance, complexity, and stability trade-offs
Identify multiplicative and non-linear relationships in your feature space
Auto-generate feature definitions, lineage, and monitoring specifications
Example Output
Feature Impact Analysis:
feature_name: customer_lifetime_days
statistical_significance: p-value < 0.001 ✓
business_lift: +3.2% churn reduction
stability_score: 0.92 (consistent across CV folds)
recommendation: INCLUDE in v1.0
Feature Engineering Candidates:
original_feature: purchase_amount
engineered_candidates:
1. log_purchase_amount
→ Gini importance: 12.3%, normal distribution
2. purchase_amount_percentile
→ rank-based, robust to outliers
3. purchase_per_day_ratio
→ rate normalization, business-relevant
ranked_recommendation: log_purchase_amount (best stability)
What's Included
- SKILL.md: Complete feature engineering workflow with decision trees and Claude prompts
- Data exploration template: Statistical analysis checklist with Python code examples
- Feature hypothesis generator: Guided brainstorm template for domain-specific features
- Statistical validation worksheet: Correlation, significance, and stability testing framework
- Business impact scoring matrix: ROI and lift estimation template
- Production readiness checklist: Data quality, monitoring, and deployment readiness assessment
- Feature monitoring specification: Drift detection and SLA monitoring template
- Data catalog documentation guide: Lineage, definitions, and ownership templates
Who It's For
- Data scientists building classification and regression models with limited feature budgets
- ML engineers optimizing model performance and reducing technical debt before production
- Analytics engineers designing feature tables and scalable data pipelines
- Product managers evaluating which ML improvements drive business outcomes
- Data-driven teams standardizing feature engineering practices across projects
Best For
- Reducing feature count without sacrificing model performance or adding complexity
- Discovering new features from existing raw data through systematic analysis
- Validating engineered features are statistically significant and stable across data splits
- Quantifying business impact of model improvements for stakeholder alignment
- Building reproducible workflows that scale as your data and models grow







