SkillsLib.ai

Systematic Feature Engineering & Selection for Production ML

Engineer and select high-impact features for production ML models

0.0(0 reviews)
100+ downloads
Updated Sep 2026

What You Can Do

This skill guides you through systematic feature engineering to identify and validate features that drive real model performance. You'll leverage Claude to analyze your data, generate feature hypotheses, validate statistical significance, and quantify business impact—all with workflows designed to catch pitfalls before production deployment.

Features

Data exploration workflow

Analyze distributions, correlations, and multicollinearity to identify candidate features from raw data

Feature hypothesis generation

Brainstorm domain-specific engineered features using Claude's domain knowledge

Statistical validation

Test feature significance, stability, and collinearity across data splits

Business impact scoring

Quantify expected model lift, ROI, and operational cost-benefit for each feature

Production readiness checklist

Verify data quality, monitoring, and deployment stability before going live

Feature selection framework

Rank candidates by performance, complexity, and stability trade-offs

Interaction detection

Identify multiplicative and non-linear relationships in your feature space

Data catalog templates

Auto-generate feature definitions, lineage, and monitoring specifications

Example Output

Feature Impact Analysis:

code
feature_name: customer_lifetime_days
statistical_significance: p-value < 0.001 ✓
business_lift: +3.2% churn reduction
stability_score: 0.92 (consistent across CV folds)
recommendation: INCLUDE in v1.0

Feature Engineering Candidates:

code
original_feature: purchase_amount
engineered_candidates:
  1. log_purchase_amount
     → Gini importance: 12.3%, normal distribution
  2. purchase_amount_percentile  
     → rank-based, robust to outliers
  3. purchase_per_day_ratio
     → rate normalization, business-relevant
ranked_recommendation: log_purchase_amount (best stability)

What's Included

  • SKILL.md: Complete feature engineering workflow with decision trees and Claude prompts
  • Data exploration template: Statistical analysis checklist with Python code examples
  • Feature hypothesis generator: Guided brainstorm template for domain-specific features
  • Statistical validation worksheet: Correlation, significance, and stability testing framework
  • Business impact scoring matrix: ROI and lift estimation template
  • Production readiness checklist: Data quality, monitoring, and deployment readiness assessment
  • Feature monitoring specification: Drift detection and SLA monitoring template
  • Data catalog documentation guide: Lineage, definitions, and ownership templates

Who It's For

  • Data scientists building classification and regression models with limited feature budgets
  • ML engineers optimizing model performance and reducing technical debt before production
  • Analytics engineers designing feature tables and scalable data pipelines
  • Product managers evaluating which ML improvements drive business outcomes
  • Data-driven teams standardizing feature engineering practices across projects

Best For

  • Reducing feature count without sacrificing model performance or adding complexity
  • Discovering new features from existing raw data through systematic analysis
  • Validating engineered features are statistically significant and stable across data splits
  • Quantifying business impact of model improvements for stakeholder alignment
  • Building reproducible workflows that scale as your data and models grow

You might also like

Experiment Design & Statistical Analysis for Research Engineers
$45
Experiment Design & Statistical Analysis for Research Engineers

You can design statistically valid experiments with proper power analysis, choose the right statistical tests for your data type, analyze results while controlling for multiple comparisons, and generate publication-ready reports with accurate interpretation of findings. Claude helps you avoid common statistical pitfalls and ensures your experimental claims are well-supported by evidence.

Analytics Documentation Generator
$25
Analytics Documentation Generator

This skill automatically documents your entire analytics infrastructure by analyzing data sources, transformations, and outputs. You'll generate production-ready data dictionaries with field definitions, lineage maps showing data flow across systems, and transformation documentation that explains logic and dependencies. Save weeks of manual documentation work while keeping your analytics stack discoverable as it evolves.

BI Data Quality Investigator
$30
BI Data Quality Investigator

You'll systematically diagnose data quality problems by developing structured root cause analysis frameworks, calculating the true business impact, and creating reproducible validation tests. This skill walks you through hypothesis-driven investigation, data lineage analysis, and remediation planning — turning data issues into documented fixes and preventive measures.

RAG Pipeline Optimization with Claude
$40
RAG Pipeline Optimization with Claude

You can systematically evaluate and improve your RAG pipelines using Claude as a design partner. You'll analyze retrieval quality, identify bottlenecks in your embedding and chunking strategies, and receive actionable recommendations to reduce hallucinations and improve context relevance. By the end, you'll have a data-driven optimization plan tailored to your specific use case and performance metrics.

Data Quality Test Framework Builder
$35
Data Quality Test Framework Builder

You'll build comprehensive data quality test suites that validate transformations, detect anomalies, and document standards across your dbt and SQL pipelines. This skill generates production-ready test configurations, anomaly detection protocols, and validation rules that catch data issues before they impact analytics.

HEOR Evidence Synthesis & Dossier Builder
$35
HEOR3.6(5)
HEOR Evidence Synthesis & Dossier Builder

You can rapidly compile, organize, and format disparate health economic evidence—from clinical trials to cost-effectiveness analyses—into structured, regulatory-compliant dossiers. This skill maps your evidence to specific payer and HTA requirements, automatically generates evidence hierarchies, and produces submission-ready dossier outlines with formatting that meets regulatory standards for NICE, EUnetHTA, and other major bodies.

Payer Evidence Synthesis & HTA Builder
$30
Payer4.0(3)
Payer Evidence Synthesis & HTA Builder

You can structure comprehensive health technology assessments (HTAs) that organize clinical evidence, economic analyses, and regulatory considerations into evidence-based coverage recommendations. This skill helps you synthesize clinical trial data, health economic models, and real-world evidence into clear, defensible payer coverage determinations. You'll generate professional HTA reports that align with major frameworks like ICER, CADTH, and NICE standards.

ML Infrastructure Failure Analysis & Optimization
$25
ML Infrastructure Failure Analysis & Optimization

This skill helps you systematically diagnose failures in distributed ML training and serving infrastructure. You provide system logs, metrics, and error traces, and Claude performs structured root-cause analysis to identify the underlying issue—whether it's resource exhaustion, distributed system deadlock, data pipeline corruption, or model serving misconfiguration. You get a detailed diagnosis with remediation steps ranked by likelihood and implementation effort.

$35.00