
Growth Experimentation & Metric Analysis Framework
Design and analyze growth experiments with statistical rigor
What You Can Do
This skill helps you design statistically rigorous experiments, calculate exact sample sizes based on your confidence and power requirements, and interpret results with confidence intervals and significance testing. You'll move beyond guesswork to measure what actually drives growth. It provides frameworks for hypothesis formation, experiment design, result interpretation, and prioritizing your experiment backlog by expected impact and effort.
Features
Determine required sample sizes based on baseline rate, desired lift, confidence level (95% or 99%), and statistical power (80% or 90%). Accounts for multiple variants and adjusts for peeking.
Structure experiments with clear control/variant specifications, success criteria, confounding variables, and recommended runtime. Includes randomization strategy and data collection checklist.
Calculate p-values, confidence intervals, and effect sizes (Cohen's h, lift %) from your results. Determines whether improvements are real or due to chance with 95%+ confidence.
Design and analyze factorial experiments (2x2 grids) and multi-armed tests. Calculates interaction effects and helps identify winning combinations.
Structure growth hypotheses with clear problem statement, testable prediction, success metric, and expected uplift. Validates logic before you run expensive experiments.
Rank your experiment backlog by impact/effort matrix. Combines expected lift, sample size required, and implementation effort to create a roadmap.
Translate statistics into clear actions: deploy, iterate, or abandon. Explains what confidence intervals mean and when to trust results despite noise.
Calculate interim decision points for long experiments. Stop early if results are conclusive instead of waiting for full duration.
Example Output
Example 1: Sample Size Calculation Based on your hypothesis (convert 25% to 28%, 95% confidence, 80% power): You need 5,240 users per variant. With 1,200 daily visitors split 50/50, this takes 9 days. Risk: novelty effect could inflate variant lift in days 1-3; recommend collecting 14 days of data.
Example 2: Results Interpretation Control: 1,205/4,890 conversions (24.6%, CI: 23.2%-26.1%) Variant: 1,456/4,920 conversions (29.6%, CI: 28.1%-31.1%) Statistical significance: p < 0.001 ✓ Lift: +20.3% (real, not noise). Cohen's h = 0.110 (small-to-medium effect). Recommendation: Deploy to 25% of users, monitor for 7 days, then full rollout.
Example 3: Experiment Prioritization High-Impact/Low-Effort (Run immediately): Checkout confirmation copy, Email frequency optimization. High-Impact/High-Effort (Q3 roadmap): Pricing model test, AI recommendation engine. Low-Impact/Quick-Win: Button color, Header text variation.
What's Included
- Hypothesis Formulation Template: Structured format for defining clear, testable hypotheses with baseline rates, expected lift, and success metrics.
- Statistical Analysis Toolkit: Functions for power analysis, sample sizing, significance testing, confidence intervals, and effect size calculation.
- Experiment Design Checklist: Risk mitigation guide covering confounding variables, novelty effects, seasonality, and recommended experiment duration.
- Portfolio Prioritization Matrix: Framework for ranking experiments by potential business impact vs. implementation effort and sample size required.
- Results Interpretation Guide: Reference for reading p-values, confidence intervals, effect sizes, and translating statistics into go/no-go decisions.
- Statistical Glossary: Definitions of key concepts (power, p-value, Cohen's h, confidence interval) and when to use each statistical test.
Who It's For
- Product Manager
- Growth Manager or Growth Hacker
- Data Analyst or Analytics Manager
- Startup Founder or CEO
- Conversion Rate Optimization Specialist
Best For
- A/B testing strategy and experimental design
- Sample size and statistical power calculations
- Interpreting experiment results and lift quantification
- Prioritizing experiment backlogs by impact and effort
- Determining statistical significance and confidence intervals



