
Experiment Design for A/B Testing & Causal Inference
Design statistically sound A/B tests and causal experiments in minutes
What You Can Do
This skill guides you through rigorous experiment design from hypothesis to analysis plan. You'll compute the statistical power needed for your tests, identify confounding variables to control for, design optimal randomization strategies, and generate pre-registered analysis plans that prevent p-hacking and publication bias. Whether you're running a product A/B test or investigating causal effects in observational data, you get structured, validated designs that account for real-world constraints.
Features
instantly computes required sample sizes for your target effect sizes, significance levels, and baseline conversion rates
ready-to-use frameworks for A/B tests, multivariate tests, quasi-experimental designs, and causal inference studies
systematically walks through potential confounding variables and recommends statistical or experimental controls
helps you define primary metrics, secondary metrics, and guardrail metrics with clear success criteria and sensitivity thresholds
produces detailed, pre-registered analysis plans that specify primary hypotheses, subgroup analyses, and robustness checks upfront
recommends stratification, blocking, or matching approaches optimized for your sample size and constraints
shows how statistical power changes across a range of effect sizes and sample sizes to stress-test your assumptions
guides you through directed acyclic graphs (DAGs), adjustment sets, and validity threats for causal claims
Example Output
Example 1: Power Analysis for Product A/B Test
Inputs: 5% baseline conversion rate, 10% relative lift target, 80% power, 95% significance
Output:
Required sample size per group: 15,472 users
Total required users: 30,944
Expected runtime (at 5,000 users/day): 6.2 days
MDE at current N=10k per group: 12.6% relative lift
Example 2: A/B Test Design Checklist
- ✅ Hypothesis: Larger checkout button increases conversion by 8–12%
- ✅ Primary metric: Checkout completion rate (target 5% lift)
- ✅ Secondary metrics: AOV, return rate, support tickets
- ✅ Guardrails: Session duration, cart abandonment (no >5% change)
- ✅ Randomization: Stratified by device type + geography
- ✅ Excluded: New users (<1 day old), bot-flagged sessions
- ✅ Duration: Fixed 2 weeks, stopping rule disabled
- ✅ Analysis: Intent-to-treat (all randomized users)
Example 3: Causal Inference DAG
Advertising Spend → Revenue
↓
Market Sentiment → Revenue
↓
Competitor Price
Confounders to adjust for: Market Sentiment
Adjustment set: {Market Sentiment}
Causal claim validity: Strong (RCT not feasible, but DAG-guided design mitigates bias)
What's Included
- SKILL.md: Complete experiment design methodology with decision trees and validation workflows
- A/B Test Design Template: Pre-built structure with hypothesis, success metrics, sample size calculator, and stopping rules
- Causal Inference Checklist: Confounder identification, DAG construction, and validity threat assessment
- Metric Definition Worksheet: Define primary, secondary, and guardrail metrics with sensitivity thresholds
- Analysis Plan Template: Pre-registered analysis specifications (SAGE framework: Specific, Appropriate, Generalized, Efficient)
- Sample Size Reference Card: Lookup tables and formulas for common effect sizes and significance levels
- Randomization Strategy Guide: Stratification, blocking, and matching best practices
Who It's For
- Product managers — designing feature A/B tests with confidence in statistical rigor
- Data scientists — planning causal inference studies and observational analyses
- UX researchers — conducting user experiments with proper controls and power analysis
- Growth teams — running conversion optimization tests with validated designs
- Academic and healthcare researchers — ensuring experiment designs meet publication and regulatory standards
Best For
- Calculating required sample size — determining if you need 1,000 or 100,000 users for statistical significance
- Preventing measurement bias — identifying confounders and designing controls to isolate causal effects
- Meeting publication standards — creating pre-registered analysis plans that prevent p-hacking and HARKing
- Designing experiments for observational data — constructing causal DAGs and selecting adjustment sets when randomization isn't possible
- Stress-testing assumptions — running sensitivity analyses to see how power degrades if your effect size or baseline rates are wrong







