
Experiment Design Framework
Design rigorous experiments with statistical power and built-in bias guardrails
What You Can Do
You can create statistically rigorous experimental designs from hypothesis to success metrics. Claude helps you calculate proper sample sizes with power analysis, identify hidden biases and confounds, and define clear decision criteria before data collection begins. This prevents p-hacking, ensures adequate statistical power, and produces defensible experimental protocols.
Features
Determines required sample sizes based on effect size, significance level, and desired statistical power using analytic formulas
Identifies confounding variables, selection bias, regression to the mean, and allocation concealment issues early in design
Structures primary and secondary outcomes, statistical thresholds, stopping rules, and decision criteria before running the experiment
Generates designs for randomized controlled trials, A/B tests, factorial experiments, regression discontinuity, and quasi-experimental approaches
Applies Bonferroni, false discovery rate (FDR), or sequential testing corrections to maintain error rates across multiple comparisons
Creates a structured pre-registration document with design details and analysis plan to prevent selective reporting
Tests how robust your conclusions are under different assumptions, effect sizes, and violation of design assumptions
Specifies primary and secondary analyses, subgroup comparisons, and decision rules for handling missing data or protocol deviations
Example Output
Example 1: A/B Test Power Analysis
Experiment: Feature rollout A/B test
Control conversion rate: 5%
Expected effect size: 15% relative lift (5.75% treatment rate)
Significance level (α): 0.05
Statistical power (1-β): 0.80
Required sample size per group: 2,847
Total sample needed: 5,694 users
Expected duration (at 500 daily users): 11-12 days
Bias risks identified:
- Selection bias if users self-select into control/treatment
→ Mitigate: Assign at point of feature exposure
- Novelty effect in treatment group
→ Mitigate: Run test for ≥7 days to allow adaptation
Example 2: Pre-registration Checklist
✓ Primary hypothesis stated before data collection
✓ Sample size justified by power calculation
✓ Primary outcome metric explicitly defined
✓ Statistical test specified (z-test, t-test, Bayesian, etc.)
✓ Significance level set (α = 0.05)
✓ Multiple testing correction applied (n=2 outcomes → Bonferroni α = 0.025)
✓ Stopping rule defined (fixed sample vs. sequential)
✓ Exclusion criteria specified in advance
✓ Potential confounds identified and mitigation strategy noted
What's Included
- SKILL.md: Full experiment design framework with workflows for every phase
- Power analysis worksheet: Templates for calculating sample sizes for common study designs
- Bias identification checklist: Systematic guide to confounds, selection bias, and design flaws
- Experimental design templates: Ready-to-customize designs for RCTs, A/B tests, factorial designs, and quasi-experiments
- Pre-registration template: Structured document for locking in your design and analysis plan
- Success metrics framework: Template for defining primary/secondary outcomes, statistical thresholds, and decision rules
Who It's For
- Product managers designing A/B tests and feature experiments
- Data scientists and statisticians planning research studies
- Academic researchers designing rigorous protocols
- Growth and marketing teams testing campaign hypotheses
- Quality assurance and engineering teams designing controlled experiments
Best For
- Designing A/B tests with adequate statistical power
- Planning clinical trials or academic research with proper controls
- Pre-registering experiments to prevent bias and p-hacking
- Calculating sample sizes for studies and surveys
- Identifying confounds and hidden biases before running experiments







