
A/B Testing: Statistical Rigor for CRO Specialists
Design statistically sound A/B tests with proper sample sizing and rigorous analysis
What You Can Do
You'll learn to design A/B tests that reliably detect real changes while minimizing false positives. This skill helps you calculate proper sample sizes, determine statistical power requirements, and interpret results using rigorous frequentist methods — ensuring every test delivers actionable insights and saves your team from chasing false signals.
Features
Compute optimal sample sizes based on baseline conversion rate, minimum detectable effect, significance level, and statistical power.
Determine the probability your test will detect a real effect given your sample size, baseline rate, and effect threshold.
Correctly interpret confidence intervals and understand what they reveal about the true effect size in your population.
Identify and adjust for inflated false positive rates when running multiple simultaneous tests or sequential peeks.
Translate business goals into statistical effect sizes (Cohen's h, lift percentage) to inform your test design.
Learn what p-values actually measure, common misinterpretations to avoid, and how to contextualize them with effect sizes.
Calculate how long your test should run to reach statistical power targets and detect the minimum effect size you care about.
Make confident launch decisions by separating statistical significance from practical significance and business impact.
Example Output
Example 1: Sample Size Calculation
You ask: "We have a 3% baseline conversion rate. We want to detect a 0.5% absolute lift (to 3.5%). What sample size do we need with 80% power and 5% significance?"
Claude returns:
- Minimum sample size per variant: 9,223 users
- Total test duration: ~5 days (at 2,000 users/day)
- Power curve visualization: Shows how power increases with sample size
- Trade-offs: "If you reduce power to 70%, you need only 6,800 users per variant (3-4 days)."
Example 2: Interpreting Ambiguous Results
You submit test data:
- Control: 100/3,500 conversions (2.86%)
- Variant: 120/4,100 conversions (2.93%)
- P-value: 0.18
Claude diagnoses:
- ✓ Not statistically significant (p > 0.05)
- ✓ Confidence interval for lift: [-1.2%, +4.8%]
- ✓ Test was underpowered; need 12,000 users per variant to detect 0.5% lift reliably
- ✓ Recommendation: Continue test or redesign with larger effect threshold
Example 3: Multiple Comparison Correction
You say: "I'm running 5 A/B tests simultaneously on different pages."
Claude advises:
- ✓ Uncorrected false positive rate: 22.6%
- ✓ Bonferroni-corrected threshold: p < 0.01 (not 0.05)
- ✓ Alternative: FDR control for less conservative approach
- ✓ Best practice: Prioritize 1-2 tests and run others sequentially
What's Included
- Sample Size Calculator Template: Pre-built formulas and lookup tables for conversion rate tests, continuous metrics, and count data scenarios.
- Statistical Power Curves: Visual relationships between sample size, effect size, power, and significance level to guide experiment design.
- Multiple Comparison Checklist: Step-by-step guidance on identifying when correction is needed and how to apply Bonferroni, Holm, or FDR methods.
- Confidence Interval Interpretation Guide: Templates for correctly stating what confidence intervals mean and translating them into business decisions.
- P-Value and Significance Reference: Common misconceptions about p-values, correct interpretations, and how to avoid misuse in reports.
- Effect Size Threshold Lookup: Industry benchmarks and guidance on what constitutes a practically meaningful effect in your domain.
Who It's For
- CRO Specialists
- Product Managers
- Growth Engineers
- Marketing Analysts
- UX Researchers
Best For
- Designing experiments before launch
- Calculating required sample sizes and test duration
- Interpreting test results and detecting false positives
- Determining practical vs statistical significance
- Troubleshooting inconclusive test results







