
A/B Test Statistical Analyzer
Validate A/B test results with statistical rigor and confidence intervals
What You Can Do
You can submit A/B test results and receive statistical validation that separates genuine conversion improvements from random noise. This skill calculates p-values, confidence intervals, statistical power, and required sample sizes—giving you the data-backed confidence to either declare a winner, extend your test, or kill a variant without wasting resources on false positives.
Features
Determines if observed lift is statistically significant at standard confidence levels (95%, 99%)
Calculates whether your test has sufficient power to detect the effect size you care about, identifying underpowered tests before declaring results
Provides bounds around estimated conversion lift so you know the true range of expected improvement
Checks if you've collected enough data and recommends minimum sample sizes for future tests
Warns when results are at risk of Type I or Type II errors based on test design
Translates statistical findings into business-friendly language and decision frameworks
Identifies when to stop tests early or continue based on statistical thresholds
Quantifies expected revenue or engagement impact based on observed conversion differences
Example Output
Input: Control: 2,400 visitors, 180 conversions (7.5%) | Variant: 2,400 visitors, 210 conversions (8.75%)
Output:
- P-value: 0.032 (statistically significant at 95% confidence)
- Observed lift: +16.7%
- 95% confidence interval: [1.2%, 32.1%]
- Statistical power: 62% (underpowered for reliable decision)
- Recommendation: Continue test to 4,800 visitors per variant for 80% power
- Business impact: At current conversion, this lift = +$12K/month (assuming $50 AOV)
Input: Control: 5,000 visitors, 150 conversions (3%) | Variant: 5,000 visitors, 152 conversions (3.04%)
Output:
- P-value: 0.87 (NOT statistically significant)
- Observed lift: +1.3% (likely noise)
- 95% confidence interval: [-1.8%, +4.5%]
- Statistical power: 8% (severely underpowered)
- Recommendation: Kill this variant. Effect size too small to detect reliably even with 10x sample size.
- Next steps: Test variants with predicted 5%+ lift or restructure hypothesis
What's Included
- SKILL.md: Core instruction file with statistical validation framework and decision rules
- A/B Test Data Template: Standardized format for submitting control/variant data (conversions, visitors, duration)
- Statistical Significance Checklist: Step-by-step validation checklist before declaring test winners
- Power Analysis Worksheet: Planning template for determining required sample sizes before launch
- Stakeholder Communication Guide: Pre-written language for explaining p-values, confidence intervals, and statistical limitations to non-technical teams
Who It's For
- Conversion Rate Optimization (CRO) Specialists — Validate test results and defend decisions with statistical evidence
- E-commerce Product Managers — Determine when A/B test data justifies feature rollouts or design changes
- Growth & Experimentation Teams — Calculate power requirements and avoid launching underpowered tests
- Data Analysts & Statisticians — Quickly assess test quality and identify methodological risks
- Marketing Managers — Understand statistical confidence limits of conversion tests before making budget decisions
Best For
- Pre-launch test planning — Calculate required sample sizes and duration before running experiments
- Results validation — Confirm statistical significance and identify false positives before celebrating wins
- Underpowered test diagnosis — Assess whether observed differences are meaningful or require more data
- Stakeholder communication — Translate p-values and confidence intervals into business-friendly decision language
- Test termination decisions — Determine when to kill losing variants or extend winners for higher confidence
- Lift estimation & ROI projection — Quantify expected revenue impact of winning variants







