
A/B Test Statistical Validator
Validate A/B test results with statistical rigor and prevent false positives
What You Can Do
You can confidently evaluate whether your A/B test results are statistically significant before implementing winners. The skill calculates p-values, z-scores, confidence intervals, and validates sample sizes against baseline variance to identify underpowered tests. You'll receive executive-ready recommendations on whether to stop testing, run longer, or implement with confidence—eliminating guesswork from conversion optimization decisions.
Features
Determines true statistical significance vs. random variation
Quantifies the range of true conversion lift with precision
Confirms your test has adequate power to detect meaningful effects
Calculates expected conversion improvement and business impact
Flags underpowered tests and sequential testing risks
Extends analysis to A/B/n tests with proper statistical adjustments
Identifies external factors and lurking variables affecting results
Recommends implementation, extended testing, or rejection with reasoning
Example Output
Example 1: Statistically Significant Result
Test Data: Control 2,000 visitors, 180 conversions (9%). Variant 2,000 visitors, 220 conversions (11%).
✓ P-value: 0.0234 (below 0.05 threshold) ✓ Confidence Interval: 1.2% to 2.8% lift (95% confidence) ✓ Statistical Power: 82% (adequate) ✓ Recommendation: Implement variant. Result is robust with sufficient sample size.
Example 2: Underpowered Test
Test Data: Control 500 visitors, 35 conversions (7%). Variant 500 visitors, 42 conversions (8.4%).
⚠ P-value: 0.089 (above 0.05 threshold) ⚠ Statistical Power: 35% (UNDERPOWERED) ⚠ Recommendation: Run test longer. Current sample detects only large effects. Need ~3,200 visitors per variant for 80% power at this baseline.
Example 3: False Positive Risk
Test Data: 10 variants tested simultaneously, one shows p = 0.04.
⚠ Multiple Comparisons Adjusted P-value: 0.40 (fails correction) ⚠ Risk: 40% chance of false positive without Bonferroni adjustment ✓ Recommendation: Use sequential A/B testing or decrease alpha threshold to 0.005 for 10 comparisons.
What's Included
- SKILL.md instruction file with full statistical validation framework:
- Statistical Calculator Template: Copy-paste formulas for p-value, z-score, and confidence interval computation
- Test Readiness Checklist: Pre-test validation checklist for sample size, duration, and power requirements
- Winner Selection Decision Matrix: Framework for interpreting results across significance thresholds and business context
- Executive Summary Template: Stakeholder-ready report format with risk callouts and recommendations
Who It's For
- Conversion Rate Optimization (CRO) specialists validating A/B test results before implementation
- E-commerce product managers deciding between variant winners and extended testing
- Data analysts documenting statistical rigor for compliance or QA review
- Marketing directors presenting test conclusions to executive leadership
- Growth team leads preventing costly false-positive deployments
Best For
- Validating A/B test statistical significance before winner implementation
- Calculating confidence intervals and conversion lift ranges
- Detecting underpowered tests requiring extended sample collection
- Recommending optimal test duration and sample size for future experiments
- Assessing multi-variant (A/B/n) test results with multiple comparison correction
- Documenting statistical reasoning for compliance, audit, or stakeholder review







