SkillsLib.ai

A/B Testing: Statistical Rigor for CRO Specialists

Design statistically sound A/B tests with proper sample sizing and rigorous analysis

3.7(3 reviews)
100+ downloads
Updated Oct 2026

What You Can Do

You'll learn to design A/B tests that reliably detect real changes while minimizing false positives. This skill helps you calculate proper sample sizes, determine statistical power requirements, and interpret results using rigorous frequentist methods — ensuring every test delivers actionable insights and saves your team from chasing false signals.

Features

Sample Size Calculator

Compute optimal sample sizes based on baseline conversion rate, minimum detectable effect, significance level, and statistical power.

Statistical Power Analysis

Determine the probability your test will detect a real effect given your sample size, baseline rate, and effect threshold.

Confidence Interval Interpretation

Correctly interpret confidence intervals and understand what they reveal about the true effect size in your population.

Multiple Comparison Correction

Identify and adjust for inflated false positive rates when running multiple simultaneous tests or sequential peeks.

Effect Size Estimation

Translate business goals into statistical effect sizes (Cohen's h, lift percentage) to inform your test design.

P-Value Demystification

Learn what p-values actually measure, common misinterpretations to avoid, and how to contextualize them with effect sizes.

Test Duration Optimization

Calculate how long your test should run to reach statistical power targets and detect the minimum effect size you care about.

Go/No-Go Decision Framework

Make confident launch decisions by separating statistical significance from practical significance and business impact.

Example Output

Example 1: Sample Size Calculation

You ask: "We have a 3% baseline conversion rate. We want to detect a 0.5% absolute lift (to 3.5%). What sample size do we need with 80% power and 5% significance?"

Claude returns:

  • Minimum sample size per variant: 9,223 users
  • Total test duration: ~5 days (at 2,000 users/day)
  • Power curve visualization: Shows how power increases with sample size
  • Trade-offs: "If you reduce power to 70%, you need only 6,800 users per variant (3-4 days)."

Example 2: Interpreting Ambiguous Results

You submit test data:

  • Control: 100/3,500 conversions (2.86%)
  • Variant: 120/4,100 conversions (2.93%)
  • P-value: 0.18

Claude diagnoses:

  • ✓ Not statistically significant (p > 0.05)
  • ✓ Confidence interval for lift: [-1.2%, +4.8%]
  • ✓ Test was underpowered; need 12,000 users per variant to detect 0.5% lift reliably
  • ✓ Recommendation: Continue test or redesign with larger effect threshold

Example 3: Multiple Comparison Correction

You say: "I'm running 5 A/B tests simultaneously on different pages."

Claude advises:

  • ✓ Uncorrected false positive rate: 22.6%
  • ✓ Bonferroni-corrected threshold: p < 0.01 (not 0.05)
  • ✓ Alternative: FDR control for less conservative approach
  • ✓ Best practice: Prioritize 1-2 tests and run others sequentially

What's Included

  • Sample Size Calculator Template: Pre-built formulas and lookup tables for conversion rate tests, continuous metrics, and count data scenarios.
  • Statistical Power Curves: Visual relationships between sample size, effect size, power, and significance level to guide experiment design.
  • Multiple Comparison Checklist: Step-by-step guidance on identifying when correction is needed and how to apply Bonferroni, Holm, or FDR methods.
  • Confidence Interval Interpretation Guide: Templates for correctly stating what confidence intervals mean and translating them into business decisions.
  • P-Value and Significance Reference: Common misconceptions about p-values, correct interpretations, and how to avoid misuse in reports.
  • Effect Size Threshold Lookup: Industry benchmarks and guidance on what constitutes a practically meaningful effect in your domain.

Who It's For

  • CRO Specialists
  • Product Managers
  • Growth Engineers
  • Marketing Analysts
  • UX Researchers

Best For

  • Designing experiments before launch
  • Calculating required sample sizes and test duration
  • Interpreting test results and detecting false positives
  • Determining practical vs statistical significance
  • Troubleshooting inconclusive test results

You might also like

Logic Pro Session Architect
$35
Logic Pro3.3(3)
Logic Pro Session Architect

You can architect and optimize your Logic Pro sessions for professional mixing, routing, and production efficiency. Claude guides you through session design decisions, routing strategies, and workflow optimization to reduce setup time and improve mix consistency across projects.

Funnel Optimization Analysis & Strategy
$35
Funnel Optimization Analysis & Strategy

You can diagnose exactly where prospects drop off in your funnel, quantify the impact of each bottleneck, and receive prioritized optimization strategies backed by conversion metrics. This skill analyzes your funnel architecture, identifies friction points, and delivers actionable experiments (A/B tests, copy changes, UX improvements) ranked by expected revenue impact.

Social Media Strategy & Engagement Optimizer
$20
Social Media Strategy & Engagement Optimizer

This skill helps you develop comprehensive social media strategies tailored to your audience demographics and platform dynamics. You'll optimize posting schedules, identify high-performing content themes, and create actionable engagement plans that drive measurable growth and community building across all major social platforms.

Competitive Intelligence Analyst
$30
Competitive Intelligence Analyst

You can quickly structure competitor research, map competitive positioning across markets, and identify white space opportunities. The skill synthesizes company data, product analysis, and market trends into actionable strategic recommendations that inform your go-to-market strategy and product roadmap.

AI-Powered Nurture Campaign Architect
$30
Nurture3.5(4)
AI-Powered Nurture Campaign Architect

You'll architect multi-stage nurture campaigns that automatically segment audiences, personalize messaging, and optimize send timing based on behavior and engagement data. The skill generates complete campaign strategies with A/B testing frameworks, content variations, and actionable analytics dashboards. Get data-driven recommendations to improve open rates, click-through rates, and conversion metrics at every stage.

DMP Audience Architecture & Segment Optimization
$35
DMP3.8(5)
DMP Audience Architecture & Segment Optimization

You can architect sophisticated audience segments within your DMP, optimize them for campaign performance, and map audiences across marketing channels. This skill helps you classify customer data into actionable segments, define activation rules, analyze segment quality metrics, and recommend optimization strategies based on your business objectives and data constraints.

Marketing Stack Integration Architect
$25
Marketing Stack Integration Architect

This skill helps you design, document, and troubleshoot integrations between your marketing platforms. You get API mapping strategies, workflow automation guidance, and ready-to-use integration templates for popular tools like Salesforce, HubSpot, Marketo, and Zapier. Whether you're building native connections or orchestrating multi-step workflows, this skill guides you through requirements gathering, authentication setup, and data validation.

PMP Deal Structure & Optimization for Programmatic Media
$35
PMP Deal Structure & Optimization for Programmatic Media

You can design sophisticated PMP deal structures tailored to specific advertiser needs, audience segments, and performance targets. This skill analyzes deal parameters—floor pricing, impression volume, audience data, and performance history—to recommend optimization strategies that maximize yield while maintaining advertiser satisfaction. You'll get detailed deal frameworks with pricing tiers, audience definitions, and performance benchmarks.

$30.00