SkillsLib.ai

A/B Testing: Statistical Rigor for CRO Specialists

Design statistically sound A/B tests with proper sample sizing and rigorous analysis

3.7(3 reviews)
100+ downloads
Updated Sep 2026

What You Can Do

You'll learn to design A/B tests that reliably detect real changes while minimizing false positives. This skill helps you calculate proper sample sizes, determine statistical power requirements, and interpret results using rigorous frequentist methods — ensuring every test delivers actionable insights and saves your team from chasing false signals.

Features

Sample Size Calculator

Compute optimal sample sizes based on baseline conversion rate, minimum detectable effect, significance level, and statistical power.

Statistical Power Analysis

Determine the probability your test will detect a real effect given your sample size, baseline rate, and effect threshold.

Confidence Interval Interpretation

Correctly interpret confidence intervals and understand what they reveal about the true effect size in your population.

Multiple Comparison Correction

Identify and adjust for inflated false positive rates when running multiple simultaneous tests or sequential peeks.

Effect Size Estimation

Translate business goals into statistical effect sizes (Cohen's h, lift percentage) to inform your test design.

P-Value Demystification

Learn what p-values actually measure, common misinterpretations to avoid, and how to contextualize them with effect sizes.

Test Duration Optimization

Calculate how long your test should run to reach statistical power targets and detect the minimum effect size you care about.

Go/No-Go Decision Framework

Make confident launch decisions by separating statistical significance from practical significance and business impact.

Example Output

Example 1: Sample Size Calculation

You ask: "We have a 3% baseline conversion rate. We want to detect a 0.5% absolute lift (to 3.5%). What sample size do we need with 80% power and 5% significance?"

Claude returns:

  • Minimum sample size per variant: 9,223 users
  • Total test duration: ~5 days (at 2,000 users/day)
  • Power curve visualization: Shows how power increases with sample size
  • Trade-offs: "If you reduce power to 70%, you need only 6,800 users per variant (3-4 days)."

Example 2: Interpreting Ambiguous Results

You submit test data:

  • Control: 100/3,500 conversions (2.86%)
  • Variant: 120/4,100 conversions (2.93%)
  • P-value: 0.18

Claude diagnoses:

  • ✓ Not statistically significant (p > 0.05)
  • ✓ Confidence interval for lift: [-1.2%, +4.8%]
  • ✓ Test was underpowered; need 12,000 users per variant to detect 0.5% lift reliably
  • ✓ Recommendation: Continue test or redesign with larger effect threshold

Example 3: Multiple Comparison Correction

You say: "I'm running 5 A/B tests simultaneously on different pages."

Claude advises:

  • ✓ Uncorrected false positive rate: 22.6%
  • ✓ Bonferroni-corrected threshold: p < 0.01 (not 0.05)
  • ✓ Alternative: FDR control for less conservative approach
  • ✓ Best practice: Prioritize 1-2 tests and run others sequentially

What's Included

  • Sample Size Calculator Template: Pre-built formulas and lookup tables for conversion rate tests, continuous metrics, and count data scenarios.
  • Statistical Power Curves: Visual relationships between sample size, effect size, power, and significance level to guide experiment design.
  • Multiple Comparison Checklist: Step-by-step guidance on identifying when correction is needed and how to apply Bonferroni, Holm, or FDR methods.
  • Confidence Interval Interpretation Guide: Templates for correctly stating what confidence intervals mean and translating them into business decisions.
  • P-Value and Significance Reference: Common misconceptions about p-values, correct interpretations, and how to avoid misuse in reports.
  • Effect Size Threshold Lookup: Industry benchmarks and guidance on what constitutes a practically meaningful effect in your domain.

Who It's For

  • CRO Specialists
  • Product Managers
  • Growth Engineers
  • Marketing Analysts
  • UX Researchers

Best For

  • Designing experiments before launch
  • Calculating required sample sizes and test duration
  • Interpreting test results and detecting false positives
  • Determining practical vs statistical significance
  • Troubleshooting inconclusive test results

You might also like

Funnel Optimization Analysis & Strategy
$35
Funnel Optimization Analysis & Strategy

You can diagnose exactly where prospects drop off in your funnel, quantify the impact of each bottleneck, and receive prioritized optimization strategies backed by conversion metrics. This skill analyzes your funnel architecture, identifies friction points, and delivers actionable experiments (A/B tests, copy changes, UX improvements) ranked by expected revenue impact.

Guerrilla Tactics Campaign Designer
$25
Guerrilla Tactics Campaign Designer

You create structured guerrilla marketing campaigns that deliver outsized impact without big budgets. The skill designs tactics from concept through execution plan, identifies and mitigates legal and reputational risks, and stress-tests ideas for feasibility before you invest time or money. You get actionable playbooks with timelines, resource estimates, contingency plans, and success metrics.

Brand Name Generation Strategy
$35
Brand Name Generation Strategy

Create compelling brand names that align with your market positioning, target audience, and business goals. You'll receive 15-25 strategically-sound candidates, each with linguistic analysis, competitive positioning assessment, and cultural implications reviewed. Every recommendation includes domain availability status, trademark conflict identification, and actionable launch strategy guidance.

AI-Powered Nurture Campaign Architect
$30
Nurture3.5(4)
AI-Powered Nurture Campaign Architect

You'll architect multi-stage nurture campaigns that automatically segment audiences, personalize messaging, and optimize send timing based on behavior and engagement data. The skill generates complete campaign strategies with A/B testing frameworks, content variations, and actionable analytics dashboards. Get data-driven recommendations to improve open rates, click-through rates, and conversion metrics at every stage.

Programmatic Attribution Analyzer
$35
Programmatic Attribution Analyzer

Analyze historical customer journeys across multiple touchpoints to determine each channel's true ROI contribution. You can model different attribution approaches, identify the most effective conversion paths, and generate data-driven recommendations for reallocating programmatic media budgets to maximize return on ad spend.

Marketing Stack Integration Architect
$25
Marketing Stack Integration Architect

This skill helps you design, document, and troubleshoot integrations between your marketing platforms. You get API mapping strategies, workflow automation guidance, and ready-to-use integration templates for popular tools like Salesforce, HubSpot, Marketo, and Zapier. Whether you're building native connections or orchestrating multi-step workflows, this skill guides you through requirements gathering, authentication setup, and data validation.

SEO Content Audit & Optimization Strategy
$35
SEO Content Audit & Optimization Strategy

Analyze your existing web content against SEO best practices and competitor strategies to find quick wins and growth opportunities. Claude evaluates keyword alignment, content structure, search intent match, and technical factors—then generates a prioritized action plan ranked by traffic potential and implementation effort. You get data-driven recommendations to improve organic search visibility and drive measurable traffic gains.

PMP Deal Structure & Optimization for Programmatic Media
$35
PMP Deal Structure & Optimization for Programmatic Media

You can design sophisticated PMP deal structures tailored to specific advertiser needs, audience segments, and performance targets. This skill analyzes deal parameters—floor pricing, impression volume, audience data, and performance history—to recommend optimization strategies that maximize yield while maintaining advertiser satisfaction. You'll get detailed deal frameworks with pricing tiers, audience definitions, and performance benchmarks.

$30.00