SkillsLib.ai

Experiment Design for A/B Testing & Causal Inference

Design statistically sound A/B tests and causal experiments in minutes

0.0(0 reviews)
100+ downloads
Updated Oct 2026

What You Can Do

This skill guides you through rigorous experiment design from hypothesis to analysis plan. You'll compute the statistical power needed for your tests, identify confounding variables to control for, design optimal randomization strategies, and generate pre-registered analysis plans that prevent p-hacking and publication bias. Whether you're running a product A/B test or investigating causal effects in observational data, you get structured, validated designs that account for real-world constraints.

Features

Statistical power and sample size calculator

instantly computes required sample sizes for your target effect sizes, significance levels, and baseline conversion rates

Experiment design templates

ready-to-use frameworks for A/B tests, multivariate tests, quasi-experimental designs, and causal inference studies

Confounder identification and control

systematically walks through potential confounding variables and recommends statistical or experimental controls

Metric selection framework

helps you define primary metrics, secondary metrics, and guardrail metrics with clear success criteria and sensitivity thresholds

Analysis plan generator

produces detailed, pre-registered analysis plans that specify primary hypotheses, subgroup analyses, and robustness checks upfront

Randomization strategy advisor

recommends stratification, blocking, or matching approaches optimized for your sample size and constraints

Power sensitivity analysis

shows how statistical power changes across a range of effect sizes and sample sizes to stress-test your assumptions

Causal inference checklist

guides you through directed acyclic graphs (DAGs), adjustment sets, and validity threats for causal claims

Example Output

Example 1: Power Analysis for Product A/B Test

Inputs: 5% baseline conversion rate, 10% relative lift target, 80% power, 95% significance

Output:

code
Required sample size per group: 15,472 users
Total required users: 30,944
Expected runtime (at 5,000 users/day): 6.2 days
MDE at current N=10k per group: 12.6% relative lift

Example 2: A/B Test Design Checklist

  • ✅ Hypothesis: Larger checkout button increases conversion by 8–12%
  • ✅ Primary metric: Checkout completion rate (target 5% lift)
  • ✅ Secondary metrics: AOV, return rate, support tickets
  • ✅ Guardrails: Session duration, cart abandonment (no >5% change)
  • ✅ Randomization: Stratified by device type + geography
  • ✅ Excluded: New users (<1 day old), bot-flagged sessions
  • ✅ Duration: Fixed 2 weeks, stopping rule disabled
  • ✅ Analysis: Intent-to-treat (all randomized users)

Example 3: Causal Inference DAG

code
Advertising Spend → Revenue
      ↓
  Market Sentiment → Revenue
      ↓
  Competitor Price

Confounders to adjust for: Market Sentiment
Adjustment set: {Market Sentiment}
Causal claim validity: Strong (RCT not feasible, but DAG-guided design mitigates bias)

What's Included

  • SKILL.md: Complete experiment design methodology with decision trees and validation workflows
  • A/B Test Design Template: Pre-built structure with hypothesis, success metrics, sample size calculator, and stopping rules
  • Causal Inference Checklist: Confounder identification, DAG construction, and validity threat assessment
  • Metric Definition Worksheet: Define primary, secondary, and guardrail metrics with sensitivity thresholds
  • Analysis Plan Template: Pre-registered analysis specifications (SAGE framework: Specific, Appropriate, Generalized, Efficient)
  • Sample Size Reference Card: Lookup tables and formulas for common effect sizes and significance levels
  • Randomization Strategy Guide: Stratification, blocking, and matching best practices

Who It's For

  • Product managers — designing feature A/B tests with confidence in statistical rigor
  • Data scientists — planning causal inference studies and observational analyses
  • UX researchers — conducting user experiments with proper controls and power analysis
  • Growth teams — running conversion optimization tests with validated designs
  • Academic and healthcare researchers — ensuring experiment designs meet publication and regulatory standards

Best For

  • Calculating required sample size — determining if you need 1,000 or 100,000 users for statistical significance
  • Preventing measurement bias — identifying confounders and designing controls to isolate causal effects
  • Meeting publication standards — creating pre-registered analysis plans that prevent p-hacking and HARKing
  • Designing experiments for observational data — constructing causal DAGs and selecting adjustment sets when randomization isn't possible
  • Stress-testing assumptions — running sensitivity analyses to see how power degrades if your effect size or baseline rates are wrong

You might also like

Experiment Design & Statistical Analysis for Research Engineers
$45
Experiment Design & Statistical Analysis for Research Engineers

You can design statistically valid experiments with proper power analysis, choose the right statistical tests for your data type, analyze results while controlling for multiple comparisons, and generate publication-ready reports with accurate interpretation of findings. Claude helps you avoid common statistical pitfalls and ensures your experimental claims are well-supported by evidence.

Analytics Documentation Generator
$25
Analytics Documentation Generator

This skill automatically documents your entire analytics infrastructure by analyzing data sources, transformations, and outputs. You'll generate production-ready data dictionaries with field definitions, lineage maps showing data flow across systems, and transformation documentation that explains logic and dependencies. Save weeks of manual documentation work while keeping your analytics stack discoverable as it evolves.

BI Data Quality Investigator
$30
BI Data Quality Investigator

You'll systematically diagnose data quality problems by developing structured root cause analysis frameworks, calculating the true business impact, and creating reproducible validation tests. This skill walks you through hypothesis-driven investigation, data lineage analysis, and remediation planning — turning data issues into documented fixes and preventive measures.

Data Quality Test Framework Builder
$35
Data Quality Test Framework Builder

You'll build comprehensive data quality test suites that validate transformations, detect anomalies, and document standards across your dbt and SQL pipelines. This skill generates production-ready test configurations, anomaly detection protocols, and validation rules that catch data issues before they impact analytics.

HEOR Evidence Synthesis & Dossier Builder
$35
HEOR3.6(5)
HEOR Evidence Synthesis & Dossier Builder

You can rapidly compile, organize, and format disparate health economic evidence—from clinical trials to cost-effectiveness analyses—into structured, regulatory-compliant dossiers. This skill maps your evidence to specific payer and HTA requirements, automatically generates evidence hierarchies, and produces submission-ready dossier outlines with formatting that meets regulatory standards for NICE, EUnetHTA, and other major bodies.

Payer Evidence Synthesis & HTA Builder
$30
Payer4.0(3)
Payer Evidence Synthesis & HTA Builder

You can structure comprehensive health technology assessments (HTAs) that organize clinical evidence, economic analyses, and regulatory considerations into evidence-based coverage recommendations. This skill helps you synthesize clinical trial data, health economic models, and real-world evidence into clear, defensible payer coverage determinations. You'll generate professional HTA reports that align with major frameworks like ICER, CADTH, and NICE standards.

Setup Agent Tail
$45
Monitoring4.4(48)
Setup Agent Tail

This skill detects your project framework (Vite, Next.js, plain Node, or monorepo) and automatically configures agent-tail to pipe dev server and browser console logs into unified log files. You'll get a proposed configuration tailored to your setup, install agent-tail with the correct plugins, and have logs immediately available for AI agents to consume and analyze.

ML Infrastructure Failure Analysis & Optimization
$25
ML Infrastructure Failure Analysis & Optimization

This skill helps you systematically diagnose failures in distributed ML training and serving infrastructure. You provide system logs, metrics, and error traces, and Claude performs structured root-cause analysis to identify the underlying issue—whether it's resource exhaustion, distributed system deadlock, data pipeline corruption, or model serving misconfiguration. You get a detailed diagnosis with remediation steps ranked by likelihood and implementation effort.

$35.00