SkillsLib.ai

Experiment Design Framework

Design rigorous experiments with statistical power and built-in bias guardrails

0.0(0 reviews)
100+ downloads
Updated Sep 2026

What You Can Do

You can create statistically rigorous experimental designs from hypothesis to success metrics. Claude helps you calculate proper sample sizes with power analysis, identify hidden biases and confounds, and define clear decision criteria before data collection begins. This prevents p-hacking, ensures adequate statistical power, and produces defensible experimental protocols.

Features

Power analysis and sample size calculation

Determines required sample sizes based on effect size, significance level, and desired statistical power using analytic formulas

Bias identification and mitigation

Identifies confounding variables, selection bias, regression to the mean, and allocation concealment issues early in design

Success metrics framework

Structures primary and secondary outcomes, statistical thresholds, stopping rules, and decision criteria before running the experiment

Experimental design templates

Generates designs for randomized controlled trials, A/B tests, factorial experiments, regression discontinuity, and quasi-experimental approaches

Multiple testing correction

Applies Bonferroni, false discovery rate (FDR), or sequential testing corrections to maintain error rates across multiple comparisons

Pre-registration checklist

Creates a structured pre-registration document with design details and analysis plan to prevent selective reporting

Sensitivity analysis

Tests how robust your conclusions are under different assumptions, effect sizes, and violation of design assumptions

Data analysis plan

Specifies primary and secondary analyses, subgroup comparisons, and decision rules for handling missing data or protocol deviations

Example Output

Example 1: A/B Test Power Analysis

code
Experiment: Feature rollout A/B test
Control conversion rate: 5%
Expected effect size: 15% relative lift (5.75% treatment rate)
Significance level (α): 0.05
Statistical power (1-β): 0.80

Required sample size per group: 2,847
Total sample needed: 5,694 users
Expected duration (at 500 daily users): 11-12 days

Bias risks identified:
- Selection bias if users self-select into control/treatment
  → Mitigate: Assign at point of feature exposure
- Novelty effect in treatment group
  → Mitigate: Run test for ≥7 days to allow adaptation

Example 2: Pre-registration Checklist

code
✓ Primary hypothesis stated before data collection
✓ Sample size justified by power calculation
✓ Primary outcome metric explicitly defined
✓ Statistical test specified (z-test, t-test, Bayesian, etc.)
✓ Significance level set (α = 0.05)
✓ Multiple testing correction applied (n=2 outcomes → Bonferroni α = 0.025)
✓ Stopping rule defined (fixed sample vs. sequential)
✓ Exclusion criteria specified in advance
✓ Potential confounds identified and mitigation strategy noted

What's Included

  • SKILL.md: Full experiment design framework with workflows for every phase
  • Power analysis worksheet: Templates for calculating sample sizes for common study designs
  • Bias identification checklist: Systematic guide to confounds, selection bias, and design flaws
  • Experimental design templates: Ready-to-customize designs for RCTs, A/B tests, factorial designs, and quasi-experiments
  • Pre-registration template: Structured document for locking in your design and analysis plan
  • Success metrics framework: Template for defining primary/secondary outcomes, statistical thresholds, and decision rules

Who It's For

  • Product managers designing A/B tests and feature experiments
  • Data scientists and statisticians planning research studies
  • Academic researchers designing rigorous protocols
  • Growth and marketing teams testing campaign hypotheses
  • Quality assurance and engineering teams designing controlled experiments

Best For

  • Designing A/B tests with adequate statistical power
  • Planning clinical trials or academic research with proper controls
  • Pre-registering experiments to prevent bias and p-hacking
  • Calculating sample sizes for studies and surveys
  • Identifying confounds and hidden biases before running experiments

You might also like

Experiment Design & Statistical Analysis for Research Engineers
$45
Experiment Design & Statistical Analysis for Research Engineers

You can design statistically valid experiments with proper power analysis, choose the right statistical tests for your data type, analyze results while controlling for multiple comparisons, and generate publication-ready reports with accurate interpretation of findings. Claude helps you avoid common statistical pitfalls and ensures your experimental claims are well-supported by evidence.

Analytics Documentation Generator
$25
Analytics Documentation Generator

This skill automatically documents your entire analytics infrastructure by analyzing data sources, transformations, and outputs. You'll generate production-ready data dictionaries with field definitions, lineage maps showing data flow across systems, and transformation documentation that explains logic and dependencies. Save weeks of manual documentation work while keeping your analytics stack discoverable as it evolves.

Agentdb Vector Search
$20
RAG4.1(34)
Agentdb Vector Search

You can build production-grade vector search systems that retrieve semantically similar documents in sub-millisecond time using AgentDB's optimized HNSW indexing. The skill enables you to implement RAG pipelines, semantic search engines, and intelligent knowledge bases with configurable embedding dimensions, distance metrics (cosine, Euclidean, dot product), and similarity thresholds—all with built-in quantization and caching for massive performance gains.

BI Data Quality Investigator
$30
BI Data Quality Investigator

You'll systematically diagnose data quality problems by developing structured root cause analysis frameworks, calculating the true business impact, and creating reproducible validation tests. This skill walks you through hypothesis-driven investigation, data lineage analysis, and remediation planning — turning data issues into documented fixes and preventive measures.

RAG Pipeline Optimization with Claude
$40
RAG Pipeline Optimization with Claude

You can systematically evaluate and improve your RAG pipelines using Claude as a design partner. You'll analyze retrieval quality, identify bottlenecks in your embedding and chunking strategies, and receive actionable recommendations to reduce hallucinations and improve context relevance. By the end, you'll have a data-driven optimization plan tailored to your specific use case and performance metrics.

Data Quality Test Framework Builder
$35
Data Quality Test Framework Builder

You'll build comprehensive data quality test suites that validate transformations, detect anomalies, and document standards across your dbt and SQL pipelines. This skill generates production-ready test configurations, anomaly detection protocols, and validation rules that catch data issues before they impact analytics.

Analytics Report Builder: Executive-Ready Data Storytelling
$40
Analytics Report Builder: Executive-Ready Data Storytelling

You can convert complex datasets and business metrics into polished, executive-ready reports that stakeholders trust and act on. Claude generates data-driven narratives, executive summaries, actionable insights, and visualization recommendations tailored to your audience's priorities. Your reports will tell a cohesive story that connects metrics to business outcomes, eliminating confusion and accelerating decision-making.

ML Infrastructure Failure Analysis & Optimization
$25
ML Infrastructure Failure Analysis & Optimization

This skill helps you systematically diagnose failures in distributed ML training and serving infrastructure. You provide system logs, metrics, and error traces, and Claude performs structured root-cause analysis to identify the underlying issue—whether it's resource exhaustion, distributed system deadlock, data pipeline corruption, or model serving misconfiguration. You get a detailed diagnosis with remediation steps ranked by likelihood and implementation effort.

$35.00