SkillsLib.ai

Experiment Design & Statistical Analysis for Research Engineers

Design sound experiments and analyze results with statistical rigor

0.0(0 reviews)
100+ downloads
Updated Oct 2026

What You Can Do

You can design statistically valid experiments with proper power analysis, choose the right statistical tests for your data type, analyze results while controlling for multiple comparisons, and generate publication-ready reports with accurate interpretation of findings. Claude helps you avoid common statistical pitfalls and ensures your experimental claims are well-supported by evidence.

Features

Experiment design templates with power analysis

specify sample size, effect size, and significance level to get sample size recommendations

Statistical test selection

get recommendations for t-tests, ANOVA, chi-square, Mann-Whitney U, and others based on your data structure

Effect size calculation and interpretation

compute Cohen's d, odds ratios, and other effect measures with plain-English explanation

Multiple comparison correction

apply Bonferroni, FDR, and other adjustments when testing multiple hypotheses

Assumption checking workflows

verify normality, homogeneity of variance, independence, and other test prerequisites

Result visualization guidance

describe charts that best communicate findings (forest plots, violin plots, confidence intervals)

Publication-ready report templates

structure methods, results, and discussion sections with proper statistical reporting

Example Output

Example 1: Experiment Design

You ask: "Design an A/B test to detect a 10% improvement in click-through rate with 95% confidence and 80% power."

Claude returns:

  • Required sample size per group: 1,847
  • Total participants: 3,694
  • Recommended test: Two-proportion z-test
  • Effect size (Cohen's h): 0.21
  • Study design considerations and critical assumptions

Example 2: Statistical Analysis

You provide: "Control group: mean = 42.3 (SD = 8.1, n=50), Treatment: mean = 46.8 (SD = 7.9, n=50)"

Claude analyzes:

  • Test selection: Independent samples t-test
  • Results: t(98) = 2.34, p = 0.021
  • Effect size (Cohen's d): 0.66 (medium effect)
  • 95% CI for difference: [0.73, 8.23]
  • Interpretation: Statistically significant with meaningful practical improvement

Example 3: Report Generation

You request: "Generate methods and results sections for a two-arm experiment testing treatment vs. control."

Claude provides complete text including study design, randomization, statistical methods with justification, results with test statistics and p-values, and interpretation aligned with pre-specified hypotheses.

What's Included

  • SKILL.md: Full skill with experiment design workflow, statistical test decision trees, and analysis templates
  • Experiment design templates: Power analysis checklist, sample size guidance, study protocol template
  • Statistical analysis workflows: Decision trees for test selection, assumption-checking procedures, interpretation guidelines
  • Report templates: Methods section template, results reporting structure, discussion framework
  • Statistical test reference: When to use t-tests, ANOVA, chi-square, Mann-Whitney U, Kruskal-Wallis, and effect size formulas
  • Assumption checking guides: Procedures for verifying normality, homogeneity, independence, and sphericity
  • Effect size interpretation: Cohen's conventions and practical significance thresholds across research domains

Who It's For

  • Research engineers — Design experiments to validate algorithm improvements and system changes
  • Data scientists — Run rigorous A/B tests and interpret results with statistical confidence
  • Product managers — Conduct user experiments and make data-driven product decisions
  • ML engineers — Validate model improvements against baselines with proper statistical testing
  • Academic researchers — Design studies and report results following publication standards

Best For

  • Designing A/B tests and experiments — Specify your hypothesis and constraints, get sample size and test recommendations
  • Analyzing experimental results — Input summary statistics or raw data, get statistical test results and interpretation
  • Choosing statistical tests — Describe your data type and research question, get recommendations with clear justification
  • Multiple hypothesis testing — Learn when and how to apply Bonferroni, FDR, and other adjustments
  • Writing statistical methodology — Generate publication-ready methods and results sections with proper reporting conventions

You might also like

Analytics Documentation Generator
$25
Analytics Documentation Generator

This skill automatically documents your entire analytics infrastructure by analyzing data sources, transformations, and outputs. You'll generate production-ready data dictionaries with field definitions, lineage maps showing data flow across systems, and transformation documentation that explains logic and dependencies. Save weeks of manual documentation work while keeping your analytics stack discoverable as it evolves.

Agentdb Vector Search
$20
RAG4.1(34)
Agentdb Vector Search

You can build production-grade vector search systems that retrieve semantically similar documents in sub-millisecond time using AgentDB's optimized HNSW indexing. The skill enables you to implement RAG pipelines, semantic search engines, and intelligent knowledge bases with configurable embedding dimensions, distance metrics (cosine, Euclidean, dot product), and similarity thresholds—all with built-in quantization and caching for massive performance gains.

BI Data Quality Investigator
$30
BI Data Quality Investigator

You'll systematically diagnose data quality problems by developing structured root cause analysis frameworks, calculating the true business impact, and creating reproducible validation tests. This skill walks you through hypothesis-driven investigation, data lineage analysis, and remediation planning — turning data issues into documented fixes and preventive measures.

RAG Pipeline Optimization with Claude
$40
RAG Pipeline Optimization with Claude

You can systematically evaluate and improve your RAG pipelines using Claude as a design partner. You'll analyze retrieval quality, identify bottlenecks in your embedding and chunking strategies, and receive actionable recommendations to reduce hallucinations and improve context relevance. By the end, you'll have a data-driven optimization plan tailored to your specific use case and performance metrics.

Data Quality Test Framework Builder
$35
Data Quality Test Framework Builder

You'll build comprehensive data quality test suites that validate transformations, detect anomalies, and document standards across your dbt and SQL pipelines. This skill generates production-ready test configurations, anomaly detection protocols, and validation rules that catch data issues before they impact analytics.

Analytics Report Builder: Executive-Ready Data Storytelling
$40
Analytics Report Builder: Executive-Ready Data Storytelling

You can convert complex datasets and business metrics into polished, executive-ready reports that stakeholders trust and act on. Claude generates data-driven narratives, executive summaries, actionable insights, and visualization recommendations tailored to your audience's priorities. Your reports will tell a cohesive story that connects metrics to business outcomes, eliminating confusion and accelerating decision-making.

HEOR Evidence Synthesis & Dossier Builder
$35
HEOR3.6(5)
HEOR Evidence Synthesis & Dossier Builder

You can rapidly compile, organize, and format disparate health economic evidence—from clinical trials to cost-effectiveness analyses—into structured, regulatory-compliant dossiers. This skill maps your evidence to specific payer and HTA requirements, automatically generates evidence hierarchies, and produces submission-ready dossier outlines with formatting that meets regulatory standards for NICE, EUnetHTA, and other major bodies.

ML Infrastructure Failure Analysis & Optimization
$25
ML Infrastructure Failure Analysis & Optimization

This skill helps you systematically diagnose failures in distributed ML training and serving infrastructure. You provide system logs, metrics, and error traces, and Claude performs structured root-cause analysis to identify the underlying issue—whether it's resource exhaustion, distributed system deadlock, data pipeline corruption, or model serving misconfiguration. You get a detailed diagnosis with remediation steps ranked by likelihood and implementation effort.

$45.00