
Benchmark Design & Analysis for Research Engineers
Design statistically valid benchmarks and interpret results
What You Can Do
You can design rigorous, reproducible benchmarks from first principles, establishing proper controls, baselines, and sample sizes. Claude analyzes performance results using appropriate statistical tests, calculates confidence intervals, and identifies whether observed differences are meaningful or noise. You'll produce peer-review-ready methodology sections and reports that satisfy the rigor demands of academic and production environments.
Features
establish proper controls, baselines, and experimental structure
choose appropriate tests (t-tests, ANOVA, Mann-Whitney U) based on data characteristics
calculate required sample sizes to detect meaningful differences
generate plots for trend analysis, distributions, and comparisons
quantify uncertainty and effect sizes with proper bounds
create detailed experiment logs and environment specifications
identify statistically significant deviations from baseline
structure multi-condition experiments with proper statistical framing
Example Output
Example 1: Benchmark Design Plan
## Benchmark Specification
- Primary metric: Throughput (ops/sec)
- Baseline: Current implementation
- Conditions: 3 optimizations vs baseline
- Sample size: 100 runs per condition (n=400 total)
- Warmup: 10 runs per condition
- Statistical test: One-way ANOVA + post-hoc Tukey HSD
- Confidence level: 95%
Example 2: Statistical Analysis Output
Results:
✓ Optimization A: +12.3% (95% CI: [10.1%, 14.5%], p<0.001)
✓ Optimization B: +8.7% (95% CI: [6.2%, 11.2%], p<0.001)
✗ Optimization C: +1.2% (95% CI: [-1.3%, 3.7%], p=0.34) — not significant
Effect sizes: Cohen's d = 1.42 (large), 0.98 (medium), 0.11 (negligible)
Example 3: Anomaly Report
Run 47 flagged: Throughput = 8,234 ops/sec (3.2σ below mean)
Likely cause: system noise (thermal throttling detected in logs)
Recommendation: Exclude from final analysis (outlier removal justified)
What's Included
- SKILL.md: Full benchmark design & analysis workflow
- Benchmark Design Template: Controls checklist, baseline definition, sample size calculator
- Statistical Analysis Checklist: Test selection flowchart, assumptions validation, post-hoc procedures
- Report Template: Methodology, results table, visualization descriptions, discussion framework
- Reproducibility Log: Hardware specs, software versions, environment variables, seed records
- Anomaly Detection Worksheet: Outlier justification, exclusion criteria, sensitivity analysis
Who It's For
- Machine learning researchers validating model improvements and comparing architectures
- Systems engineers benchmarking infrastructure changes and optimization techniques
- Algorithm researchers proving theoretical improvements with empirical results
- Performance engineers detecting regressions and quantifying optimization wins
- Academic researchers publishing peer-reviewed experimental work with statistical rigor
Best For
- Designing rigorous A/B tests and controlled experiments with proper statistical framing
- Validating performance claims with confidence intervals and effect size reporting
- Detecting real improvements vs. noise and random variation in benchmarks
- Creating reproducible benchmark suites with documented controls and methodology
- Writing experimental methodology sections that satisfy publication and peer-review standards







