
Summarize Experiment
Extract and summarize experiment results into lightweight markdown reports
4.3(24 reviews)100+ downloadsUpdated Sep 2026Verified SafeSecurity VerifiedThis skill was analyzed by our AI security scanner for harmful content including data exfiltration, system manipulation, credential theft, and prompt injection. No threats were detected.
What You Can Do
This skill parses experiment metadata and outputs to extract critical metrics from completed machine learning experiments. It reads experiment_summary.yaml to identify runs, pulls final training loss from SLURM logs, extracts accuracy metrics from inspect-ai evaluation files, and generates a clean summary.md report—all while logging the extraction process for reproducibility.
Features
Parse experiment_summary.yaml to identify fine-tuned and control runs with their hyperparameters
Extract final training loss from SLURM stdout logs automatically
Pull accuracy and performance metrics from inspect-ai .eval files
Generate summary.md with structured metrics and run metadata
Log all extraction steps to logs/summarize-experiment.log for audit trails
Support partial experiment results—handle incomplete runs gracefully
Map evaluation tasks and epochs to their corresponding results
Example Output
Example summary.md output:
Experiment Summary: bert-classification-v2
Run Status
| Run Name | Type | Model | Status | Training Loss | Accuracy |
|---|---|---|---|---|---|
| run-001-lr0.001 | fine-tuned | bert-base | Completed | 0.245 | 0.924 |
| run-002-lr0.0001 | fine-tuned | bert-base | Completed | 0.198 | 0.931 |
| control-baseline | control | bert-base | Completed | — | 0.847 |
Key Findings
- Best accuracy: 0.931 (run-002-lr0.0001, learning rate 0.0001)
- Improvement over baseline: +8.4 percentage points
- Training converged in all runs
Log entry example:
code
[2024-01-15 14:32:01] Parsing experiment_summary.yaml...
[2024-01-15 14:32:02] Found 3 runs (2 fine-tuned, 1 control)
[2024-01-15 14:32:03] Extracted training loss from run-001: 0.245
[2024-01-15 14:32:04] Extracted accuracy from eval/logs/run-002.eval: 0.931
[2024-01-15 14:32:05] summary.md generated successfully
What's Included
- `summarize-experiment.md` instruction file with full workflow documentation:
- YAML parsing template for reading experiment_summary.yaml structures:
- Python extraction script (parse_eval_log.py) for inspect-ai eval file parsing:
- summary.md template with markdown table and findings sections:
- SLURM log extraction patterns and regular expressions:
Who It's For
- Machine learning researchers conducting hyperparameter tuning experiments
- ML engineers documenting model fine-tuning results for team review
- Research scientists tracking multiple experimental runs for publications
- AI labs automating experiment result documentation workflows
- Data scientists creating reproducible experiment reports
Best For
- Summarizing multi-run fine-tuning experiments with varied hyperparameters
- Extracting metrics from completed inspect-ai evaluation workflows
- Generating quick reference documents comparing control vs. fine-tuned models
- Automating post-experiment documentation after run-experiment completion
- Creating audit trails of experiment execution and metric extraction
You might also like
$36.00$45.00







