RigorPilot Research Skills
Auditable deep learning research with reproducibility and scientific rigor
What You Can Do
RigorPilot is a comprehensive skill suite for deep learning research, providing auditable repository analysis, README-first reproducibility workflows, bounded exploration with evidence-based ranking, and publication-ready scientific documentation. Each skill enforces rigor through documented assumptions, comparability tracking, and explicit decision checkpoints. Use these skills to understand ML repository internals, faithfully reproduce published results, systematically explore novel ideas with proper baselines, and generate standardized research artifacts.
Features
Read-only deep inspection of deep learning repositories, mapping model architectures, training/inference entrypoints, configs, insertion points, and flagging suspicious implementation patterns without modification.
Faithful reproduction prioritizing README guidance, selecting minimal trustworthy targets (inference before evaluation before training), and recording all deviations, assumptions, and environment state.
Bounded experimental work on frozen branches or checkpoints. Gates ideas with explicit scoring, runs smoke tests, and ranks candidates by real evidence before committing to full execution.
Automatic setup of datasets, checkpoints, dependencies, and caches required for selected reproduction or exploration targets without silent protocol changes.
Generates standardized outputs including SUMMARY.md, SCIENTIFIC_CHANGELOG.md, COMPARABILITY_REPORT.md, ANNOTATED_README.md, and machine-readable status.json for publication-quality evidence.
Records all executed commands, environment variables, code modifications, human decision points, and reproducibility assumptions with conservative patch governance rules.
Loads contextual principles for research safety, deep learning experiments, comparability boundaries, and patch impact to ensure decisions meet scientific standards and align with repository intent.
Example Output
Example 1: Reproducing Baseline Results
I want to reproduce the baseline inference results from the README. What's the smallest target and what do I need to set up?
Claude reads the README, extracts documented inference commands, bootstraps required datasets and checkpoints, executes the minimal target with full logging, and produces SUMMARY.md, COMMANDS.md, SCIENTIFIC_CHANGELOG.md, and COMPARABILITY_REPORT.md showing exact environment, deviations, and metrics reproduction.
Example 2: Exploring Architectural Variants
I have a frozen current_research branch with baseline results. Can you propose and smoke-test 3 architectural variants and rank them by expected impact?
Claude analyzes the repository, gates candidate ideas against the frozen SOTA reference and evaluation method, ranks each by cost and success likelihood, executes one as a smoke test, collects evidence, and documents all decisions with rollback instructions in explore_outputs/.
What's Included
- ai-research-reproduction skill: End-to-end workflow for README-first trusted reproduction with minimal target selection, auditable execution, and standardized documentation
- ai-research-explore skill: Bounded candidate exploration with idea gating, evidence ranking, and scientific comparability tracking for novelty assessment
- analyze-project skill: Read-only repository analysis for architecture, entrypoints, config relationships, and suspicious patterns
- Research governance policies: Embedded references for research rigor, deep learning experiments, patch safety, and continuous learning guidance
- Standardized artifact templates: Pre-configured markdown and JSON formats for SUMMARY.md, SCIENTIFIC_CHANGELOG.md, COMPARABILITY_REPORT.md, and status.json
- Orchestration helpers: Python utilities and scripts for artifact writing, experiment orchestration, and lesson store integration
Who It's For
- Deep Learning Researchers
- ML Paper Authors
- Research Engineers
- Benchmark and Evaluation Specialists
Best For
- Reproducible research workflows
- Auditable experiment exploration
- Scientific rigor and documentation
- Repository-grounded ML analysis
- Publication-ready evidence generation
