
Evaluate
Review Claude's execution steps and provide structured feedback
What You Can Do
You can trace Claude's decision-making process step-by-step through the Agent Execution Loop, seeing exactly which skills were matched, what was expected, which tools were called, how results were verified, and what was learned. Generate an interactive HTML review page where you write targeted feedback next to each execution step, then feed that structured feedback back into Claude for iterative improvements.
Features
view MATCH, THINK, ACT, VERIFY, LEARN phases with actual values and tool calls
write feedback directly next to each step and export structured comments
understand why a specific skill was chosen or why none matched
see which tools were invoked, their order, and whether they ran in parallel or sequence
review what checks passed or failed and identify audit issues
track what was updated in the skill or why no update occurred
Claude automatically parses your comments per step and applies targeted fixes
capture execution context from the current session without external logging
Example Output
MATCH Step: Selected 'Professional Writing' skill because user asked for a client email. No matching skills found for technical documentation.
THINK Step: Expected output: formal tone, 3-paragraph structure, call-to-action. Verification criteria: tone check via sentiment analysis, paragraph count validation.
ACT Step: Called email-generator tool (sequential), then grammar-check tool (parallel). Generated 285 words, processed through Hemingway API.
VERIFY Step: Passed tone check (formal: 92%), failed paragraph count (2 instead of 3). Skill audit: no security issues detected.
LEARN Step: Updated skill prompt to enforce paragraph structure with numbered examples.
What's Included
- SKILL.md instruction file with Agent Execution Loop definitions:
- HTML template with embedded CSS/JavaScript for interactive feedback collection:
- Reflection worksheet for mapping MATCH, THINK, ACT, VERIFY, LEARN steps:
- Feedback parser logic to extract and apply per-step improvements:
- Timestamp-based file naming convention for managing multiple evaluations:
Who It's For
- Professional writers refining client deliverables and optimizing writing workflows
- Content strategists auditing content generation processes and skill effectiveness
- AI prompt engineers debugging complex multi-step agent behaviors
- Learning & development specialists analyzing how Claude processes training content
- Quality assurance teams validating consistency and accuracy of automated writing tasks
Best For
- Troubleshooting unsatisfactory writing outputs with precise, actionable feedback
- Debugging skill selection logic when the wrong writing style was applied
- Auditing tool chains to optimize writing workflows and eliminate redundant steps
- Validating verification checkpoints in content quality assurance pipelines
- Iteratively improving writing skills through structured evaluation cycles



