RunComfy Agent Skills
30 Claude skills for RunComfy model endpoints (image, video, audio, vision)
What You Can Do
Access 30+ RunComfy model endpoints directly from Claude Code, including image-to-video, video-editing, AI image generation, AI video generation, and music composition. Each skill wraps a production endpoint with schema validation and intent-based routing. Use the RunComfy CLI to generate, inpaint, outpaint, and transform media from ACE Step music (tag-driven, $0.0002/s) to video and image models. Skills route automatically based on your request intent.
Features
Automatic endpoint selection based on what you ask for (music gen request routes to ACE Step, video request routes to video endpoint)
Single skill exposes multiple routes (ACE Step includes text-to-audio, inpaint, outpaint, allowing generation, editing, and extension in one skill)
ACE Step music uses comma-separated tags for composition control (genre, mood, instruments, BPM) instead of free-text prompts
RunComfy pricing significantly cheaper than competing APIs (ACE Step $0.0002/s vs ElevenLabs Music $0.0083/s, 27x savings)
Every endpoint has explicit input/output schema with field-level validation before hitting the RunComfy API
Skills default to open-weights models (Apache 2.0) like ACE Step, with commercial alternatives available for specific use cases
Chain generation, inpaint, and outpaint calls to build complex media pipelines (draft cheap, polish premium)
Example Output
Tag-driven instrumental music generation:
runcomfy run acestep-ai/ace-step/text-to-audio \
--input '{"tags": "lo-fi hip-hop, mellow, vinyl crackle, rhodes piano, soft drums, 75 BPM", "lyrics": "[inst]", "duration": 90}' \
--output-dir ./out
Generates a 90-second lo-fi hip-hop track with the specified vibe and instrumentation.
Audio inpainting (regenerate a section):
runcomfy run acestep-ai/ace-step/audio-inpaint \
--input '{"audio": "https://your-cdn.example/original-track.mp3", "tags": "indie pop, breakdown, piano only, soft, no drums", "start_time": 20, "end_time": 40, "lyrics": "[inst]"}' \
--output-dir ./out
Replaces the 20 to 40 second segment of an existing track with a new bridge matching the specified vibe.
Audio outpainting (extend a track):
runcomfy run acestep-ai/ace-step/audio-outpaint \
--input '{"audio": "https://your-cdn.example/hook-30s.mp3", "tags": "indie pop, electric guitar, drums, build-up before chorus, fade outro", "extend_before_duration": 30, "extend_after_duration": 60, "lyrics": "[inst]"}' \
--output-dir ./out
Turns a 30-second hook into a 2-minute track by adding a 30-second intro and 60-second outro.
What's Included
- 30 production skills: Full catalog of image-to-video, video-edit, image generation, video generation, music generation, and specialized routers
- RunComfy CLI integration: Each skill uses the runcomfy CLI to invoke endpoints with schema-validated input and automatic output handling
- Multi-route endpoints: Skills like ACE Step expose 4+ routes (text-to-audio base, text-to-audio 1.5, audio-inpaint, audio-outpaint) in one package
- Schema validation and error handling: Every endpoint has explicit input schema, field validation, type checking, and exit codes for debugging
- Detailed documentation: Each skill includes route selection guide, schema tables, usage examples, prompting tips, and cost breakdowns
- Security and privacy defaults: Token management via ~/.config/runcomfy/token.json (0600), HTTPS-only API calls, no shell-injection surface
Who It's For
- AI developers and prompt engineers
- Content creators (music, video, image)
- Game developers and indie studios
- Batch processing and automation builders
- Cost-conscious teams scaling media generation
Best For
- High-volume, cost-sensitive content generation
- Multi-modal media pipelines (generate, edit, extend)
- Automatic model selection based on user intent
- Production-grade schema validation and error handling
- Open-weights model pipelines



