RunComfy Agent Skills
Run 29 AI generation models via Claude with the RunComfy CLI
What You Can Do
Access 29 production-ready Claude agent skills that wrap RunComfy's AI models for image generation, video creation, audio synthesis, and avatar videos. Skills include music generation (ACE Step with inpaint/outpaint), AI avatars and lip-sync (OmniHuman, Seedance, HappyHorse), image editing (inpainting, outpainting, face-swap, relight), video generation and extension, and upscaling. All skills are CLI-driven via the runcomfy command, with built-in schema validation, error handling, and examples for each model.
Features
ACE Step for tag-driven music composition, inpainting, and outpainting. Cost-effective alternative to ElevenLabs Music (27x cheaper at $0.0002/s), supports 50+ languages and structured lyrics with section markers.
Route across OmniHuman (audio-driven full-body), Wan 2-7 (with audio_url field), HappyHorse (in-pass audio), and Seedance v2 (multi-modal cinematic). Pick the right model for UGC, dubbed video, or styled characters.
Multiple image models including Flux 2, GPT Image 2, Nano Banana 2, and specialized editors for inpainting, outpainting, face-swap, and relight. Support for text-to-image, image-to-image, and style transfer workflows.
Generate videos from prompts or extend existing clips bidirectionally. Models like HappyHorse, Seedance, and specialized outpainting endpoints for seamless video composition and looping content.
Face-swap, relight, and character animation tools. Edit facial expressions, lighting, and emotions in both images and videos while maintaining identity and quality.
Skills intelligently route user intent to the optimal model. Whether generating music, creating avatars, or editing images, each skill picks the right endpoint with documented prompting patterns.
CLI-driven, repeatable workflows optimized for cost-sensitive batches. Set seeds for reproducibility, control durations and resolutions, and iterate cheaply before final rendering.
Each skill ships with detailed schema tables, prompting tips, common patterns, comparison matrices, and troubleshooting guides. 29 skills with 100% coverage of model capabilities.
Example Output
Generate lo-fi hip-hop with ACE Step
runcomfy run acestep-ai/ace-step/text-to-audio \
--input '{
"tags": "lo-fi hip-hop, mellow, vinyl crackle, rhodes piano, soft drums, 75 BPM",
"lyrics": "[inst]",
"duration": 90
}' \
--output-dir ./out
Generates a 90-second royalty-free lo-fi track with vinyl character at $0.018. Seed for reproducibility or randomize for variations.
Create a talking-head presenter video
runcomfy run bytedance/omnihuman/api \
--input '{
"image_url": "https://your-cdn.example/presenter.jpg",
"audio_url": "https://your-cdn.example/voiceover.mp3"
}' \
--output-dir ./out
Feed one portrait image and one voiceover MP3, get back a video where the subject speaks naturally with synchronized mouth and gestures.
Extend a music track bidirectionally
runcomfy run acestep-ai/ace-step/audio-outpaint \
--input '{
"audio": "https://your-cdn.example/hook-30s.mp3",
"tags": "indie pop, electric guitar, drums, build-up before chorus, fade outro",
"extend_before_duration": 30,
"extend_after_duration": 60
}' \
--output-dir ./out
Transform a 30-second hook into a 2-minute arrangement by adding intro and outro sections in a single call.
What's Included
- 29 Production-Ready Skills: Fully implemented Claude agent skills that wrap 15+ AI models, from music to avatars to image editing, each with validated schemas and error handling.
- RunComfy CLI Integration: All skills use the open-source `runcomfy` CLI as the execution layer, with auto-polling, file management, and transparent error reporting via standardized exit codes.
- Model-Specific Documentation: Each skill includes detailed schema tables, prompting tips, common patterns, and comparison matrices to help users pick the right endpoint and parameters.
- Reproduction & Iteration Support: Seed fields for deterministic output, duration/resolution controls, and cost estimates per call. Designed for batch workflows and cost-sensitive iteration.
- Security & Privacy Guidance: Token storage best practices, input boundary documentation, and mitigation strategies for third-party content and steganography risks.
- Troubleshooting & Error Codes: Standardized exit codes (0 success, 64 bad args, 65 schema mismatch, 69 upstream error, 75 retryable, 77 auth), with links to troubleshooting guides.
Who It's For
- AI Engineers & Developers
- Content Creators & Video Producers
- Product Teams Building AI Features
- Automation & Data Science Teams
- AI Research & Prototyping Teams
Best For
- Batch Content Generation
- Creative Automation Workflows
- Multi-Modal AI Pipelines
- Cost-Sensitive Prototyping
- Reproducible AI Workflows