RunComfy AI Media Skills
AI music, video, and image generation suite for Claude agents using RunComfy
What You Can Do
Access RunComfy's full suite of AI media models directly from Claude. Generate music with ACE Step (tag-driven composition, inpainting, outpainting at $0.0002/s), create talking-head and avatar videos with OmniHuman, Wan 2-7, HappyHorse, and Seedance v2. Edit images, generate videos, perform face-swaps, and inpaint media. Skills classify user intent automatically and pick the right model, invoke via RunComfy CLI, and retrieve outputs into your workflow.
Features
Tag-driven composition (genre, mood, instruments, BPM), multilingual lyric support (50+ languages via ACE Step 1.5), audio inpainting (regenerate time ranges), audio outpainting (extend before/after), 5-240 seconds per call. Cost: $0.0002-0.0003/s (27x cheaper than ElevenLabs Music).
Routes across OmniHuman (audio-driven full-body from portrait), Wan 2-7 (scene-controlled with audio_url), HappyHorse (text-to-video or image-to-video with in-pass audio), and Seedance v2 Pro (cinematic multi-modal). Automatic intent classification picks the right model.
Multiple models including Flux Kontext, GPT Image Edit, Nano Banana 2, and nano-banana-edit. Generate, edit, and manipulate images with straightforward JSON schemas.
Text-to-video, image-to-video, video inpainting, video editing. Models include HappyHorse 1.0 (Arena top-ranked), Seedance v2, Kling 3.0, Seedance v2 Minimax, and WAN 2-7.
Skills analyze user requests (cost-sensitive vs premium, photoreal vs stylized, cinematic vs simple, script vs audio file) and automatically select the optimal model and endpoint. No manual model ID selection needed.
Fixed seeds for reproducibility, detailed cost tables per endpoint, cheap-draft patterns (iterate with ACE Step base then polish), and explicit pricing guidance throughout.
Token stored with mode 0600, JSON-only input (no shell injection), input boundary validation, outbound endpoint allowlist, guidance on third-party content risks and lyrics provenance.
Unified command interface (`runcomfy run <model_id> --input '...' --output-dir ./out`). One-line install, automatic polling, automatic file downloads, structured exit codes (0-77) for reliable error handling.
Example Output
Generate background music for a game loop:
runcomfy run acestep-ai/ace-step/text-to-audio \
--input '{"tags": "lo-fi hip-hop, mellow, vinyl crackle, rhodes piano, soft drums, 75 BPM", "lyrics": "[inst]", "duration": 90}' \
--output-dir ./out
ACE Step generates a 90-second loopable background track with consistent groove and seamless loop structure. Cost: approximately 1.8 cents.
Create a talking-head video from existing voiceover:
runcomfy run bytedance/omnihuman/api \
--input '{"image_url": "https://your-cdn.example/presenter.jpg", "audio_url": "https://your-cdn.example/voiceover.mp3"}' \
--output-dir ./out
OmniHuman syncs the portrait to the audio with natural mouth movements and gestures. One call, no prompt required, output ready for immediate use.
Extend a short musical hook into a complete track:
runcomfy run acestep-ai/ace-step/audio-outpaint \
--input '{"audio": "https://your-cdn.example/hook-30s.mp3", "tags": "indie pop, electric guitar, drums, build-up, fade outro", "extend_before_duration": 30, "extend_after_duration": 60}' \
--output-dir ./out
ACE Step outpainting adds 30 seconds of intro and 60 seconds of outro to the original hook, creating a full 2-minute production-ready track.
What's Included
- 30+ AI Skills: ACE Step music generation, OmniHuman avatar, Wan 2-7 video, HappyHorse talking-head, Seedance cinematic, plus image generation, video editing, inpainting, face-swap, and specialized models.
- Documented Prompting Patterns: Tag combinations for music, lyric markers for structured vocals, scene descriptions for video, model-specific tips (e.g., quote scripts for HappyHorse, match tags for inpaint blending).
- Intent Classification Logic: Built-in classification to analyze whether user has pre-recorded audio vs needs generation, photoreal vs stylized, single-shot vs cinematic, and route to optimal model automatically.
- Cost Analysis & Comparison Tables: Pricing breakdown per endpoint, ACE Step vs ElevenLabs Music tradeoff matrix, cheap-draft patterns, and model-specific cost estimates with Gaussian distributions.
- Security & Privacy Guidance: Best practices for token storage, input boundary analysis (JSON prevents shell injection), third-party content risk assessment, and lyrics provenance confirmation.
- Exit Code Reference: Structured error codes (0 success, 64 bad args, 65 schema mismatch, 69 upstream error, 75 retryable timeout, 77 auth failed) for reliable error handling in automation.
Who It's For
- Content Creators
- Game & Interactive Developers
- Video Producers & Editors
- Music Producers & Sound Designers
- Marketing & UGC Creative Teams
Best For
- Batch media generation at scale
- Multi-language dubbed content pipelines
- Iterative creative workflows and drafting
- Cost-sensitive background media production
- Avatar and talking-head video workflows





