
Elevenlabs Transcribe
Transcribe audio & video files with speaker diarization using ElevenLabs API
What You Can Do
You can transcribe audio and video files into accurate text with automatic speaker diarization (identifying who said what), support for 32+ languages, and detection of audio events like background noise or music. The skill accepts flexible parameters for output formatting, custom terminology, and speaker count optimization, making it ideal for podcasts, interviews, meetings, lectures, and multilingual content.
Features
automatically identifies and labels different speakers in audio files
transcribes 32+ languages with ISO-639 language code customization
identifies background noise, music, silence, and other acoustic events
bias transcription toward domain-specific terms or key phrases
save transcripts as .txt files with optional timestamps at word or character level
processes mp3, wav, mp4, m4a, ogg, flac, webm and other common audio/video formats
specify 1-32 speakers for optimized diarization accuracy
choose between none, word-level, or character-level timestamps for precise timing data
Example Output
Example 1: Basic podcast transcription
[Speaker 1 - 0:00-0:15]: "Welcome to the tech podcast. Today we're discussing AI deployment strategies."
[Speaker 2 - 0:16-0:42]: "Thanks for having me. The key challenge is managing model latency in production environments."
[Audio Event: Background Music] - 0:43-0:50
[Speaker 1 - 0:51-1:05]: "Can you walk us through your approach?"
Example 2: Meeting transcription with key terms
[Speaker 1 - CEO] - 0:00-0:30: "Q3 results show strong adoption of our API infrastructure."
[Speaker 2 - CFO] - 0:31-1:15: "Revenue grew 45% YoY. The machine learning pipeline optimization reduced costs significantly."
[Audio Event: Side Conversation] - 1:16-1:25
[Speaker 3 - Engineer] - 1:26-2:10: "Our deployment strategy now includes automated canary releases and real-time monitoring."
What's Included
- elevenlabs-transcribe SKILL.md: Complete skill definition with argument parsing and API integration
- ElevenLabs Scribe v2 API wrapper: Handles authentication, file streaming, and response formatting
- Parameter templates: Pre-built examples for language codes, speaker count configurations, and custom keyterm lists
- Output formatting guide: Markdown templates for speaker-labeled transcripts with timestamps
- Error handling checklist: Common issues (missing API key, unsupported formats, rate limits) and solutions
Who It's For
- Podcast producers & audio engineers — Generate searchable transcripts with speaker labels for distribution and SEO
- Researchers & academics — Transcribe interviews, lectures, and focus group discussions with speaker identification
- Meeting facilitators & note-takers — Convert team meetings, client calls, and webinars into timestamped transcripts
- Content creators & journalists — Rapidly transcribe raw audio from field recordings, interviews, and multimedia projects
- Software engineers & DevOps teams — Integrate transcription into CI/CD workflows for accessibility and documentation
Best For
- Multilingual audio transcription across 32+ languages with language-specific optimization
- Identifying and separating multiple speakers in interviews, panels, meetings, and podcasts
- Detecting non-speech audio events (music, noise, silence) for post-production analysis
- Domain-specific transcription (medical, legal, technical) using custom keyterm biasing
- Batch processing audio/video files from multiple sources with consistent formatting and timestamps







