
Video Understand
Extract frames and transcribe video locally with ffmpeg and Whisper
What You Can Do
You can intelligently extract representative frames from videos using scene detection, keyframe analysis, or regular intervals, while simultaneously transcribing audio to text using OpenAI's Whisper model. This skill processes videos entirely offline, giving you frame-by-frame visual context and complete transcripts without relying on external APIs or cloud services.
Features
automatically identifies shot changes and extracts frames at transitions
isolates the most important frames based on visual content analysis
captures frames at regular time intervals for consistent sampling
converts video audio to text using offline speech recognition
control maximum frames extracted to balance detail vs. file size
choose from tiny to large models based on accuracy needs
extract visual content without transcription for faster processing
structured metadata with timestamps, frame paths, and transcription data
Example Output
Scene Detection with Transcription
{
"video": "interview.mp4",
"duration_seconds": 180,
"frames_extracted": 12,
"frames": [
{"timestamp": 0.0, "mode": "scene", "path": "frame_0000.jpg"},
{"timestamp": 15.3, "mode": "scene", "path": "frame_0001.jpg"}
],
"transcription": {
"full_text": "Welcome to the interview. Today we're discussing video production techniques...",
"segments": [
{"start": 0.0, "end": 5.2, "text": "Welcome to the interview."},
{"start": 5.2, "end": 12.1, "text": "Today we're discussing video production techniques..."}
]
}
}
Keyframe-Only Output
✓ Extracted 8 keyframes from video.mp4
✓ Identified critical visual moments
✓ Saved frames: frame_0000.jpg, frame_0001.jpg, ...
What's Included
- SKILL.md: complete setup and CLI reference
- understand_video.py: main script with scene detection, keyframe, and interval extraction modes
- Whisper integration: optional local transcription with configurable model sizes
- CLI flags & options: control frame limits, models, output formats, and processing modes
- JSON schema templates: structured output format for frame metadata and transcription segments
Who It's For
- Video editors — quickly understand video structure and content before detailed editing
- Content creators — extract key moments and auto-generate captions locally
- Researchers — analyze video datasets offline without API costs or privacy concerns
- Accessibility specialists — transcribe video audio to create subtitles and transcripts
- Post-production teams — preview shot sequences and identify editing points visually
Best For
- Transcribing video interviews and podcasts to searchable text
- Extracting thumbnail candidates and key frames for previews
- Analyzing shot composition and scene transitions
- Creating subtitle files from video audio tracks
- Cataloging video library content without external dependencies







