
Ai Multimodal
Transcribe audio, analyze images, extract documents, and generate images via Google Gemini
What You Can Do
You can process multimedia content across multiple formats: transcribe audio files up to 9.5 hours with timestamps and speaker identification, analyze images for objects, text, and visual relationships, extract structured data from PDF documents including tables and charts, and generate images from text prompts with refinement capabilities. This skill unifies fragmented multimodal workflows into one coherent Claude-powered interface, eliminating context-switching between different tools.
Features
Convert speech to text with timestamps, summarize audio content, identify speakers, and analyze music or environmental sounds from files up to 9.5 hours long
Perform OCR, detect objects with bounding boxes, segment pixels, answer questions about images, and compare multiple images (up to 3,600) in a single request
Detect scenes, analyze temporal sequences, answer questions with frame-level precision, and process YouTube URLs and local videos up to 6 hours
Process native PDFs up to 1,000 pages, extract tables and forms, analyze charts and diagrams, and convert visual data into structured formats
Create images from text prompts, edit existing images, refine compositions, and iterate on designs without leaving Claude
Access Gemini 2.5 and 2.0 models with context windows up to 2M tokens for handling large documents and complex analyses
Handle multiple files simultaneously and extract data from entire document sets in one workflow
Example Output
Audio Transcription Example:
File: meeting_recording.mp3 (45 minutes)
Transcript with timestamps:
[00:00] Speaker A: "Let's review Q4 targets..."
[02:15] Speaker B: "I suggest we prioritize the API integration..."
Summary: Meeting covered budget allocation, timeline adjustments, and resource planning. Key decision: move launch date to March 15.
Image Analysis Example:
Input: Screenshot of dashboard
Detected Objects: 3 charts, 2 tables, 1 title, 4 metric cards
Extracted Text: "Monthly Revenue: $245K | Growth: +18%"
Caption: Professional sales dashboard showing revenue trends and performance metrics by region.
Document Extraction Example:
Input: Multi-page PDF contract
Extracted Data:
- Table 1: Payment Terms (rows: 8, columns: 5) → CSV format
- Signature blocks identified: 3 locations
- Key clauses extracted: Liability, Indemnification, Termination
What's Included
- SKILL.md instruction file with full API reference and capability documentation:
- Audio Processing Template: Scripts for transcription workflows, speaker detection setup, and batch audio analysis
- Image Analysis Checklist: OCR extraction best practices, multi-image comparison patterns, object detection configuration
- Document Extraction Framework: PDF parsing workflows, table extraction pipelines, form data structuring examples
- Image Generation Workflow: Prompt engineering guidelines, iterative refinement patterns, composition templates
Who It's For
- Document Analysts — Extract data from PDFs, contracts, and compliance documents at scale without manual entry
- Content Creators — Generate images, transcribe video/podcast content, and extract visual elements from existing media
- Researchers — Analyze scientific papers, charts, and figures; process lecture recordings and conference presentations
- Data Engineers — Build pipelines that ingest multimodal data, normalize structured extraction, and feed downstream systems
- Product Managers — Analyze user feedback videos, extract UI screenshots, and generate design assets from text descriptions
Best For
- Batch transcription of long-form audio (podcasts, webinars, interviews, meetings)
- Optical character recognition (OCR) and text extraction from images and PDFs
- Visual data extraction (tables, charts, diagrams) from documents and screenshots
- Multi-image comparison and relationship analysis across dozens of files
- Text-to-image generation with iterative refinement for marketing and design workflows
- Video scene analysis and temporal question-answering for research and documentation






