SkillsLib.ai

Ai Multimodal

Transcribe audio, analyze images, extract documents, and generate images via Google Gemini

3.9(25 reviews)
100+ downloads
Updated Oct 2026
Verified SafeSecurity VerifiedThis skill was analyzed by our AI security scanner for harmful content including data exfiltration, system manipulation, credential theft, and prompt injection. No threats were detected.

What You Can Do

You can process multimedia content across multiple formats: transcribe audio files up to 9.5 hours with timestamps and speaker identification, analyze images for objects, text, and visual relationships, extract structured data from PDF documents including tables and charts, and generate images from text prompts with refinement capabilities. This skill unifies fragmented multimodal workflows into one coherent Claude-powered interface, eliminating context-switching between different tools.

Features

Audio Transcription & Analysis

Convert speech to text with timestamps, summarize audio content, identify speakers, and analyze music or environmental sounds from files up to 9.5 hours long

Image Understanding

Perform OCR, detect objects with bounding boxes, segment pixels, answer questions about images, and compare multiple images (up to 3,600) in a single request

Video Processing

Detect scenes, analyze temporal sequences, answer questions with frame-level precision, and process YouTube URLs and local videos up to 6 hours

Document Extraction

Process native PDFs up to 1,000 pages, extract tables and forms, analyze charts and diagrams, and convert visual data into structured formats

Image Generation

Create images from text prompts, edit existing images, refine compositions, and iterate on designs without leaving Claude

Multi-Model Support

Access Gemini 2.5 and 2.0 models with context windows up to 2M tokens for handling large documents and complex analyses

Batch Processing

Handle multiple files simultaneously and extract data from entire document sets in one workflow

Example Output

Audio Transcription Example:

code
File: meeting_recording.mp3 (45 minutes)
Transcript with timestamps:
[00:00] Speaker A: "Let's review Q4 targets..."
[02:15] Speaker B: "I suggest we prioritize the API integration..."
Summary: Meeting covered budget allocation, timeline adjustments, and resource planning. Key decision: move launch date to March 15.

Image Analysis Example:

code
Input: Screenshot of dashboard
Detected Objects: 3 charts, 2 tables, 1 title, 4 metric cards
Extracted Text: "Monthly Revenue: $245K | Growth: +18%"
Caption: Professional sales dashboard showing revenue trends and performance metrics by region.

Document Extraction Example:

code
Input: Multi-page PDF contract
Extracted Data:
- Table 1: Payment Terms (rows: 8, columns: 5) → CSV format
- Signature blocks identified: 3 locations
- Key clauses extracted: Liability, Indemnification, Termination

What's Included

  • SKILL.md instruction file with full API reference and capability documentation:
  • Audio Processing Template: Scripts for transcription workflows, speaker detection setup, and batch audio analysis
  • Image Analysis Checklist: OCR extraction best practices, multi-image comparison patterns, object detection configuration
  • Document Extraction Framework: PDF parsing workflows, table extraction pipelines, form data structuring examples
  • Image Generation Workflow: Prompt engineering guidelines, iterative refinement patterns, composition templates

Who It's For

  • Document Analysts — Extract data from PDFs, contracts, and compliance documents at scale without manual entry
  • Content Creators — Generate images, transcribe video/podcast content, and extract visual elements from existing media
  • Researchers — Analyze scientific papers, charts, and figures; process lecture recordings and conference presentations
  • Data Engineers — Build pipelines that ingest multimodal data, normalize structured extraction, and feed downstream systems
  • Product Managers — Analyze user feedback videos, extract UI screenshots, and generate design assets from text descriptions

Best For

  • Batch transcription of long-form audio (podcasts, webinars, interviews, meetings)
  • Optical character recognition (OCR) and text extraction from images and PDFs
  • Visual data extraction (tables, charts, diagrams) from documents and screenshots
  • Multi-image comparison and relationship analysis across dozens of files
  • Text-to-image generation with iterative refinement for marketing and design workflows
  • Video scene analysis and temporal question-answering for research and documentation

You might also like

Ui Styling
$35
Frontend4.1(31)
Ui Styling

You can rapidly prototype and build accessible user interfaces by combining shadcn/ui's pre-built Radix UI components with Tailwind CSS utility-first styling. This skill enables you to implement complex UI patterns, responsive layouts, dark mode themes, and design systems while maintaining full type safety and accessibility standards. Generate visual designs, posters, and branded materials alongside functional interface components.

Ui Ux Pro Max
$35
Ui Ux Pro Max

You can plan, build, and optimize complete UI/UX solutions for websites, dashboards, e-commerce platforms, and mobile apps. The skill provides curated design patterns (glassmorphism, minimalism, neumorphism, bento grids), 21 color palettes, 50 font pairings, 20+ chart types, and implementation guidance for 8 development stacks. Whether you're reviewing existing code, improving accessibility, or creating animations, you get framework-specific recommendations and responsive design strategies.

Gemini Cli
$25
Gemini Cli

This skill lets you leverage Gemini CLI as a complementary AI tool within your Claude-driven workflows. You can request code reviews from a different AI perspective, access current web information through Google Search integration, analyze complex codebase architectures, run parallel code generation tasks, and handle specialized operations like test suite generation or documentation—all while Claude coordinates the orchestration and results integration.

Aesthetic
$35
Web3.3(21)
Aesthetic

You can build visually stunning interfaces by systematically applying visual hierarchy, color theory, and typography principles. Analyze inspiration from design platforms, generate and iterate on design images until they meet aesthetic standards, and create comprehensive design systems that ensure your interfaces are not only beautiful but also accessible and functionally sound.

Systematic Prompt Evaluation & Optimization
$35
Systematic Prompt Evaluation & Optimization

Systematically evaluate your Claude prompts using quantitative scoring frameworks to identify weaknesses in clarity, task alignment, and output quality. You'll receive detailed optimization recommendations with A/B testing guidance to improve prompt effectiveness and consistency across different use cases and input variations.

Llm Cost Optimizer
$35
Llm Cost Optimizer

You can identify inefficiencies in your LLM API spending and receive prioritized recommendations with validated ROI estimates. The skill performs static analysis on your current prompts and configurations, examines dynamic usage patterns to pinpoint high-cost workflows, and runs comparative modeling across optimization scenarios. You'll get implementation roadmaps that quantify financial impact for each change.

Prompt Engineer
$45
Prompt Engineer

You can transform underperforming or inconsistent prompts into reliable, production-ready specifications that work across diverse inputs and use cases. This skill guides you through diagnosing prompt failures, applying evidence-based optimization techniques (chain-of-thought reasoning, few-shot examples, output constraints), testing variants systematically, and documenting final prompts for team-wide reproducibility. You'll move beyond trial-and-error to engineering prompts that scale reliably across different model versions and maintain quality with thousands of inputs.

$28.00$35.00