SkillsLib.ai

Pdf Reading

Read, extract, and inspect PDF content from disk

4.6(47 reviews)
1,000+ downloads
Updated Sep 2026
Verified SafeSecurity VerifiedThis skill was analyzed by our AI security scanner for harmful content including data exfiltration, system manipulation, credential theft, and prompt injection. No threats were detected.

What You Can Do

You can read PDFs directly from disk, extract text and embedded content, identify document structure (text-based vs. scanned), list images and attachments, and choose the right extraction strategy based on document type. This skill helps you understand what's in a PDF before processing it, diagnose extraction issues, and pull specific content like tables, forms, and embedded files.

Features

Content inventory

quickly assess page count, file size, PDF version, and metadata using pdfinfo diagnostics

Text extraction detection

determine if a PDF is text-based or scanned before attempting extraction

Image and attachment listing

identify embedded images, attachments, and supplementary files within the PDF

Font and encoding analysis

diagnose text extraction problems caused by broken font encoding or font substitution

Page rasterization

render specific pages as images for visual inspection of layout, charts, and forms

Table and form detection

locate and analyze structured data, forms, and embedded tables

Document type classification

automatically identify whether a PDF is text-heavy, scanned, a slide deck, form-based, or data-heavy

Example Output

Example 1: Quick content inventory of a multi-page report

code
File: annual_report.pdf
- Pages: 48 | Size: 12.5 MB | PDF version: 1.7
- Text extractable: ✓ (text-based PDF)
- Images found: 23 | Attachments: 2
- Fonts: Standard (no encoding issues detected)

Example 2: Extracting structured data from a form PDF

code
Document type: Form-based PDF
- Form fields identified: 14 (text, checkbox, dropdown)
- Embedded tables: 3
- Recommended extraction: pdfplumber for field mapping + pypdf for form data

Example 3: Diagnosing a scanned document

code
File: scanned_contract.pdf
- Text extractable: ✗ (image-based scan)
- OCR fallback recommended: Yes
- Resolution: 300 DPI (suitable for OCR)

What's Included

  • SKILL.md: Complete PDF reading workflow with diagnostic commands and strategy selection
  • REFERENCE.md: Advanced techniques (pypdfium2 rendering, pdfplumber table extraction, OCR fallback, corrupted PDF handling)
  • Diagnostic checklist: Quick reference for running content inventories
  • Document type decision tree: Framework for choosing extraction methods based on PDF structure
  • Python snippet library: Ready-to-use code for common extraction tasks

Who It's For

  • Content strategists and editors — audit PDF libraries and extract content for reuse
  • Data analysts and researchers — extract tables, datasets, and structured information from reports
  • Legal and compliance professionals — review and extract data from contracts, forms, and regulatory documents
  • Product managers — inventory and assess documentation, user guides, and technical specifications
  • Document management specialists — diagnose and prepare PDFs for processing pipelines

Best For

  • Content extraction and migration — pulling text and structured data from PDFs into other formats
  • Document audits and inventory — understanding what's in a large collection of PDFs
  • Scanned document diagnosis — detecting OCR requirements and image quality issues
  • Form and field extraction — identifying and mapping form data within PDFs
  • Visual inspection workflows — rendering pages as images to review layout, charts, and formatting

You might also like

File Reading
$45
Backend4.4(50)
File Reading

When a user uploads a file to Claude, you can intelligently detect its type and read it using the appropriate method. Instead of blindly running cat on binary files or loading massive CSVs into context, this skill routes each file type to the right tool, reading only what's needed to answer the user's question. You'll extract text from PDFs, parse structured data from CSVs and JSON, process images, and decompress archives—all without wasting context or producing garbage output.

Research Data Curation Workflow
$30
Research Data Curation Workflow

Automate the curation of research datasets by generating standardized metadata, performing compliance verification, and producing quality assessments. This skill standardizes your data documentation process, ensures datasets meet regulatory requirements, and prepares them for repository publication—reducing manual work and improving consistency across your research organization.

Pdf
$25
Backend4.4(48)
Pdf

You can perform end-to-end PDF operations including extracting text and metadata from documents, merging multiple PDFs into a single file, splitting PDFs by page ranges, filling fillable forms programmatically, and adding text overlays to non-fillable PDFs. This skill handles both simple read operations and complex document workflows, making it essential for document processing, form automation, and PDF batch operations.

Mechanical License Compliance Auditor
$45
Mechanical4.0(34)
Mechanical License Compliance Auditor

You can systematically verify licensing compliance across music compositions, calculate accurate mechanical royalties against current statutory rates, and generate audit-ready documentation that satisfies internal controls and external audit demands. This skill consolidates manual spreadsheet workflows into a structured, repeatable process with built-in error detection and cross-referencing against NMPA rates, licensing agreements, and regulatory requirements.

Music Licensing Rights Clearance Advisor
$25
Licensing4.2(11)
Music Licensing Rights Clearance Advisor

You can systematically identify all licensing requirements for a musical work across multiple territories, coordinate approvals from rights holders and collection societies, and generate legally compliant clearance letters and license terms. This skill maps composition rights (mechanical, performance, synchronization) and sound recording rights (master use, digital performance) while accounting for territory-specific regulations like EU directives, UK post-Brexit rules, and jurisdiction-specific collection society requirements.

Xlsx
$45
Xlsx

This skill lets you handle spreadsheet files as primary inputs and outputs. You can open and read existing Excel or CSV files, edit them by adding columns, computing formulas, and applying formatting, clean up messy or malformed data into proper spreadsheets, create new spreadsheets from scratch or from other data sources, and convert between tabular file formats. Every output file is delivered error-free and production-ready.

Research Data Documentation Generator
$45
Research Data Documentation Generator

You can create comprehensive, standards-compliant documentation for research datasets including data dictionaries, metadata records, and institutional compliance reports. The skill generates structured documentation following FAIR principles (Findable, Accessible, Interoperable, Reusable), validates against disciplinary standards like DDI, DataCite, or Dublin Core, and produces documentation your institution and publishers require for data archival and discovery.

Product Self Knowledge
$25
Tier 24.0(28)
Product Self Knowledge

This skill prevents you from providing incorrect information about Anthropic's products by routing you to official documentation. You'll verify details about Claude API (models, rate limits, function calling, batch processing), Claude Code (installation, Node.js requirements, MCP integration), and Claude.ai (subscription tiers, feature limits) before answering user questions. Use it whenever you're uncertain about product specs, pricing, or capabilities.

$35.00