
Pdf Reading
Read, extract, and inspect PDF content from disk
What You Can Do
You can read PDFs directly from disk, extract text and embedded content, identify document structure (text-based vs. scanned), list images and attachments, and choose the right extraction strategy based on document type. This skill helps you understand what's in a PDF before processing it, diagnose extraction issues, and pull specific content like tables, forms, and embedded files.
Features
quickly assess page count, file size, PDF version, and metadata using pdfinfo diagnostics
determine if a PDF is text-based or scanned before attempting extraction
identify embedded images, attachments, and supplementary files within the PDF
diagnose text extraction problems caused by broken font encoding or font substitution
render specific pages as images for visual inspection of layout, charts, and forms
locate and analyze structured data, forms, and embedded tables
automatically identify whether a PDF is text-heavy, scanned, a slide deck, form-based, or data-heavy
Example Output
Example 1: Quick content inventory of a multi-page report
File: annual_report.pdf
- Pages: 48 | Size: 12.5 MB | PDF version: 1.7
- Text extractable: ✓ (text-based PDF)
- Images found: 23 | Attachments: 2
- Fonts: Standard (no encoding issues detected)
Example 2: Extracting structured data from a form PDF
Document type: Form-based PDF
- Form fields identified: 14 (text, checkbox, dropdown)
- Embedded tables: 3
- Recommended extraction: pdfplumber for field mapping + pypdf for form data
Example 3: Diagnosing a scanned document
File: scanned_contract.pdf
- Text extractable: ✗ (image-based scan)
- OCR fallback recommended: Yes
- Resolution: 300 DPI (suitable for OCR)
What's Included
- SKILL.md: Complete PDF reading workflow with diagnostic commands and strategy selection
- REFERENCE.md: Advanced techniques (pypdfium2 rendering, pdfplumber table extraction, OCR fallback, corrupted PDF handling)
- Diagnostic checklist: Quick reference for running content inventories
- Document type decision tree: Framework for choosing extraction methods based on PDF structure
- Python snippet library: Ready-to-use code for common extraction tasks
Who It's For
- Content strategists and editors — audit PDF libraries and extract content for reuse
- Data analysts and researchers — extract tables, datasets, and structured information from reports
- Legal and compliance professionals — review and extract data from contracts, forms, and regulatory documents
- Product managers — inventory and assess documentation, user guides, and technical specifications
- Document management specialists — diagnose and prepare PDFs for processing pipelines
Best For
- Content extraction and migration — pulling text and structured data from PDFs into other formats
- Document audits and inventory — understanding what's in a large collection of PDFs
- Scanned document diagnosis — detecting OCR requirements and image quality issues
- Form and field extraction — identifying and mapping form data within PDFs
- Visual inspection workflows — rendering pages as images to review layout, charts, and formatting







