
File Reading
Detect file types and read uploads without binary garbage or context overflow
What You Can Do
When a user uploads a file to Claude, you can intelligently detect its type and read it using the appropriate method. Instead of blindly running cat on binary files or loading massive CSVs into context, this skill routes each file type to the right tool, reading only what's needed to answer the user's question. You'll extract text from PDFs, parse structured data from CSVs and JSON, process images, and decompress archives—all without wasting context or producing garbage output.
Features
inspect extension and MIME type to route correctly
use stat before reading to avoid context overflow on large files
read text content without binary garbage or formatting bloat
load only rows needed, not entire 100MB datasets
extract metadata and visual content when applicable
list contents and decompress selectively without unpacking everything
read and validate without line-wrapping display issues
DOCX, XLSX, ebooks, and plaintext use optimized extraction
Example Output
Example 1: CSV file uploaded
user: I uploaded a sales report. How many rows are there?
clyde: I'll check the file size and row count without loading the whole thing.
$ stat -c '%s bytes' /mnt/user-data/uploads/sales_2024.csv
47283945 bytes
$ wc -l /mnt/user-data/uploads/sales_2024.csv
125847 lines
→ Your CSV has approximately 125,847 rows of data (47 MB).
Example 2: PDF uploaded
user: Can you summarize the key findings in the research paper I uploaded?
clyde: I'll extract the text from the PDF.
$ file /mnt/user-data/uploads/research.pdf
PDF document, version 1.4
[PDF text extraction → first 2000 tokens of readable content]
→ The paper demonstrates that X correlates with Y under conditions Z...
Example 3: ZIP archive uploaded
user: What's in this backup file?
clyde: I'll list the archive contents.
$ unzip -l /mnt/user-data/uploads/backup.zip | head -20
→ Archive contains 342 files: configs/, data/, logs/, reports/...
What's Included
- SKILL.md: core routing logic and protocol for each file type
- File type dispatch table: extension-to-tool mapping for PDFs, CSVs, DOCX, XLSX, JSON, images, archives, ebooks
- Stat & sampling checklist: when to check file size before reading
- Context-safe extraction templates: snippets for each format to avoid overflow
- Fallback protocol: how to handle unknown or corrupted file types
Who It's For
- Data analysts — quickly sample large CSVs and tabular exports without context waste
- Researchers — extract and summarize text from academic PDFs and documents
- Engineers — inspect config files, logs, and structured data in archives
- Product managers — read uploaded reports, spreadsheets, and presentation exports
- Content reviewers — process batches of images, ebooks, and document formats
Best For
- Reading uploaded files without seeing binary garbage or raw ZIP bytes
- Sampling large datasets (CSVs, JSON logs) to answer specific questions quickly
- Extracting text from multi-page PDFs for summarization or analysis
- Inspecting archive contents before deciding what to decompress
- Processing mixed file types in a single upload batch







