
Api Data Extractor
Extract, transform, and load data from REST and GraphQL APIs with authentication and pagination
What You Can Do
You can pull structured data from any REST or GraphQL endpoint, automatically managing authentication (API keys, OAuth, Basic Auth), pagination across multiple pages, rate limit compliance, and transient failure retries. Claude transforms raw API responses into normalized, schema-validated datasets deduplicated and formatted for immediate downstream analysis or storage.
Features
automatically detects and handles both API types
supports API keys, OAuth tokens, Basic Auth, and custom headers
manages cursor-based, offset, and page-number pagination schemes
respects API quotas and implements exponential backoff retries
validates extracted data against your predefined data schemas
gracefully handles 429, 503, timeout, and network errors
flattens nested structures, renames fields, converts types, deduplicates
reports success rates, row counts, and failed requests with diagnostics
Example Output
Example 1: REST API extraction with pagination
Extracted 2,847 user records from /api/v1/users endpoint
- Applied pagination: 20 records/page × 143 pages
- Schema validation: 100% success (0 failures)
- Deduplicated: 12 duplicate entries removed
- Output: users_2024-01-15.json (2.3 MB)
Example 2: GraphQL transformation
Query: { users { id name email posts { id title } } }
Extracted 1,543 users with nested posts flattened
Transformation applied:
- Nested posts array → separate user_posts table
- email field renamed to contact_email
- created_at timestamp → ISO 8601 format
Output: users.csv (143 KB), user_posts.csv (856 KB)
Example 3: Rate-limited API with retries
Stripe API extraction (100 req/sec limit)
- Initial attempt: 3,200 requests over 45 seconds
- Rate limit hit: 429 responses at request 1,847
- Backoff retry: resumed after 30s delay
- Final success: 5,120 transaction records extracted
- Duration: 2 min 15 sec (with intelligent delays)
What's Included
- api-data-extractor SKILL.md: instruction file with activation signals and best practices
- REST API template: pre-built configuration for common REST endpoints with pagination patterns
- GraphQL query template: example queries and response transformation logic
- Authentication config checklist: secure credential handling patterns for API keys, OAuth, tokens
- Error handling reference: retry strategies, rate limit detection, and failure logging
- Schema validation framework: JSON schema templates for validating extracted datasets
Who It's For
- Backend engineers — building data pipelines and ETL workflows
- Data engineers — extracting datasets from SaaS APIs for warehouses
- API integration specialists — connecting third-party services and syncing data
- Product analysts — pulling user and usage data from analytics or business APIs
- DevOps/Site reliability engineers — automating data collection for monitoring and reporting
Best For
- Multi-page data extraction — handling pagination at scale
- Rate-limited API integrations — respecting quotas and implementing backoff logic
- Data transformation workflows — normalizing and flattening nested API responses
- Schema validation pipelines — ensuring data quality before storage or analysis
- Error-resilient batch jobs — recovering from transient failures and retrying intelligently






