
Llm Cost Optimizer
Reduce LLM API costs by 30-60% through prompt compression, caching, and model optimization
What You Can Do
You can identify inefficiencies in your LLM API spending and receive prioritized recommendations with validated ROI estimates. The skill performs static analysis on your current prompts and configurations, examines dynamic usage patterns to pinpoint high-cost workflows, and runs comparative modeling across optimization scenarios. You'll get implementation roadmaps that quantify financial impact for each change.
Features
Identifies redundant tokens and restructures prompts to maintain quality while reducing input size by 20-40%
Recommends where prompt caching delivers maximum ROI based on your repetition patterns and API pricing
Analyzes task complexity and suggests cheaper model alternatives (e.g., Claude 3.5 Haiku vs. Opus) with quality tradeoff assessment
Identifies opportunities to group requests and reduce per-call overhead by 15-35%
Analyzes 2+ weeks of API logs to surface cost drivers and high-impact optimization targets
Quantifies monthly and annual savings for each optimization, prioritized by implementation effort
Provides step-by-step guidance for deploying changes without breaking production systems
Establishes metrics to ensure optimizations maintain output quality and performance standards
Example Output
Example 1: Prompt Compression
Original prompt: 1,200 tokens → Optimized: 720 tokens
CompressionRatio: 40% reduction
Method: Removed redundant context, converted instructions to structured format
MonthlySavings: $180 (at 0.003/1K tokens)
Example 2: Model Tier Switch
Current: 10,000 API calls/month using Claude 3.5 Opus
Recommendation: Switch 60% of calls to Claude 3.5 Haiku (for classification/extraction tasks)
ProjectedSavings: $420/month (42% reduction in token costs)
QualityRisk: Low — Haiku scores 94% accuracy on your existing test set
Example 3: Caching Impact
High-repetition workflow identified: 500 calls/week with identical 300-token system prompt
Caching opportunity: Cache the system prompt + conversation history
Savings: 150,000 tokens/month × $0.0008 = $120/month
Implementation effort: 30 minutes (add caching headers to API calls)
What's Included
- SKILL.md: Complete optimization framework with decision trees for each optimization vector
- Cost Analysis Template: Spreadsheet to input your API usage data and calculate baselines
- Prompt Compression Checklist: Step-by-step guide to identify and remove redundant tokens
- Model Selection Matrix: Capability and cost comparison chart across Claude models (3.5 Haiku, Sonnet, Opus)
- Implementation Roadmap Template: Prioritized action plan with effort estimates and safety checkpoints
- ROI Validation Worksheet: Track pre/post-optimization costs to measure actual savings
Who It's For
- AI/Prompt Engineers — Responsible for prompt design and API integration optimization
- ML/AI Team Leads — Managing LLM infrastructure costs at scale
- Finance/Operations Managers — Identifying cost reduction opportunities in AI spend
- Product Managers — Building LLM-powered features with budget constraints
- DevOps/Platform Engineers — Implementing caching and batching infrastructure
Best For
- Analyzing API billing data to identify cost drivers and inefficiencies
- Redesigning prompts for token efficiency without losing quality
- Evaluating trade-offs between model tiers and task-specific performance
- Planning infrastructure changes (caching, batching, queue optimization)
- Quantifying ROI before implementing cost-reduction changes







