Caveman Mode
Cut token usage 65% by talking like a caveman
What You Can Do
Drastically reduce token consumption in Claude interactions by using ultra-compressed communication that preserves full technical accuracy. Caveman mode strips articles, filler, and hedging while keeping all technical substance, code, and meaning intact. Choose from multiple intensity levels (lite, full, ultra) and linguistic variants (wenyan-lite, wenyan-full, wenyan-ultra). Includes cavecrew subagent delegation that compresses agent output by 60%, memory file compression for CLAUDE.md files, and real-time token savings tracking.
Features
Strip articles, filler, and hedging while preserving technical accuracy. Measured 65% token reduction with full substance retention.
Choose from lite (professional tight), full (classic caveman), ultra (stripped conjunctions), or wenyan variants for classical Chinese compression.
Spawn caveman-compressed investigator, builder, or reviewer agents. Tool results return 60% smaller, extending main context duration across long sessions.
Compress CLAUDE.md, todos, and preferences while preserving code blocks, URLs, file paths, and all technical content exactly.
Real-time stats showing actual token savings, estimated overhead, and net efficiency gains with honest warnings for negative-savings scenarios.
Preserves user's dominant language and technical terms verbatim. Compresses style, not language, with classical Chinese mode.
Active every response until explicitly turned off. No filler drift. Auto-triggers when token efficiency requested.
Example Output
Extending session context
How can I run longer Claude sessions without losing context?
Response: Three ways. Lite mode drops filler but keeps grammar. Cavecrew split work to subagents with 60% smaller results. Compress CLAUDE.md with /caveman-compress saves 40-60% of memory file.
High-volume token optimization
We run 100 Claude analyses daily and token costs are high.
Response: Caveman full mode in batch scripts. Measured 65% output token savings, zero accuracy loss. Use cavecrew subagents for investigation phases. Compress memory at session start. Backend savings: ~$30-40 per 1000 analyses.
Mode intensity comparison
Show me the same technical answer in different caveman intensities.
Response: Full mode: "Component re-renders because new object ref created each render. Wrap in useMemo." Ultra mode: "New object ref each render triggers re-render. useMemo fixes." Wenyan-full: "每繪新生對象參照,故重繪;以 useMemo 包之則免。"
What's Included
- Caveman skill: Main mode with full, lite, ultra, and wenyan intensity levels. Auto-activates on user request or token-efficiency context.
- Cavecrew system: Three subagent presets: cavecrew-investigator (locate code), cavecrew-builder (1-2 file edits), cavecrew-reviewer (diff audit). All output compressed.
- Caveman-compress utility: Compress natural language memory files while preserving code blocks, URLs, paths, and technical content exactly. Creates human-readable backups.
- Caveman-stats tracker: Real-time token usage metrics, estimated savings, per-turn overhead, and net efficiency with honest warnings for negative-savings scenarios.
- Decision guides and documentation: When to use caveman vs vanilla agents, chaining patterns (locate-fix-verify), and compression rules that guarantee accuracy preservation.
Who It's For
- Claude Code users managing context limits
- Engineering teams optimizing token budgets
- Researchers running high-volume LLM analysis
- Developers in long multi-hour coding sessions
Best For
- Extending main context duration in long sessions
- Reducing token costs for batch processing
- Maintaining accuracy while cutting verbosity
- Memory file compression and archival