headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Headroom is an MCP server and context compression layer that reduces token consumption for AI agents by compressing tool outputs, logs, files, and RAG chunks before they reach the LLM. It achieves 20% token reduction for coding agents and 60-95% for JSON payloads while preserving answer quality through reversible compression (CCR). The server runs locally, offers three integration modes (library, proxy, MCP server), and includes cross-agent memory sharing and output token reduction features. It supports major AI coding assistants including Claude Code, Codex, Cursor, Aider, and Cline.
Key Features
Use Cases
- 01Reduce token costs for daily coding agent sessions with Claude Code, Codex, or Cursor by 20-60%
- 02Compress large tool outputs from code search (100+ results), SRE incident logs, or codebase exploration before LLM processing
- 03Share compressed context and memory across multiple AI agents (Claude, Gemini, Grok) with automatic deduplication
- 04Enable retrieval of original uncompressed content when the LLM needs full detail via CCR
- 05Cut output token costs by steering model verbosity and reducing reasoning effort on routine tool result turns
- 06Integrate compression into existing LangChain, LiteLLM, Vercel AI SDK, or custom Python/TypeScript applications
Related MCP Servers
View moremarkitdown
Python tool for converting files and office documents to Markdown.
firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
prompts.chat
f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
langflow
Langflow is a powerful tool for building and deploying AI-powered agents and workflows.
headroom — FAQ
What is Headroom?+
Headroom is an MCP server and compression layer that reduces LLM token usage by compressing tool outputs, logs, files, and RAG results before they reach the model. It runs locally, preserves answer accuracy, and supports reversible compression so originals can be retrieved on demand.
How do I install Headroom as an MCP server?+
Install via 'uv tool install --python 3.13 headroom-ai[all]' or 'pip install headroom-ai[all]', then run 'headroom mcp install' to configure it for MCP clients. For manual setup, add the server to your MCP client config pointing to the 'headroom mcp serve' command.
Which AI clients work with Headroom?+
Headroom works as an MCP server with any MCP client (Claude Code, Codex, Cline, Continue). It also wraps Claude Code, Cursor, Aider, Copilot CLI, Goose, OpenHands, Grok CLI, and others via 'headroom wrap <agent>' or transparent proxy mode.
Do I need API keys or external services?+
No external services or API keys are required for compression—all processing runs locally on your machine. Your prompts and files never leave your system during compression. You still need your own LLM provider API keys (Anthropic, OpenAI, etc.) for the actual model calls.
Is Headroom free and open source?+
Yes, Headroom is fully open source under the Apache 2.0 license. The entire codebase is free to use, and there is an optional managed team offering for enterprises needing centralized deployment and support.
What compression ratios can I expect?+
Token reduction varies by content type: 20-30% for code search results, 40-60% for logs and codebase exploration, and 60-95% for repetitive JSON. Prose and already-dense content compress minimally. Run 'headroom savings' against your own traffic for personalized metrics.
How do I install headroom?+
Open the source repository on GitHub and follow its README. headroom is a mcp server — MCP Agents Market links you directly to the official repo.
Is headroom free?+
headroom is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.