</>MCP Agents Market
MCP Server

headroom

by headroomlabs-ai74.3kPythonUpdated 2026-10-02

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

Claude CodeCodexCursorClineContinueAiderGoose

Headroom is an MCP server and context compression layer that reduces token consumption for AI agents by compressing tool outputs, logs, files, and RAG chunks before they reach the LLM. It achieves 20% token reduction for coding agents and 60-95% for JSON payloads while preserving answer quality through reversible compression (CCR). The server runs locally, offers three integration modes (library, proxy, MCP server), and includes cross-agent memory sharing and output token reduction features. It supports major AI coding assistants including Claude Code, Codex, Cursor, Aider, and Cline.

Key Features

MCP server exposing headroom_compress, headroom_retrieve, and headroom_stats tools for any MCP client
Content-aware compression routing: SmartCrusher for JSON (60-95% reduction), CodeCompressor for source code (AST-aware), Kompress-v2-base ML model for text
Reversible compression (CCR) that caches originals locally for on-demand retrieval via headroom_retrieve
Cross-agent memory with automatic deduplication shared across Claude, Codex, Gemini, and Grok
Output token reduction through verbosity steering and effort routing (reduces model response tokens by 31.7% estimated)
Live-zone compression that preserves provider KV-cache by only compressing new content, never rewriting frozen context
One-command agent wrapping for Claude Code, Cursor, Aider, Copilot CLI, Cline, Continue, Goose, and others
Failure learning with 'headroom learn' that mines failed sessions and auto-generates corrections to CLAUDE.md or AGENTS.md

Use Cases

  • 01Reduce token costs for daily coding agent sessions with Claude Code, Codex, or Cursor by 20-60%
  • 02Compress large tool outputs from code search (100+ results), SRE incident logs, or codebase exploration before LLM processing
  • 03Share compressed context and memory across multiple AI agents (Claude, Gemini, Grok) with automatic deduplication
  • 04Enable retrieval of original uncompressed content when the LLM needs full detail via CCR
  • 05Cut output token costs by steering model verbosity and reducing reasoning effort on routine tool result turns
  • 06Integrate compression into existing LangChain, LiteLLM, Vercel AI SDK, or custom Python/TypeScript applications

Related MCP Servers

View more

headroom — FAQ

What is Headroom?+

Headroom is an MCP server and compression layer that reduces LLM token usage by compressing tool outputs, logs, files, and RAG results before they reach the model. It runs locally, preserves answer accuracy, and supports reversible compression so originals can be retrieved on demand.

How do I install Headroom as an MCP server?+

Install via 'uv tool install --python 3.13 headroom-ai[all]' or 'pip install headroom-ai[all]', then run 'headroom mcp install' to configure it for MCP clients. For manual setup, add the server to your MCP client config pointing to the 'headroom mcp serve' command.

Which AI clients work with Headroom?+

Headroom works as an MCP server with any MCP client (Claude Code, Codex, Cline, Continue). It also wraps Claude Code, Cursor, Aider, Copilot CLI, Goose, OpenHands, Grok CLI, and others via 'headroom wrap <agent>' or transparent proxy mode.

Do I need API keys or external services?+

No external services or API keys are required for compression—all processing runs locally on your machine. Your prompts and files never leave your system during compression. You still need your own LLM provider API keys (Anthropic, OpenAI, etc.) for the actual model calls.

Is Headroom free and open source?+

Yes, Headroom is fully open source under the Apache 2.0 license. The entire codebase is free to use, and there is an optional managed team offering for enterprises needing centralized deployment and support.

What compression ratios can I expect?+

Token reduction varies by content type: 20-30% for code search results, 40-60% for logs and codebase exploration, and 60-95% for repetitive JSON. Prose and already-dense content compress minimally. Run 'headroom savings' against your own traffic for personalized metrics.

How do I install headroom?+

Open the source repository on GitHub and follow its README. headroom is a mcp server — MCP Agents Market links you directly to the official repo.

Is headroom free?+

headroom is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.

Related searches