Headroom

Local context compression for LLM apps and coding agents. Runs on your machine, keeps originals for retrieval, and works with OpenAI-compatible clients.

Screenshot of Headroom website

Headroom compresses the material an AI agent reads before it reaches the model. It's for developers whose LLM apps or coding assistants spend context space on repetitive tool results, logs, files, and retrieved documents. Compression runs on your machine, and no prompt or file content goes to an external service for compression. Requests still go to your chosen LLM provider.

Compression is reversible. Headroom caches originals locally and gives the model a retrieval tool so it can fetch details omitted from the shortened context. Its JSON compressor uses statistical analysis to retain errors, unusual values, and boundary items. Log compression keeps failures while removing routine passing output; search results are ranked by relevance. Source code compression preserves signatures and collapses function bodies, but it's opt-in. It also handles prose, Git diffs, and images.

You can use Headroom as a local proxy, a Python or TypeScript library, or an MCP server. It connects with Claude Code, Codex, Cursor, and other coding agents, plus frameworks such as LangChain and Vercel AI SDK. Provider support includes OpenAI, Anthropic, Google, and Bedrock, with broader compatibility through LiteLLM and OpenAI-compatible clients.

Savings depend on repetition: dense prose offers less room to shrink than recurring JSON records or verbose logs. Headroom also preserves stable prompt prefixes for provider caching, stores memory across conversations, and compresses context shared between agents. It's open source under Apache 2.0.

Similar to Headroom