
Headroom compresses the material an AI agent reads before it reaches the model. It's for developers whose LLM apps or coding assistants spend context space on repetitive tool results, logs, files, and retrieved documents. Compression runs on your machine, and no prompt or file content goes to an external service for compression. Requests still go to your chosen LLM provider.
Compression is reversible. Headroom caches originals locally and gives the model a retrieval tool so it can fetch details omitted from the shortened context. Its JSON compressor uses statistical analysis to retain errors, unusual values, and boundary items. Log compression keeps failures while removing routine passing output; search results are ranked by relevance. Source code compression preserves signatures and collapses function bodies, but it's opt-in. It also handles prose, Git diffs, and images.
You can use Headroom as a local proxy, a Python or TypeScript library, or an MCP server. It connects with Claude Code, Codex, Cursor, and other coding agents, plus frameworks such as LangChain and Vercel AI SDK. Provider support includes OpenAI, Anthropic, Google, and Bedrock, with broader compatibility through LiteLLM and OpenAI-compatible clients.
Savings depend on repetition: dense prose offers less room to shrink than recurring JSON records or verbose logs. Headroom also preserves stable prompt prefixes for provider caching, stores memory across conversations, and compresses context shared between agents. It's open source under Apache 2.0.
Claim this page with an email at docs.headroomlabs.ai. Headroom gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Headroom?Promote it
Something wrong or outdated on this page?
3.8KUpdated 2 months agoMIT
Linux · Docker · Web#Hybrid search#Knowledge graphs#LM Studio integration
SimpleMem is a Python memory system for developers building AI agents that need to recall earlier conversations and stored media. You can self-host it with Docker or use its hosted text-memory service. The code is open source under the MIT license.
3.7KUpdated 5 months agoApache-2.0
#Agent Skills#Code execution#Context compression
8.5KUpdated 5 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
31.3KUpdated 1 day agoApache-2.0
Docker#Hybrid search#Knowledge graphs#llama.cpp backend
13.3KUpdated 1 day agoApache-2.0
Docker#Agent Skills#Hybrid search#MCP
4.5KUpdated 2 months agoApache-2.0
Docker · Web#Hybrid search#Knowledge graphs#MCP
Acontext turns AI agent sessions into Markdown skill files that agents can reuse on later tasks. It's for developers who want agents to retain successful approaches, past mistakes, and user preferences in memory they can inspect and correct. The full stack can run on your own infrastructure, and a hosted cloud service is also available. It's open source under Apache 2.0.
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
Graphiti is a self-hosted Python framework for developers building AI agents that need to remember changing facts. It builds knowledge graphs from conversations, structured records and unstructured text, so an agent can query current information or recover what was true earlier. It's open source under Apache 2.0.
EverOS is a self-hosted memory runtime and Python library for developers building AI agents that need to remember across sessions and applications. It stores conversations, files, and task histories as readable Markdown, with local SQLite and LanceDB indexes for retrieval. The code uses the Apache 2.0 license; a managed cloud service is also available.
M Flow is a self-hosted memory engine for developers building AI agents and applications that need to recall earlier conversations, facts, and workflows. Its distinguishing feature is how it selects context: vector search finds possible matches, then a knowledge graph ranks them by the evidence connecting them to the query. It's open source under Apache 2.0 and runs as a Python library or a Docker service.