Player not loading? Watch on YouTube
Headroom compresses context between an AI agent and the model API. This review examines its input-token savings and compares them with Caveman's approach to shortening model output. The speaker describes Headroom as open source under Apache 2.0 and introduces installation through pip with all extras.
In the speaker's log test, 423 lines became 15, a reported 96% reduction. Both errors and the warning survived. The speaker says this compression ran without a model call or network access. That result concerns one noisy log, not a guaranteed reduction in API costs.
The pipeline routes content to different compression strategies. JSON compression retains anomalies while collapsing repetitive arrays. An AST pass preserves code signatures and imports but folds function bodies. Logs receive statistical filtering; prose uses a small language model to rank meaningful tokens. A cache aligner stabilizes the beginning of the context for provider caching.
Headroom stores original content behind a hash for later retrieval. The speaker warns that retrieval adds another API round trip and can outweigh savings on small payloads. A short recent tool call saved 0% in a separate test: by default, Headroom protects the last four messages and user-written input.
The review also cites published token reductions of 47% to 92% across several tasks and unchanged math accuracy of 0.87. These are reported benchmark results, not proof that compression preserves every answer.