HOW AI AGENTS CAN USE FEWER TOKENS: A PRACTICAL INTRODUCTION TO HEADROOM
Learn where context compression is actually useful, how to integrate Headroom into an app or agent, and why a high compression rate does not necessarily translate into cost savings.


Author:
Dmitrii Vasilev
Category:
articles
Publication date:
Intro
Imagine an AI agent needs to find the cause of an incident in an event log. The log contains thousands of identical messages indicating normal operation, dozens of technical fields, and only a few relevant errors that actually explain what went wrong.
The agent receives the entire log, sends it to the model, and then rereads specific sections several more times. The task may be completed correctly, but the team still pays to process a large amount of almost useless context. A larger context window does not solve the problem: the data fits into the request, but processing it still takes time and may still incur costs.
One way to reduce these costs is to shrink tool outputs before they are added to the request sent to the model. That is the purpose of the open-source project Headroom.
This article focuses on a specific release: Headroom v0.34.0, commit `9fd5ae3`, released on August 5, 2026. This distinction matters because the project is actively evolving, so the capabilities of main, older documentation, and the version covered here may differ.
Why Agent Context Grows So Quickly
In a regular chat, most of the context comes from the conversation itself. An agent also accumulates tool outputs:
- An API returns hundreds of objects with the same fields.
- A monitoring system produces a large log in which most events are repetitive.
- A web tool returns HTML along with menus, navigation, and technical markup.
- A RAG system retrieves several similar fragments from the same document.
- A coding agent reads entire files for the sake of a single function or signature.
- Results from previous calls are included in subsequent requests again.
At the level of a single request, the difference may seem small. But an agentic task usually consists of a chain of steps: the model calls a tool, examines the response, requests additional data, and only then produces the final result. If a large response remains in the history, some of that unnecessary context gets processed again.
A large context window is like a spacious folder for documents. Being able to put a thousand pages into it does not make reading those pages free. What's more, when repetitions and technical noise accumulate, it becomes harder for the model to spot a rare error, important date, or unusual value.
The goal of optimization, therefore, is not simply to send fewer tokens. The aim is to retain enough information for the agent to make the right decision without forcing it to compensate for compression by rereading the data.

A large tool response can increase costs more than once: it remains in the history and is processed again in subsequent requests to the model.
What Headroom Does
In simplified form, Headroom sits here:
tool response → Headroom → compressed context → model
Here, compression does not mean creating a ZIP archive or generating one universal summary. Headroom identifies the type of content and applies an appropriate transformation: structurally compacting JSON, reducing repetitive logs, preserving important parts of code, or shortening ordinary text.
The potential benefits appear in three areas:
- The model receives fewer input tokens.
- More room remains in the context window for useful information.
- Large tool responses are less likely to interfere with subsequent agent steps.
Reduced latency is possible, but not guaranteed: local compression takes time too.
There is an important limitation. Headroom only processes content that is routed through its library, proxy, wrapper, or custom MCP tools. It does not connect to an arbitrary agent and automatically intercept all of its tool responses.
Headroom also works well alongside tools for structural code navigation, such as CodeGraph*. CodeGraph helps the agent identify relevant parts of a project in advance, so it does not have to read unnecessary files. Headroom comes in later, once the data has already been retrieved but before it is sent to the model. The two approaches are complementary rather than competing: one reduces unnecessary reading, while the other reduces the amount of context that has already been retrieved.

CodeGraph and Headroom reduce context at different stages.
Sources: CodeGraph and Headroom v0.34.0
* For more details about how to improve code search with CodeGraph read the article.
Where Headroom Can Be Useful
The best candidate for a pilot is not simply any large request, but a repeatable scenario in which one or more tools regularly return large amounts of structured, technical, or duplicated content.
|
Scenario |
Potential benefit |
What to verify |
|
Incident diagnosis |
Reduce repetitive normal events |
Are errors, event sequences, and anomalies preserved? |
|
Large API responses |
Represent repetitive objects and duplicates more compactly |
Are rare values, outliers, and important fields preserved? |
|
Documents and web pages |
Remove navigation noise, markup, and repetition |
Are dates, terms, sources, and caveats still present? |
|
RAG and enterprise search |
Avoid passing several nearly identical fragments |
Have answer completeness and attribution been preserved? |
|
Working with code |
Preserve the structural overview without full function bodies |
Is the original needed for precise edits and verification? |
Logs and Monitoring
Logs often contain long sequences of health checks, successful requests, and identical metrics. Against this background, the cause of an incident may consist of only a few lines among tens of thousands.
Headroom aims to preserve errors and anomalies, but it does so using heuristics. A pilot should therefore use a log with a known root cause. The evaluation should focus not only on the final answer, but also on the evidence: which events the agent connected and whether it selected a prominent but misleading signal.
APIs and Structured Data
JSON arrays are particularly well suited to compression when they contain many similar objects. For example, a tool might return 500 orders even though the task only requires the overall picture, a few representative records, and exceptions.
In Headroom v0.34.0, such data is processed through structural transformation, deduplication, and selection. This does not mean the tool will always simply remove identical fields or guarantee that a rare value is preserved. If exceptions matter more than typical entries, that needs to be part of the quality evaluation.
Documents, HTML, and Search Results
Research agents often receive useful text alongside menus, repeated blocks, markup, and technical page fragments. Reducing this noise allows the model to process more sources within the same context window.
The main risk is losing conditions and caveats. A sentence such as "the feature is available only if three conditions are met" should not become "the feature is available" after overly aggressive processing. For contracts, medical materials, and other documents where wording matters, compressed output should not be used as the sole source.
RAG and Knowledge Bases
In a RAG system, it usually makes more sense to improve retrieval itself first: tune ranking, remove duplicates, and limit the number of fragments. Headroom can be the next layer if even relevant search results remain too large.
Here, it is important to evaluate not only answer accuracy but also attribution: can the agent identify the document and fragment on which its conclusion is based?
Coding Agents
For supported languages, Headroom can preserve imports, signatures, types, and the overall syntactic structure while reducing function bodies and comments. This is useful for getting an initial overview of a large code fragment.
However, this representation does not replace a dependency graph and is not always appropriate before making edits. If the agent needs to modify a specific branch, check exception handling, or preserve formatting, it will likely need the original source code.
Why Different Types of Data Are Processed Differently
The word compression in Headroom covers several mechanisms:
- JSON and structured responses are transformed based on their schema, repetitions, and characteristic elements.
- Logs, search results, and HTML undergo specialized processing designed to address the types of noise typical of each format.
- Source code can be reduced with syntax awareness in modes where CodeAware is enabled.
- Ordinary text can, by default, be processed by the local Kompress model, which selects portions of the content; an external endpoint can be configured instead, and if it is unavailable, the transformation operates in fail-open mode.
Kompress is not a separate cloud LLM that reads the user's question and generates a new summary. In v0.34.0, it is a local token-classification model; it does not select content specifically for the agent's current question. A semantically important detail can therefore still be removed.
In the standard proxy/cache mode, Headroom tries not to rewrite the stable portion of an existing request and primarily works with newly added content. This helps preserve the prefix that a model provider may be able to cache. When the library or /v1/compress endpoint is called directly, the exact boundary depends on the integration parameters.

The data type determines how content is processed; important details are preserved heuristically, so quality must be validated.
Sources: ContentRouter and CacheAligner.
How to Integrate Headroom
The integration method depends on which part of the system you control.
|
Method |
When to use it |
What passes through Headroom |
Separate process required |
|
Python library |
You control the application code |
Only data passed to compress() |
No |
|
TypeScript SDK |
The application is written in TypeScript |
Messages passed to the SDK; the SDK communicates with a local proxy |
Yes |
|
Proxy |
The client allows you to change the base URL |
Supported requests and content explicitly routed through the proxy |
Yes |
|
MCP |
The agent supports the Model Context Protocol |
Explicit calls to headroom_compress, headroom_retrieve, and headroom_stats |
MCP process |
|
Agent wrapper |
You use a supported coding agent |
Session traffic launched through the wrapper |
Wrapper launches the proxy |
|
Framework adapter |
The application is built on a supported framework |
Context handled by the specific adapter |
Depends on the integration |
The simplest way to evaluate the tool alongside an existing application is to use a local proxy. To make the setup reproducible, pin the version:
uv tool install --python 3.13 "headroom-ai[proxy]==0.34.0" headroom --version
This command was tested in an isolated environment on August 11, 2026; the CLI returned headroom, version 0.34.0. The standard [proxy] installation does not include the tree-sitter dependencies required by CodeAware, so the extra [code] is needed for that use case—for example, headroom-ai[proxy,code]==0.34.0. In v0.34.0, the documentation, configuration classes, and executable CLI differ in their default CodeAware behavior, so when working with code, it is safer to explicitly pass --code-aware or --no-code-aware. The [all] package installs the full set of optional features, but it is usually unnecessary for an initial pilot.
The proxy can then be started with:
HEADROOM_BEACON=off headroom proxy --port 8787
In another terminal, point the application to the local endpoint. For example, for a client that supports the standard OpenAI environment variable:
OPENAI_BASE_URL=http://127.0.0.1:8787/v1 your-app
So the statement “no code changes required” is only true under one condition: the client already allows you to change the base URL through configuration. The routing path still changes:
application → local Headroom proxy → model provider
The wrapper automates this setup for supported coding agents. The command looks like headroom wrap <agent>, but behind the scenes it starts a local proxy and changes the session configuration. In v0.34.0, the wrapper can also connect Serena for code navigation; this can be disabled with --code-memory none. Persistent settings for supported tools can be reverted with headroom unwrap <tool>.
MCP works differently. The headroom mcp serve command provides the agent with explicit tools for compression, retrieving the original, and viewing statistics:
headroom mcp serve
Results from other MCP servers do not automatically pass through Headroom. The agent must explicitly call headroom_compress, or the host application must orchestrate such a call. Statistics are available through headroom_stats, while the stored original can be retrieved through headroom_retrieve if a CCR marker was created for the specific call.
A Practical Six-Step Pilot
There is no need to connect Headroom to all agent traffic at once. A more reliable approach is to start with a single, well-defined scenario.
- Find an expensive response. Choose a tool that regularly returns a large JSON response, log, HTML page, or document.
- Record the baseline. Measure the cost of the full task, execution time, number of model calls, and answer quality without Headroom.
- Choose an integration. A library works well for your own application, a proxy for an existing SDK, and explicit tools for an MCP agent.
- Define your original-data policy. Decide where originals will be stored, who can retrieve them, and how long they need to be retained.
- Repeat the same task. Keep the model, prompt, tools, and input data identical.
- Compare complete sessions. Compressing a single tool response does not by itself prove that the overall task became cheaper.
It is useful to define a success criterion before the pilot. For example: median cost should decrease, the task success rate should not noticeably decline, and the number of repeated reads and false conclusions should not increase.
How to Retrieve the Original via CCR
For scenarios where details are critical, Headroom offers CCR — Compress, Cache, Retrieve.
In an integration where CCR mode is enabled and a retrieval path is configured, the mechanism works somewhat like a temporary storage vault:
- Headroom compresses the content.
- The original is stored in the CCR store.
- The compressed version receives an identifier.
- When necessary, the agent calls headroom_retrieve to retrieve the original data.
This is useful as a safety net, but it is not a guarantee of “lossless compression.” Retrieval is performed using a hash and returns the complete stored original rather than a semantically searched result. It is only possible if the integration created a marker and retrieval tool, the record has not expired, the storage is available, and the agent understands that the compressed data is insufficient.
Technical note. By default, the main CCR store uses a local SQLite file at ~/.headroom/ccr_store.db; the path can be changed through HEADROOM_WORKSPACE_DIR or HEADROOM_CCR_SQLITE_PATH. A memory backend is also available, but it does not survive a process restart. Retention periods differ between the proxy, standalone MCP server, and some compressors, so there is no universal TTL. For an autonomous agent, the TTL should be set deliberately and cleanup should be verified. In v0.34.0, the standalone /v1/compress endpoint operates without a CCR marker and does not store the original by default. To enable a recoverable mode, you need config.mode="ccr", access to /v1/retrieve, and the headroom_retrieve tool on the calling application's side. A direct Python call to compress() does not add a retrieval tool automatically. In other words, calling compression alone is not enough: the entire retrieval path needs to be tested.
Every original-data retrieval adds latency and increases the context again. If the agent does this after almost every compression, the chosen profile is too aggressive, or the scenario is simply a poor fit for Headroom.
Headroom and Prompt Caching: Different Ways to Save
Prompt caching at the model provider reduces the cost of repeatedly processing an identical prefix. Headroom reduces the amount of context being sent. These mechanisms can work together, but the effectiveness of one does not guarantee the effectiveness of the other.
If the stable part of the request remains unchanged while Headroom compresses only the new suffix, caching can continue to work. If the transformed content becomes part of the prefix or looks different from one request to the next, the cache hit rate may decrease.
In Headroom v0.34.0, the CacheAligner component is disabled by default. When enabled, it only detects potentially unstable content and generates a warning. Despite its name and some older descriptions, it does not reorder or rewrite parts of the system prompt.
A practical calculation should account for all cost categories:
session cost = regular input + cache read + cache write + output + additional costs
For example, for the model selected in the protocol, gpt-5.6-luna, OpenAI separately charges for cache writes. According to the official pricing as of August 12, 2026, under the standard short-context tier, the rates are $0.20 per million regular input tokens, $0.02 per million cached input tokens, $0.25 per million cache-write tokens, and $1.20 per million output tokens. Prices should be checked again before running real-world tests.
Compression therefore should not be evaluated based solely on the input tokens column. It is important to track whether cache writes, original-data retrievals, or total task execution time have increased.

CCR provides a safety net against loss of details, while prompt caching reduces the cost of repeated input; the mechanisms are independent.
Sources: CCR Headroom and OpenAI Prompt Caching.
What Actually Stays Local
Compression and original-data storage can take place on the user's computer or within their own infrastructure. However, if the agent uses a cloud-based model, the compressed context is still sent to the provider. “Local processing” in Headroom does not mean that the entire session stays on the local machine.
Before deployment, four things should be checked:
- Dependencies and models. Packages are downloaded from a package registry, while Kompress and tokenizer artifacts may be required during setup or on the first run. In an isolated environment, they should be downloaded in advance.
- Original-data storage. You need clearly defined paths, access permissions, TTLs, cleanup procedures, and behavior after a restart.
- Network requests. In addition to the model provider, there may be initial artifact downloads and an external Kompress endpoint if one is configured.
- Telemetry. Local statistics and external data transmission are separate mechanisms.
The last point is particularly important. The v0.34.0 proxy documentation states that the external beacon has been removed, but the source code for the same tag contains a separate HEADROOM_BEACON mechanism that is enabled by default. It is supposed to send aggregate data without prompts, code, or file paths, but in sensitive environments it is reasonable to explicitly set HEADROOM_BEACON=off or DO_NOT_TRACK=1. The HEADROOM_OFFLINE=1 variable disables Headroom's auxiliary network requests, including the beacon, checks, and downloads, but it does not disable calls to the cloud model provider. Actual network traffic should still be verified.
Compression is not the same as anonymization. Personal, commercial, and other sensitive data still require separate rules for redaction, access, and storage.
When You Don't Need Headroom
An additional layer is not always justified. If possible, improve the data source first:
- Add server-side log filtering.
- Limit API fields.
- Use pagination.
- Remove duplicate metadata.
- Return structured errors.
- Reduce repeated reads.
- Improve RAG ranking.
- Use CodeGraph or another navigation tool to avoid reading unnecessary files.
These changes eliminate noise before it enters the context and are generally more predictable than semantic compression.
Headroom may not be worthwhile if tool responses are short, every detail is critical, the environment does not allow local processes, or compression takes roughly as long as the main request. Frequent calls to headroom_retrieve are another warning sign: the compressed representation is systematically insufficient.
Finally, Headroom does not fix a poorly designed agent. If an agent unnecessarily calls the same tool multiple times or repeatedly reads huge files in full, compression reduces the consequences but does not eliminate the underlying cause.
Conclusion
Headroom is best viewed as an additional layer for specific, expensive scenarios rather than a universal cost-saving switch.
A good starting point is a single repeatable workflow: log analysis, a large API response, a RAG workflow, or code research. For that workflow, establish a baseline, choose the appropriate integration, define rules for retrieving originals, and repeat the same task several times.
The key metric is not the maximum compression percentage for an individual JSON response, but the cost of successfully completing the task while maintaining quality. The greatest benefits typically come from combining several measures: precise retrieval of relevant data, well-designed tools, controlled compression, and properly measured caching.
Key Sources
- Headroom v0.34.0 Release
- README and Repository for v0.34.0
- Quickstart v0.34.0
- Proxy Documentation v0.34.0
- MCP Integration
- CCR: Storing and Retrieving the Original
- ContentRouter, SmartCrusher, CodeCompressor, and KompressCompressor
- CacheAligner Implementation and Configuration Values
- Beacon Telemetry and Aggregate Data
- Headroom Benchmark Documentation
- OpenAI: GPT-5.6 Luna, Prompt Caching, and Pricing