What this series covers. It assumes you have already chosen your model, and that the same model serves the whole conversation. It is about keeping a long agent conversation inside its context window without losing what the agent was told, and what that does to the prompt-cache bill; it is not a guide to choosing a model or to the cheapest way to run an agent.
If you only read one paragraph: as long as the conversation fits the model's window, leave it alone. The cache already serves about 94% of every prompt, and nothing we tried beat that. The exception is a model with a price line below its window, like gpt-6-luna, which doubles its rates on every call above 272K tokens: there, treat the line as the window and compact below it, which cut the bill by 38–46% in our test. Once the conversation outgrows the window, each of the built-in strategies either loses a good share of what the agent was told (between a quarter and four fifths of the facts in our test) or still ends up over the limit. The trick is to start compacting before you hit the real input limit, and we explain how to tell further down. Part 2 is about the strategies we built to do better.
Why compaction is tricky
An agent sends its whole conversation to the model on every call. Compaction trims that conversation so it still fits the window and each call carries fewer tokens. Simple enough, except that two things work against you.
The first is that prompt caching only works on an unchanged prefix. Providers cache the prompt from the first token forward, and if the start of this call's prompt is byte for byte the same as the last call's, that part is billed as a cache read, roughly a tenth of the normal price. On the models we tested that is $0.02 per million tokens instead of $0.20. The moment a strategy edits something, everything after that edit is back at full price.
The second is that compaction throws information away, and the agent can't use what it no longer has. The damage is uneven. Losing the turn where the user stated a requirement hurts far more than losing a chatty acknowledgement, and a strategy that keeps a value but drops the line saying what it belongs to has kept nothing useful at all.
So we judge every strategy on three things at once: does the prompt fit the window, does the agent still know what it was told, and what did the whole conversation cost once the cache is taken into account.
The conversation we test on
All the numbers here come from one scripted conversation, run through a real agent (the Agent Framework harness agent, with real tool calls and real replies). We shaped it like a typical retrieval-heavy session: lookups return large documents, only a few values inside them matter, and the facts the agent needs at the end are scattered across the whole history.
The window is simulated. The model's real limit is far larger, so every prompt actually goes through, and the benchmark simply disqualifies (DQ) any strategy whose prompt ever went past 120,000 tokens. That is why the uncompacted control is disqualified once the conversation is one and a half or three times the window. We still report its cost, because it tells you what a model with no size limit would charge.
The built-in strategies, and where they break
Agent Framework ships seven compaction strategies in agent_framework._compaction, and the harness agent installs one of them for you. Here is what each one did to our test conversation on gpt-5.6-luna, measured on agent-framework-core 1.20.0. Costs are per conversation, averaged over five seeds, with the summarizer's own calls included. DQ marks a cell where every seed went over the window.
| Strategy | What it does | Facts kept, of 53 conversation at 0.9× / 1.5× / 3× the window | Cost per conversation at 0.9× / 1.5× / 3× the window | Where it breaks |
|---|---|---|---|---|
| No compaction reference | Sends the whole conversation on every call. | 53 / 53 / 53 | $0.068 / $0.150 DQ / $0.443 DQ | Past the window no real model could run it. Its cost is what a model with no size limit would pay. |
| Context harness default | Works in two steps against a budget it derives from the window. Once the prompt reaches half of that budget it collapses the older tool results, and once it reaches 80% it drops the oldest groups. | 27 / 15 / 16 | $0.090 / $0.125 / $0.267 | Dropping the oldest groups removes the requirements and the correction first, because they are the oldest turns. |
| Truncation | Once the prompt passes a token threshold, it drops the oldest groups until the prompt is under a lower threshold. | 35 / 18 / 17 | $0.076 / $0.137 / $0.237 | The same problem: it drops the oldest groups whatever they contain, so the values survive while the turns that explain them are gone. |
| Sliding | Keeps only the most recent groups, a fixed number of them, and drops everything older. | 8 / 8 / 8 | $0.076 / $0.121 / $0.236 | It changes the start of the prompt on every call, so the cache hit rate is about 1%, the worst we saw. |
| Tool | Replaces each older tool result with a one-line summary and keeps the last four intact. | 50 / 47 / 43 | $0.095 / $0.178 DQ / $0.493 DQ | It never touches user text, so once that alone outgrows the window it goes over on every seed. |
| Selective | Removes the older tool calls together with their results and keeps the last four. | 50 / 47 / 44 | $0.100 / $0.195 DQ / $0.522 DQ | It has the same limit, and a removed result is gone for good. |
| Summarization | Asks an LLM to summarise the older groups into one summary message that links back to them. | 20 / 17 / 17 | $0.138 / $0.213 / $0.408 | A five-sentence summary can't hold 48 codes. This model kept 17–20 of 53; gpt-6-luna did better at 30–47. Re-summarising also rewrites the start of the prompt on every call, so the cache hit rate is about 1% and you pay the summarizer on top. Out of the box it is worse still: the default summarizer input budget of 8,000 tokens skips any tool result bigger than that, so it never compacted this conversation at all until we raised the budget to the window. |
| Token | Runs a list of other strategies in turn until the prompt is under a token budget, and if none of them gets there, drops the oldest groups. | 10–29 / 11–24 / 8–20 | $0.058–0.129 / $0.083–0.263 / $0.148–0.430 | It always fits the prompt and never keeps the facts: every ordering we tried ends in the oldest-first fallback. |
Four patterns explain all of it:
- Age is the wrong order to evict in. The oldest turns hold the task, the requirements and the constraints, so truncating them first does the most damage you could possibly do.
- Re-deciding every turn burns the cache. A rule like "over 80%, cut to 50%" rewrites the history a little differently each time it trips, so the cached prefix is lost from the head again and again. Truncation and the harness default manage 87–90% cache hits where not compacting manages 94–96%, and that difference is billed at ten times the price.
- Tool-only strategies can't shrink the rest. They keep facts well, but once user text alone passes the window they have nothing left to cut.
- Values without their labels are useless. Truncation left codes in the prompt whose labelling turn was gone, and the model used none of them.
And inside the window none of them pays off, because the cache already serves around 94% of the prompt and any edit costs more than it saves. The exception is a model that charges more above a certain prompt size, like gpt-6-luna above 272K tokens. The next section covers it.
Where not compacting breaks, and when to start
Not compacting is the right default inside the window. Here is when that stops being true. As the conversation grows, three things go wrong, roughly in this order.
- Cost does not grow in step with the conversation. Every call re-sends everything said so far, so a conversation twice as long makes twice as many calls, each over a prompt twice the size: roughly four times the tokens, not two. Cached tokens are cheap, but they are not free. On gpt-5.6-luna the uncompacted conversation cost $0.068 at 0.9 of the window, $0.150 at one and a half times it and $0.443 at three times it, with 94–98% of every prompt served from cache the whole way. A long-context price line bends the curve further: every token in a call above the line costs double, cached ones included. On gpt-6-luna the same conversation cost $0.313 at three times the window, because the calls above 272K tokens were billed at the long rates.
- Answers get worse long before the limit. We asked the agent the closing questions with nothing compacted, on a model that could hold the whole conversation. Inside the window it answered them all correctly. At one and a half times the window it got 97–100% right, and at three times only 84–90%. The composed strategy's compacted prompt scored 94–100% at that same size, with a quarter of the tokens. A long context is not just more expensive; past a certain size the model reads it less reliably.
- Then the call is refused. At the model's real input limit the request fails outright. In the benchmark, that is the disqualification you see wherever the conversation is one and a half or three times the size the model accepts. On a real deployment it is an error, and by then it is too late to recover anything that was never written down.
So the real question is when to start, and our answer is: before the conversation gets anywhere near the limit, with enough margin that the model can still read everything it is asked to summarise.
- It fits with room to spare. Don't compact. Every edit costs more in cache than it saves, and the best compacting strategies we measured landed within 8% of not compacting either way, which is inside the seed-to-seed noise.
- It will outgrow the real input limit. Compact, and trigger early. Our strategies ask the agent to write down what it learned at 60% of the budget, where the budget is the model's real input limit minus the reply you reserve, not the advertised window. A record written at 95% comes after the model has stopped reading the oldest lookups carefully, and one asked for past the limit never comes at all.
- There is a price line below the window. Treat the line as your window. On gpt-6-luna we triggered at 20% of a 1M window, about 200K, so no call ever crossed 272K, and the bill fell by 38–46% on a conversation that fitted the window comfortably.
- Recall over a long session matters to you. Compact even if the cost case is marginal. At 3.0 the compacted prompt answered better than the whole conversation did.
Which strategy to compact with, and how each one behaves, is the subject of Part 2. The settings behind "60% of the budget" are in Part 3.
The lab tool: cachebench
Every number in this article comes from cachebench, a benchmark we built on top of Agent Framework. We built it because the obvious measurements kept lying to us: fewer tokens didn't mean a lower bill, and a fact still sitting in the prompt didn't mean the agent had used it. It is open source as maf-cachebench, and its documentation covers every option and how to read the table it prints.
What it gives you
- Two harnesses.
cachebenchreplays byte-identical scripted conversations, comparable across providers.cachebench_livedrives a real agent, one model at a time. - All 20 strategies: the seven built-ins in their harness configurations, token-budget compositions, and ours, selectable by name.
- Any provider the framework supports: Foundry projects, Azure OpenAI resources (Chat Completions or Responses), OpenRouter, Mistral and Ollama. Pricing includes a cache-write premium for models that bill one.
- Records you can trust later. Every run appends one JSON line per strategy and seed, with its full configuration, so tables can be re-rendered and compared across runs without paying again.
--dry-runprints the plan, sizes and token estimate first. - Loud failures. A strategy that goes over the window is disqualified rather than averaged in, and throttling and connection retries are counted and reported.
Running it
# one extra per provider: openai, foundry, mistral, ollama
pip install "maf-cachebench[foundry]"
cachebench_live foundry:gpt-5.6-luna --agent harness \
--strategies none,tool_and_user_summary_anchored \
--summarizer-provider foundry:gpt-5.6-luna \
--fill 1.5 --context-window 120000 --repeats 5 \
--price-input 0.20 --price-cached 0.02 --price-output 1.20 \
--results-jsonl runs/my-cell.jsonl
# re-render later, or pool several runs into one comparison
cachebench_live --from-jsonl runs/
The grid behind these articles (six cells, 20 strategies, five seeds each) and the long-context runs are published with every record and the scripts that produced them, so any table here can be re-rendered with cachebench_live --from-jsonl without paying again.
Next: Part 2: Compaction that keeps the cache. The three rules we built our strategies on, how the composed strategy works, which strategy to use when, and what changes when a model charges more above 272K tokens. Part 3: Using the compaction strategies is the hands-on guide, with the settings and the code.
Further reading
- Part 2: Compaction that keeps the cache · this series
The strategies we built, how they win, and which one to use for which case. - Part 3: Using the compaction strategies · this series
What each strategy does, its settings and counters, and the code to attach it to an agent. - maf-compaction · PyPI
The package with the strategies from this series:pip install maf-compaction. - maf-cachebench · PyPI
The benchmark behind every number in this series:pip install maf-cachebench. - Running the benchmark · maf-extensions on GitHub
Every option, how to read the table it prints, and the flags that say a row is not measuring what its name claims. - Compaction · Microsoft Learn, Agent Framework documentation
The official reference for the built-in strategies in .NET, Python and Go, and how the harness agent wires them in. - Chat History Storage Patterns in Microsoft Agent Framework · Agent Framework blog, April 2026
Service-managed against client-managed history, and why compaction becomes your job in the second. - Managing Chat History for Large Language Models · Agent Framework blog, November 2024
The earlier Semantic Kernel take: message-count, token-limit and summarizing reducers. - Prompt caching · OpenAI API documentation
How prefix caching works and why summarization, compaction or truncation can reset it. - TokenPilot: Cache-Efficient Context Management for LLM Agents · arXiv, 2026
Research on the same tension: compaction that keeps prompt prefixes stable. - Micro-compaction: amortizing context compression in agent loops · DEV Community
Compaction in another agent framework, and an honest note that its cache cost was never priced. - Context compaction in agent frameworks · DEV Community, CrabTalk
A survey of how eight agent frameworks compact context.
Measured 5–6 October 2026 on agent-framework-core 1.20.0 (Foundry client 1.14.0, OpenAI client 1.15.0), gpt-5.6-luna and gpt-6-luna, Agent Framework harness agent, simulated 120K window, five seeds per cell. Costs are per conversation at list prices, summarizer calls included, closing questions excluded. gpt-5.6-luna's long-context surcharge above 272K input tokens is not modelled; it would only raise the uncompacted control's cost when the conversation is three times the window.
