Recap of Part 1. Compaction edits the prompt, and every edit makes the cached text behind it billable at full price again, so while the conversation fits its window, leaving it alone is cheapest, unless the model has a price line inside that window. Once it outgrows the window, Agent Framework's built-in strategies either lose much of what the agent was told or still go over.
What this series covers. It assumes you have already chosen your model, and that the same model serves the whole conversation. It is about keeping a long agent conversation inside its context window without losing what the agent was told, and what that does to the prompt-cache bill; it is not a guide to choosing a model or to the cheapest way to run an agent.
If you only read one paragraph: once the conversation outgrows the window by half, one strategy we built kept every fact inside a 120K window on every seed of both models, and did it 22% cheaper than a model with an unlimited window on gpt-6-luna and 18–35% cheaper on four seeds out of five on gpt-5.6-luna. At three times the window it still held on every gpt-6-luna seed, at a third of the unlimited cost. And on a model that charges double above 272K tokens, compacting below that line cut the bill by 38–46%.
Three rules, four strategies
We built our strategies around three rules. Each one is a fix for a failure we measured in Part 1.
- Compact the same way every time. Decide what to shrink by where a message sits in the conversation, not by how full the window happens to be right now. That way every call sends the same bytes for the old part of the conversation and the cache keeps working.
- Compact each part only once. Once an old stretch of the conversation is shrunk, leave it alone. Later calls only touch newer messages, so the start of the prompt stays cached.
- Remove the least useful things first. Drop bulky tool output before anything else, and the user's own words last or never. And never delete the only copy of a fact. If nothing else holds it, keep it.
Anchored
anchored · anchored_min_gain · anchored_no_assistantKeeps a fixed head and tail word for word and collapses the band between them: tool results first, then tool calls, assistant narration last, user turns never. min_gain skips collapses too small to repay the cache they break (below 29% of the tokens after the edit).
Wins Inside the window: anchored_min_gain kept 53/53 and cost 6% less than not compacting on gpt-6-luna.
Loses Past it, it keeps every user turn and cannot fit.
Tool-summary anchored
tool_summary_anchoredAt 60% of the budget, the agent is made to call a special tool whose only job is to write down the facts from the tool results so far: the codes, the names, the numbers. We call that note the record. Later passes drop the results the record covers; a result the record missed is kept whole, never silently dropped. It repeats for new tool work by default.
Wins At 1.5× on gpt-6-luna: 53/53, 32% under an unlimited model.
Loses When the model writes an incomplete record, or when user text alone fills the window.
User-summary anchored
user_summary_anchoredSummarises the oldest band of user turns with a separate summarizer call, keeping the first and last user turns. Summaries stand once written and are not re-summarised, so the prefix stays stable.
Wins Keeps every fact.
Loses Alone, it leaves the tool results, so it disqualified on every seed past the window.
Tool-and-user summary anchored
tool_and_user_summary_anchoredThe two halves above, run in order on one shared trigger, with a last-resort chain for when both have run and the prompt is still over budget. The only strategy that held past the window.
Wins Everywhere past the window.
Watch An incomplete record at 3× can still push one seed over.
How the composed strategy works
The composed strategy runs before every model call and does as little as it can get away with: nothing below the trigger, tool results first, user text only if that wasn't enough, and heavier rewriting only when the prompt wouldn't fit at all.
Two design choices matter more than the rest. The record is written by the agent itself, in its own turn, so it is part of the cached conversation rather than an outside rewrite, and the agent gets to name the values it will need. And nothing is removed without a replacement. A tool result the record missed stays whole, and if that leaves no room the strategy stops over budget and says so rather than guessing. Every seed that was disqualified still had every fact in its prompt.
Which strategy to use
Pick by how far the conversation is going to outgrow the model's window. The numbers come from our final grid: all 20 strategies, a 120K window, five seeds each, on gpt-5.6-luna and gpt-6-luna, measured on agent-framework-core 1.20.0.
| Conversation vs window | Use | Cost per conversation gpt-5.6-luna / gpt-6-luna | Measured |
|---|---|---|---|
| Fits (up to ~0.9) | No compaction | $0.068 / $0.036 | The best strategies that kept all the facts landed within 8% of this on gpt-5.6-luna and within 3% on gpt-6-luna. The cache already serves ~94% of the prompt, so there is nothing to win. |
| Up to ~1.5× | tool_ | $0.143 / $0.058 unlimited: $0.150 / $0.075 | 53/53 on every seed of both models, never over the window. 22% under an unlimited model on gpt-6-luna. On gpt-5.6-luna it was 18–35% under on four seeds out of five, and 93% over on the one seed where the model's records left five lookups uncovered. The mean hides that, so we show both. |
| Up to ~1.5×, and the model's single record covers every lookup | tool_ | $0.308 / $0.051 | Cheapest on gpt-6-luna, 32% under an unlimited model. On gpt-5.6-luna the records were incomplete and 2 seeds out of 5 went over. |
| ~3× and beyond | tool_ | $0.19–0.46 / $0.10–0.11 unlimited: $0.443 / $0.313 | On gpt-6-luna, 53/53 inside the window on every seed at a third of an unlimited model's cost. On gpt-5.6-luna three seeds came in 30–56% under, one 13% over, and one was disqualified with every fact still kept (the cost range is the four that held). The gpt-6-luna unlimited figure includes its long-context rates above 272K. |
| Fits a 1M window, but crosses the model's long-context price line (gpt-6-luna: 272K) | tool_ | — / $0.209 and $0.236 no compaction: $0.384 | Both stayed under 272K on every call and came in 46% and 38% cheaper with all 53 facts kept. Details below. |
| Anything that must be recalled later | Avoid truncation, sliding window, token-budget and plain summarization | $0.083–0.263 / $0.045–0.132 at 1.5× the window | They fit, but past the window they keep only 8–24 of 53 facts. Plain summarization does better on gpt-6-luna (38 of 53 at 1.5×) and not on gpt-5.6-luna (17), costs 20% more than not compacting, and gets about 1% of its prompt from the cache. |
gpt-6-luna costs about half of gpt-5.6-luna on every row, even with its cache-write premium charged, and ranks the strategies the same way. Both models also answer less reliably from a whole 360K conversation (84–90% on the first answer) than from the composed strategy's compacted one (94–100%), so compaction can help recall, not just cost.
The long-context price line
Some models charge more for big requests. On gpt-6-luna a request with more than 272,000 input tokens is billed entirely at the long-context rates, not just the part over the line, and every token in it costs more, cached ones included.
| gpt-6-luna, per 1M tokens | Input | Cached | Cache write | Output |
|---|---|---|---|---|
| Request up to 272K tokens | $0.10 | $0.01 | $0.125 | $0.50 |
| Request over 272K tokens | $0.20 | $0.02 | $0.25 | $0.75 |
This changes the usual answer. Inside a window, compaction normally doesn't pay, because breaking the cache costs more than the smaller prompt saves. But once every call above the line costs double, keeping the prompt under it is worth a few cache breaks.
We tested it with a 1M window and a conversation of about 400K tokens, which fits the window but sits well over the line. The strategies were set to start compacting at 20% of the window, about 200K tokens.
| gpt-6-luna, ~400K conversation | Cost per conversation | Largest prompt | Input billed at long rates | Facts kept | First answer |
|---|---|---|---|---|---|
| No compaction | $0.384 | 385K | 54% | 53 / 53 | 83% |
| tool_ | $0.236 −38% | 204K | 0% | 53 / 53 | 92% |
| tool_ | $0.209 −46% | 211K | 0% | 53 / 53 | 96% |
- Both strategies stayed under the line on every call. Every one of their five seeds cost less than every seed without compaction.
- Most of the saving comes from the price line. Without the surcharge, not compacting would have cost $0.271 and the strategies would only have saved 13–23%.
- Answers got better too. The compacted prompts scored 92–96% on the first answer, against 83% from the full conversation.
Our rule of thumb: if your model has a price line, treat the line as your window, and set the trigger comfortably below it so no call ever crosses.
Next: Part 3: Using the compaction strategies. What each strategy does to the conversation, the settings and counters, and the code to attach each one to an agent with the maf-compaction package. How these were measured: Part 1: Compaction and the prompt cache covers the test conversation, the built-in strategies these are compared against, and the cachebench tool.
Further reading
- Part 1: Compaction and the prompt cache · this series
Why compaction is tricky, the built-in strategies and where they break, and the cachebench tool. - Part 3: Using the compaction strategies · this series
What each strategy does, its settings and counters, and the code to attach it to an agent. - maf-compaction · PyPI
The package with the strategies from this series:pip install maf-compaction. - The strategies, documented · maf-extensions on GitHub
Each strategy's mechanism, defaults and failure modes, with the source and tests beside it. - maf-cachebench · PyPI
The benchmark behind every number in this series:pip install maf-cachebench. - The records behind these numbers · maf-extensions on GitHub
One JSON line per strategy and seed for the grid and the long-context runs, with the scripts that produced them. - Compaction · Microsoft Learn, Agent Framework documentation
The official reference for the built-in strategies in .NET, Python and Go, and how the harness agent wires them in. - Chat History Storage Patterns in Microsoft Agent Framework · Agent Framework blog, April 2026
Service-managed against client-managed history, and why compaction becomes your job in the second. - Managing Chat History for Large Language Models · Agent Framework blog, November 2024
The earlier Semantic Kernel take: message-count, token-limit and summarizing reducers. - Prompt caching · OpenAI API documentation
How prefix caching works and why summarization, compaction or truncation can reset it. - TokenPilot: Cache-Efficient Context Management for LLM Agents · arXiv, 2026
Research on the same tension: compaction that keeps prompt prefixes stable. - Micro-compaction: amortizing context compression in agent loops · DEV Community
Compaction in another agent framework, and an honest note that its cache cost was never priced. - Context compaction in agent frameworks · DEV Community, CrabTalk
A survey of how eight agent frameworks compact context.
The strategy comparison was measured 5–6 October 2026 on agent-framework-core 1.20.0, gpt-5.6-luna and gpt-6-luna, Agent Framework harness agent, simulated 120K window, five seeds per cell. The long-context section was measured 4 October 2026 on core 1.16.0, gpt-6-luna, 1M window, the surcharge above 272K input tokens priced per request; the strategies' own code is the same in both. Costs are per conversation at list prices. The records are published with maf-cachebench.
