This article was written with the help of AI.

AGENT FRAMEWORK · COMPACTION · PART 2 OF 3

Compaction that keeps the cache

The compaction strategies we built for Microsoft Agent Framework: the rules behind them, how the composed strategy works, where each one wins, and which to use for which case.

Recap of Part 1. Compaction edits the prompt, and every edit makes the cached text behind it billable at full price again, so while the conversation fits its window, leaving it alone is cheapest, unless the model has a price line inside that window. Once it outgrows the window, Agent Framework's built-in strategies either lose much of what the agent was told or still go over.

What this series covers. It assumes you have already chosen your model, and that the same model serves the whole conversation. It is about keeping a long agent conversation inside its context window without losing what the agent was told, and what that does to the prompt-cache bill; it is not a guide to choosing a model or to the cheapest way to run an agent.

If you only read one paragraph: once the conversation outgrows the window by half, one strategy we built kept every fact inside a 120K window on every seed of both models, and did it 22% cheaper than a model with an unlimited window on gpt-6-luna and 18–35% cheaper on four seeds out of five on gpt-5.6-luna. At three times the window it still held on every gpt-6-luna seed, at a third of the unlimited cost. And on a model that charges double above 272K tokens, compacting below that line cut the bill by 38–46%.

Three rules, four strategies

We built our strategies around three rules. Each one is a fix for a failure we measured in Part 1.

  • Compact the same way every time. Decide what to shrink by where a message sits in the conversation, not by how full the window happens to be right now. That way every call sends the same bytes for the old part of the conversation and the cache keeps working.
  • Compact each part only once. Once an old stretch of the conversation is shrunk, leave it alone. Later calls only touch newer messages, so the start of the prompt stays cached.
  • Remove the least useful things first. Drop bulky tool output before anything else, and the user's own words last or never. And never delete the only copy of a fact. If nothing else holds it, keep it.

Anchored

anchored · anchored_min_gain · anchored_no_assistant

Keeps a fixed head and tail word for word and collapses the band between them: tool results first, then tool calls, assistant narration last, user turns never. min_gain skips collapses too small to repay the cache they break (below 29% of the tokens after the edit).

Wins Inside the window: anchored_min_gain kept 53/53 and cost 6% less than not compacting on gpt-6-luna.

Loses Past it, it keeps every user turn and cannot fit.

Tool-summary anchored

tool_summary_anchored

At 60% of the budget, the agent is made to call a special tool whose only job is to write down the facts from the tool results so far: the codes, the names, the numbers. We call that note the record. Later passes drop the results the record covers; a result the record missed is kept whole, never silently dropped. It repeats for new tool work by default.

Wins At 1.5× on gpt-6-luna: 53/53, 32% under an unlimited model.

Loses When the model writes an incomplete record, or when user text alone fills the window.

User-summary anchored

user_summary_anchored

Summarises the oldest band of user turns with a separate summarizer call, keeping the first and last user turns. Summaries stand once written and are not re-summarised, so the prefix stays stable.

Wins Keeps every fact.

Loses Alone, it leaves the tool results, so it disqualified on every seed past the window.

Tool-and-user summary anchored

tool_and_user_summary_anchored

The two halves above, run in order on one shared trigger, with a last-resort chain for when both have run and the prompt is still over budget. The only strategy that held past the window.

Wins Everywhere past the window.

Watch An incomplete record at 3× can still push one seed over.

How the composed strategy works

The composed strategy runs before every model call and does as little as it can get away with: nothing below the trigger, tool results first, user text only if that wasn't enough, and heavier rewriting only when the prompt wouldn't fit at all.

Decision flow of the composed strategy Before each call: if the prompt is under the trigger, send it unchanged. Otherwise the tool half records new tool results and drops what the record covers. If still over the trigger, the user half summarises the oldest band of user turns. If the prompt is over the input budget, a last-resort chain runs: merge records, merge user summaries, rewrite records harder, anchored fallback, and finally a loud disqualification. Before each model call Over the trigger? 60% of budget no Send unchanged: cache intact yes 1 · Tool half: record, then drop New tool results no record covers? The agent is made to call the record tool and write a compact record of them. The next pass drops the results the record covers; a result it missed stays whole (never silently lost). Still over trigger? after the record no Send yes 2 · User half: summarise the oldest band A summarizer condenses the oldest user turns; the first and last stay word for word. Earlier summaries are kept as they are, not re-summarised. Over the budget? window − reply no Send yes 3 · Last-resort chain stops, and sends, once under its target a · Merge records all records into one; old ones retired b · Merge user summaries the same, for the user half c · Rewrite records harder up to 2 attempts, asking for more compression; kept only if smaller. A refused rewrite is not asked again. d · Anchored fallback sheds narration, oldest first e · Still over: fail loudly disqualified, every fact still there Hysteresis Once started, the chain works down to a target below the budget: it removes 29% of the tokens behind its first edit, so the next turn doesn't push straight back over. Its decisions are kept for the whole run.
Figure 1. The composed strategy's decision flow. Each step touches the prompt only when the one before it wasn't enough, so most calls go out unchanged and keep their cache. In our long test (a 30K window, with the conversation 6.5 times the window) two details, not re-asking for rewrites the model had refused and a bit of hysteresis, cut the chain's cost in half, from $0.71 to $0.36 per conversation with all 197 facts kept.

Two design choices matter more than the rest. The record is written by the agent itself, in its own turn, so it is part of the cached conversation rather than an outside rewrite, and the agent gets to name the values it will need. And nothing is removed without a replacement. A tool result the record missed stays whole, and if that leaves no room the strategy stops over budget and says so rather than guessing. Every seed that was disqualified still had every fact in its prompt.

Which strategy to use

Pick by how far the conversation is going to outgrow the model's window. The numbers come from our final grid: all 20 strategies, a 120K window, five seeds each, on gpt-5.6-luna and gpt-6-luna, measured on agent-framework-core 1.20.0.

Conversation vs windowUseCost per conversation
gpt-5.6-luna / gpt-6-luna
Measured
Fits (up to ~0.9)No compaction$0.068 / $0.036The best strategies that kept all the facts landed within 8% of this on gpt-5.6-luna and within 3% on gpt-6-luna. The cache already serves ~94% of the prompt, so there is nothing to win.
Up to ~1.5×tool_and_user_summary_anchored$0.143 / $0.058
unlimited: $0.150 / $0.075
53/53 on every seed of both models, never over the window. 22% under an unlimited model on gpt-6-luna. On gpt-5.6-luna it was 18–35% under on four seeds out of five, and 93% over on the one seed where the model's records left five lookups uncovered. The mean hides that, so we show both.
Up to ~1.5×, and the model's single record covers every lookuptool_summary_anchored$0.308 / $0.051Cheapest on gpt-6-luna, 32% under an unlimited model. On gpt-5.6-luna the records were incomplete and 2 seeds out of 5 went over.
~3× and beyondtool_and_user_summary_anchored$0.19–0.46 / $0.10–0.11
unlimited: $0.443 / $0.313
On gpt-6-luna, 53/53 inside the window on every seed at a third of an unlimited model's cost. On gpt-5.6-luna three seeds came in 30–56% under, one 13% over, and one was disqualified with every fact still kept (the cost range is the four that held). The gpt-6-luna unlimited figure includes its long-context rates above 272K.
Fits a 1M window, but crosses the model's long-context price line (gpt-6-luna: 272K)tool_summary_anchored or tool_and_user_summary_anchored, triggered below the line— / $0.209 and $0.236
no compaction: $0.384
Both stayed under 272K on every call and came in 46% and 38% cheaper with all 53 facts kept. Details below.
Anything that must be recalled laterAvoid truncation, sliding window, token-budget and plain summarization$0.083–0.263 / $0.045–0.132
at 1.5× the window
They fit, but past the window they keep only 8–24 of 53 facts. Plain summarization does better on gpt-6-luna (38 of 53 at 1.5×) and not on gpt-5.6-luna (17), costs 20% more than not compacting, and gets about 1% of its prompt from the cache.

gpt-6-luna costs about half of gpt-5.6-luna on every row, even with its cache-write premium charged, and ranks the strategies the same way. Both models also answer less reliably from a whole 360K conversation (84–90% on the first answer) than from the composed strategy's compacted one (94–100%), so compaction can help recall, not just cost.

The long-context price line

Some models charge more for big requests. On gpt-6-luna a request with more than 272,000 input tokens is billed entirely at the long-context rates, not just the part over the line, and every token in it costs more, cached ones included.

gpt-6-luna, per 1M tokensInputCachedCache writeOutput
Request up to 272K tokens$0.10$0.01$0.125$0.50
Request over 272K tokens$0.20$0.02$0.25$0.75

This changes the usual answer. Inside a window, compaction normally doesn't pay, because breaking the cache costs more than the smaller prompt saves. But once every call above the line costs double, keeping the prompt under it is worth a few cache breaks.

We tested it with a 1M window and a conversation of about 400K tokens, which fits the window but sits well over the line. The strategies were set to start compacting at 20% of the window, about 200K tokens.

gpt-6-luna, ~400K conversationCost per conversationLargest promptInput billed at long ratesFacts keptFirst answer
No compaction$0.384385K54%53 / 5383%
tool_and_user_summary_anchored$0.236 −38%204K0%53 / 5392%
tool_summary_anchored$0.209 −46%211K0%53 / 5396%
  • Both strategies stayed under the line on every call. Every one of their five seeds cost less than every seed without compaction.
  • Most of the saving comes from the price line. Without the surcharge, not compacting would have cost $0.271 and the strategies would only have saved 13–23%.
  • Answers got better too. The compacted prompts scored 92–96% on the first answer, against 83% from the full conversation.

Our rule of thumb: if your model has a price line, treat the line as your window, and set the trigger comfortably below it so no call ever crosses.

Further reading

The strategy comparison was measured 5–6 October 2026 on agent-framework-core 1.20.0, gpt-5.6-luna and gpt-6-luna, Agent Framework harness agent, simulated 120K window, five seeds per cell. The long-context section was measured 4 October 2026 on core 1.16.0, gpt-6-luna, 1M window, the surcharge above 272K input tokens priced per request; the strategies' own code is the same in both. Costs are per conversation at list prices. The records are published with maf-cachebench.

What is context costing your agents?

We measure compaction, caching and recall on your own workloads, and tune agent context for cost and accuracy.