This article was written with the help of AI.

AGENT FRAMEWORK · COMPACTION · PART 3 OF 3

Using the compaction strategies

A hands-on guide to the strategies from Part 2: what each one does to the conversation, the settings that matter, the counters to read afterwards, and the code to attach each one to a Microsoft Agent Framework agent.

Recap of Part 1 and Part 2. Every compaction edit costs prompt cache, and past the window the built-in strategies lose facts or still go over. Our strategies decide by position and drop tool output before the user's words, and the composed one kept every fact at 1.5 times the window for less than an unlimited window would cost.

What this series covers. It assumes you have already chosen your model, and that the same model serves the whole conversation. It is about keeping a long agent conversation inside its context window without losing what the agent was told, and what that does to the prompt-cache bill; it is not a guide to choosing a model or to the cheapest way to run an agent.

If you only read one paragraph: every strategy is a plain object that takes a token budget and a tokenizer. You hand the same object to the agent twice, once for the pass before each model call and once for the pass after each run. The record-based strategies also need a tool and a middleware, and the user-summary one needs a second chat client to write the summaries. Everything else is a setting, and every setting has a default we measured.

Before you start

We built these strategies to address the shortcomings of what Microsoft Agent Framework offers out of the box. They ship as maf-compaction, an open-source package in our maf-extensions repository. It depends on Agent Framework's core package alone, and it is pinned to one minor release of it, because it builds on a private compaction module of the framework. It is experimental: importing it emits a MafCompactionExperimentalWarning, and releases before 1.0 may change the API.

pip install maf-compaction
# the examples below also use these two
pip install tiktoken agent-framework-openai

Three things are the same for every strategy.

A tokenizer. Every decision is sized in tokens: the budget, the triggers, how much of a tool result to keep. The strategies accept anything with a count_tokens(text) method, and you want exact counts for the model you actually run. The framework's CharacterEstimatorTokenizer works, but it guesses from character counts, and on our workload that guess was off by a factor of two.

import tiktoken

class Tokenizer:
    """Exact token counts. The strategies size every decision with this."""

    def __init__(self, encoding: str = "o200k_base") -> None:
        self._encoding = tiktoken.get_encoding(encoding)

    def count_tokens(self, text: str) -> int:
        return len(self._encoding.encode(text))

tokenizer = Tokenizer()

A budget. max_input_tokens is the ceiling the prompt has to stay under. Set it to the model's real input limit minus the output you reserve, not to the advertised context window. On GPT-5-class deployments the two differ by 128,000 tokens, and a budget set to the bigger number puts every trigger above what the service will accept. We learned that one the hard way.

CONTEXT_WINDOW = 128_000   # the model's real input limit
MAX_OUTPUT = 4_000         # reserved for the reply
BUDGET = CONTEXT_WINDOW - MAX_OUTPUT

Two phases, one object. Agent Framework runs compaction at two points. Before each model call, the agent's compaction_strategy runs inside the client, on the messages about to be sent. After each run, a CompactionProvider runs on the stored history. Give the same strategy object to both. The easiest way is the harness agent, which takes a strategy for each phase and wires the history provider for you:

from agent_framework import create_harness_agent
from agent_framework_openai import OpenAIChatClient

client = OpenAIChatClient(model_id="gpt-6-luna")

agent = create_harness_agent(
    client,
    name="assistant",
    agent_instructions="Answer from the lookups you make. Quote values exactly.",
    tools=[lookup_deployment],
    max_context_window_tokens=CONTEXT_WINDOW,
    max_output_tokens=MAX_OUTPUT,
    before_compaction_strategy=strategy,
    after_compaction_strategy=strategy,
    tokenizer=tokenizer,
)

session = agent.create_session()
reply = await agent.run("Which deployment serves region EU-2?", session=session)

With a plain Agent you wire the same thing by hand. Two details matter. The provider's before_strategy stays None, because the before phase travels as the agent's own compaction_strategy. And require_per_service_call_history_persistence is on, so the history is written after every model call inside a tool loop, not only at the end of the run. The record-based strategies depend on that: their record is a tool result, and it has to be in the stored history before the next pass can act on it.

from agent_framework import Agent, CompactionProvider, InMemoryHistoryProvider

history = InMemoryHistoryProvider()
agent = Agent(
    client=client,
    name="assistant",
    instructions="You answer from the lookups you make. Quote values exactly.",
    tools=[lookup_deployment],
    context_providers=[
        history,
        CompactionProvider(
            before_strategy=None,
            after_strategy=strategy,
            tokenizer=tokenizer,
            history_source_id=history.source_id,
        ),
    ],
    compaction_strategy=strategy,
    require_per_service_call_history_persistence=True,
)

Every example below builds strategy. The agent code stays the same unless a section says otherwise.

Anchored

AnchoredCompactionStrategy keeps a fixed number of message groups at the start and at the end of the conversation word for word. Everything between them is the band. In the band it shortens tool results, oldest first, to a size that depends only on where the result sits. Only if shortening is not enough does it remove whole groups, and assistant narration goes last. A message group is one user turn, one assistant reply, or one tool call with its result.

Head, band and tail in the anchored strategy A conversation drawn as a row of message groups. The first three groups are the head and are kept word for word. The last four are the tail and are also kept word for word. The groups between them form the band, where each tool result is shortened to a share of the budget divided by its position in the band, so the oldest keeps the most and each later one keeps less. User turns in the band are never touched. Head · keep_head_groups = 3 system user reply Band · shortened by position tool 1 keeps ½ user tool 2 keeps ⅓ reply tool 3 keeps ¼ user tool 4 keeps ⅕ Tail · keep_tail_groups = 4 user reply tool 5 user The share is the result's own position, so it never changes as the conversation grows keep = band_share × max_input_tokens ÷ (position + 1) band_share = 0.25 by default Kept word for word: the head, the tail, every user turn, and anything another strategy marked as preserved. Removed only when shortening was not enough: whole tool groups oldest first, then assistant replies, each replaced by a one-line marker. A shortened result is never shortened again: the band is re-read on every pass, and a second trim would break the cache for nothing.
Figure 1. The anchored layout. The oldest tool result in the band keeps the biggest slice and each later one a smaller slice, because the share is read off the result's own position. The same conversation compacts to the same bytes on every later call, which is what keeps the cached prefix intact.

Here is the strategy with its defaults spelled out.

from maf_compaction import AnchoredCompactionStrategy

strategy = AnchoredCompactionStrategy(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,
    keep_head_groups=3,            # task and requirements, never touched
    keep_tail_groups=4,            # the working set, never touched
    band_share=0.25,               # oldest banded result keeps 25% of the budget
    keep_tokens=None,              # or a fixed token count per shortened result
    collapse_assistant_text=True,  # allow dropping narration as the last resort
)
SettingDefaultWhat it does
keep_head_groups3Groups at the start kept word for word. The system prompt and the first exchange carry the task; every deleting strategy we measured threw these away first.
keep_tail_groups4Recent groups kept word for word. Too small and the model loses the thread; too large and every new turn shifts a big block and re-bills it.
band_share0.25Share of the budget the oldest banded tool result may keep. Lower it to shorten harder. It cannot rescue a payload of small results: if results are already under their allowance, the strategy does nothing, and that is correct.
keep_tokensNoneA fixed number of tokens to keep per shortened result, split between its head and its tail. Overrides the position rule.
collapse_assistant_textTrueWhether assistant replies in the band may be dropped when shedding tool groups was not enough. Set False when the model restates tool values in its replies, since those replies may be the only surviving copy.

Use it when tool results are large and the conversation stays near the window. It is cheap to run, needs no extra model calls, and kept all 53 planted facts at a 272K window. Don't expect it to hold a conversation that outgrows the window, though. It keeps every user turn, so once user text alone fills the budget it stops with the prompt over the limit rather than delete what the user said.

Minimum-gain anchored

Every edit costs something even when it removes nothing useful. The provider re-reads everything from the edit to the end of the prompt at full price on the next call. MinimumGainAnchoredCompactionStrategy is the anchored strategy with one addition: before it changes anything, it adds up what the collapse would save and compares that with the tokens behind the earliest edit. If the saving is under min_gain_fraction of those tokens, it declines and leaves the conversation as it found it.

The default, 0.29, is the break-even for a model whose cached tokens cost a tenth of fresh ones and a conversation with about twenty turns left. Fewer turns ahead means fewer calls to repay the edit, so raise the fraction for short conversations and lower it for long ones. The floor does not apply when the prompt is already over the budget; there, shortening is what keeps the conversation sendable at all.

from maf_compaction import (
    MinimumGainAnchoredCompactionStrategy,
)

strategy = MinimumGainAnchoredCompactionStrategy(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,
    min_gain_fraction=0.29,   # decline a collapse saving under 29% of what follows
)

# after a run: how many collapses it refused
print(strategy.declined_collapses)

Use it instead of plain anchored whenever the conversation fits the window. It was the best in-window result we measured, 53 of 53 facts at 6% under not compacting on gpt-6-luna. It doesn't make compaction cheaper than no compaction; what it does is stop you paying for edits that can't pay for themselves.

Tool-summary anchored

ToolResultAnchoredSummarizationCompactionStrategy is the only one that carries information forward instead of discarding it. It makes the agent write down what it learned from its tool results, in its own turn, and then drops the results that record covers. The record is a tool result, so the provider issued it, the cache keeps it, and the strategies that shed assistant prose never touch it.

It takes four pieces working together, and you need all four. The strategy decides what to drop. The tool is what the model calls to write the record, and it just echoes the text straight back. The gate keeps that tool inert until it is asked for. The middleware is what does the asking: it watches the prompt size on the way out of every call, and when the prompt passes the trigger it pins the next call to the tool.

How a record is asked for, written and used Four steps. One: the middleware sees the prompt over sixty percent of the budget after a call, arms the gate, and pins the next call to the recall tool. Two: the model calls the tool and writes the record; the tool returns the text, and the record lands in history as a tool result. Three: on the next pass, the strategy drops every tool group in front of the record whose distinctive values the record quotes, at least eighty percent of them; an uncovered group is kept whole and the strategy asks for one more record for it. Four: new tool work after the record re-arms the trigger. If no record ever arrives by ninety percent of the budget, the anchored fallback shortens the band instead. 1 · The middleware asks After a call, the prompt is over trigger_fraction (60%) of the budget. The middleware arms the gate and pins the next call to the recall tool. 2 · The model writes the record The tool's description is the whole prompt: quote every value that could not be guessed, aim at 2,000 tokens. The tool echoes the text back as a tool result. 3 · The strategy drops what the record covers Next pass: a tool group in front of the record is dropped when the record quotes 80% of its distinctive values. The record itself is preserved from every strategy. 4 · Repeat for new tool work New tool groups after the newest record re-arm the trigger, so later batches get a record of their own. On by default (repeat_records=True). Uncovered: kept whole, asked again A group the record did not quote is never dropped. It is marked preserved, and the strategy asks for one more record aimed at it. No record by fallback_fraction (90%) The model may never call the tool. Past this line the anchored fallback shortens the band instead, and the run counts a fallback. The middleware sends no message of its own: a message appended by middleware would be stored and replayed to the user, so the tool's description carries the whole instruction. The record arrives one call late by construction, which is why the trigger sits well below the fallback line.
Figure 2. The record handshake. The decision to ask is made on the way out of one call and acted on in the next, and the drop happens on the pass after that. Nothing is removed without a replacement the agent wrote itself.
from maf_compaction import (
    AnchoredCompactionStrategy,
    RecallGate,
    ToolResultAnchoredSummarizationCompactionStrategy,
    ToolResultRecallMiddleware,
    make_recall_tool,
)

strategy = ToolResultAnchoredSummarizationCompactionStrategy(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,
    keep_head_groups=3,
    keep_tail_groups=4,
    trigger_fraction=0.6,      # ask for a record past 60% of the budget
    fallback_fraction=0.9,     # give up waiting for one past 90%
    coverage_share=0.8,        # covered when 80% of a group's values are quoted
    fallback=AnchoredCompactionStrategy(
        max_input_tokens=BUDGET, tokenizer=tokenizer
    ),
)

gate = RecallGate()
recall_tool = make_recall_tool(gate, target_tokens=2_000)  # length to aim at

recall = ToolResultRecallMiddleware(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,                       # the same tokenizer as the strategy
    arm=gate.arm,
    trigger_fraction=strategy.trigger_fraction,
    record_max_tokens=4_000,                   # hard cap on the pinned call's reply
    repeat_records=True,                       # a record per new batch of tool work
    reforce=strategy.take_reforce,             # may ask again for uncovered groups
)

agent = create_harness_agent(
    client,
    name="assistant",
    agent_instructions="Answer from the lookups you make. Quote values exactly.",
    tools=[lookup_deployment, recall_tool],    # registered like any other tool
    middleware=[recall],
    max_context_window_tokens=CONTEXT_WINDOW,
    max_output_tokens=MAX_OUTPUT,
    before_compaction_strategy=strategy,
    after_compaction_strategy=strategy,
    tokenizer=tokenizer,
)
SettingDefaultWhat it does
trigger_fraction0.6Share of the budget at which the middleware pins a call to the recall tool. Early on purpose: the record degrades with the amount of tool output the model must read, and an edit repays itself over the turns that follow it.
fallback_fraction0.9Share of the budget past which the strategy stops waiting for a record and hands the conversation to the anchored fallback. Must be above the trigger.
coverage_share0.8Share of a tool group's distinctive values (tokens of four or more characters with a digit) the record must quote before the group may be dropped. 1.0 keeps a whole group over one reformatted value; much below 0.5 lets two quoted values delete six.
fallbackanchoredThe strategy run when no record arrives, or when a record did not free enough. Build it yourself so its band_share matches what you measured; the default takes the anchored defaults.
target_tokens (tool)2,000Length stated in the tool's description. The only bound that makes the model plan to fit.
record_max_tokens (middleware)4,000Cap on the pinned call's reply. A cap truncates rather than shortens, and a truncated tool call loses its arguments, so keep it well above the target. Truncations are counted.
repeat_records (middleware)TrueAsk again once there is new tool work after the newest record. Off, the strategy records once and every later result sits in the prompt uncompacted; at three times the window that disqualified every seed. Turn it off only for a conversation that ends soon after it first outgrows the window.
max_groups_before_record (middleware)NoneForce a record every N tool groups regardless of size. Use it for models whose records cover only part of what they read; asking for less per record is what buys coverage back.

After a run, the counters say what happened. On the strategy: records_in_conversation (how many records stand), groups_kept_uncovered (tool groups a record failed to quote, kept whole), fallbacks_used (no record ever came) and fallbacks_after_record (a record came but did not free enough). On the middleware: forced_calls, records_forced, records_volunteered (the model called the tool uninvited; the gate let it answer but recorded nothing) and records_truncated.

Three things go wrong in practice. Pinning the model's tool_choice anywhere else breaks it, because the model has to be free to call the recall tool on the call the middleware pins. The strategy and the middleware must share one tokenizer and one budget, or they will disagree about when a record is due. And a model that writes incomplete records compacts almost nothing rather than lose facts. That is the failure we want, but it shows up as a prompt that keeps growing, so read groups_kept_uncovered before blaming the strategy, and set max_groups_before_record if it is high.

Use it when tool output is the bulk of the conversation and the model writes good records. It was the cheapest strategy we measured at 1.5 times the window on gpt-6-luna, 32% under an unlimited model with every fact kept, and the cheapest way to stay under a long-context price line. It can't help when user text alone fills the window; that is what the next two are for.

User-summary anchored

UserTurnAnchoredSummarizationCompactionStrategy works on the other half of the conversation. Past its trigger, it takes the user turns between a fixed head and a fixed tail, sends them to a summarizer, and puts the result back in their place as one user message. The replaced turns are linked to the summary and excluded, using the framework's own replace-and-link mechanism, so a history store can still show what was summarised. It never reads a tool call, a tool result or an assistant reply.

It needs a chat client of its own to write the summary. Its output replaces the user's words and is trusted from then on like any other message, so point it only at a service you trust as much as the main model.

from maf_compaction import (
    SUMMARY_MODE_RECOMPACT,
    UserTurnAnchoredSummarizationCompactionStrategy,
)

summarizer = OpenAIChatClient(model_id="gpt-6-luna")  # trusted; may be the same

strategy = UserTurnAnchoredSummarizationCompactionStrategy(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,
    client=summarizer,
    keep_head_user_turns=1,         # the task, kept word for word
    keep_tail_user_turns=1,         # the live request, kept word for word
    trigger_fraction=0.8,           # later than the record strategy's 0.6
    min_band_share=0.1,             # only when the band is worth 10% of the prompt
    summary_mode=SUMMARY_MODE_RECOMPACT,
)

# after a run
print(strategy.user_compactions, strategy.user_messages_replaced)
print(strategy.user_passes_declined)
SettingDefaultWhat it does
keep_head_user_turns1User turns at the start kept word for word. The first turn carries the task; truncation was measured leaving 29 facts in a prompt the model could no longer use because it had lost them.
keep_tail_user_turns1User turns at the end kept word for word. A model answering a summary of the question it was just asked answers a different question.
trigger_fraction0.8Share of the budget before anything happens. Later than the record strategy's 0.6 because this edit sits just behind the head and breaks nearly the whole cached prefix.
min_band_share0.1The band must be worth this share of the prompt before a pass runs. Without it the strategy fired once per turn for the rest of a run: 31 passes on one seed, each rewriting the prefix to free almost nothing. At 0.1 a pass happens about once per 11% of prompt growth.
summary_moderecompactrecompact re-reads its own previous summary and replaces it with one covering everything. boundary leaves earlier summaries standing and compacts only what is newer, so the prefix up to the newest summary never changes. fold is boundary mode that collapses the standing summaries into one once they are worth the break.
prompt, fold_promptbuilt inWhat the summarizer is asked for. The default asks for a much shorter third-person account that keeps every requirement, correction and preference and quotes identifiers, codes, paths and numbers verbatim.

Use it when the conversation is mostly the user's own text and the exact wording of old turns doesn't matter. On its own it won't hold a tool-heavy conversation, because it can't touch the tool results; on our test it was disqualified on every seed past the window. It exists to be composed.

The composed strategy

ToolResultAndUserTurnAnchoredSummarizationCompactionStrategy runs the two summarising strategies over one conversation, the record half first. It owns no selection rule of its own; what it adds is an order and a last resort. Both halves are judged against one line, the record half's trigger, and the user half acts only if the record half was not enough, because its edit is the one that breaks the cache. While a record has been asked for and may still arrive, the user half waits for it. Only while the prompt is over the budget with both halves spent does the chain run: merge the records, fold the user summaries, rewrite the record harder, the anchored fallback, and finally a loud failure with every fact still in the prompt. Part 2 draws the flow.

Build it from the two halves. The user half runs in boundary mode, so a summary stands once written, and remembers two summarizer requests rather than one, because a pass over the budget may ask for a band and then a fold.

from maf_compaction import (
    SUMMARY_MODE_BOUNDARY,
    ToolResultAndUserTurnAnchoredSummarizationCompactionStrategy,
)

record_half = ToolResultAnchoredSummarizationCompactionStrategy(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,
    trigger_fraction=0.6,
    fallback=AnchoredCompactionStrategy(
        max_input_tokens=BUDGET, tokenizer=tokenizer
    ),
)
user_half = UserTurnAnchoredSummarizationCompactionStrategy(
    max_input_tokens=BUDGET,            # must equal the record half's
    tokenizer=tokenizer,
    client=summarizer,
    summary_mode=SUMMARY_MODE_BOUNDARY,
    remembered_requests=2,
)
strategy = ToolResultAndUserTurnAnchoredSummarizationCompactionStrategy(
    tokenizer=tokenizer,
    tool_results=record_half,
    user_turns=user_half,
    harder_attempts=2,          # rewrite-harder attempts in the chain
    chain_gain_fraction=0.29,   # once started, remove this share behind its edit
)

gate = RecallGate()
recall_tool = make_recall_tool(gate)
recall = ToolResultRecallMiddleware(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,
    arm=gate.arm,
    trigger_fraction=record_half.trigger_fraction,
    repeat_records=True,
    reforce=record_half.take_reforce,   # the half inside the composition
)

The agent is wired exactly as for the record strategy: recall_tool in tools, recall in middleware, and strategy for both compaction phases. Notice that the middleware is wired to the record half inside the composition. A composition is not an instance of its parts, so code that checks isinstance(strategy, ToolResultAnchoredSummarizationCompactionStrategy) before installing the middleware installs nothing, and the row then runs with no record and no error. We made exactly that mistake. Look for the nested half instead; the package exports the lookup it uses itself:

from maf_compaction import find_nested_strategy

# the strategy itself, or the record half inside a composition; None if there is none
record_half = find_nested_strategy(
    strategy, ToolResultAnchoredSummarizationCompactionStrategy
)
SettingDefaultWhat it does
user_trigger_fractionthe record half'sThe line the user half is judged at. Defaults to the record half's trigger so the two cannot drift apart; pass a value only if you want two lines.
harder_attempts2How many times the chain may ask the summarizer to rewrite the standing record shorter, each attempt asking for more compression. A rewrite is kept only if it is smaller; a refused rewrite is not asked again.
chain_gain_fraction0.29Hysteresis. Once the chain starts, it keeps going until it has removed this share of the tokens behind its first edit, so the next turn does not push straight back over the budget. Zero stops it exactly at the budget. This setting halved the composed strategy's cost in the long test.
merge_promptbuilt inWhat the summarizer is asked for when merging records into one.

Every counter from both halves is readable on the composed object, plus one per chain step: records_merged, user_summaries_merged, record_rewrites, last_resort_fallbacks, and user_passes_waited for the passes the user half held while a record was due. A row where last_resort_fallbacks is zero never needed the chain; a row where it is not tells you the first three steps were not enough.

Use it for any conversation that is going to outgrow the window and whose facts you will need later. It was the only strategy that held past the window on every seed of both models at 1.5 times the window; at three times it held on every seed on gpt-6-luna and on four out of five on gpt-5.6-luna. It pays both halves' costs, an agent turn for the record and a summarizer call for the summary, so inside the window it is not cheaper than doing nothing, unless a price line sits inside that window.

Which one, and the mistakes to avoid

  • The conversation fits the window. Do not compact, or use minimum-gain anchored if you want a guard against growth. Nothing beats the cache here, unless the model has a price line inside the window, which is the last case below.
  • Tool output is the bulk and it will outgrow the window. Tool-summary anchored. Check groups_kept_uncovered on your model first; if it stays high, bound the ask with max_groups_before_record or move to the composed strategy.
  • It will outgrow the window and you cannot predict the shape. The composed strategy.
  • The model charges more above a size line. Treat the line as your window: set max_input_tokens below it and the record trigger comfortably under that. On gpt-6-luna's 272K line that cut the bill by 38% to 46%.

And the mistakes we made, so you don't have to. For each one: what we did, what it caused, and what to do instead.

  • A budget set to the advertised window.
    What we did. We set max_input_tokens to the context window on the model card.
    What it caused. The model refuses a prompt that leaves no room for its answer, so the real limit is that window minus the output you let it generate. Our prompts were under budget and over the limit at the same time, and every trigger, being a fraction of the budget, fired late.
    Do instead. Set the budget to the real input limit minus your output reservation, and let every trigger derive from that number.
  • Different tokenizers in different places.
    What we did. The strategy, the middleware and the agent each counted tokens their own way.
    What it caused. They disagreed about how full the prompt was, so the record was asked for at the wrong moment: too early by one count, too late by another.
    Do instead. Build one tokenizer and pass that same object to all three.
  • Only one compaction phase wired.
    What we did. We wired the strategy into the before phase only.
    What it caused. The before phase compacts what is sent and the after phase compacts what is stored, so the stored history kept growing uncompacted. A record the agent wrote inside a tool loop never landed in that history either, so the next call had no record to lean on.
    Do instead. Give the same strategy object to both phases, and turn on per-service-call history persistence.
  • Installing the recall middleware by isinstance.
    What we did. We installed the middleware only when the strategy was an instance of the record strategy.
    What it caused. A composed strategy is not an instance of its record half, so the check failed and the middleware never went in. Nothing asked the agent to write records, and a strategy that refuses to drop an unrecorded group then dropped nothing.
    Do instead. Find the record half inside the composition with find_nested_strategy and install the middleware for that.
  • Repeats off on a long conversation.
    What we did. We turned record repeats off, so the agent wrote one record and never another.
    What it caused. One record covers only what existed when it was written. Everything after it was never recorded and, because the strategy refuses to drop an unrecorded group, never shortened either, so the prompt grew until it went over.
    Do instead. Leave repeat_records on, which is the default, so new tool work after the newest record gets a record of its own.
  • A summarizer we had not checked on the user's words.
    What we did. We gave the user half a summarizer without reading what it made of the user's turns.
    What it caused. Its output replaces those turns as one user message, so whatever it dropped or reworded became what the user had said for the rest of the conversation.
    Do instead. Use a summarizer you would trust to restate the requirements, and read a few of its summaries against the turns they replaced before relying on it.

Further reading

  • Part 1: Compaction and the prompt cache · this series
    Why compaction is tricky, the built-in strategies and where they break, and the cachebench tool.
  • Part 2: Compaction that keeps the cache · this series
    The three rules, the composed strategy's decision flow, and which strategy wins where.
  • maf-compaction · PyPI
    The package with the strategies from this series: pip install maf-compaction.
  • The strategies, documented · maf-extensions on GitHub
    Each strategy's mechanism, defaults and failure modes, with the source and tests beside it.
  • maf-cachebench · PyPI
    The benchmark behind every number in this series: pip install maf-cachebench.
  • Compaction · Microsoft Learn, Agent Framework documentation
    The official reference for the built-in strategies, the compaction provider, and how the harness agent wires the two phases.
  • Chat History Storage Patterns in Microsoft Agent Framework · Agent Framework blog, April 2026
    Service-managed against client-managed history, and why compaction becomes your job in the second.
  • Prompt caching · OpenAI API documentation
    How prefix caching works and why an edit resets it from that point on.

Settings and defaults as of maf-compaction 0.1.0, measured against agent-framework-core 1.20.0. The strategies are experimental and depend on a private compaction module of Agent Framework; re-read them against the version you install.

What is context costing your agents?

We measure compaction, caching and recall on your own workloads, and tune agent context for cost and accuracy.