Sets when an Agent compacts its conversation and when large tool results
are moved out of the model context. By default, the agent compacts before a
request would exceed about 32,000 tokens and saves tool results larger than
64 KiB to disk, leaving a short preview and a deputy://tool-result/...
reference in the context. The model can read the rest with the
deputy_read_tool_result tool, and you can with
Agent$resolve_tool_result().
Arguments
- max_tokens
Estimated context size, in tokens, that triggers compaction.
NULLturns automatic compaction off.- compact_to
After compaction, the recent turns that are kept take up about this fraction of
max_tokens. Must be between 0 and 1.- fallback
What to do if the model can't write the summary.
"error"(the default) stops with an error and leaves the conversation unchanged."text"uses a plain summary built from the start of each turn instead.- max_tool_result_bytes
Size, in bytes, above which a tool result is saved to disk and replaced in the model context by a preview and a reference. Compaction applies the same limit to tool results and tool call arguments in the turns it summarises.
NULLturns this off.- offload_dir
Directory for saved tool results; each session gets its own subdirectory. A relative path is resolved against the R working directory when the policy is created.
NULLuses Deputy's user cache directory.- summary_fallback_chats
A list of ellmer Chats to try, in order, if the agent's own Chat fails to write the summary during automatic compaction with a transient error. Each must have no turns or tools. They are only used for summaries and don't change the agent's Chat; manual
$compact()doesn't use them.- max_tool_result_image_bytes
Maximum total size of the images kept in one tool result's model context, 2 MiB by default. Inline images count their encoded size; images given by URL count only the URL. Images over the limit stay in the saved result.
NULLremoves the limit.- max_tool_result_images
Maximum number of images kept in one tool result's model context, 4 by default.
0moves all images out;NULLremoves the limit. Image limits are separate frommax_tool_result_bytes.- estimator
How to measure the context when the provider can't count tokens (some gateways answer the counting request with HTTP 404).
"auto"(the default) asks the provider first and otherwise estimates: the usage the provider reported for its latest response, plus an estimate of what was added since, including a longer system prompt or new tools. Without reported usage, everything is estimated. Estimates assume three bytes of text per token, a fixed amount per image and per document page, so they run high for typical text."provider"uses only the provider's count, so automatic compaction doesn't run when the provider can't count.
Details
Automatic compaction runs at the start of a run and between rounds of tool
calls. Summary requests count toward the run's UsageLimits. The summary is
appended to the system prompt and the model context keeps only the recent
turns; the removed turns stay available from Agent$get_turns() and in
saved sessions. If summarising fails or is cancelled, the conversation is
left as it was. Agent$last_compaction() describes the latest compaction,
including every summary attempt.
The "compaction_start" run event records estimate_source,
"provider" or "estimate". Usage reported before a compaction or a
$microcompact(), or saved in a session, isn't reused until the provider
reports usage for the smaller context. A token-counting request that fails
with HTTP 404, 405 or 501 isn't repeated for the same provider and base URL.
The policy is read-only: read fields with $, and create a new policy to
change one (the example shows how). The Chats in summary_fallback_chats
are copied, so later changes to your Chat objects don't affect the policy.
Saved sessions don't include the policy; $load_session() keeps the
loading agent's policy.
Examples
policy <- ContextPolicy(max_tokens = 16000, fallback = "text")
policy$max_tokens
#> [1] 16000
settings <- S7::props(policy)
settings$max_tokens <- 24000L
do.call(ContextPolicy, settings)
#> <ContextPolicy>
#> compact at: 24000 tokens
#> compact to: 50%
#> estimator: auto
#> fallback: text
#> summary fallback Chats: 0
#> offload above: 65536 bytes