Repository navigation
[Agent refactor] Context management: boundaries, compression, and token budgeting #1439
Description
Activity
- addedtype: enhancementNew feature or requestNew feature or request
on Mar 12, 2026 或许在此之前我们可以构建一个上下文的结构图,这可以让截断和压缩都更能正确的找到其边界。
Perhaps before this, we could construct a structural diagram of the context, which would allow both truncation and compression to more accurately identify their boundaries.
Good idea — the boundary rules should be visual before we touch code.
1. Context Window Regions
Everything sent to the LLM consumes a single context window budget. The regions stack in this order:
Loadinggraph TB A["<b>System Prompt</b> (fixed, cached)<br/>identity · rules · skills · memory · dynamic context · summary"] B["<b>Conversation History</b> (compressible)<br/>message groups 1 ... N — see diagram below"] C["<b>Current User Message</b>"] D["<b>Tool Definitions</b> (provider-injected)"] E["<b>Output Reserve</b> (max_tokens)"] A --> B --> C --> D --> E style A fill:#dbeafe,stroke:#2563eb style B fill:#fef3c7,stroke:#d97706 style C fill:#d1fae5,stroke:#059669 style D fill:#e0e7ff,stroke:#6366f1 style E fill:#fce4ec,stroke:#e11d48history_budget = context_window − system_tokens − current_msg_tokens − tool_def_tokens − max_tokensOnly the yellow region (conversation history) is compressible. Everything else is fixed overhead or reserved.
2. Message Groups and Safe Boundaries
In the session history, messages form atomic groups per conversation turn. The session stores all intermediate messages (
AddFullMessageinloop.go:1130for assistant+tool_calls,loop.go:1264for tool results).Three group types exist:
Loadinggraph TB subgraph Turn1["Turn 1 — Simple Exchange"] U1["user: hello"]:::user --> A1["assistant: hi"]:::asst end subgraph Turn2["Turn 2 — Tool Call"] U2["user: search X"]:::user --> A2["assistant + TC: search"]:::tc A2 -->|"atomic"| TR2["tool: result"]:::tool TR2 -->|"atomic"| AF2["assistant: found X"]:::asst end subgraph Turn3["Turn 3 — Chained Tool Calls"] U3["user: save and notify"]:::user --> A3["assistant + TC: save"]:::tc A3 -->|"atomic"| TR3["tool: saved"]:::tool TR3 -->|"atomic"| A3b["assistant + TC: notify"]:::tc A3b -->|"atomic"| TR3b["tool: sent"]:::tool TR3b -->|"atomic"| AF3["assistant: done"]:::asst end A1 ==>|"SAFE"| U2 AF2 ==>|"SAFE"| U3 AF3 ==>|"SAFE"| MORE["..."] classDef user fill:#dbeafe,stroke:#2563eb classDef asst fill:#d1fae5,stroke:#059669 classDef tc fill:#fef3c7,stroke:#d97706 classDef tool fill:#e5e7eb,stroke:#6b7280Boundary rule: a SAFE cut point (thick arrow) exists only after a final
assistant(notool_calls) and before the nextuser. Theatomicedges inside tool sequences must never be broken — splitting them creates orphanedtool_call/tool_resultpairs that trigger provider 400 errors (Anthropic, DeepSeek, Zhipu).To find the nearest safe cut point at a given index:
- Scan backward — look for a
usermessage whose predecessor is anassistantwithouttool_calls(or history start) - If none found, scan forward with the same rule
- Cut before the
usermessage at the boundary
3. Boundary-Aware Compression Flow
All three compression paths (proactive, emergency, async) should use the same safe-boundary detection:
Loadingflowchart TD START["Build messages for LLM"] --> EST["Estimate total tokens<br/>(system + history + current + tools)"] EST --> CHECK{"total exceeds<br/>context_window − max_tokens ?"} CHECK -->|"No"| CALL["Call LLM"] CHECK -->|"Yes"| PRO["Proactive: summarize oldest<br/>groups at safe boundary"] PRO --> REBUILD["Rebuild messages"] REBUILD --> EST CALL --> RESP{"Response?"} RESP -->|"Success"| SAVE["Save to session"] RESP -->|"Context error"| FORCE["forceCompression:<br/>truncate at safe boundary"] FORCE --> REBUILD2["Rebuild messages"] REBUILD2 --> CALL SAVE --> POST{"history_tokens<br/>above soft threshold ?"} POST -->|"No"| DONE["Done"] POST -->|"Yes"| ASYNC["Async: summarizeSession<br/>split at safe boundary"] ASYNC --> DONE style PRO fill:#fff3cd,stroke:#ffc107 style FORCE fill:#f8d7da,stroke:#dc3545 style ASYNC fill:#d1ecf1,stroke:#0dcaf0Key change from the current code: all three boxes (
forceCompression,summarizeSession, and the proposed proactive check) use the same boundary-finding logic to select cut points, instead of slicing at arbitrary indices likemidorlen-4.- Scan backward — look for a
所以我们需要四个上下文Builder吗?抽取一个公共的上下文Builder接口,为系统提示词、对话历史、用户消息、工具定义,四个区块创建各自的具体的Builder。还是全部使用同一个通用的Builder会更好?
So do we need four context Builders? Extract a common context Builder interface and create specific Builders for the four blocks: system prompt, conversation history, user message, and tool definitions. Or would it be better to use the same generic Builder for all?
I'd lean toward neither — the issue isn't how we build each region, but that we never budget for them. Let me explain.
Construction is already separated
Region Current code Where System Prompt BuildSystemPromptWithCache()+buildDynamicContext()context.go:493History GetHistory()→sanitizeHistoryForProvider()context.go:548User Message Message{Role:"user", Content:msg}context.go:563Tool Definitions ToolRegistry.ToProviderDefs()registry.go:261Tool definitions are already a separate parameter in
Chat(ctx, messages, toolDefs, model, opts)atloop.go:950— they never flow throughBuildMessages().Why 4 Builders adds complexity without proportional benefit
A common interface would need to unify
string,[]Message, and[]ToolDefinition— fundamentally different output types. And only history needs complex logic (compression, boundary detection, summarization). The other three regions are fixed-size pass-through data, meaning 3 out of 4 Builder implementations would be trivial wrappers — adding concepts without matching value.From the refactor guide: "if a behavior can be expressed without a new abstraction, do not add one".
The actual gap: no budget awareness before calling LLM
runAgentLoop (loop.go:757) ├─ BuildMessages(history, ...) → messages ← no size check └─ runLLMIteration (loop.go:876) ├─ ToProviderDefs() → toolDefs ← no size check └─ Chat(messages, toolDefs, max_tokens) ← send blind └─ context error → forceCompression ← reactive fix onlyWe assemble everything, send it to the LLM, and only compress after the provider rejects it with a 400. What we need is a proactive check.
Proposed: budget-aware functions, no new types
Each region's token cost is computed independently (addressing the "separate concerns" intuition), but as pure functions — not interface methods:
history_budget = context_window − estimateTokens(system_prompt) − estimateTokens(current_msg) − estimateToolDefsTokens(tool_defs) − max_tokensThe proactive check slots in between
BuildMessages()andChat():Loadingflowchart TD BUILD["BuildMessages + ToProviderDefs"] --> EST["Estimate each region's token cost"] EST --> CALC["history_budget =<br/>context_window minus fixed costs"] CALC --> CHECK{"history<br/>within budget?"} CHECK -->|"Yes"| CALL["Call LLM"] CHECK -->|"No"| COMP["Compress at safe boundary"] COMP --> REBUILD["Rebuild messages"] REBUILD --> CHECK CALL --> RESP{"Response?"} RESP -->|"OK"| SAVE["Save + async summarize"] RESP -->|"Context error"| FORCE["forceCompression fallback"] FORCE --> REBUILD2["Rebuild"] REBUILD2 --> CALL style COMP fill:#fff3cd,stroke:#ffc107 style FORCE fill:#f8d7da,stroke:#dc3545The proactive path (yellow) prevents most context errors before they happen. The emergency path (red) stays as a fallback for edge cases where estimation undershoots.
How this fits the refactor
- Makes the implicit constraint (context must fit) explicit and preventive — "if an existing boundary can be made explicit, do that first"
- Adds no new types or interfaces — "do not introduce a new concept unless it is strictly necessary"
- No restructuring of
ContextBuilderor the provider call path — purely additive
One prerequisite: the budget calculation needs a correct
context_windowvalue. Currentlyinstance.go:227setsContextWindow: maxTokens, which conflates the model's context window with the output generation limit. These should be separated (a config field with sensible default) before the budget check can be meaningful.Reacted by 美電球 and is-Xiaoen- added 4 commits that reference this issue
on Mar 13, 2026 Yeah, the core proposals (boundary separation, proactive budget check, Turn-based compression) are covered by #1490 now. Closing this one.
- added a commit that references this issue
on Apr 14, 2026
Context
This addresses track 6 of the agent refactor (#1216):
I've been working in this area through the session persistence track (#732, #1170) and spent time reading the compression and context-building code. Below is what I found, and a proposal for how to clarify these boundaries.
Current state
Context management is currently spread across three locations with implicit boundaries between them:
context.go:BuildMessages()— assembles system prompt (cached static + dynamic) + summary + history + current message into[]Message. RunssanitizeHistoryForProvider()to drop orphaned tool pairs at read time.loop.go:maybeSummarize()— checks two conditions after each turn:len(history) > SummarizeMessageThreshold(default 20) orestimateTokens(history) > ContextWindow * SummarizeTokenPercent / 100. If either is true, fires a background goroutine to runsummarizeSession().loop.go:forceCompression()— called reactively when the LLM returns a context-window error. Drops the oldest 50% of conversation messages, appends an emergency note to the system prompt.There is no explicit model of how much context space is available, what fills it, or when compression should happen relative to the actual budget.
Specific problems
1. ContextWindow defaults to MaxTokens
In
instance.go:227:MaxTokensis the max output tokens (default 32768 indefaults.go:33), passed to the LLM as themax_tokensrequest parameter (loop.go:930). ButContextWindowshould represent the model's input capacity — typically 128K+ for modern models.Setting
ContextWindow = maxTokensmeans:maybeSummarizethreshold =32768 * 75 / 100 = 24576estimated tokensmax_tokensto a large value, summarization never triggers at allPR #556 identified the same issue.
2. forceCompression can orphan tool pairs
forceCompression()slices conversation atmid = len(conversation) / 2(loop.go:1355) without checking whether the cut falls between an assistant message withToolCallsand its matchingtoolresult messages.The read-path defense (
sanitizeHistoryForProvideratcontext.go:577) catches orphaned pairs at query time, but the stored session history remains corrupted — tool messages without their matching assistant predecessor, or assistant messages with tool_calls but no results following. PR #665 identified this gap.3. Compression is reactive, not proactive
forceCompressiononly runs after the LLM already rejected the request with a context-window error (loop.go:1009-1027). This means:A proactive check before the LLM call would prevent this entirely for the common case.
4. Token estimation undercounts
estimateTokens()(loop.go:1691) only countsutf8.RuneCountInString(m.Content). It ignores:ToolCalls— function name + JSON arguments can be substantial (complex tool args easily add thousands of tokens)BuildMessages, not included in the history estimateThe summarization threshold check in
maybeSummarizecompares againstContextWindowusing this undercount, so the check is weaker than intended in both directions.Proposal
The goal is to clarify existing implicit boundaries, not introduce new abstractions. This follows the refactor's "minimum concepts" rule — no new types unless the current code cannot be clarified without them.
A. Separate context_window from max_tokens in config
Add
context_windowas an explicit field inAgentDefaults. Default to 0, meaning "fall back to a safe default" (e.g. 131072). This letsContextWindowandMaxTokensserve their actual distinct purposes:MaxTokens→ max output tokens per LLM callContextWindow→ total input capacity of the modelA follow-up improvement could auto-detect context window from the provider, but that's not needed for the initial fix.
B. Compute the available history budget explicitly
After building the system prompt, we know the fixed overhead. The available space for history becomes a simple subtraction:
This
historyBudgetreplaces the currentContextWindow * SummarizeTokenPercent / 100as the compression threshold. No new types — just making the arithmetic explicit and correct.C. Proactive pre-call check
Before calling the LLM in
runLLMIteration, estimate the total token cost of the assembledmessagesslice. If it exceedscontextWindow - reserve, run summarization before the call.forceCompressionstays as a last-resort fallback for cases where the estimate was too low. But it should stop being the primary compression path.D. Tool-pair-aware truncation
When
forceCompressionorsummarizeSessiontruncates history, ensure the cut point does not fall inside a[assistant+tool_calls, tool_result, ...]group. If it does, move the cut before the group start. This gives write-path protection to complement the existing read-path sanitization.E. Include ToolCalls in token estimation
Extend
estimateTokensto account form.ToolCalls— serialize function name and arguments into the character count. A rough estimate is still better than ignoring them.What this does NOT propose
ContextBudgetstruct, noContextManagerinterface)SessionStoreorBuildMessagesAPI signaturesThis is a boundary-clarification and correctness track, not a feature expansion.
Relationship to other tracks
Intentionally independent. Context management operates on
[]providers.Messageand integer token counts. It does not depend on the Agent abstraction (track 1), EventBus (track 3), persona assembly (track 4), or capability model (track 5). If the AgentLoop lifecycle changes (track 2/3), the call sites may shift, but the budget logic itself is unaffected.Related issues and PRs
I'd be happy to take this on if it fits the refactor direction. My plan would be to start with a working note in
docs/agent-refactor/context.mdcovering the boundary definitions, then follow up with implementation PRs targeting therefactor/agentbranch.