Skip to content

Percentage-of-context-window budgets scale the wrong way: ToolSearch preload never engages, and the always-on context warning never fires, on large windows #12029

Description

@yiliang114

Part of #12028.

Two budgets in the codebase are expressed as a percentage of the model's context window. Both guard against costs that have nothing to do with the window size, so both stop working exactly as windows get large — and the project is moving towards large windows.

1. tools.toolSearch.threshold (default 10% of the window)

packages/core/src/core/client.ts:1766-1772 computes budget = floor(contextWindow * percent / 100), and packages/core/src/tools/tool-registry.ts:1073-1079 reveals the entire deferred set if it fits (all-or-nothing by design).

docs/design/toolsearch-preload-threshold.md explains the tradeoff being made, and it is a real one:

"every mid-session reveal rewrites the function-declaration list, which sits at the front of the tools→system→messages prefix, so a single ToolSearch load invalidates the entire prompt KV cache. For a small deferred set the deferral saves little and the cache damage plus the extra ToolSearch round-trip make it a net loss."

So the gate is asking: is the deferred set big enough that hiding it beats one prefix rebuild? Neither quantity depends on the context window. The window only appears because the 10% figure was carried over from Claude Code's ENABLE_TOOL_SEARCH=auto, where the same number is used under a different cost model (the doc says as much at lines 65-70).

The doc never discusses what the percentage means at 256K or 1M. The only sizing guard is an upper clamp at 100% (client.ts:1758-1764).

Observed on a 1M-context model: budget = 100,000 tokens, deferred pool ≈ 4.2k, so every deferred tool is revealed at session start. Ten tools that are shouldDefer: true upstream were all present in the request:

cron_create 1,103 · web_fetch 693 · web_search 499 · send_message 498 · monitor 475 · create_sub_session 409 · read_mcp_resource 191 · cron_delete 116 · task_stop 102 · cron_list 80 — 4,166 tokens, 19.4% of the built-in tool block, on every request of every session.

Deferral is not merely weakened here; at this window size it can never engage.

2. MEMORY_CONTEXT_WARNING_RATIO (15% of the window)

packages/core/src/config/config.ts:322, warning at :4702-4727:

Warning: Loaded always-on context (QWEN.md context files + auto-memory) uses about ${estimatedTokens} tokens, more than 15% of this model's ${contextWindowSize} token context window.

At 1M that warning fires at 150,000 tokens. The same session carried 15,400 tokens of context files — 33% of all its non-conversation context — and printed nothing. The warning is silent precisely where always-on context is cheapest to accumulate and most expensive in aggregate.

Suggestion

Express both as an absolute token budget, or as min(percent-of-window, absolute cap) so small-window behaviour is preserved. For the preload gate the natural cap is a small multiple of what one prefix rebuild costs, which is a function of the prefix size — not the window.

Worth stating explicitly in the docs either way: threshold: 0 is not a free win. It trades a fixed per-request cost (the deferred schemas) for an occasional prefix rebuild, so it pays off only when the deferred tools are rarely used in a session. Today that tradeoff is invisible to anyone who has not read the design doc.

中文说明

属于 #12028。

代码里有两个预算是按"模型上下文窗口的百分比"表达的。它们防范的成本都与窗口大小无关,因此窗口越大越失效——而项目正在转向大窗口。

1. tools.toolSearch.threshold(默认窗口的 10%)

packages/core/src/core/client.ts:1766-1772 计算 budget = floor(contextWindow * percent / 100),packages/core/src/tools/tool-registry.ts:1073-1079 在装得下时把整个延迟集合一次性揭示(全有全无,属于设计)。

docs/design/toolsearch-preload-threshold.md 解释了这个权衡,而且它是真实存在的:

每次会话中途的揭示都会重写函数声明列表,而它位于 tools→system→messages 前缀的最前面,因此一次 ToolSearch 加载会让整段 prompt KV 缓存失效。对于较小的延迟集合,延迟本身省不下多少,缓存损失加上额外一次 ToolSearch 往返反而是净亏。

也就是说这个门限在问:延迟集合是否大到足以让"隐藏它"胜过"一次前缀重建"? 这两个量都不依赖上下文窗口。窗口之所以出现在公式里,是因为 10% 这个数字来自 Claude Code 的 ENABLE_TOOL_SEARCH=auto,而那边的成本模型不同(设计文档 65-70 行自己说明了这一点)。

文档从未讨论过 256K 或 1M 下这个百分比意味着什么。唯一的尺寸保护是 100% 的上限(client.ts:1758-1764)。

在 1M 窗口模型上实测:预算 = 100,000 token,延迟池约 4.2k,于是所有延迟工具在会话开始时被全部揭示。十个上游标记为 shouldDefer: true 的工具全部出现在请求中:

cron_create 1,103 · web_fetch 693 · web_search 499 · send_message 498 · monitor 475 · create_sub_session 409 · read_mcp_resource 191 · cron_delete 116 · task_stop 102 · cron_list 80 —— 4,166 token,占内置工具块的 19.4%,每个会话的每一轮都在付。

在这个窗口尺寸下,延迟加载不是被削弱,而是根本无法生效。

2. MEMORY_CONTEXT_WARNING_RATIO(窗口的 15%)

packages/core/src/config/config.ts:322,告警在 :4702-4727。1M 窗口下这条告警要到 150,000 token 才触发。同一个会话携带了 15,400 token 的上下文文件——占其全部非对话上下文的 33%——却什么都没提示。常驻上下文最容易悄悄堆积、累计代价最高的场景,恰恰是告警沉默的场景。

建议

两者都改成绝对 token 预算,或 min(窗口百分比, 绝对上限),以保留小窗口下的行为。对预加载门限而言,自然的上限是"一次前缀重建的代价"的某个倍数,而那是前缀大小的函数,不是窗口的函数。

无论如何,文档里都应写清楚:threshold: 0 不是白捡的收益。它用"偶尔一次前缀重建"换掉"每轮固定的延迟 schema 成本",只有在会话很少用到延迟工具时才划算。今天这个权衡对没读过设计文档的人是不可见的。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    category/coreCore engine and logicmodel/long-contextpriority/P2Medium - Moderately impactful, noticeable problemroadmap/context-performanceRoadmap: Context and performancescope/memoryMemory and context managementscope/settingsSettings and preferencesscope/token-managementToken handling and limitsstatus/ready-for-humanSpecified but requires human judgment to implement; not suitable for an autonomous agenttype/bugSomething isn't working as expected

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions