You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(core): the eager tool surface is a hand-maintained static list — let something choose it, without invalidating the prompt prefix #12326
Part of #12028. Complements #12029 (which bounds the preload budget); this is about who chooses the resident set in the first place. No dependency on open work — the measurement ruler it needs landed in #12119.
What happens today
tools.eager is the one knob that materially shrinks the resident tool surface: in the #12028 sample it takes built-in tool schemas from 21,461 to ~5,868 tokens, more than every upstream code change on this umbrella put together. But it is a hand-written list, and it is static:
The allowlist is snapshotted once in PermissionManager.initialize() (permissions/permission-manager.ts:353), and a tool's registration status is decided when it is registered (config.ts → registerLazyTool). The setting is requiresRestart.
A demoted tool is not disabled: it stays registered, listed in /tools, and loadable through tool_search.
So the mechanism already in the codebase is static starting set + model-driven top-up. The dynamic half exists. What is missing is that nothing chooses the starting set — every deployment enumerates tool names by hand, or (much more commonly) never sets the key at all and pays for the whole surface on every request.
Why this cannot simply become per-turn
Function declarations sit at the very front of the prompt prefix, so changing the declared set mid-session rewrites the prefix and invalidates its KV cache. From the #12028 sample (prefix ≈ 42k tokens, implicit cache at 20% of full price):
Deferred pool
Turns before preloading is the cheaper choice
4.2k (bundled built-ins only)
~40
8k
~21
20k (MCP-heavy)
~8
Demotion saves on the order of ¥0.01 per cached turn, while a single mid-session reveal costs ¥0.28–0.40. That is why the design deliberately decides once at session start and thereafter only adds — see docs/design/toolsearch-preload-threshold.md.
The contrast worth noting: a skill's paths: activation is turn-level, because the skill listing lives in history[0] rather than in the tools block, so changing it does not break the tools → system prefix (the reasoning is in the comment at core/environmentContext.ts). The same "adapt to the session" idea is cheap one layer up and expensive here. This is a placement constraint, not a missing capability.
What could choose the set
Four options, cheapest first. None of them changes the declared set mid-session.
Option
Cost
Cache impact
Notes
A. Project-scope tools.eager
zero code — already possible
none
Still a hand-written list, but one per project instead of one globally. Worth documenting rather than building.
B. Named presets (e.g. tools.eagerProfile: "file-work")
small
none
#7625's fork profiles already ship named tool-restriction presets as .qwen/fork-profiles/<name>.md; reuse that shape rather than inventing a second file format. Turns "enumerate 40 tool names" into "pick one word".
C. Usage-driven, computed at session start
medium
none
Adaptive across sessions, static within one — which is exactly the shape the cache constraint allows. ToolCallEvent.function_name already exists (telemetry/types.ts:212), so per-tool call counts need no new instrumentation.
D. First-turn intent classification
medium
none (nothing is cached before the first request)
The only point where "pick a set from intent" is safe. A wrong guess costs one tool_search round trip at best, and at worst a tool the model never thinks to look for — a silent quality loss, not an error.
Recommendation: C, with B as the manual override. C puts the adaptation between sessions instead of inside one, so it gets the benefit without touching the prefix-cache constraint at all.
Questions to settle before implementation
Where usage state lives. A project-scope file under .qwen/? It has to survive /clear and be per-project, since the right set differs by codebase.
Cold start. The first session in a project has no history. Fall back to the current behaviour (everything eager), or to a preset?
Starvation, and it is the serious one. A tool that is never offered is never called, so a usage-driven set can lock itself in: absence of calls is not evidence the tool was not needed. Options: a floor set that is always eager, periodic re-offering, or counting tool_search reveals as usage (a reveal is direct evidence the model wanted something it was not given).
How a wrong set is observed.tool_search calls per session is the natural signal — the tracking(core): non-conversation context token governance #12028 analysis put the break-even at ≤2 reveals per session. Anything above that means the computed set is wrong, and it should be visible in telemetry rather than inferred from complaints.
Explicitly not proposed
Mid-session promotion or demotion of declared tools. The cache arithmetic above says no.
Per-turn classification after the first turn, for the same reason.
Idle cost (the input tokens of a session that asks one question and calls no tool) falls to the same order as a hand-written allowlist achieves, without a deployment writing one.
tool_search calls per session stay at or below ~2 on a representative task set; above that the saving is cancelled by prefix rebuilds.
No drop in tool-call success rate or task completion on that task set. This needs the evaluation harness the umbrella still lacks — worth stating plainly, because without it the starvation risk in question 3 is unfalsifiable.
Part of #12028. Complements #12029 (which bounds the preload budget); this is about who chooses the resident set in the first place. No dependency on open work — the measurement ruler it needs landed in #12119.
What happens today
tools.eageris the one knob that materially shrinks the resident tool surface: in the #12028 sample it takes built-in tool schemas from 21,461 to ~5,868 tokens, more than every upstream code change on this umbrella put together. But it is a hand-written list, and it is static:PermissionManager.initialize()(permissions/permission-manager.ts:353), and a tool's registration status is decided when it is registered (config.ts→registerLazyTool). The setting isrequiresRestart./tools, and loadable throughtool_search.So the mechanism already in the codebase is static starting set + model-driven top-up. The dynamic half exists. What is missing is that nothing chooses the starting set — every deployment enumerates tool names by hand, or (much more commonly) never sets the key at all and pays for the whole surface on every request.
Why this cannot simply become per-turn
Function declarations sit at the very front of the prompt prefix, so changing the declared set mid-session rewrites the prefix and invalidates its KV cache. From the #12028 sample (prefix ≈ 42k tokens, implicit cache at 20% of full price):
Demotion saves on the order of ¥0.01 per cached turn, while a single mid-session reveal costs ¥0.28–0.40. That is why the design deliberately decides once at session start and thereafter only adds — see
docs/design/toolsearch-preload-threshold.md.The contrast worth noting: a skill's
paths:activation is turn-level, because the skill listing lives inhistory[0]rather than in the tools block, so changing it does not break the tools → system prefix (the reasoning is in the comment atcore/environmentContext.ts). The same "adapt to the session" idea is cheap one layer up and expensive here. This is a placement constraint, not a missing capability.What could choose the set
Four options, cheapest first. None of them changes the declared set mid-session.
tools.eagertools.eagerProfile: "file-work").qwen/fork-profiles/<name>.md; reuse that shape rather than inventing a second file format. Turns "enumerate 40 tool names" into "pick one word".ToolCallEvent.function_namealready exists (telemetry/types.ts:212), so per-tool call counts need no new instrumentation.tool_searchround trip at best, and at worst a tool the model never thinks to look for — a silent quality loss, not an error.Recommendation: C, with B as the manual override. C puts the adaptation between sessions instead of inside one, so it gets the benefit without touching the prefix-cache constraint at all.
Questions to settle before implementation
.qwen/? It has to survive/clearand be per-project, since the right set differs by codebase.tool_searchreveals as usage (a reveal is direct evidence the model wanted something it was not given).maxPreloadTokens, or is that cap only about preloading deferred tools?tool_searchcalls per session is the natural signal — the tracking(core): non-conversation context token governance #12028 analysis put the break-even at ≤2 reveals per session. Anything above that means the computed set is wrong, and it should be visible in telemetry rather than inferred from complaints.Explicitly not proposed
Acceptance
tool_searchcalls per session stay at or below ~2 on a representative task set; above that the saving is cancelled by prefix rebuilds.中文说明
属于 #12028。与 #12029 互补(那个管的是预加载预算的上限),本单管的是常驻集合最初由谁来选。不依赖任何在途工作——它需要的度量尺已随 #12119 落地。
现状
tools.eager是唯一能实质缩小常驻工具面的开关:在 #12028 的样本里,它把内置工具 schema 从 21,461 降到约 5,868 token,比这条伞 issue 下所有上游代码改动加起来还多。但它是一份手写名单,而且是静态的:PermissionManager.initialize()里快照一次(permissions/permission-manager.ts:353),工具的注册状态在注册时就定了(config.ts→registerLazyTool)。该设置requiresRestart。/tools、仍可通过tool_search加载。也就是说代码里已有的机制是静态起点集合 + 模型驱动的按需补齐,动态那一半是有的。缺的是没有任何东西来选这个起点——每个部署要么手工枚举工具名,要么(更常见)根本不设这个键,于是每一次请求都为整个工具面付费。
为什么不能直接做成"每轮动态"
函数声明位于 prompt 前缀的最前面,会话中途修改声明集合会重写前缀并作废其 KV 缓存。按 #12028 样本(前缀约 42k token,隐式缓存为全价 20%):
降级每轮(命中缓存)省下的量级是 ¥0.01,而一次会话中途的揭示要付 ¥0.28–0.40。这正是该设计刻意"会话开始定一次、之后只增不减"的原因——见
docs/design/toolsearch-preload-threshold.md。值得对照的一点:skill 的
paths:门控确实是 turn-level 的,因为 skill 清单位于history[0]而不在 tools 段,改它不会打断 tools → system 的前缀(理由写在core/environmentContext.ts的注释里)。同一个"随会话自适应"的想法,在上一层很便宜,在这一层很贵。这是位置带来的约束,不是能力缺失。可以由谁来选这个集合
四个选项,由便宜到贵。它们都不在会话中途改变声明集合。
tools.eagertools.eagerProfile: "file-work").qwen/fork-profiles/<name>.md的形式提供了命名工具限制预设;复用那个形状,不要再造一种文件格式。把"枚举 40 个工具名"变成"选一个词"。ToolCallEvent.function_name已经存在(telemetry/types.ts:212),按工具的调用次数不需要新埋点。tool_search往返,坏的情况是模型压根想不到去找某个工具——静默的质量损失,不是报错。建议:选 C,以 B 作为手动覆盖。 C 把自适应放在会话之间而不是会话之内,因此完全不触碰前缀缓存这条约束。
实现前需要定的问题
.qwen/下的项目级文件?它必须能跨/clear存活,并且按项目区分——不同代码库的正确集合不同。tool_search的揭示计为用量(一次揭示正是"模型想要它没被给到"的直接证据)。maxPreloadTokens的 token 上限约束,还是那个上限只管延迟工具的预加载?tool_search调用次数是自然信号——tracking(core): non-conversation context token governance #12028 的分析把平衡点定在每会话 ≤2 次揭示。超过这个数就说明集合算错了,而这应当能从遥测里直接读到,而不是靠用户抱怨倒推。明确不提议的
验收
tool_search调用次数保持在约 2 次以内;超过就意味着省下的量被前缀重建抵消掉了。