Skip to content

feat(core): the eager tool surface is a hand-maintained static list — let something choose it, without invalidating the prompt prefix #12326

Description

@yiliang114

Part of #12028. Complements #12029 (which bounds the preload budget); this is about who chooses the resident set in the first place. No dependency on open work — the measurement ruler it needs landed in #12119.

What happens today

tools.eager is the one knob that materially shrinks the resident tool surface: in the #12028 sample it takes built-in tool schemas from 21,461 to ~5,868 tokens, more than every upstream code change on this umbrella put together. But it is a hand-written list, and it is static:

  • The allowlist is snapshotted once in PermissionManager.initialize() (permissions/permission-manager.ts:353), and a tool's registration status is decided when it is registered (config.ts → registerLazyTool). The setting is requiresRestart.
  • A demoted tool is not disabled: it stays registered, listed in /tools, and loadable through tool_search.

So the mechanism already in the codebase is static starting set + model-driven top-up. The dynamic half exists. What is missing is that nothing chooses the starting set — every deployment enumerates tool names by hand, or (much more commonly) never sets the key at all and pays for the whole surface on every request.

Why this cannot simply become per-turn

Function declarations sit at the very front of the prompt prefix, so changing the declared set mid-session rewrites the prefix and invalidates its KV cache. From the #12028 sample (prefix ≈ 42k tokens, implicit cache at 20% of full price):

Deferred pool Turns before preloading is the cheaper choice
4.2k (bundled built-ins only) ~40
8k ~21
20k (MCP-heavy) ~8

Demotion saves on the order of ¥0.01 per cached turn, while a single mid-session reveal costs ¥0.28–0.40. That is why the design deliberately decides once at session start and thereafter only adds — see docs/design/toolsearch-preload-threshold.md.

The contrast worth noting: a skill's paths: activation is turn-level, because the skill listing lives in history[0] rather than in the tools block, so changing it does not break the tools → system prefix (the reasoning is in the comment at core/environmentContext.ts). The same "adapt to the session" idea is cheap one layer up and expensive here. This is a placement constraint, not a missing capability.

What could choose the set

Four options, cheapest first. None of them changes the declared set mid-session.

Option Cost Cache impact Notes
A. Project-scope tools.eager zero code — already possible none Still a hand-written list, but one per project instead of one globally. Worth documenting rather than building.
B. Named presets (e.g. tools.eagerProfile: "file-work") small none #7625's fork profiles already ship named tool-restriction presets as .qwen/fork-profiles/<name>.md; reuse that shape rather than inventing a second file format. Turns "enumerate 40 tool names" into "pick one word".
C. Usage-driven, computed at session start medium none Adaptive across sessions, static within one — which is exactly the shape the cache constraint allows. ToolCallEvent.function_name already exists (telemetry/types.ts:212), so per-tool call counts need no new instrumentation.
D. First-turn intent classification medium none (nothing is cached before the first request) The only point where "pick a set from intent" is safe. A wrong guess costs one tool_search round trip at best, and at worst a tool the model never thinks to look for — a silent quality loss, not an error.

Recommendation: C, with B as the manual override. C puts the adaptation between sessions instead of inside one, so it gets the benefit without touching the prefix-cache constraint at all.

Questions to settle before implementation

  1. Where usage state lives. A project-scope file under .qwen/? It has to survive /clear and be per-project, since the right set differs by codebase.
  2. Cold start. The first session in a project has no history. Fall back to the current behaviour (everything eager), or to a preset?
  3. Starvation, and it is the serious one. A tool that is never offered is never called, so a usage-driven set can lock itself in: absence of calls is not evidence the tool was not needed. Options: a floor set that is always eager, periodic re-offering, or counting tool_search reveals as usage (a reveal is direct evidence the model wanted something it was not given).
  4. Interaction with Percentage-of-context-window budgets scale the wrong way: ToolSearch preload never engages, and the always-on context warning never fires, on large windows #12029. Should the computed set be capped in tokens by maxPreloadTokens, or is that cap only about preloading deferred tools?
  5. How a wrong set is observed. tool_search calls per session is the natural signal — the tracking(core): non-conversation context token governance #12028 analysis put the break-even at ≤2 reveals per session. Anything above that means the computed set is wrong, and it should be visible in telemetry rather than inferred from complaints.

Explicitly not proposed

Acceptance

  • Idle cost (the input tokens of a session that asks one question and calls no tool) falls to the same order as a hand-written allowlist achieves, without a deployment writing one.
  • tool_search calls per session stay at or below ~2 on a representative task set; above that the saving is cancelled by prefix rebuilds.
  • No drop in tool-call success rate or task completion on that task set. This needs the evaluation harness the umbrella still lacks — worth stating plainly, because without it the starvation risk in question 3 is unfalsifiable.
中文说明

属于 #12028。与 #12029 互补(那个管的是预加载预算的上限),本单管的是常驻集合最初由谁来选。不依赖任何在途工作——它需要的度量尺已随 #12119 落地。

现状

tools.eager 是唯一能实质缩小常驻工具面的开关:在 #12028 的样本里,它把内置工具 schema 从 21,461 降到约 5,868 token,比这条伞 issue 下所有上游代码改动加起来还多。但它是一份手写名单,而且是静态的:

  • 白名单在 PermissionManager.initialize() 里快照一次(permissions/permission-manager.ts:353),工具的注册状态在注册时就定了(config.ts → registerLazyTool)。该设置 requiresRestart。
  • 被降级的工具不是被禁用:它仍然注册、仍出现在 /tools、仍可通过 tool_search 加载。

也就是说代码里已有的机制是静态起点集合 + 模型驱动的按需补齐,动态那一半是有的。缺的是没有任何东西来选这个起点——每个部署要么手工枚举工具名,要么(更常见)根本不设这个键,于是每一次请求都为整个工具面付费。

为什么不能直接做成"每轮动态"

函数声明位于 prompt 前缀的最前面,会话中途修改声明集合会重写前缀并作废其 KV 缓存。按 #12028 样本(前缀约 42k token,隐式缓存为全价 20%):

延迟工具池 预加载更划算所需轮数
4.2k(仅内置工具) 约 40
8k 约 21
20k(MCP 较多) 约 8

降级每轮(命中缓存)省下的量级是 ¥0.01,而一次会话中途的揭示要付 ¥0.28–0.40。这正是该设计刻意"会话开始定一次、之后只增不减"的原因——见 docs/design/toolsearch-preload-threshold.md。

值得对照的一点:skill 的 paths: 门控确实是 turn-level 的,因为 skill 清单位于 history[0] 而不在 tools 段,改它不会打断 tools → system 的前缀(理由写在 core/environmentContext.ts 的注释里)。同一个"随会话自适应"的想法,在上一层很便宜,在这一层很贵。这是位置带来的约束,不是能力缺失。

可以由谁来选这个集合

四个选项,由便宜到贵。它们都不在会话中途改变声明集合。

选项 成本 缓存影响 备注
A. 项目级 tools.eager 0 代码,今天即可 无 仍是手写名单,但每个项目一份而不是全局一份。值得写进文档,而不是去实现。
B. 命名预设(如 tools.eagerProfile: "file-work") 小 无 #7625 的 fork profiles 已经以 .qwen/fork-profiles/<name>.md 的形式提供了命名工具限制预设;复用那个形状,不要再造一种文件格式。把"枚举 40 个工具名"变成"选一个词"。
C. 用量驱动,在会话开始时计算 中 无 跨会话自适应、会话内静态——正是缓存约束允许的形状。ToolCallEvent.function_name 已经存在(telemetry/types.ts:212),按工具的调用次数不需要新埋点。
D. 首轮意图分类 中 无(第一次请求之前没有任何缓存) 这是"按意图选集合"唯一安全的时机。猜错的代价:好的情况是多一次 tool_search 往返,坏的情况是模型压根想不到去找某个工具——静默的质量损失,不是报错。

建议:选 C,以 B 作为手动覆盖。 C 把自适应放在会话之间而不是会话之内,因此完全不触碰前缀缓存这条约束。

实现前需要定的问题

  1. 用量状态存在哪。 .qwen/ 下的项目级文件?它必须能跨 /clear 存活,并且按项目区分——不同代码库的正确集合不同。
  2. 冷启动。 项目里的第一个会话没有历史。回退到当前行为(全部 eager),还是回退到某个预设?
  3. 饿死问题,这是真正严重的一条。 从未被提供的工具永远不会被调用,所以用量驱动的集合会自我锁定:没有调用记录不等于这个工具不需要。可选解法:一个永远 eager 的地板集合、定期重新提供、或把 tool_search 的揭示计为用量(一次揭示正是"模型想要它没被给到"的直接证据)。
  4. 与 Percentage-of-context-window budgets scale the wrong way: ToolSearch preload never engages, and the always-on context warning never fires, on large windows #12029 的关系。 计算出的集合是否也要受 maxPreloadTokens 的 token 上限约束,还是那个上限只管延迟工具的预加载?
  5. 集合选错了怎么被发现。 每会话的 tool_search 调用次数是自然信号——tracking(core): non-conversation context token governance #12028 的分析把平衡点定在每会话 ≤2 次揭示。超过这个数就说明集合算错了,而这应当能从遥测里直接读到,而不是靠用户抱怨倒推。

明确不提议的

验收

  • 空载成本(一个只问一句、不调用任何工具的会话所发出的输入 token)降到与手写白名单同一量级,而无需部署方去写那份名单。
  • 在一组有代表性的任务集上,每会话 tool_search 调用次数保持在约 2 次以内;超过就意味着省下的量被前缀重建抵消掉了。
  • 在同一任务集上,工具调用成功率与任务完成率不下降。这需要这条伞 issue 至今仍缺的评测框架——有必要明说,因为没有它,第 3 条的饿死风险是无法被证伪的。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions