Skip to content

feat(core): assemble the tool-policy and example sections of the system prompt from the resident tool set #12032

Description

@yiliang114

Part of #12028.

The mismatch

Which tool schemas reach the model is highly configurable — shouldDefer, tools.eager, tools.visible, tools.disabled, permissions.deny, the ToolSearch preload. The system prompt's description of those tools is not: it is assembled the same way regardless.

Measured on the default base prompt (20,801 chars; no git, no sandbox, interactive, generic model):

section chars notes
## Using Your Tools 4,031 ~66% of its lines name a specific tool
# Examples 3,283 entirely [tool_call: …] transcripts

45.7% of all lines in the base prompt name a specific tool. Roughly 7.3 KB is structurally tool-specific.

So a session that defers or denies, say, grep_search and agent still ships their policy bullets and their worked examples — and the prompt actively instructs the model to use tools it has not been given. The token waste is the smaller half of the problem; the correctness half is that prompt and tool set can silently disagree.

Why this is tractable

Tool names in the prompt are interpolated from ToolNames (packages/core/src/tools/tool-names.ts) rather than hard-coded, and the sections are already bullet-organised. The assembly point is buildDefaultBasePrompt (packages/core/src/core/prompts.ts:359-513), which already does conditional assembly for several flags — todoWriteEnabled gates its own bullets at :374-385, codeModeOnly swaps the whole Using-Your-Tools body at :310-331, and keepCodingInstructions: false drops ## Software Engineering Tasks at :369-372.

Extending that to "emit a tool's policy bullet and examples only if the tool is in the current declaration set" is the same pattern applied per tool instead of per feature flag.

Why it is worth doing

  • It makes prompt/tool-set consistency a property of the mechanism rather than something a deployment has to maintain by hand.
  • It benefits every configuration that trims tools, not just one profile — so it does not need a new opt-in knob.
  • It removes the main reason an integrator would fork the system prompt via --system-prompt / QWEN_SYSTEM_MD. That matters: prompts.ts carries roughly 6,349 chars (~30%) of safety and behavioural boundaries — denied-tool-call bypass rules, hook-context-is-not-user-input, the dangerous-action classes — and a forked copy silently stops tracking upstream changes to them, with no test failing.

Caveats

  • Mid-session tool reveals already call setTools() and invalidate the prefix cache; making the system prompt depend on the resident tool set means a reveal would additionally change the system block. The prompt is cached behind the tools block, so this needs to be measured rather than assumed. Regenerating the prompt only at session start (leaving a reveal to change the tools block alone) is the conservative option.
  • getActionsSection() (prompts.ts:714-744) is unconditional today and should stay that way; the safety boundaries are not per-tool.
中文说明

属于 #12028。

失配点

哪些工具 schema 会到达模型是高度可配置的——shouldDefer、tools.eager、tools.visible、tools.disabled、permissions.deny、ToolSearch 预加载。但系统提示词里对这些工具的描述不是:无论如何装配方式都一样。

按默认基础提示词测量(20,801 字符;无 git、无沙箱、交互模式、通用模型):

段落 字符 说明
## Using Your Tools 4,031 约 66% 的行提到具体工具名
# Examples 3,283 全部是 [tool_call: …] 记录

基础提示词中 45.7% 的行提到了具体工具名,结构上与工具绑定的约 7.3 KB。

于是一个把 grep_search、agent 改为延迟或禁用的会话,仍然会发送它们的策略条目和完整示例——而且提示词还在主动指引模型去使用它并没有拿到的工具。浪费 token 是问题较小的一半;较大的一半是提示词与工具集会静默地不一致。

为什么可行

提示词里的工具名是从 ToolNames(packages/core/src/tools/tool-names.ts)插值进去的,不是硬编码字符串,而且这些段落本来就是按条目组织的。装配点在 buildDefaultBasePrompt(packages/core/src/core/prompts.ts:359-513),它已经在按若干开关做条件装配——todoWriteEnabled 在 :374-385 门控自己的条目,codeModeOnly 在 :310-331 整体替换 Using Your Tools 的正文,keepCodingInstructions: false 在 :369-372 删掉 ## Software Engineering Tasks。

把它扩展成"仅当该工具在当前声明集合中时才生成它的策略条目和示例",只是把同一个模式从"按 feature flag"改成"按工具"。

为什么值得做

  • 它让"提示词与工具集一致"成为机制的性质,而不是每个部署方需要手工维护的东西。
  • 它惠及所有裁剪了工具的配置,而不只是某一个 profile,因此不需要新增 opt-in 开关。
  • 它消除了集成方 fork 系统提示词(--system-prompt / QWEN_SYSTEM_MD)的主要理由。这一点很重要:prompts.ts 里约有 6,349 字符(约 30%)是安全与行为边界——被拒工具调用不得绕路、hook 注入内容不算用户输入、危险操作四分类——fork 出来的副本会静默地停止跟踪上游对这些条款的修改,而且不会有任何测试失败。

注意

  • 会话中途的工具揭示本来就会调用 setTools() 并让前缀缓存失效;让系统提示词依赖常驻工具集意味着一次揭示还会额外改变 system 块。提示词缓存在 tools 块之后,因此这一点需要实测而非假设。保守做法是只在会话开始时生成提示词,中途揭示只改 tools 块。
  • getActionsSection()(prompts.ts:714-744)目前是无条件生成的,应保持不变;安全边界不是按工具划分的。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    category/coreCore engine and logicpriority/P3Low - Minor, cosmetic, nice-to-fix issuesscope/corescope/token-managementToken handling and limitsstatus/ready-for-humanSpecified but requires human judgment to implement; not suitable for an autonomous agenttype/feature-requestNew feature or enhancement request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions