Skip to content

perf(core): the built-in tool descriptions and schemas are the largest block of non-conversation context, and have no size tracking #12054

Description

@yiliang114

Part of #12028.

The gap

#12028 itemises non-conversation context for one session on a 1M-context model. Built-in tools are the largest category by a wide margin:

category tokens share of non-conversation
Built-in tools 21,461 45.9%
Context (QWEN.md) files 15,400 33.0%
System prompt 5,253 11.2%
Skills (listing) 4,620 9.9%
MCP tools 0 —
Non-conversation total 46,734

Every other category already has a tracking issue with a mechanism attached:

The size of the tool descriptions and schemas themselves has nothing. Not the deferral mechanism (#12029), not the leak (#11814) — what each tool costs when it is resident.

Six query angles across issues and PRs (tool schema tokens, shorten tool description, reduce tool schema tokens, slim built-in tools, tool description too long, and title-scoped variants) returned no matching item. That is absence of evidence rather than proof, but there is no tracking issue either way, which is the actionable part.

Where the mass is

From #12028's per-tool table, the top of the block:

workflow 3,829 · agent 3,613 · run_shell_command 1,495 · cron_create 1,103 · report_findings 1,016 · record_artifact 842 · update_goal 783 · ask_user_question 751 · web_fetch 693 · edit 612 · exit_plan_mode 592 · read_file 585 · web_search 499 · send_message 498

The two largest alone are 7,442 tokens — 34.7% of the block, 15.9% of all non-conversation context. Both are description-heavy rather than parameter-heavy: agent carries a subagent-type catalogue plus delegation policy, workflow a full authoring contract.

workflow is gated behind isWorkflowsEnabled(), which defaults to false on current main (packages/core/src/config/config.ts:9061: return this.workflowsEnabled ?? false). It appears in the sample because that session enabled it — so it is not dead weight on a default install. It is however the most expensive opt-in in the product: enabling one experimental feature costs more than run_shell_command, edit, read_file, write_file, grep_search and glob combined (4,080 for all six, per #12028).

What is actually askable

Not "make descriptions short". Three specific classes:

  1. Restatement of what the schema already says. Prose repeating a parameter's name, type or enum back at the model. The schema is machine-read; the duplicate sentence costs tokens and adds nothing.
  2. Worked examples embedded in descriptions. Several tools carry multi-example transcripts. feat(core): assemble the tool-policy and example sections of the system prompt from the resident tool set #12032 makes exactly this observation about # Examples in the system prompt (3,283 chars, entirely [tool_call: …]); the copies inside tool descriptions have the same shape and raise the same question — does the model need N examples or one?
  3. Overlapping tools that could merge. tracking(core): non-conversation context token governance #12028 notes seven file tools total 4,080 tokens. Whether some of those are one tool with a mode parameter is a design question rather than a trimming one, and should be answered separately from (1) and (2).

What must not be trimmed

The discipline #12032 applies to prompts.ts applies here unchanged. Load-bearing text in tool descriptions includes:

  • call boundaries and the dangerous-operation classes
  • run_shell_command's shell-quoting rules and the "prefer dedicated tools over cat/grep/sed" policy
  • the denied-tool-call rule — a denial must not be routed around via shell indirection, generated scripts, aliases, symlinks, config changes or hooks
  • the "explain a mutating command before running it" requirement

These read as verbose but are behavioural constraints. Cutting them to hit a token target trades a measurable saving for an unmeasurable regression, and no test in CI would fail — the same failure mode #12032 identifies for a forked system prompt.

Why this is not a pure win

Trimming descriptions risks tool-selection recall. Any change needs a before/after on task success, not only on token count, which is why the acceptance list below carries a recall gate alongside the size gate.

Blocked on the ruler

There is currently no trustworthy instrument for the before/after. #12033 documents that /context's categories do not sum to its reported total, and that per-category estimates have never been compared against provider inputTokens. A "reduced by N tokens" claim sourced from /context today is not defensible. Either land #12033 first, or measure directly against provider usage in the PR.

Suggested acceptance

  • a per-tool token table for a default install, measured against provider inputTokens rather than estimated
  • a measurable reduction in the built-in tool block for a default configuration
  • no regression in tool-selection recall or task success on a fixed task set
  • the preserved safety / call-boundary text named explicitly in the PR, so a reviewer can check nothing behavioural was cut
中文说明

属于 #12028。

空白在哪: #12028 的样本里,内置工具是非对话上下文的最大项(21,461 token,占 45.9%),远超上下文文件(15,400 / 33.0%)、系统提示词(5,253 / 11.2%)和 skill 清单(4,620 / 9.9%)。其余每一类都已经有带机制的跟踪单——系统提示词 #12032、上下文文件 #12030、大窗口下延迟机制失效 #12029、禁用后 schema 仍发送 #11814、尺子本身 #12033。唯独「工具描述与 schema 本身有多大」这件事没有任何单:不是延迟机制、也不是泄漏,而是每个工具在常驻时的单价。

已用六个查询角度(issue + PR)搜索,无匹配项。这属于「查不到证据」而非「证明不存在」,但无论如何目前确实没有跟踪单,这才是可行动的部分。

质量集中在哪: 最大的两项 workflow 3,829 与 agent 3,613 合计 7,442 token,占该块 34.7%、占全部非对话上下文 15.9%。两者都是描述重、参数轻:agent 带一整套 subagent 类型目录与委派策略,workflow 带完整的编写契约。

workflow 由 isWorkflowsEnabled() 门控,当前 main 上默认为 false(config.ts:9061:return this.workflowsEnabled ?? false)。它出现在样本里是因为该会话显式启用了,所以默认安装下它不是死重。但它是全产品最贵的一个 opt-in:启用一个实验特性的成本,超过 run_shell_command、edit、read_file、write_file、grep_search、glob 六个工具之和(4,080)。

真正能要求的是三类,不是「把描述改短」:

  1. 重复 schema 已表达的内容——把参数名、类型、枚举再用散文说一遍。schema 是机器读的,那句重复的散文只花 token、不增信息。
  2. 描述里内嵌的示例——若干工具带多例转录。feat(core): assemble the tool-policy and example sections of the system prompt from the resident tool set #12032 对系统提示词的 # Examples(3,283 字符,全是 [tool_call: …])提的正是同一条,工具描述里的副本形态相同、问题也相同:模型需要 N 个例子还是 1 个?
  3. 可合并的重叠工具——tracking(core): non-conversation context token governance #12028 指出七个文件类工具合计 4,080 token。其中某几个是否应为「一个工具 + mode 参数」是设计问题而非裁剪问题,应与前两类分开回答。

不能裁的部分: #12032 对 prompts.ts 的纪律在这里同样适用。工具描述里的承重文本包括:调用边界与危险操作分类;run_shell_command 的 shell 引号规则与「优先用专用工具而非 cat/grep/sed」策略;被拒工具调用不得绕路(不得通过 shell 间接、生成脚本、别名、软链、改配置或 hook 规避);以及「改动型命令执行前先解释」的要求。这些读起来冗长,但都是行为约束。为了达标去砍它们,是用可测量的节省换不可测量的回退,而且 CI 里没有任何测试会失败——正是 #12032 为 fork 系统提示词指出的同一种失效模式。

为什么这不是纯收益: 裁描述有工具选择召回率的风险。任何改动都需要任务成功率上的前后对比,而不只是 token 数,所以下面的验收清单在尺寸门之外还带了一道召回门。

卡在尺子上: 目前没有可信的前后对比工具。#12033 记录了 /context 的分类之和与其报告总数不闭合,且各分类估算从未与 provider 的 inputTokens 对照过。今天拿 /context 出的「降低了 N token」站不住。要么先落地 #12033,要么在 PR 里直接对 provider usage 量测。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions