Skip to content

tracking(core): non-conversation context token governance #12028

Description

@yiliang114

Non-conversation context — the system prompt, built-in tool schemas, context (QWEN.md) files and the skill listing — is sent and paid for on every request. On a large-context model this block easily dwarfs the conversation itself without anyone noticing, because it surfaces as a small percentage of the window.

This issue tracks the individual findings; each gets its own issue below.

Evidence

A /context detail sample from an interactive session on a 1M-context model, taken one turn in:

category tokens share of non-conversation
Built-in tools 21,461 45.9%
Context (QWEN.md) files 15,400 33.0%
System prompt 5,253 11.2%
Skills (listing) 4,620 9.9%
MCP tools 0 —
Non-conversation total 46,734
Messages 614

98.7% of that request's input was prefix, while the window read as only 6.5% used. That gap is the theme of this issue: the health metric everyone looks at is a ratio against the window, and the costs are absolute.

The built-in tool block, itemised:

tool tokens tool tokens
workflow 3,829 web_fetch 693
agent 3,613 edit 612
run_shell_command 1,495 exit_plan_mode 592
cron_create 1,103 read_file 585
report_findings 1,016 web_search 499
record_artifact 842 send_message 498
update_goal 783 monitor 475
ask_user_question 751 write_file 446

plus create_sub_session 409, tool_search 375, get_goal 365, enter_plan_mode 343, grep_search 301, list_agents 292, notebook_edit 291, zoom_image 279, glob 266, record_source 220, read_mcp_resource 191, cron_delete 116, task_stop 102, cron_list 80.

Two things stand out. The largest single entry, workflow at 3,829 tokens, is a feature that is off by default upstream (isWorkflowsEnabled, packages/core/src/config/config.ts:8637) — enabling one experimental feature costs more than the six file tools combined. And the seven tools this session actually needed to do file work — run_shell_command, edit, read_file, write_file, grep_search, glob, tool_search — total 4,080 tokens, 19% of the block.

The same mis-parameterization shows up twice in the code, and both instances get less protective as windows grow:

  • tools.toolSearch.threshold — 10% of the window. At 1M that is a 100,000-token preload budget against a ~4.2k deferred pool, so nothing ever stays deferred.
  • MEMORY_CONTEXT_WARNING_RATIO — 15% of the window. At 1M the always-on context warning fires at 150,000 tokens, so 15,400 tokens of context files draws no warning at all.

Findings

  1. Percentage-of-window budgets scale the wrong way. Both knobs above protect against costs that scale with the deferred set or the context files, not with the window.
  2. Every active extension's context file is unconditionally resident. Nine extension QWEN.md files totalled 9,989 tokens — 21% of all non-conversation context — concatenated into every request regardless of relevance. There is no path-gating, no per-extension budget, no truncation. The same extensions already ship skills, which cost 100–200 tokens each in the listing and load their bodies on demand. (Correction: the same sample measured 84 skills at 4,620 tokens, about 55 each.)
  3. The system prompt is ~46% tool-specific but is assembled independently of which tools are resident. ## Using Your Tools (4,031 chars) is ~66% tool policy; # Examples (3,283 chars) is entirely [tool_call: …] transcripts. Deferring or denying a tool leaves its policy and examples in the prompt, which both wastes tokens and points the model at tools it was not given.
  4. /context accounting does not close. Categories sum to 47,348 against a reported 65,267. Token governance cannot be driven from numbers that do not add up. (Correction 2026-09-17: the 1.378 ratio is not CJK undercounting — the display path uses the CJK-aware estimator for every category. The gap is structural: messages is derived as total − apiCachedTokens, and the startup prelude belongs to no category. See fix(cli): /context category breakdown does not close (skills listing unattributed, messages derived from a cache subtraction, startup prelude uncounted) #12033.)
  5. tools.disabled does not always stop the schema being sent (tools.disabled removes zoom_image from the registry but its schema is still sent to the model #11814).

Status (2026-10-07)

The 46,734-token sample above and the 2026-09-29 first-main comparison at fa4a4c92ce (17,794 → 13,956) are historical measurements with different configurations, not refreshed baselines. The direct list below has 19 issues: 12 completed, 1 closed as not planned, and 6 open. Five open items are on the main path; #12235 is a P3 follow-up on hold.

No implementation PR under this tracker is open any more. Everything that was open on 2026-10-04 has merged:

PR Merge commit Merged (UTC) First release containing it Remaining acceptance
#12580 context-first prompt 1fb5a7152270 10-04 07:19 v0.25.0 Representative net-token acceptance; scoped paired arms only (ON 31,441 vs OFF 31,437 provider tokens). Closed #12579.
#13158 opt-in selector skip (#13003) 537464132303 10-04 07:49 v0.25.0 Default off. Paired memory quality, full cost and P95 latency. #13003 stays open.
#13246 /context estimate clamp 149cb3380c16 10-04 08:08 v0.25.0 None. Closed #13239; per-tool detail mismatch stays in #12235.
#13033 defer Agent/Goal declarations by default 8e3f3923b8c8 10-06 04:13 v0.25.1-preview.0 only This changes default behavior and landed before the #12333(b) quality gate. Natural-discovery recall and net savings are unmeasured; #12326 stays open.
#12901 Responses nested tool_call.arguments + target-named errors b54d49c073c7 10-06 14:10 v0.25.1-preview.0 only Refs #12889. The original news prompt on the reporter's Responses provider has not been rerun; #12889 stays open.

v0.25.0 does not contain #13033 or #12901; no stable release does yet.

Open main-path items and what each one is waiting on:

Related, tracked separately: #12889 (provider acceptance, above), #13003 (opt-in experiment, above), #13004 (window-bounded phase 1 up as #13571, default-off; process exit and replay deferred).

Acceptance order is unchanged: #12333(b) runner consumes a settings overlay → matched recall/task-success and full-cost runs → #12326 decision (now including whether #13033's default holds) → refreshed baseline → default-enable decisions. Memory experiments remain off by default. Deployment-side savings additionally require upgrading to a release that contains these PRs and configuring tools.eager; neither is a code change in this repository.

Historical memory evidence; withdrawn implementation

The earlier #13158 cooldown/window/cursor/drain/failure-bound work, including a4568de26e49, e5f51b4b8d98, 2662afbb9485, 4b8b45c11e14 and bbfeaac17e3d, is historical and absent from the current retained experiment. Its controlled checks do not accept #13004 or establish current source behavior. The retained three-turn OFF/ON sample at b5d948cd44 recorded captured-process input 77,125 → 61,655, primary input 47,908 → 47,838, and requests 13 → 8. ON followed OFF with shared provider cache; this is not causal billing evidence or a current-main comparison. Earlier-head Goal and Agent samples likewise include pay-on-use regressions and incomplete/confounded billing attribution.

Evidence remains in the original queue, tail, write-recovery, failure-bound and tool-cost reports with their limitations.

Sub-issues

中文说明

非对话上下文——系统提示词、内置工具 schema、QWEN.md 上下文文件和 skill 清单——每轮请求都会完整发送并计费。在大上下文模型上,这一块很容易在无人察觉的情况下远超对话本身,因为它只表现为窗口里一个很小的百分比。

本 issue 用于统一跟进,下列每条各自开 issue。

数据

某次交互会话开场第一轮的 /context detail,模型上下文窗口 1M:

类别 token 占非对话
内置工具 21,461 45.9%
上下文(QWEN.md)文件 15,400 33.0%
系统提示词 5,253 11.2%
Skills(清单) 4,620 9.9%
MCP 工具 0 —
非对话合计 46,734
消息 614

这轮请求 98.7% 的输入是前缀,而窗口只显示用了 6.5%。这个落差正是本 issue 的主题:大家看的健康指标是"占窗口的比例",而成本是绝对值。

内置工具块逐项:

工具 token 工具 token
workflow 3,829 web_fetch 693
agent 3,613 edit 612
run_shell_command 1,495 exit_plan_mode 592
cron_create 1,103 read_file 585
report_findings 1,016 web_search 499
record_artifact 842 send_message 498
update_goal 783 monitor 475
ask_user_question 751 write_file 446

以及 create_sub_session 409、tool_search 375、get_goal 365、enter_plan_mode 343、grep_search 301、list_agents 292、notebook_edit 291、zoom_image 279、glob 266、record_source 220、read_mcp_resource 191、cron_delete 116、task_stop 102、cron_list 80。

有两点值得注意。占比最大的 workflow(3,829 token)在上游是默认关闭的功能(isWorkflowsEnabled,packages/core/src/config/config.ts:8637)——开一个实验性功能的代价超过六个文件工具之和。而这个会话真正用于文件工作的七个工具——run_shell_command、edit、read_file、write_file、grep_search、glob、tool_search——合计 4,080 token,只占该块的 19%。

同一种参数化错误在代码里出现了两次,且都随着窗口变大而更不起作用:

  • tools.toolSearch.threshold —— 窗口的 10%。1M 窗口对应 100,000 token 的预加载预算,而延迟池只有约 4.2k,于是没有任何工具能保持延迟。
  • MEMORY_CONTEXT_WARNING_RATIO —— 窗口的 15%。1M 窗口下,常驻上下文的警告要到 150,000 token 才触发,因此 15,400 token 的上下文文件完全不会告警。

结论

  1. 按窗口百分比计算预算的方向是反的。 上面两个开关防范的成本分别随"延迟工具集"和"上下文文件"增长,与窗口大小无关。
  2. 每个已启用 extension 的上下文文件都无条件常驻。 9 个 extension 的 QWEN.md 合计 9,989 token,占全部非对话上下文的 21%,无论会话在做什么都会被拼进每轮请求。没有路径门控、没有单个 extension 的预算、也不截断。而这些 extension 本身已经带了 skill——清单里每个只要 100–200 token,正文按需加载。(更正:同一样本实测 84 个 skill 共 4,620 token,平均约 55。)
  3. 系统提示词约 46% 与具体工具绑定,但其装配与"哪些工具实际常驻"无关。 ## Using Your Tools(4,031 字符)约 66% 是工具策略,# Examples(3,283 字符)全部是 [tool_call: …] 记录。把工具改为延迟或禁用之后,它对应的策略和示例仍留在提示词里——既浪费 token,又会指引模型去用它并没有拿到的工具。
  4. /context 的账不闭合。 分类相加 47,348,报告总数 65,267。账都对不上,就没法用它驱动 token 治理。(2026-09-17 更正:1.378 这个比值不是 CJK 低估——展示路径所有分类用的都是 CJK-aware 估算器。缺口是结构性的:messages 由 total − apiCachedTokens 推导,且启动 prelude 不属于任何分类。详见 fix(cli): /context category breakdown does not close (skills listing unattributed, messages derived from a cache subtraction, startup prelude uncounted) #12033。)
  5. tools.disabled 并不总能阻止 schema 被发送(tools.disabled removes zoom_image from the registry but its schema is still sent to the model #11814)。

子 issue

见上方直接清单。#12047 由 #12066 解决,#12033 由 #12119 解决;#12048 的实验性修复 #12273 在真实 Bailian 验证后关闭,该 issue 按 not planned 收口。内置工具体积跟踪 #12054 已由 #12142、#12323、#12532 完成。

当前进展(2026-10-07)

上方 46,734 token 样本,以及 2026-09-29 在 fa4a4c92ce 上的首个主请求 17,794 → 13,956,都是不同配置下的历史测量,不是刷新后的基线。直接清单共 19 个 issue:12 个完成、1 个按 not planned 关闭、6 个开放。开放项中 5 个属于主线,#12235 为暂缓的 P3 后续项。

本 tracker 下已没有开放的实现 PR。 10-04 时仍开放的 PR 已全部合并:

PR 合并提交 合并时间(UTC) 首个包含它的版本 剩余验收
#12580 上下文优先提示词 1fb5a7152270 10-04 07:19 v0.25.0 代表性净 token 验收;目前只有限定配对(ON 31,441 / OFF 31,437 provider token)。已关闭 #12579。
#13158 selector 跳过实验(#13003) 537464132303 10-04 07:49 v0.25.0 默认关闭。记忆质量、完整成本、P95 延迟配对;#13003 保持开放。
#13246 /context 估算修正 149cb3380c16 10-04 08:08 v0.25.0 无。已关闭 #13239;逐工具明细差异留在 #12235。
#13033 Agent/Goal 声明默认延迟 8e3f3923b8c8 10-06 04:13 仅 v0.25.1-preview.0 改变默认行为,且在 #12333(b) 质量门之前合入。 自然发现召回率与净节省尚未测量;#12326 保持开放。
#12901 Responses 嵌套 tool_call.arguments 与带目标名的报错 b54d49c073c7 10-06 14:10 仅 v0.25.1-preview.0 Refs #12889。尚未在报告人的 Responses 提供商上重跑原始新闻提示词;#12889 保持开放。

v0.25.0 不含 #13033 和 #12901,目前也没有包含它们的正式版。

主线开放项及各自等待的条件:

相关但单独跟踪:#12889(提供商验收,见上)、#13003(开关实验,见上)、#13004(按窗口约束的第一阶段已提为 #13571,默认关闭;进程退出与重放延后)。

验收顺序不变:#12333(b) 的 runner 实际消费 settings overlay → 相同条件下的召回/任务成功率与完整成本配对 → #12326 决策(含 #13033 的默认值是否保留)→ 刷新基线 → 默认开启决策。记忆实验继续默认关闭。部署侧要看到收益,还需要升级到包含这些 PR 的版本并配置 tools.eager;这两项都不是本仓库的代码改动。

历史证据与撤回范围

#13158 早期的冷却、窗口、游标、drain 和失败次数限制(包括 a4568de26e49、e5f51b4b8d98、2662afbb9485、4b8b45c11e14、bbfeaac17e3d)不在当前保留的实验中。其受控检查不构成 #13004 验收或当前行为证明。旧 b5d948cd44 三轮 OFF/ON 记录进程输入 77,125 → 61,655、主请求输入 47,908 → 47,838、请求数 13 → 8;ON 后运行且共享 provider 缓存,不能推导因果账单收益或当前 main 对照。旧 Agent/Goal 样本同样包含按需使用路径成本上升,以及计费归因不完整或混杂。

原始排队、尾部、写入恢复、失败次数限制及工具成本报告保留了各自边界。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions