You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Non-conversation context — the system prompt, built-in tool schemas, context (QWEN.md) files and the skill listing — is sent and paid for on every request. On a large-context model this block easily dwarfs the conversation itself without anyone noticing, because it surfaces as a small percentage of the window.
This issue tracks the individual findings; each gets its own issue below.
Evidence
A /context detail sample from an interactive session on a 1M-context model, taken one turn in:
category
tokens
share of non-conversation
Built-in tools
21,461
45.9%
Context (QWEN.md) files
15,400
33.0%
System prompt
5,253
11.2%
Skills (listing)
4,620
9.9%
MCP tools
0
—
Non-conversation total
46,734
Messages
614
98.7% of that request's input was prefix, while the window read as only 6.5% used. That gap is the theme of this issue: the health metric everyone looks at is a ratio against the window, and the costs are absolute.
Two things stand out. The largest single entry, workflow at 3,829 tokens, is a feature that is off by default upstream (isWorkflowsEnabled, packages/core/src/config/config.ts:8637) — enabling one experimental feature costs more than the six file tools combined. And the seven tools this session actually needed to do file work — run_shell_command, edit, read_file, write_file, grep_search, glob, tool_search — total 4,080 tokens, 19% of the block.
The same mis-parameterization shows up twice in the code, and both instances get less protective as windows grow:
tools.toolSearch.threshold — 10% of the window. At 1M that is a 100,000-token preload budget against a ~4.2k deferred pool, so nothing ever stays deferred.
MEMORY_CONTEXT_WARNING_RATIO — 15% of the window. At 1M the always-on context warning fires at 150,000 tokens, so 15,400 tokens of context files draws no warning at all.
Findings
Percentage-of-window budgets scale the wrong way. Both knobs above protect against costs that scale with the deferred set or the context files, not with the window.
Every active extension's context file is unconditionally resident. Nine extension QWEN.md files totalled 9,989 tokens — 21% of all non-conversation context — concatenated into every request regardless of relevance. There is no path-gating, no per-extension budget, no truncation. The same extensions already ship skills, which cost 100–200 tokens each in the listing and load their bodies on demand. (Correction: the same sample measured 84 skills at 4,620 tokens, about 55 each.)
The system prompt is ~46% tool-specific but is assembled independently of which tools are resident.## Using Your Tools (4,031 chars) is ~66% tool policy; # Examples (3,283 chars) is entirely [tool_call: …] transcripts. Deferring or denying a tool leaves its policy and examples in the prompt, which both wastes tokens and points the model at tools it was not given.
The 46,734-token sample above and the 2026-09-29 first-main comparison at fa4a4c92ce (17,794 → 13,956) are historical measurements with different configurations, not refreshed baselines. The direct list below has 19 issues: 12 completed, 1 closed as not planned, and 6 open. Five open items are on the main path; #12235 is a P3 follow-up on hold.
No implementation PR under this tracker is open any more. Everything that was open on 2026-10-04 has merged:
This changes default behavior and landed before the #12333(b) quality gate. Natural-discovery recall and net savings are unmeasured; #12326 stays open.
Related, tracked separately: #12889 (provider acceptance, above), #13003 (opt-in experiment, above), #13004 (window-bounded phase 1 up as #13571, default-off; process exit and replay deferred).
Acceptance order is unchanged: #12333(b) runner consumes a settings overlay → matched recall/task-success and full-cost runs → #12326 decision (now including whether #13033's default holds) → refreshed baseline → default-enable decisions. Memory experiments remain off by default. Deployment-side savings additionally require upgrading to a release that contains these PRs and configuring tools.eager; neither is a code change in this repository.
The earlier #13158 cooldown/window/cursor/drain/failure-bound work, including a4568de26e49, e5f51b4b8d98, 2662afbb9485, 4b8b45c11e14 and bbfeaac17e3d, is historical and absent from the current retained experiment. Its controlled checks do not accept #13004 or establish current source behavior. The retained three-turn OFF/ON sample at b5d948cd44 recorded captured-process input 77,125 → 61,655, primary input 47,908 → 47,838, and requests 13 → 8. ON followed OFF with shared provider cache; this is not causal billing evidence or a current-main comparison. Earlier-head Goal and Agent samples likewise include pay-on-use regressions and incomplete/confounded billing attribution.
Non-conversation context — the system prompt, built-in tool schemas, context (
QWEN.md) files and the skill listing — is sent and paid for on every request. On a large-context model this block easily dwarfs the conversation itself without anyone noticing, because it surfaces as a small percentage of the window.This issue tracks the individual findings; each gets its own issue below.
Evidence
A
/context detailsample from an interactive session on a 1M-context model, taken one turn in:QWEN.md) files98.7% of that request's input was prefix, while the window read as only 6.5% used. That gap is the theme of this issue: the health metric everyone looks at is a ratio against the window, and the costs are absolute.
The built-in tool block, itemised:
workflowweb_fetchagenteditrun_shell_commandexit_plan_modecron_createread_filereport_findingsweb_searchrecord_artifactsend_messageupdate_goalmonitorask_user_questionwrite_fileplus
create_sub_session409,tool_search375,get_goal365,enter_plan_mode343,grep_search301,list_agents292,notebook_edit291,zoom_image279,glob266,record_source220,read_mcp_resource191,cron_delete116,task_stop102,cron_list80.Two things stand out. The largest single entry,
workflowat 3,829 tokens, is a feature that is off by default upstream (isWorkflowsEnabled,packages/core/src/config/config.ts:8637) — enabling one experimental feature costs more than the six file tools combined. And the seven tools this session actually needed to do file work —run_shell_command,edit,read_file,write_file,grep_search,glob,tool_search— total 4,080 tokens, 19% of the block.The same mis-parameterization shows up twice in the code, and both instances get less protective as windows grow:
tools.toolSearch.threshold— 10% of the window. At 1M that is a 100,000-token preload budget against a ~4.2k deferred pool, so nothing ever stays deferred.MEMORY_CONTEXT_WARNING_RATIO— 15% of the window. At 1M the always-on context warning fires at 150,000 tokens, so 15,400 tokens of context files draws no warning at all.Findings
QWEN.mdfiles totalled 9,989 tokens — 21% of all non-conversation context — concatenated into every request regardless of relevance. There is no path-gating, no per-extension budget, no truncation. The same extensions already ship skills, which cost 100–200 tokens each in the listing and load their bodies on demand. (Correction: the same sample measured 84 skills at 4,620 tokens, about 55 each.)## Using Your Tools(4,031 chars) is ~66% tool policy;# Examples(3,283 chars) is entirely[tool_call: …]transcripts. Deferring or denying a tool leaves its policy and examples in the prompt, which both wastes tokens and points the model at tools it was not given./contextaccounting does not close. Categories sum to 47,348 against a reported 65,267. Token governance cannot be driven from numbers that do not add up. (Correction 2026-09-17: the 1.378 ratio is not CJK undercounting — the display path uses the CJK-aware estimator for every category. The gap is structural:messagesis derived astotal − apiCachedTokens, and the startup prelude belongs to no category. See fix(cli): /context category breakdown does not close (skills listing unattributed, messages derived from a cache subtraction, startup prelude uncounted) #12033.)tools.disableddoes not always stop the schema being sent (tools.disabled removes zoom_image from the registry but its schema is still sent to the model #11814).Status (2026-10-07)
The 46,734-token sample above and the 2026-09-29 first-main comparison at
fa4a4c92ce(17,794 → 13,956) are historical measurements with different configurations, not refreshed baselines. The direct list below has 19 issues: 12 completed, 1 closed as not planned, and 6 open. Five open items are on the main path; #12235 is a P3 follow-up on hold.No implementation PR under this tracker is open any more. Everything that was open on 2026-10-04 has merged:
1fb5a7152270537464132303/contextestimate clamp149cb3380c168e3f3923b8c8tool_call.arguments+ target-named errorsb54d49c073c7Refs #12889. The original news prompt on the reporter's Responses provider has not been rerun; #12889 stays open.v0.25.0 does not contain #13033 or #12901; no stable release does yet.
Open main-path items and what each one is waiting on:
Related, tracked separately: #12889 (provider acceptance, above), #13003 (opt-in experiment, above), #13004 (window-bounded phase 1 up as #13571, default-off; process exit and replay deferred).
Acceptance order is unchanged: #12333(b) runner consumes a settings overlay → matched recall/task-success and full-cost runs → #12326 decision (now including whether #13033's default holds) → refreshed baseline → default-enable decisions. Memory experiments remain off by default. Deployment-side savings additionally require upgrading to a release that contains these PRs and configuring
tools.eager; neither is a code change in this repository.Historical memory evidence; withdrawn implementation
The earlier #13158 cooldown/window/cursor/drain/failure-bound work, including
a4568de26e49,e5f51b4b8d98,2662afbb9485,4b8b45c11e14andbbfeaac17e3d, is historical and absent from the current retained experiment. Its controlled checks do not accept #13004 or establish current source behavior. The retained three-turn OFF/ON sample atb5d948cd44recorded captured-process input 77,125 → 61,655, primary input 47,908 → 47,838, and requests 13 → 8. ON followed OFF with shared provider cache; this is not causal billing evidence or a current-main comparison. Earlier-head Goal and Agent samples likewise include pay-on-use regressions and incomplete/confounded billing attribution.Evidence remains in the original queue, tail, write-recovery, failure-bound and tool-cost reports with their limitations.
Sub-issues
tools.disabledremoveszoom_imagefrom the registry but its schema is still sent — closed; fix(core): hide unavailable image zoom guidance #12271 covers the only evidenced path, and no raw schema leak was reproducedmain(fix(core): keep the deferred-tool bridge halves on the same tool #12539 shared name resolution and fingerprint check; Deferred-tool bridge: a hidden tool whose schema left context via /compress is still invocable by name (remaining half of #11321) #12569 review gate via fix(core): deferred-tool selection rules in the reminder line and a reviewed-schema gate for tool_call #13020)/compressis still invocable by name — closed by merged fix(core): deferred-tool selection rules in the reminder line and a reviewed-schema gate for tool_call #13020; the broader residual items in Deferred review findings from PR #10410: feat(core): preserve prompt cache for deferred tools #11321 remain separatemainat728c13dedid not reproduce the original information loss. The issue is still open; model-choice acceptance is unverified. Current-main reproduction report<available_skills>entry, so the 8,000-char listing trim takes its budget from the user's own skills — closed by the entry bounds in merged fix(core,cli): budget the bundled skill listing; drop OpenTUI's category heading without a provider total #13159; broader listing compression remains gated on feat(ci): the token work has no recall or task-success gate — teach the existing benchmark to compare two configurations #12333(b)/contextcategory breakdown does not close (skills listing unattributed, messages derived from a cache subtraction, startup prelude uncounted) — implemented by fix(cli): make /context categories add up to the provider total #12119/contextsubtracts a process-global cached-token count, so a daemon session can be charged another session's cache — implemented by fix(cli): prefer per-session cached token count in /context #12066/contextaccounting) — fix(cli): close the deferred /context accounting follow-ups #12540 landed the actionable fixes; residual Suggestions on hold (P3)/contextshows a conversation-sized Messages row under pre-conversation overhead when usage is estimated — closed by the OpenTUI heading fix in merged fix(core,cli): budget the bundled skill listing; drop OpenTUI's category heading without a provider total #1315936533afb82reported foreground main-model input −11.65%; selector ablation and P95 tail validation remain open. The default-off perf(memory): skip the selector after a delivered unique strong recall hit #13003 selector experiment merged in feat(memory): Opt-in selector skip for a unique new recall hit #13158; perf(memory): add a bounded cooldown after no-op extraction #13004 was withdrawn and stays open for preservation/replay design. Rollout is not complete中文说明
非对话上下文——系统提示词、内置工具 schema、
QWEN.md上下文文件和 skill 清单——每轮请求都会完整发送并计费。在大上下文模型上,这一块很容易在无人察觉的情况下远超对话本身,因为它只表现为窗口里一个很小的百分比。本 issue 用于统一跟进,下列每条各自开 issue。
数据
某次交互会话开场第一轮的
/context detail,模型上下文窗口 1M:QWEN.md)文件这轮请求 98.7% 的输入是前缀,而窗口只显示用了 6.5%。这个落差正是本 issue 的主题:大家看的健康指标是"占窗口的比例",而成本是绝对值。
内置工具块逐项:
workflowweb_fetchagenteditrun_shell_commandexit_plan_modecron_createread_filereport_findingsweb_searchrecord_artifactsend_messageupdate_goalmonitorask_user_questionwrite_file以及
create_sub_session409、tool_search375、get_goal365、enter_plan_mode343、grep_search301、list_agents292、notebook_edit291、zoom_image279、glob266、record_source220、read_mcp_resource191、cron_delete116、task_stop102、cron_list80。有两点值得注意。占比最大的
workflow(3,829 token)在上游是默认关闭的功能(isWorkflowsEnabled,packages/core/src/config/config.ts:8637)——开一个实验性功能的代价超过六个文件工具之和。而这个会话真正用于文件工作的七个工具——run_shell_command、edit、read_file、write_file、grep_search、glob、tool_search——合计 4,080 token,只占该块的 19%。同一种参数化错误在代码里出现了两次,且都随着窗口变大而更不起作用:
tools.toolSearch.threshold—— 窗口的 10%。1M 窗口对应 100,000 token 的预加载预算,而延迟池只有约 4.2k,于是没有任何工具能保持延迟。MEMORY_CONTEXT_WARNING_RATIO—— 窗口的 15%。1M 窗口下,常驻上下文的警告要到 150,000 token 才触发,因此 15,400 token 的上下文文件完全不会告警。结论
QWEN.md合计 9,989 token,占全部非对话上下文的 21%,无论会话在做什么都会被拼进每轮请求。没有路径门控、没有单个 extension 的预算、也不截断。而这些 extension 本身已经带了 skill——清单里每个只要 100–200 token,正文按需加载。(更正:同一样本实测 84 个 skill 共 4,620 token,平均约 55。)## Using Your Tools(4,031 字符)约 66% 是工具策略,# Examples(3,283 字符)全部是[tool_call: …]记录。把工具改为延迟或禁用之后,它对应的策略和示例仍留在提示词里——既浪费 token,又会指引模型去用它并没有拿到的工具。/context的账不闭合。 分类相加 47,348,报告总数 65,267。账都对不上,就没法用它驱动 token 治理。(2026-09-17 更正:1.378 这个比值不是 CJK 低估——展示路径所有分类用的都是 CJK-aware 估算器。缺口是结构性的:messages由total − apiCachedTokens推导,且启动 prelude 不属于任何分类。详见 fix(cli): /context category breakdown does not close (skills listing unattributed, messages derived from a cache subtraction, startup prelude uncounted) #12033。)tools.disabled并不总能阻止 schema 被发送(tools.disabled removes zoom_image from the registry but its schema is still sent to the model #11814)。子 issue
见上方直接清单。#12047 由 #12066 解决,#12033 由 #12119 解决;#12048 的实验性修复 #12273 在真实 Bailian 验证后关闭,该 issue 按 not planned 收口。内置工具体积跟踪 #12054 已由 #12142、#12323、#12532 完成。
当前进展(2026-10-07)
上方 46,734 token 样本,以及 2026-09-29 在
fa4a4c92ce上的首个主请求 17,794 → 13,956,都是不同配置下的历史测量,不是刷新后的基线。直接清单共 19 个 issue:12 个完成、1 个按 not planned 关闭、6 个开放。开放项中 5 个属于主线,#12235 为暂缓的 P3 后续项。本 tracker 下已没有开放的实现 PR。 10-04 时仍开放的 PR 已全部合并:
1fb5a7152270537464132303/context估算修正149cb3380c168e3f3923b8c8tool_call.arguments与带目标名的报错b54d49c073c7Refs #12889。尚未在报告人的 Responses 提供商上重跑原始新闻提示词;#12889 保持开放。v0.25.0 不含 #13033 和 #12901,目前也没有包含它们的正式版。
主线开放项及各自等待的条件:
相关但单独跟踪:#12889(提供商验收,见上)、#13003(开关实验,见上)、#13004(按窗口约束的第一阶段已提为 #13571,默认关闭;进程退出与重放延后)。
验收顺序不变:#12333(b) 的 runner 实际消费 settings overlay → 相同条件下的召回/任务成功率与完整成本配对 → #12326 决策(含 #13033 的默认值是否保留)→ 刷新基线 → 默认开启决策。记忆实验继续默认关闭。部署侧要看到收益,还需要升级到包含这些 PR 的版本并配置
tools.eager;这两项都不是本仓库的代码改动。历史证据与撤回范围
#13158 早期的冷却、窗口、游标、drain 和失败次数限制(包括
a4568de26e49、e5f51b4b8d98、2662afbb9485、4b8b45c11e14、bbfeaac17e3d)不在当前保留的实验中。其受控检查不构成 #13004 验收或当前行为证明。旧b5d948cd44三轮 OFF/ON 记录进程输入 77,125 → 61,655、主请求输入 47,908 → 47,838、请求数 13 → 8;ON 后运行且共享 provider 缓存,不能推导因果账单收益或当前 main 对照。旧 Agent/Goal 样本同样包含按需使用路径成本上升,以及计费归因不完整或混杂。原始排队、尾部、写入恢复、失败次数限制及工具成本报告保留了各自边界。