You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(ci): the token work has no recall or task-success gate — teach the existing benchmark to compare two configurations #12333
Part of #12028. This is the acceptance criterion that has no owner: every token change on this umbrella is measured for what it saves, and nothing measures what it costs in tool recall or task success. Until that exists, the largest saving available cannot responsibly be turned on.
What already exists
.github/workflows/dsw-swe-verified-release.yml — DSW Harbor Benchmark Release — is a workflow_dispatch that evaluates a published npm release against SWE-bench Verified (500 instances) and Terminal-Bench 2.0 (89 tasks), with inputs for release_tag / qwen_release_tag, per-suite limits, a single-instance selector, and execution_backend.
So task success on a standard suite is already harnessed. The gap is narrower than "build an eval harness".
What is missing
No settings overlay. The workflow installs a published release and runs it with default settings; nothing lets a run carry a settings.json. So the one question this umbrella needs answered — does task success drop when tools.eager withholds most tool schemas? — cannot be asked of it at all. An input that writes a settings file into the run's QWEN_HOME (and an A/B mode that runs the same manifest twice, with and without it) is the whole change.
No token or recall profile per run. The suite reports task outcomes. It does not report:
idle cost — input tokens of a session that asks one question and calls no tool;
total input tokens per task, which is the number that decides whether a saving is real or was merely moved into the conversation (take away grep_search and the model reaches for grep through the shell, whose output is never cached);
bad-call rate — calls naming an undeclared tool, or failing parameter validation.
None of this needs new instrumentation. /context reports the breakdown with an explicit residual since fix(cli): make /context categories add up to the provider total #12119, ToolCallEvent.function_name (packages/core/src/telemetry/types.ts) already carries per-call tool names, and tool-call errors are already events. It is aggregation and reporting.
No task set for the sessions this actually governs. SWE-bench and Terminal-Bench are coding suites. A deployment whose sessions are mostly data queries and file work will see its recall regressions somewhere those suites do not look. A stratified set of real-shaped tasks (pure Q&A, file work, shell, multi-step) plus a handful of safety cases belongs beside them — stated as a requirement here, not designed here.
Why it blocks, concretely
The tools.eager allowlist is worth 21,461 → ~5,868 tokens in the tracking(core): non-conversation context token governance #12028 sample, with no code change — more than every merged PR on this umbrella combined. Nobody should enable it in production on the strength of "the prefix got smaller".
feat(core): preserve prompt cache for deferred tools #10410 makes deferral nearly free on the cost side, which removes the economic argument against deferring everything and leaves recall as the only remaining question — sharpening the need for this rather than softening it.
Acceptance
A single dispatched run answers, for two configurations of the same commit:
reported per config
idle cost
tokens
input tokens per task
mean, p50, p95
tool_search calls per session
mean, and share of sessions above 2
bad-call rate
share of calls
task success
the suites' own pass rate
safety cases
pass/fail, must be 100%
and the comparison is legible enough that "we saved 33% of the prefix and lost nothing" is either supported or refuted by it.
Not proposed here
A new benchmark suite. The existing two plus a stratified real-task set is enough.
Part of #12028. This is the acceptance criterion that has no owner: every token change on this umbrella is measured for what it saves, and nothing measures what it costs in tool recall or task success. Until that exists, the largest saving available cannot responsibly be turned on.
What already exists
.github/workflows/dsw-swe-verified-release.yml— DSW Harbor Benchmark Release — is aworkflow_dispatchthat evaluates a published npm release against SWE-bench Verified (500 instances) and Terminal-Bench 2.0 (89 tasks), with inputs forrelease_tag/qwen_release_tag, per-suite limits, a single-instance selector, andexecution_backend.So task success on a standard suite is already harnessed. The gap is narrower than "build an eval harness".
What is missing
No settings overlay. The workflow installs a published release and runs it with default settings; nothing lets a run carry a
settings.json. So the one question this umbrella needs answered — does task success drop whentools.eagerwithholds most tool schemas? — cannot be asked of it at all. An input that writes a settings file into the run'sQWEN_HOME(and an A/B mode that runs the same manifest twice, with and without it) is the whole change.No token or recall profile per run. The suite reports task outcomes. It does not report:
grep_searchand the model reaches forgrepthrough the shell, whose output is never cached);tool_searchcalls per session — the routing-miss proxy. The tracking(core): non-conversation context token governance #12028 analysis put the break-even at ≤2 reveals per session under today's reveal semantics; feat(core): preserve prompt cache for deferred tools #10410 changes that arithmetic, but the metric stays the right one;None of this needs new instrumentation.
/contextreports the breakdown with an explicit residual since fix(cli): make /context categories add up to the provider total #12119,ToolCallEvent.function_name(packages/core/src/telemetry/types.ts) already carries per-call tool names, and tool-call errors are already events. It is aggregation and reporting.No task set for the sessions this actually governs. SWE-bench and Terminal-Bench are coding suites. A deployment whose sessions are mostly data queries and file work will see its recall regressions somewhere those suites do not look. A stratified set of real-shaped tasks (pure Q&A, file work, shell, multi-step) plus a handful of safety cases belongs beside them — stated as a requirement here, not designed here.
Why it blocks, concretely
tools.eagerallowlist is worth 21,461 → ~5,868 tokens in the tracking(core): non-conversation context token governance #12028 sample, with no code change — more than every merged PR on this umbrella combined. Nobody should enable it in production on the strength of "the prefix got smaller".Acceptance
A single dispatched run answers, for two configurations of the same commit:
tool_searchcalls per sessionand the comparison is legible enough that "we saved 33% of the prefix and lost nothing" is either supported or refuted by it.
Not proposed here
中文说明
属于 #12028。这是唯一没有归属的验收标准:这条伞 issue 下每一个 token 改动都度量了"省了多少",却没有任何东西度量它在工具召回率与任务成功率上"付出了多少"。这个东西不存在之前,收益最大的那个杠杆就不该被负责地打开。
已经存在的部分
.github/workflows/dsw-swe-verified-release.yml——DSW Harbor Benchmark Release——是一个workflow_dispatch,对已发布的 npm 版本跑 SWE-bench Verified(500 个实例)与 Terminal-Bench 2.0(89 个任务),输入项包括release_tag/qwen_release_tag、各套件的数量上限、单实例选择器与execution_backend。也就是说**"标准套件上的任务成功率"已经有 harness 了**。缺口比"建一个评测框架"窄得多。
缺的部分
没有配置覆盖。 该工作流安装一个已发布版本并以默认设置运行,没有任何途径让一次运行携带
settings.json。于是这条伞 issue 真正需要回答的那个问题——当tools.eager扣住大部分工具 schema 后,任务成功率是否下降?——根本无法向它提问。加一个把 settings 文件写进运行时QWEN_HOME的输入项(以及一个对同一 manifest 跑两遍、一遍带一遍不带的 A/B 模式),就是全部改动。每次运行没有 token 与召回画像。 套件报告任务结果,但不报告:
grep_search,模型就去用 shell 里的grep,那些输出永远不命中缓存);tool_search调用次数——路由未命中的代理指标。tracking(core): non-conversation context token governance #12028 的分析在当时的揭示语义下把平衡点定在每会话 ≤2 次;feat(core): preserve prompt cache for deferred tools #10410 改变了那笔账,但这个指标依然是对的那个;这些都不需要新埋点。
/context自 fix(cli): make /context categories add up to the provider total #12119 起报告带显式残差的分类;ToolCallEvent.function_name(packages/core/src/telemetry/types.ts)已经携带每次调用的工具名;工具调用错误本来就是事件。这是聚合与呈现的工作。没有针对"真正被治理的那些会话"的任务集。 SWE-bench 与 Terminal-Bench 是编码套件。一个会话以数据查询和文件操作为主的部署,其召回率回退会出现在那两个套件不看的地方。一组分层的真实形态任务(纯问答、文件操作、Shell、多步)外加少量安全用例,应当与它们并列——这里只作为需求点明,不在此设计。
为什么它是阻塞项
tools.eager白名单在 tracking(core): non-conversation context token governance #12028 样本里价值 21,461 → 约 5,868 token,且不需要改代码——比这条伞 issue 下所有已合 PR 加起来都多。没有人应该仅凭"前缀变小了"就在生产环境打开它。验收
一次 dispatch 的运行,对同一个 commit 的两种配置,给出:
tool_search次数并且这份对比要足够清楚,使"我们省掉了 33% 的前缀且没有损失"这句话能被它支持或推翻。
本单不提议