Skip to content

feat(ci): the token work has no recall or task-success gate — teach the existing benchmark to compare two configurations #12333

Description

@yiliang114

Part of #12028. This is the acceptance criterion that has no owner: every token change on this umbrella is measured for what it saves, and nothing measures what it costs in tool recall or task success. Until that exists, the largest saving available cannot responsibly be turned on.

What already exists

.github/workflows/dsw-swe-verified-release.yml — DSW Harbor Benchmark Release — is a workflow_dispatch that evaluates a published npm release against SWE-bench Verified (500 instances) and Terminal-Bench 2.0 (89 tasks), with inputs for release_tag / qwen_release_tag, per-suite limits, a single-instance selector, and execution_backend.

So task success on a standard suite is already harnessed. The gap is narrower than "build an eval harness".

What is missing

  1. No settings overlay. The workflow installs a published release and runs it with default settings; nothing lets a run carry a settings.json. So the one question this umbrella needs answered — does task success drop when tools.eager withholds most tool schemas? — cannot be asked of it at all. An input that writes a settings file into the run's QWEN_HOME (and an A/B mode that runs the same manifest twice, with and without it) is the whole change.

  2. No token or recall profile per run. The suite reports task outcomes. It does not report:

    • idle cost — input tokens of a session that asks one question and calls no tool;
    • total input tokens per task, which is the number that decides whether a saving is real or was merely moved into the conversation (take away grep_search and the model reaches for grep through the shell, whose output is never cached);
    • tool_search calls per session — the routing-miss proxy. The tracking(core): non-conversation context token governance #12028 analysis put the break-even at ≤2 reveals per session under today's reveal semantics; feat(core): preserve prompt cache for deferred tools #10410 changes that arithmetic, but the metric stays the right one;
    • bad-call rate — calls naming an undeclared tool, or failing parameter validation.

    None of this needs new instrumentation. /context reports the breakdown with an explicit residual since fix(cli): make /context categories add up to the provider total #12119, ToolCallEvent.function_name (packages/core/src/telemetry/types.ts) already carries per-call tool names, and tool-call errors are already events. It is aggregation and reporting.

  3. No task set for the sessions this actually governs. SWE-bench and Terminal-Bench are coding suites. A deployment whose sessions are mostly data queries and file work will see its recall regressions somewhere those suites do not look. A stratified set of real-shaped tasks (pure Q&A, file work, shell, multi-step) plus a handful of safety cases belongs beside them — stated as a requirement here, not designed here.

Why it blocks, concretely

Acceptance

A single dispatched run answers, for two configurations of the same commit:

reported per config
idle cost tokens
input tokens per task mean, p50, p95
tool_search calls per session mean, and share of sessions above 2
bad-call rate share of calls
task success the suites' own pass rate
safety cases pass/fail, must be 100%

and the comparison is legible enough that "we saved 33% of the prefix and lost nothing" is either supported or refuted by it.

Not proposed here

中文说明

属于 #12028。这是唯一没有归属的验收标准:这条伞 issue 下每一个 token 改动都度量了"省了多少",却没有任何东西度量它在工具召回率与任务成功率上"付出了多少"。这个东西不存在之前,收益最大的那个杠杆就不该被负责地打开。

已经存在的部分

.github/workflows/dsw-swe-verified-release.yml——DSW Harbor Benchmark Release——是一个 workflow_dispatch,对已发布的 npm 版本跑 SWE-bench Verified(500 个实例)与 Terminal-Bench 2.0(89 个任务),输入项包括 release_tag / qwen_release_tag、各套件的数量上限、单实例选择器与 execution_backend。

也就是说**"标准套件上的任务成功率"已经有 harness 了**。缺口比"建一个评测框架"窄得多。

缺的部分

  1. 没有配置覆盖。 该工作流安装一个已发布版本并以默认设置运行,没有任何途径让一次运行携带 settings.json。于是这条伞 issue 真正需要回答的那个问题——当 tools.eager 扣住大部分工具 schema 后,任务成功率是否下降?——根本无法向它提问。加一个把 settings 文件写进运行时 QWEN_HOME 的输入项(以及一个对同一 manifest 跑两遍、一遍带一遍不带的 A/B 模式),就是全部改动。

  2. 每次运行没有 token 与召回画像。 套件报告任务结果,但不报告:

    • 空载成本——一个只问一句、不调用任何工具的会话的输入 token;
    • 每任务总输入 token,这是判断"省下的量是真省还是只是挪进了对话"的关键数字(去掉 grep_search,模型就去用 shell 里的 grep,那些输出永远不命中缓存);
    • 每会话 tool_search 调用次数——路由未命中的代理指标。tracking(core): non-conversation context token governance #12028 的分析在当时的揭示语义下把平衡点定在每会话 ≤2 次;feat(core): preserve prompt cache for deferred tools #10410 改变了那笔账,但这个指标依然是对的那个;
    • 坏调用率——调用了未声明的工具名、或参数校验失败的比例。

    这些都不需要新埋点。/context 自 fix(cli): make /context categories add up to the provider total #12119 起报告带显式残差的分类;ToolCallEvent.function_name(packages/core/src/telemetry/types.ts)已经携带每次调用的工具名;工具调用错误本来就是事件。这是聚合与呈现的工作。

  3. 没有针对"真正被治理的那些会话"的任务集。 SWE-bench 与 Terminal-Bench 是编码套件。一个会话以数据查询和文件操作为主的部署,其召回率回退会出现在那两个套件不看的地方。一组分层的真实形态任务(纯问答、文件操作、Shell、多步)外加少量安全用例,应当与它们并列——这里只作为需求点明,不在此设计。

为什么它是阻塞项

验收

一次 dispatch 的运行,对同一个 commit 的两种配置,给出:

每种配置报告
空载成本 token
每任务输入 token 均值、p50、p95
每会话 tool_search 次数 均值,以及超过 2 次的会话占比
坏调用率 占全部调用的比例
任务成功率 套件自己的通过率
安全用例 通过/失败,必须 100%

并且这份对比要足够清楚,使"我们省掉了 33% 的前缀且没有损失"这句话能被它支持或推翻。

本单不提议

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions