Skip to content

feat(serve): batch workspace session live-state snapshots - #12513

Open
XIQIXIQIXIQI wants to merge 15 commits into
mainfrom
codex/batch-workspace-live-state
Open

XIQIXIQIXIQI wants to merge 15 commits into
mainfrom
codex/batch-workspace-live-state

Conversation

@XIQIXIQIXIQI

@XIQIXIQIXIQI XIQIXIQIXIQI commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Adds one read-only request for complete live-session snapshots from 1–20 explicitly selected registered workspaces. Each workspace keeps its own canonical identity, catalog version, and independent success or error result. The new request shares the existing single-workspace snapshot and catalog-cache reconciliation logic. A distinct capability and TypeScript SDK method let clients opt in with cancellation and timeout support.

Why it's needed

Clients showing several workspaces currently make one live-state HTTP request per workspace on every refresh. The recently added batch session catalog can read persisted history, so it is not a suitable high-frequency execution-state endpoint. This API lets clients fetch the same authoritative in-memory state in one request, while preserving the existing single-workspace route and per-session SSE.

Reviewer Test Plan

How to verify

Start a local daemon with a primary workspace, register a trusted secondary workspace, then send one batch request containing both IDs and an unknown selector. Expect two full snapshots and one explicit workspace_not_found member in input order, with Cache-Control: no-store. Confirm malformed bodies fail with HTTP 400, an untrusted or transitioning workspace fails only its own member, and the capability is advertised. Change a live session between running, waiting, and idle without advancing its catalog version; every batch snapshot should still reflect the current state. Confirm the SDK makes one authenticated native REST request and preserves abort and timeout behavior. Older daemons should continue to use the single-workspace route after callers preflight capabilities.

Evidence (Before & After)

N/A (HTTP and SDK change; no TUI change). The local daemon returned two workspace snapshots and one explicit unknown-workspace error in a single POST.

Tested on

OS Status
🍏 macOS ⚠️ Build, typecheck, bundle, focused tests, and local HTTP E2E passed; the large server test file showed unrelated intermittent failures
🪟 Windows Not tested locally
🐧 Linux ✅ dev x86_64: frozen install/build/bundle, typecheck, focused tests, and daemon HTTP E2E passed

Environment (optional)

Node 22.23.1 on macOS arm64 and dev Linux x86_64, each with an isolated daemon on a dynamically assigned loopback port. The global qwen binary was unavailable for the pre-change CLI dry-run; the absence of this route was confirmed in the current base source.

Risk & Scope

  • Main risk or tradeoff: A batch is not atomic across workspaces. Each successful member is capped at 512 KiB and an oversized member gets an explicit error.
  • Not validated / out of scope: Web Shell still polls workspaces separately until a client migration adopts the new capability; no model-driven live session was created in the HTTP E2E. Windows was not tested locally.
  • Breaking changes / migration notes: None. Clients preflight workspace_session_live_state_batch and retain the existing single-workspace path for older daemons.

Design: English · 简体中文

Linked Issues

Closes #12511

中文说明

本 PR 的改动

新增一个只读请求,一次获取 1–20 个显式选定、已注册 workspace 的完整会话运行状态快照。每个 workspace 独立返回规范化身份、目录版本及成功或错误结果。新请求与单 workspace 接口共用快照生成和目录缓存校准逻辑。独立 capability 和 TypeScript SDK 方法支持客户端按能力接入,并支持取消与超时。

为什么需要

同时展示多个 workspace 的客户端目前每次刷新都要逐 workspace 发起运行状态 HTTP 请求。近期增加的批量会话目录可能读取持久历史,不适合高频查询执行状态。新接口让客户端通过一次请求获取同样的权威内存状态,保留原单 workspace 路由和逐会话 SSE。

评审验证计划

如何验证

启动包含 primary workspace 的本地 daemon,注册可信的 secondary workspace,再用一次批量请求查询两个 ID 和一个未知 selector。预期按输入顺序得到两个完整快照和一个明确的 workspace_not_found 成员,响应包含 Cache-Control: no-store。确认无效请求体返回 HTTP 400,不可信或过渡中的 workspace 只影响自身成员,且 capability 已公布。在目录版本不变时使会话在运行、等待和空闲之间切换,每次快照都应反映当前状态。确认 SDK 只发送一次带鉴权的原生 REST 请求,并保留取消与超时行为。旧 daemon 的调用方应预检能力,继续使用单 workspace 路由。

前后证据

不适用(HTTP 与 SDK 改动,无 TUI 变化)。本地 daemon 实测一次 POST 返回两个 workspace 快照及一个未知 workspace 的明确错误。

测试平台

OS 状态
🍏 macOS ⚠️ 构建、类型检查、打包、聚焦测试和本地 HTTP E2E 通过;大型 server 测试文件出现无关的间歇性失败
🪟 Windows 本地未测试
🐧 Linux ✅ dev x86_64:锁文件安装/构建/打包、类型检查、聚焦测试与 daemon HTTP E2E 均通过

环境

在 macOS arm64 和 dev Linux x86_64 上使用 Node 22.23.1,各自启动使用动态 loopback 端口的隔离 daemon。全局 qwen 命令不可用,无法进行改动前 CLI dry-run;已从当前基线源码确认该路由不存在。

风险与范围

  • 主要风险或取舍:批量结果不保证跨 workspace 原子性;每个成功成员上限为 512 KiB,超限会返回明确错误。
  • 未验证或范围外:Web Shell 在后续客户端迁移前仍逐 workspace 轮询;HTTP E2E 未创建由模型驱动的运行会话;未在本地运行 Windows 测试。
  • 破坏性变更或迁移说明:无。客户端预检 workspace_session_live_state_batch,旧 daemon 保留现有单 workspace 路径。

设计文档:English · 简体中文

关联 Issue

Closes #12511

@github-actions github-actions Bot added the review/self-reported The linked issue was opened by the PR author (self-reported) label Sep 23, 2026
@XIQIXIQIXIQI

Copy link
Copy Markdown
Collaborator Author

E2E test report (macOS arm64, Node 22.23.1)

The bundled CLI at 212a61dc97 started an isolated loopback daemon with --port 0 --no-web. GET /capabilities advertised the new batch capability. One POST /sessions/live-state for a primary workspace, a registered trusted secondary workspace, and an unknown selector returned two complete empty snapshots plus an explicit 404 workspace_not_found member in input order, with Cache-Control: no-store. The existing single-workspace GET returned the matching snapshot. An empty workspace list returned HTTP 400 invalid_session_live_state_batch_request. No JSONL transcript was created, and the daemon stopped cleanly.

npm run build, npm run typecheck, and npm run bundle passed. Focused CLI tests passed (287 tests across four files), SDK tests passed (493 tests across two files), and the capability group in the larger server test file passed (67 tests). Two complete runs of that larger server test file each had one different, unrelated intermittent failure: a registration-capacity response missing limits and an empty session-delete request returning 200 instead of 400. The changed capability group passed in isolation. CI remains the cross-platform check.

The globally installed qwen command was unavailable, so the requested pre-change global-CLI dry-run could not run. The base source at 114be08f has neither the new route nor its capability. The HTTP E2E did not create a model-driven running session; focused bridge tests cover running/waiting/idle transitions with an unchanged catalog version, trust and generation boundaries, and the size limit.

@XIQIXIQIXIQI

Copy link
Copy Markdown
Collaborator Author

Linux validation on dev

I checked out the exact PR head 212a61dc9786aa316f9ba704c03d83c10736ad97 in an isolated directory on Linux x86_64 with Node 22.23.1. corepack pnpm install --frozen-lockfile passed, including the repository's full build and bundle prepare step; root npm run typecheck passed. Focused CLI tests passed (287/287), SDK tests passed (493/493), and the server capability group passed (67/67).

The bundled CLI started an isolated loopback daemon. One batch request returned successful primary and trusted secondary snapshots beside an explicit unknown-workspace 404, preserving order and Cache-Control: no-store. A 20-member request returned HTTP 200; 21 members returned the documented HTTP 400. The existing single-workspace route returned the matching catalog version, no JSONL transcript was created, and the daemon stopped cleanly. The 20-member empty-session request took 1.7 ms in this smoke run; this is not a load benchmark or evidence that a higher limit would be unsafe.

The limit of 20 is a conservative bound aligned with the existing batch catalog. Together with the 512 KiB successful-member cap, it permits roughly 10 MiB of successful entries per request. It is an API policy choice, not a measured threshold at which 21 workspaces fail; clients with more workspaces can split the explicit selection across requests.

@yiliang114 yiliang114 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review: no blocking findings in the implementation itself (see below), but CI Test (ubuntu-latest) is red on this head, and the failure belongs to this PR — so this needs a fix before approval.

The failure

Two existing contract tests fail at 8cc1567, both pointing at the same gap — the new route was added to the code and to qwen-serve-protocol.md, but its companion contract surfaces were not all updated:

  1. telemetry-catalog.test.ts > matches the explicit Express route registrations in both directions — expected […(74)] to have a length of 74 but got 75. The test hard-codes the registered-route count at line 106 (expect(registered).toHaveLength(74)); the new POST /sessions/live-state registration makes it 75. (The sibling hard-code in server/telemetry.test.ts was updated to 75; this one was missed.)

  2. rest-integration-docs-contract.test.ts > indexes every operation with a dedicated protocol section — expected […(76)] to deeply equal […(77)]. qwen-serve-protocol.md gained the ### \POST /sessions/live-state`section (77 protocol operations), butdocs/developers/daemon-rest-api-reference.md` has no link row for it (76 reference links).

Fixing #2 is not just adding a row to the reference table: the neighboring test keeps the published reference index in step with the OpenAPI document requires rows.length === operations.size and every row's capability/scope/sdk-method cells to equal the OpenAPI operation's extension fields. So the consistent fix is three surfaces in sync:

  • docs/developers/daemon-rest-api.openapi.json — add the POST /sessions/live-state operation, with x-qwen-capability: workspace_session_live_state_batch, the matching x-qwen-scope, x-qwen-sdk-method (getSessionsLiveState), and externalDocs pointing at the new protocol anchor.
  • docs/developers/daemon-rest-api-reference.md — add the corresponding table row (link + the three cells).
  • packages/cli/src/serve/server/telemetry-catalog.test.ts:106 — 74 → 75.

What I verified in the implementation (all clean)

  • The batch handler is strict and well-isolated: zod .strict() envelope (1–20 selectors, 4096 chars, extra keys rejected), per-member error isolation with the right status mapping, a TOCTOU re-check that fails closed to 503 when the generation is replaced mid-read (and a test that exercises exactly that), sequential reads with req.aborted/res.destroyed bail-outs, 512 KiB member cap, Cache-Control: no-store.
  • The trust model is the strict one: untrusted primary and secondary get 403 — it does not inherit the permissive persisted-catalog policy, and unknown/internal selectors never fall back to primary. Internal-workspace exclusion has its own test asserting the bridge is never read.
  • The extraction of readWorkspaceLiveState from the single-workspace handler is behavior-identical, and both routes sharing lastExposedCatalogVersions preserves the catalog-reconciliation contract (the amended 7447-region test pins this).
  • Rate-limit classification (read tier, anchored regex), telemetry attribution (handler_resolved, per-member workspace hash), capability registration, and the SDK method (one native REST request, no fan-out, signal/timeoutMs not serialized, older-daemon 404 surfaces as-is) all match the design doc and have tests.
  • The protocol and SDK docs, capabilities lists, and integration test were updated consistently.

Non-blocking notes for when you undraft: macOS intermittent failures you mentioned in the body are worth a second look if they reappear in CI, and there is no model-driven live session in the HTTP E2E (the in-memory read path is covered by the bridge mocks, so this is acceptable). The 512 KiB member check serializes twice (measure + res.json) — bounded and fine.

@doudouOUC doudouOUC left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks — this is in good shape structurally: the trust gate is strict per member (unknown/internal → 404, untrusted primary and secondary → 403, generation replaced mid-read → 503, never a fallback to primary), the shared projection/cache-invalidation contract with the single-workspace route is preserved and test-pinned, and the bounds (1–20 selectors, 4096 chars, 512 KiB per member with an explicit 413, no-store, sequential reads with disconnect bail-outs) all behave as the design doc says. The SDK method is a single native REST request with cancellation and no hidden fan-out. Two things before this can land:

  1. CI is red at 8cc1567 and the failure belongs to this PR (already flagged in the previous review round, re-verified at head): telemetry-catalog.test.ts:106 still expects 74 registered routes while the sibling server/telemetry.test.ts was bumped to 75, and the docs contract test fails because POST /sessions/live-state has a protocol section but no operation in daemon-rest-api.openapi.json (with its x-qwen-capability / x-qwen-scope / x-qwen-sdk-method cells) and no row in daemon-rest-api-reference.md.
  2. The rebase changes the numbers. Since this branch's base, main landed POST /session/:id/mcp-app/tools/call, so current main already counts 75 routes — after rebasing, the combined tree is 76, and both telemetry tests (telemetry-catalog.test.ts and telemetry.test.ts) should assert 76 with the 74/2 attribution split. The capability baselines (server.test.ts, integration-tests/cli/qwen-serve-routes.test.ts) likewise need the union with the five capabilities main gained (daemon_update, hosted_harness_private_v1, session_branch_worktree, session_startup_config, workspace_git_worktrees). Resolving the conflict as a plain "keep both" without this adjustment will land red.

Non-blocking: the member size check serializes twice (measure + res.json) — bounded and fine as-is; and the macOS intermittent failures you noted in the PR body are worth a second look if they reappear in CI.


结构上没有大问题:信任门禁按成员严格生效(未知/内部 → 404,不可信的 primary 和 secondary 都是 403,读取途中 generation 被替换 → 503,任何情况下都不回退到 primary),与单 workspace 路由共用的投影与缓存失效契约保持且有测试钉住,资源边界(1–20 个 selector、4096 字符、单成员 512 KiB 显式 413、no-store、顺序读取 + 断连退出)与设计文档一致。SDK 方法是单次原生 REST 请求,支持取消,没有隐藏的扇出。合入前有两件事:

  1. CI 在当前 head(8cc1567)是红的,且失败属于本 PR(上一轮评审已指出,我在 head 上重新核实):telemetry-catalog.test.ts:106 仍断言 74 条注册路由,而姊妹文件 server/telemetry.test.ts 已改为 75;文档契约测试失败是因为 POST /sessions/live-state 有协议章节,但 daemon-rest-api.openapi.json 中没有对应 operation(含 x-qwen-capability / x-qwen-scope / x-qwen-sdk-method 三个扩展字段),daemon-rest-api-reference.md 中也没有对应行。
  2. rebase 会改变数字。 自本分支的 base 以来,main 已合入 POST /session/:id/mcp-app/tools/call,当前 main 的路由计数已是 75——rebase 之后合并树是 76 条,两个遥测测试(telemetry-catalog.test.ts 与 telemetry.test.ts)都应断言 76(74/2 归属比例)。能力基线(server.test.ts、integration-tests/cli/qwen-serve-routes.test.ts)同样需要与 main 新增的五个能力(daemon_update、hosted_harness_private_v1、session_branch_worktree、session_startup_config、workspace_git_worktrees)取并集。如果按简单的「两边都保留」解冲突而不做这些调整,合入后依然是红的。

非阻塞:成员大小检查会序列化两次(测量 + res.json)——有界,保持现状即可;另外你在 PR 描述里提到的 macOS 间歇性失败,如果 CI 复现值得再看一眼。

@wenshao
wenshao marked this pull request as ready for review September 29, 2026 05:12
@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /resolve

Both sides added one handler-resolved legacy session telemetry route: main added POST /session/:id/mcp-app/tools/call and this branch added POST /sessions/live-state. Each independently bumped the audited catalog from 74/72 to 75/73, so Git merged the identical assertion text cleanly while the merged catalog actually holds 76 routes at 74/2. Re-audit the counts and widen the explanatory comment to cover both routes.
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

Qwen Code resolved the merge conflicts and pushed the branch update.

Root cause

main 81c260bb04 (#12258) added POST /session/:id/mcp-app/tools/call to legacySessionTelemetryRoutes (handler_resolved), bumping the audited counts 74/72 → 75/73. PR #12513 did the same for POST /sessions/live-state: two +1 route additions on one shared counter.

Semantic, not textual

The only marker was #12258's new comment (HEAD side empty). Git auto-merged the numbers — both sides made the byte-identical edit 74→75 / 72→73 — and that merge is wrong: the catalog now holds both routes. Merged telemetry.ts: 76 entries, 74 handler_resolved, 2 pre_resolved, 76 unique keys. Kept #12258's comment (widened to both routes) and both sides' tests.

it('contains 76 unique routes with the audited 74/2 attribution split', () => {
  expect(keys).toHaveLength(76);
  expect(new Set(keys).size).toBe(76);
  /* handler_resolved */ ).toHaveLength(74);
  /* pre_resolved     */ ).toHaveLength(2);   // unchanged

Load-bearing

  • total = handler_resolved + pre_resolved (76 = 74 + 2). pre_resolved stays 2 — neither new route is pre-resolved.
  • The title 76 … 74/2 is hardcoded prose: bump it with the three assertions or the name lies.
  • resolveDaemonTelemetryRoute is first-match-wins; POST /sessions/live-state (2 segments) cannot collide with POST /session/:id/… (3), so its slot at the array head is safe.

Follow-up needed — not done here (file did not conflict)

packages/cli/src/serve/server/telemetry-catalog.test.ts:106 still asserts toHaveLength(75). That drift guard builds a real Express app from registerSessionRoutes and compares it to the catalog. The merged catalog is 76 and POST /sessions/live-state registers at body level (session.ts:9464, no early return before it), so registered is 76 → the length assertion fails while toEqual(catalog) passes. Needs 75 → 76. Pre-existing on the PR branch: #12258 bumped it, the PR never touched the file, so PR HEAD asserted 74 against a 75-entry catalog.

No build, lint or tests run.

中文说明

根因:main 的 81c260bb04(#12258)向 legacySessionTelemetryRoutes 新增 handler_resolved 条目 POST /session/:id/mcp-app/tools/call,把审计计数 74/72 → 75/73;PR #12513 为 POST /sessions/live-state 做了同样改动:两侧各 +1,共用一个计数器。

语义冲突,非文本冲突:唯一标记是 #12258 新增的注释(HEAD 侧为空)。数字被 Git 自动合并——两侧做了逐字节相同的 74→75 / 72→73——而该合并是错的:目录现同时含两条路由。合并后 telemetry.ts:76 条,handler_resolved 74,pre_resolved 2,76 个唯一键。 保留了 #12258 的注释(扩写覆盖两条)及双方新增测试。

关键约束:总数 = handler_resolved + pre_resolved(76 = 74 + 2);pre_resolved 保持 2,两条新路由都非预解析。标题 76 … 74/2 是硬编码文本,须与三处断言同步更新,否则测试名失真。resolveDaemonTelemetryRoute 首次匹配即返回;POST /sessions/live-state(2 段)不会与 POST /session/:id/…(3 段)冲突,故排首位安全。

需后续修改——本次未处理(该文件未冲突):packages/cli/src/serve/server/telemetry-catalog.test.ts:106 仍断言 toHaveLength(75)。该守卫用 registerSessionRoutes 构建真实 Express 应用并与目录比对;合并后目录为 76,且 POST /sessions/live-state 注册在函数体层级(session.ts:9464,其前无提前 return),故 registered 为 76 → 长度断言失败,而 toEqual(catalog) 通过。需把 75 改为 76。此问题在 PR 分支上已存在:#12258 更新过它,PR 从未改动该文件,故 PR HEAD 用 74 断言 75 条目录。

未运行构建或测试。

@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /takeover

@qwen-code-dev-bot qwen-code-dev-bot added the autofix/takeover Summon the autofix loop to manage this PR (remove to release; needs triage+) label Sep 30, 2026
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤝 Takeover engaged: the autofix loop now manages this PR — it will address new review feedback and resolve base conflicts until the label is removed or the round cap is reached. Remove the autofix/takeover label (or comment @qwen-code /takeover stop) to release.

中文说明

🤝 已接管:autofix 循环现在管理此 PR —— 将持续处理新的评审反馈与 base 冲突,直到移除标签或达到轮次上限。移除 autofix/takeover 标签(或评论 @qwen-code /takeover stop)即可释放。

@qwen-code-dev-bot

qwen-code-dev-bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

✅ AutoFix round 9 finished — view run. See this round's report below.

中文说明

✅ AutoFix 第 9 轮已完成 —— 查看运行。本轮报告见下方。

Resolve conflicts by keeping the extracted readWorkspaceLiveState helper and
porting main's agent-host sourceType filter into it so both live-state routes
share the behavior, and reconcile the telemetry route guards to the combined
80 routes with the 78/2 attribution split.

Co-authored-by: Qwen-Coder <[email protected]>
…12513)

Index POST /sessions/live-state in the REST reference's persisted-catalog
row so the docs contract sees the protocol section, give the batch route the
single-workspace route's server-side records (routing-failure warn on
untrusted members, daemonLog error on unexpected bridge failures), make the
mid-batch disconnect guard observable with an AbortController and a
macrotask yield between members, and move the duplicated member-failure,
generation-currency, and bound literals into workspace-route-runtime for
both plural batch routes. Six tests pin the new witnesses.

Co-authored-by: Qwen-Coder <[email protected]>
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Addressed the latest review feedback (round 1/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 1/100 轮)。改动内容与我反驳保留之处如下:

Autofix round — PR #12513 (batch workspace session live-state)

Two commits this round: a merge of main (--conflict true) resolving two content conflicts, and one fix commit addressing all seven review findings.

Base-conflict resolution (68cd30c8a3)

  • packages/cli/src/serve/routes/session.ts — kept the PR's extracted readWorkspaceLiveState helper and ported main's AGENT_HOST_SESSION_SOURCE_TYPE filter (from 8c914ebe03, which hid daemon-owned agent-host sessions from the single-workspace live-state route) into the shared helper, so both live-state routes keep that behavior. Main's own test omits the hidden agent host without changing the catalog version pins it: removing the filter from the helper turns that test red (probed), restoring it returns green.
  • packages/cli/src/serve/server/telemetry.test.ts — union of both sides: 80 catalog routes with the 78/2 handler/pre-resolved attribution split (branch's 76 + main's 4 attachment-upload routes).
  • Semantic merge fallout: main had bumped the auto-merged telemetry-catalog.test.ts drift guard to 79; the combined tree registers 80 explicit routes, so the guard now asserts 80 (mutation probe: 79 → red "got 80", 80 → green). This supersedes the 75 → 76 instruction from earlier rounds, which was correct at 8cc1567 and stale after both merges.

Feedback dispositions

  • [rc:4132272367] (Critical, route-count drift guard) — Resolved. The guard asserts the combined count 80 and still asserts set equality with the catalog in both directions; no second catalog entry added, count kept a hardcoded literal. Probe: reverting the literal to 79 turns the guard red.
  • [rc:4132272373] (Critical, docs contract) — Resolved exactly as measured in the finding: a secondary link [POST /sessions/live-state](./qwen-serve-protocol.md#post-sessionslive-state) inside the existing "Persisted session catalog" row of daemon-rest-api-reference.md, with the row's capability cell extended by workspace_session_live_state_batch (selected-runtime per member, trusted-only, memory-only) and the SDK cell by getSessionsLiveState. No OpenAPI operation and no new | [ row: rest-integration-docs-contract.test.ts filters index rows by /^\| \[/ and asserts rows.length === operations.size, the OpenAPI document contains no plural /sessions/* path at all, and the shipped POST /sessions/catalog precedent is indexed exactly this way (secondary link in that row, no OpenAPI operation). The anchor post-sessionslive-state matches the contract's slug(). Baseline before the fix: indexes every operation with a dedicated protocol section red (missing link); after: green, and keeps the published reference index in step with the OpenAPI document stays green. Note for the earlier human reviews (rv:5286726086, rv:5347072572): their prescribed shape (OpenAPI operation + dedicated row + extension cells) also satisfies the tests, but the secondary-link shape is the repo's established convention for plural batch routes and the minimal change; the capability, scope, and SDK method are still published in the same row that carries the single-workspace live-state entry.
  • [rc:4132272383] (Suggestion, duplicated batch scaffolding) — Resolved. failedBatchMember, captureWorkspaceEntryCurrency (the four-condition generation-currency predicate), and the shared bounds MAX_BATCH_WORKSPACES / MAX_SELECTOR_CHARS / MAX_MEMBER_BYTES now live in packages/cli/src/serve/workspace-route-runtime.ts, beside resolveWorkspaceEntryBySelector and isGenerationClosedError; both session-catalog.ts and the live-state batch route consume them. The deliberately different trust gates (runtime.primary && !runtime.trusted vs !runtime.trusted) and read strategies (concurrency-4 pool with plumped AbortSignal vs sequential reads) stay caller-side, as the finding required. Both sibling suites (generation-replacement and transitioning-member cases on both routes) stay green.
  • [rc:4132272387] (Suggestion, dropped server-side records) — Resolved. The 403 arm now emits logSessionRoutingFailure('POST /sessions/live-state', 'untrusted_workspace', { workspaceId, workspaceCwd }) (matching the five existing call sites), and the catch's 500 arm logs the raw error via daemonLog?.error(message, error, { route, workspaceId }) before returning the generic member. The documented member bodies are unchanged, the raw error never enters error.message, and the 404 path stays silent so polling a removed workspace does not warn per poll. Two new tests pin both records.
  • [rc:4132272395] (Suggestion, unexecuted failure arms) — Resolved with the two prescribed witnesses. returns a generic 500 member for an unexpected bridge failure without leaking details (plain Error from listWorkspaceSessions; asserts the 500 session_live_state_failed member, no sessions key, no internal-message leak, healthy sibling member) and fails closed if the selected runtime starts draining during its read (beginDrain inside the read, asserting 503 workspace_runtime_unavailable — this state has state = 'draining' with the guard still open, so only the post-read re-check produces the 503). Probes: collapsing the 500 arm to return unavailable() reddens the first; deleting the post-read if (!isCurrent()) return unavailable(); reddens the second.
  • [rc:4132272406] (Suggestion, unobservable disconnect guard) — Resolved with option (b), matching the sibling catalog route: req.once('aborted') / res.once('close') (gated on !res.writableEnded) wired into an AbortController, a synchronous pre-loop abort check, one setImmediate macrotask yield between members, and a pre-res.json aborted check, with listener cleanup in finally. Reads remain sequential and bridge-memory-only per the protocol and design docs; the design doc's "synchronously from bridge memory" sentence was reworded in both languages to state the between-member yield. The pinning test destroys the server sockets during the first member's read over a real socket and asserts the third member's bridge is never read — the second member may complete because socket-close delivery lags the destroy by a turn or two (the finding's own measurement), which the test's comment records. Probe: removing the listener wiring and the yield reddens it.
  • [rc:4132272415] (Suggestion, unpinned telemetry attribution) — Resolved. The test file now mocks the daemon-tracing seam with a pass-through default (identical to the real SDK-uninitialized behavior) and asserts: one request-scoped qwen-code.daemon.session_live_state_batch.members write with value 2 (scope undefined — outside the per-member span), one ...batch.member span per member, and each span's qwen-code.workspace.hash equal to hashDaemonWorkspace of that member's cwd in request order. Probes: deleting the attribute write reddens it; removing the per-member span reddens it.
  • [rv:5351024966] (CHANGES_REQUESTED, review gap) — The disclosed gap was that integration-tests/cli/qwen-serve-routes.test.ts had never been executed. It now ran locally against the bundled CLI: 42/42 passed, including the workspace_session_live_state_batch capability assertion.
  • [rv:5286726086], [rv:5347072572] (COMMENTED) — Both reviews' blocking substance (red CI on the route-count guard and the docs contract, and the post-merge recount to the union with main's routes and capabilities) is resolved as above. Their non-blocking notes (double serialization, macOS flakiness watch) need no change.
  • Issue-level E2E reports and the serve A/B comment — informational, no action.

Verification

  • npm run build — passed (also covers the merge resolution and the extraction's types)
  • npm run typecheck — passed
  • npm run lint — passed
  • npx prettier --check on all touched files — passed after --write reformatted the reference table (mandatory re-pad: the file is prettier-gated and the capability cell grew the column) and the test file
  • npx vitest run src/serve/multi-workspace-sessions.test.ts src/serve/routes/session-catalog.test.ts src/serve/server/telemetry-catalog.test.ts src/serve/server/telemetry.test.ts src/serve/rest-integration-docs-contract.test.ts (packages/cli) — 318 passed
  • npx vitest run src/serve/server.test.ts (packages/cli) — 1369 passed (capability baselines, union with main's five new capabilities)
  • npx vitest run cli/qwen-serve-routes.test.ts (integration-tests, after npm run bundle, QWEN_SANDBOX=false) — 42 passed
  • Mutation probes (each mutated, observed red, restored, re-greened): route-count literal 80→79; agent-host filter removed from the shared helper; 403 routing warn removed; 500 daemonLog.error removed; 500 arm collapsed to unavailable(); post-read currency re-check deleted; disconnect listener wiring + yield removed; members attribute write removed; per-member span removed. Each probe reddened exactly its witness test.
  • Pre-round failure evidence for the gate: on the pre-round head 2a80c8dd, this round's changed tests fail — telemetry-catalog.test.ts (asserts 80, registered 76), telemetry.test.ts (asserts 80/78/2, catalog has 76), and in multi-workspace-sessions.test.ts the two logging tests and the disconnect test (no logging calls, no cancellation wiring existed).

npm run generate:settings-schema was not needed: no settings source changed. Self-review was not armed for this round (Self-review: off).

中文说明

Autofix 本轮处理 — PR #12513(批量 workspace 会话运行状态路由)

本轮共两个提交:一次按 --conflict true 要求合并 main 并解决两处内容冲突,以及一个处理全部七条评审发现的修复提交。

基线冲突解决(68cd30c8a3)

  • packages/cli/src/serve/routes/session.ts —— 保留本 PR 抽取的 readWorkspaceLiveState 辅助函数,并把 main 的 AGENT_HOST_SESSION_SOURCE_TYPE 过滤(来自 8c914ebe03,该提交使单 workspace 运行状态路由不再暴露 daemon 自有的 agent-host 会话)移植进这个共享辅助函数,使两条运行状态路由都保持该行为。main 自带的测试 omits the hidden agent host without changing the catalog version 钉住了这一点:探针实测从辅助函数中移除该过滤会让该测试变红,恢复后回到绿色。
  • packages/cli/src/serve/server/telemetry.test.ts —— 取两侧并集:80 条目录路由,归属比例为 78/2(handler_resolved/pre_resolved)(分支的 76 条 + main 新增的 4 条附件上传路由)。
  • 合并的语义性后续:main 已把自动合并的 telemetry-catalog.test.ts 漂移守卫推进到 79;合并后的树实际注册 80 条显式路由,因此该守卫现在断言 80(变异探针:改为 79 → 红,报 "got 80";改回 80 → 绿)。这取代了前几轮评审给出的 75 → 76 指示——该指示在 8cc1567 上是正确的,两次合并后已过期。

评审发现处置

  • [rc:4132272367](Critical,路由计数漂移守卫) —— 已解决。守卫断言合并后的计数 80,并继续在两个方向上与目录断言集合相等;未新增第二条目录条目,计数保持硬编码字面量。探针:把字面量改回 79 会使守卫变红。
  • [rc:4132272373](Critical,文档契约) —— 完全按该发现的实测结论解决:在 daemon-rest-api-reference.md 已有的「Persisted session catalog」行内加入次级链接 [POST /sessions/live-state](./qwen-serve-protocol.md#post-sessionslive-state),该行的 capability 单元格补充 workspace_session_live_state_batch(按成员的 selected-runtime、仅可信、仅内存),SDK 单元格补充 getSessionsLiveState。不新增 OpenAPI operation、不新增 | [ 行:rest-integration-docs-contract.test.ts 用 /^\| \[/ 过滤索引行并断言 rows.length === operations.size;OpenAPI 文档中完全没有任何复数形式的 /sessions/* 路径;已合入的 POST /sessions/catalog 先例正是以这种方式索引的(该行内的次级链接,无 OpenAPI operation)。锚点 post-sessionslive-state 与契约测试的 slug() 一致。修复前基线:indexes every operation with a dedicated protocol section 红(缺链接);修复后变绿,且 keeps the published reference index in step with the OpenAPI document 保持绿色。给前两轮人工评审(rv:5286726086、rv:5347072572)的说明:他们给出的形态(OpenAPI operation + 独立行 + 扩展字段单元格)同样能通过测试,但次级链接形态是仓库对复数批量路由的既定惯例、且改动最小;capability、scope 与 SDK 方法仍然发布在承载单 workspace 运行状态条目的同一行中。
  • [rc:4132272383](Suggestion,批量脚手架重复) —— 已解决。failedBatchMember、captureWorkspaceEntryCurrency(四条件的 generation 现势性判定)以及共享上限 MAX_BATCH_WORKSPACES / MAX_SELECTOR_CHARS / MAX_MEMBER_BYTES 现在位于 packages/cli/src/serve/workspace-route-runtime.ts,就在 resolveWorkspaceEntryBySelector 与 isGenerationClosedError 旁边;session-catalog.ts 与运行状态批量路由都改为消费它们。两处有意不同的信任门禁(runtime.primary && !runtime.trusted 对 !runtime.trusted)和读取策略(并发 4 的读取池并下穿 AbortSignal 对顺序读取)按发现要求保留在调用方。两条姊妹路由的测试套件(各自的 generation 替换与 transitioning 成员用例)保持绿色。
  • [rc:4132272387](Suggestion,丢失的服务端记录) —— 已解决。403 分支现在发出 logSessionRoutingFailure('POST /sessions/live-state', 'untrusted_workspace', { workspaceId, workspaceCwd })(与现有五处调用点一致);catch 的 500 分支在返回泛化成员之前,通过 daemonLog?.error(message, error, { route, workspaceId }) 记录原始错误。有文档约束的成员响应体不变,原始错误绝不进入 error.message,404 路径保持沉默,以免轮询已移除 workspace 的客户端每次轮询都产生一条警告。两个新测试钉住这两条记录。
  • [rc:4132272395](Suggestion,从未执行的失败分支) —— 已解决,补齐规定的两个见证用例。returns a generic 500 member for an unexpected bridge failure without leaking details(listWorkspaceSessions 抛普通 Error;断言 500 session_live_state_failed 成员、无 sessions 键、不泄露内部消息、兄弟成员健康)与 fails closed if the selected runtime starts draining during its read(在读取中调用 beginDrain,断言 503 workspace_runtime_unavailable——该状态下 state = 'draining' 且 guard 未关闭,因此只有读取后复检能产生这个 503)。探针:把 500 分支塌缩为 return unavailable() 会让第一个变红;删除读取后的 if (!isCurrent()) return unavailable(); 会让第二个变红。
  • [rc:4132272406](Suggestion,不可观测的断连守卫) —— 按方案 (b) 解决,与姊妹 catalog 路由一致:req.once('aborted') / res.once('close')(以 !res.writableEnded 为条件)接入 AbortController,循环前的同步中止检查,成员之间让出一个 setImmediate 宏任务,res.json 前的中止检查,并在 finally 中移除监听器。成员读取仍按协议与设计文档保持顺序执行、仅读 bridge 内存;设计文档中「从 bridge 内存同步读取」一句已在两种语言中改写,说明成员之间的让出。钉住该行为的测试通过真实 socket 在第一个成员读取期间销毁服务端 socket,并断言第三个成员的 bridge 从未被读取——第二个成员可能仍会完成,因为 socket close 事件的投递比 destroy 晚一到两个轮次(该发现自带的实测结论),测试注释中记录了这一点。探针:移除监听器接线与让出会让其变红。
  • [rc:4132272415](Suggestion,遥测归属未被钉住) —— 已解决。测试文件现在以穿透默认实现 mock 了 daemon-tracing 接缝(与真实 SDK 未初始化时的行为完全一致),并断言:一次请求作用域的 qwen-code.daemon.session_live_state_batch.members 写入、值为 2(scope 为 undefined,即在按成员 span 之外);每个成员一个 ...batch.member span;每个 span 的 qwen-code.workspace.hash 按请求顺序等于对应成员 cwd 的 hashDaemonWorkspace。探针:删除该属性写入会变红;移除按成员 span 会变红。
  • [rv:5351024966](CHANGES_REQUESTED,审查缺口) —— 所披露的缺口是 integration-tests/cli/qwen-serve-routes.test.ts 从未被执行。现在已基于打包后的 CLI 在本地运行:42/42 通过,包括 workspace_session_live_state_batch 能力断言。
  • [rv:5286726086]、[rv:5347072572](COMMENTED) —— 两份评审的阻塞性实质内容(路由计数守卫与文档契约导致的 CI 红,以及合并后与 main 新增路由和能力取并集的重新计数)已按上述解决。其非阻塞备注(双重序列化、macOS 抖动观察)无需改动。
  • issue 级 E2E 报告与 serve A/B 评论 —— 信息性内容,无需处理。

验证

  • npm run build —— 通过(同时覆盖合并冲突解决与抽取改动的类型)
  • npm run typecheck —— 通过
  • npm run lint —— 通过
  • 对所有改动文件执行 npx prettier --check —— 在 --write 重排参考文档表格后通过(强制重排:该文件受 prettier 门禁约束,capability 单元格变长撑宽了列)并格式化了测试文件
  • npx vitest run src/serve/multi-workspace-sessions.test.ts src/serve/routes/session-catalog.test.ts src/serve/server/telemetry-catalog.test.ts src/serve/server/telemetry.test.ts src/serve/rest-integration-docs-contract.test.ts(packages/cli)—— 318 通过
  • npx vitest run src/serve/server.test.ts(packages/cli)—— 1369 通过(能力基线,与 main 新增的五个能力取并集)
  • npx vitest run cli/qwen-serve-routes.test.ts(integration-tests,在 npm run bundle 之后,QWEN_SANDBOX=false)—— 42 通过
  • 变异探针(每次变异后确认变红、恢复后回到绿色):路由计数字面量 80→79;从共享辅助函数中移除 agent-host 过滤;移除 403 路由警告;移除 500 的 daemonLog.error;把 500 分支塌缩为 unavailable();删除读取后现势性复检;移除断连监听器接线与让出;移除成员计数属性写入;移除按成员 span。每个探针恰好使其对应的见证测试变红。
  • 供验证门禁使用的本轮前失败证据:在本轮前的 head 2a80c8dd 上,本轮改动的测试会失败——telemetry-catalog.test.ts(断言 80,实际注册 76)、telemetry.test.ts(断言 80/78/2,目录为 76),以及 multi-workspace-sessions.test.ts 中的两个日志测试和断连测试(当时不存在日志调用与取消接线)。

未运行 npm run generate:settings-schema:没有 settings 源文件变更。本轮未启用轮内自审(Self-review: off)。

Base-conflict check · 基分支冲突检查: conflicted with main — resolved in this push. · 与 main 有冲突——已在本次推送中解决。

🧵 Resolved all 7 selected review thread(s). · 已关闭全部选中的 7 条评审线程。

Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.7

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Addressed the latest review feedback (round 2/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 2/100 轮)。改动内容与我反驳保留之处如下:

Autofix round summary — PR #12513

All six inline suggestions from the round-2 review are addressed in commit 39c2624 (fix(serve): align live-state scope docs and pin batch route guards (#12513)): one documentation correction, one new documentation contract test, and three new route tests that separately pin the three previously unwitnessed guards. The two convergence-deferred items in the review body were recorded by the reviewer as not requested this round and are untouched.

Findings and dispositions

R2-1 (rc:4152008465) — contradictory scope labels in the grouped catalog row — implemented

docs/developers/daemon-rest-api-reference.md: the Persisted session catalog row's capability cell now gives both live-state operations one scope —

`session_info`, `session_list`; persisted-workspace; `session_catalog_batch`; persisted-workspace per member; `workspace_session_live_state`, `workspace_session_live_state_batch`; selected-runtime (batch: per member), trusted-only, memory-only

Verified against the code before editing: readWorkspaceLiveState reads only bridge memory (getSessionCatalogVersion / listWorkspaceSessions / invalidateSessionLists), and both live-state routes go through trusted-only resolution — matching the protocol doc's "trusted-only, memory-only snapshot" (qwen-serve-protocol.md:316) and "the batch uses the same trusted, memory-only snapshot and catalog-version semantics as the single-workspace route" (:318). The row's first cell stays an Area name, so the OpenAPI row-count contract is undisturbed. Prettier re-aligned the rest of the table's padding (the file is prettier-gated and HEAD is prettier-clean); after whitespace normalization the only content change is this one cell.

R2-2 (rc:4152008469) — grouped rows' capability/SDK cells unpinned — implemented

rest-integration-docs-contract.test.ts gains pins the grouped reference rows to registered capabilities and SDK methods: for every Area-named row (the shape the OpenAPI row check cannot see), each backticked snake_case token in the "Capability and scope" cell must be a key of SERVE_CAPABILITY_REGISTRY, and each backticked DaemonClient.<name> / WorkspaceDaemonClient.<name> token in the "TypeScript SDK" cell must exist on the class prototype. Candidate selection is deliberate — scope prose ("persisted-workspace per member") and "raw REST" notes share those cells and are not backticked.

Probe: corrupting `workspace_session_live_state_batch` in the doc fails the new test with Persisted session catalog: \workspace_session_live_state_batched` is not a registered serve capability`; restored and re-green.

R2-3 (rc:4152008472) — envelope assertion only reachable via CI integration jobs — verified, no source change

Ran the exact gate locally against a freshly built bundle: the full integration-tests/cli/qwen-serve-routes.test.ts suite (42/42, which includes the qwen serve — capabilities envelope describe asserting workspace_session_live_state_batch is advertised by a real spawned daemon), and that case alone (1 passed). The CI integration jobs on the merge head remain the final gate; this confirms the bundled CLI's envelope carries the tag.

R2-4 (rc:4152008477) — post-loop abort guard unwitnessed — implemented

New test skips the response when the client disconnects during the last member read: a one-member batch over a real server; the member read is held on a gate inside the withDaemonSpan mock while the sockets are destroyed, and released only after the server-side response close is observed — so the post-loop if (controller.signal.aborted) return; (not the in-loop check) decides. Asserts the response json write never happens while the member read itself completes (listCalls is [PRIMARY_CWD] either way, as documented).

Probe: deleting the post-loop guard turns exactly this test red; stops reading members once the client disconnects mid-batch and the pre-handler test stay green — the in-loop and post-loop arms are separately pinned. Guard restored.

R2-5 (rc:4152008481) — !generation pre-read guard unwitnessed — implemented

New test returns 503 for a member whose entry lost its current generation while preserving a healthy member: clears entry.current and calls registry.blockReplacement(entry, 'apply failed') — the exact state the trust reconciler produces at workspace-trust-reconciler.ts:230-231 — keeping the entry resolvable, then asserts the 503 workspace_runtime_unavailable member carries no sessions, the secondary bridge is never read, and the healthy member is preserved.

Probe: deleting if (!generation || !isCurrent()) return unavailable(); turns exactly this test red (the whole request 500s instead of returning an isolated 503 member); the other 17 batch tests stay green. Guard restored.

R1-6 (rc:4152008490) — pre-handler abort guard unwitnessed — implemented

New test skips every member read when the client disconnects before the handler runs: an outer express app parses the body while the socket is alive, destroys the socket, and dispatches to the real serve app only after the response close is delivered (the serve app's own express.json() skips the re-parse because the body is already read), so the handler observes an already-destroyed response. Asserts zero bridge reads on either workspace.

The first version of this test destroyed the socket before the body was parsed, so the request 400'd ahead of the guard and the probe stayed green — caught by the probe, exactly as intended. With the corrected construction the probe is deterministically red (2/2 runs): deleting if (req.aborted || res.destroyed) abort(); lets the handler run and read (primaryBridge.listCalls becomes ['/work/primary']), while with the guard both bridges stay unread. Guard restored.

Verification

  • npm run build — passed (needed COREPACK_HOME=/tmp/corepack-cache; the sandbox denies the default corepack cache path under $HOME/.cache)
  • npm run typecheck — passed
  • npm run lint — passed
  • vitest run src/serve/multi-workspace-sessions.test.ts (packages/cli) — 185 passed
  • vitest run src/serve/multi-workspace-sessions.test.ts -t 'batch workspace session live-state route' — 18 passed (run twice for flakiness; stable)
  • vitest run src/serve/rest-integration-docs-contract.test.ts (packages/cli) — 14 passed
  • vitest run src/serve/capabilities-docs-contract.test.ts src/serve/rate-limit.test.ts src/serve/server/telemetry-catalog.test.ts src/serve/server/telemetry.test.ts (packages/cli) — 116 passed
  • vitest run src/serve/server.test.ts (packages/cli) — 1369 passed
  • npm run bundle — passed
  • QWEN_SANDBOX=false vitest run --root ./integration-tests cli/qwen-serve-routes.test.ts — 42 passed; -t 'capabilities envelope' — 1 passed
  • Mutation probes (each restored and re-verified green after the red run):
    • delete the !generation || !isCurrent() pre-read guard → only the new R2-5 test fails (17 batch tests green)
    • delete the post-loop if (controller.signal.aborted) return; → only the new R2-4 test fails (mid-batch disconnect and pre-handler tests green)
    • delete the if (req.aborted || res.destroyed) abort(); pre-handler guard → only the new R1-6 test fails, deterministically (2/2 runs)
    • corrupt one capability token in the doc → only the new grouped-row contract test fails
中文说明

Autofix 本轮总结 — PR #12513

第 2 轮评审的 6 条行内建议全部在提交 39c2624(fix(serve): align live-state scope docs and pin batch route guards (#12513))中处理完毕:一处文档更正、一个新的文档契约测试、以及三个分别钉住此前无见证守卫的新路由测试。评审正文中按收敛姿态延后的两条由评审者记录为本轮不要求处理,未触碰。

发现与处置

R2-1(rc:4152008465)——分组目录行中的 scope 标签自相矛盾——已实现

docs/developers/daemon-rest-api-reference.md:Persisted session catalog 行的 capability 单元格现在给两个 live-state operation 统一的 scope(单元格文本见上方英文部分)。编辑前已对照代码验证:readWorkspaceLiveState 只读 bridge 内存(getSessionCatalogVersion / listWorkspaceSessions / invalidateSessionLists),两条 live-state 路由都走 trusted-only 解析——与协议文档「trusted-only, memory-only snapshot」(qwen-serve-protocol.md:316)及「批量路由使用与单 workspace 路由相同的可信、仅内存快照与目录版本语义」(:318)一致。该行首单元格保持为 Area 名称,OpenAPI 行数契约不受影响。表格其余部分因 Prettier 重排了填充空格(该文件受 prettier 门禁且 HEAD 是 prettier 干净的);按空白归一化后唯一的内容变化就是这一个单元格。

R2-2(rc:4152008469)——分组行的 capability/SDK 单元格无钉住——已实现

rest-integration-docs-contract.test.ts 新增 pins the grouped reference rows to registered capabilities and SDK methods:对每个以 Area 名称开头的行(OpenAPI 行检查看不到的形状),「Capability and scope」单元格中每个反引号 snake_case token 必须是 SERVE_CAPABILITY_REGISTRY 的键;「TypeScript SDK」单元格中每个反引号 DaemonClient.<name> / WorkspaceDaemonClient.<name> token 必须存在于对应类原型上。候选筛选是有意为之——scope 文字(「persisted-workspace per member」)与「raw REST」说明共用这些单元格且不带反引号。

探针:把文档中的 `workspace_session_live_state_batch` 改坏后,新测试以 Persisted session catalog: \workspace_session_live_state_batched` is not a registered serve capability` 失败;已恢复并复绿。

R2-3(rc:4152008472)——能力集断言只能由 CI 集成任务触达——已验证,无源码改动

在本地对新构建的 bundle 跑了完整的 integration-tests/cli/qwen-serve-routes.test.ts(42/42,其中包含断言真实启动的 daemon 公布 workspace_session_live_state_batch 的 qwen serve — capabilities envelope 块),并单独跑了该用例(1 passed)。合入 head 上的 CI 集成任务仍是最终门禁;本地结果确认打包后 CLI 的能力集携带该标签。

R2-4(rc:4152008477)——循环后中止守卫无见证——已实现

新测试 skips the response when the client disconnects during the last member read:单成员批量、真实 server;成员读取在 withDaemonSpan mock 内被一个 gate 挂起,同时销毁 socket,直到观察到服务端响应 close 后才放行——因此由循环后的 if (controller.signal.aborted) return;(而非循环内检查)决定结果。断言响应的 json 写入从未发生,而成员读取本身完成(无论有无守卫 listCalls 都是 [PRIMARY_CWD],如注释所述)。

探针:删除循环后守卫时恰好只有该测试变红;stops reading members once the client disconnects mid-batch 与 pre-handler 测试保持绿色——循环内与循环后两个分支被分别钉住。守卫已恢复。

R2-5(rc:4152008481)——!generation 读取前守卫无见证——已实现

新测试 returns 503 for a member whose entry lost its current generation while preserving a healthy member:清空 entry.current 并调用 registry.blockReplacement(entry, 'apply failed')——正是 trust reconciler 在 workspace-trust-reconciler.ts:230-231 产生的状态——同时保持条目可解析,然后断言 503 workspace_runtime_unavailable 成员不带 sessions、secondary bridge 从未被读取、健康成员不受影响。

探针:删除 if (!generation || !isCurrent()) return unavailable(); 时恰好只有该测试变红(整个请求变成 500,而不是返回隔离的 503 成员);其余 17 个批量测试保持绿色。守卫已恢复。

R1-6(rc:4152008490)——handler 前中止守卫无见证——已实现

新测试 skips every member read when the client disconnects before the handler runs:一个外层 express app 在 socket 存活时先解析请求体,销毁 socket,并只在响应 close 投递之后才分发给真实的 serve app(serve app 自己的 express.json() 因请求体已读取而跳过重复解析),于是 handler 观察到的是一个已销毁的响应。断言两个 workspace 的 bridge 读取次数都为零。

该测试的第一版在请求体解析之前就销毁了 socket,请求在守卫之前就 400 了,探针保持绿色——被探针如实抓获。改用正确构造后探针确定性变红(2/2 次运行):删除 if (req.aborted || res.destroyed) abort(); 后 handler 会执行并读取(primaryBridge.listCalls 变为 ['/work/primary']),而有守卫时两个 bridge 都保持未读。守卫已恢复。

验证

  • npm run build —— 通过(需要 COREPACK_HOME=/tmp/corepack-cache;沙箱拒绝 $HOME/.cache 下默认的 corepack 缓存路径)
  • npm run typecheck —— 通过
  • npm run lint —— 通过
  • vitest run src/serve/multi-workspace-sessions.test.ts(packages/cli)—— 185 通过
  • vitest run src/serve/multi-workspace-sessions.test.ts -t 'batch workspace session live-state route' —— 18 通过(为排查抖动跑了两次,稳定)
  • vitest run src/serve/rest-integration-docs-contract.test.ts(packages/cli)—— 14 通过
  • vitest run src/serve/capabilities-docs-contract.test.ts src/serve/rate-limit.test.ts src/serve/server/telemetry-catalog.test.ts src/serve/server/telemetry.test.ts(packages/cli)—— 116 通过
  • vitest run src/serve/server.test.ts(packages/cli)—— 1369 通过
  • npm run bundle —— 通过
  • QWEN_SANDBOX=false vitest run --root ./integration-tests cli/qwen-serve-routes.test.ts —— 42 通过;-t 'capabilities envelope' —— 1 通过
  • 变异探针(每个探针在变红后都已恢复并重新验证为绿):
    • 删除 !generation || !isCurrent() 读取前守卫 → 只有新的 R2-5 测试失败(17 个批量测试绿)
    • 删除循环后的 if (controller.signal.aborted) return; → 只有新的 R2-4 测试失败(mid-batch 断开与 pre-handler 测试绿)
    • 删除 if (req.aborted || res.destroyed) abort(); 这一 handler 前守卫 → 只有新的 R1-6 测试失败,且确定(2/2 次运行)
    • 改坏文档中的一个 capability token → 只有新的分组行契约测试失败

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 6 selected review thread(s). · 已关闭全部选中的 6 条评审线程。

Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.7

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🔀 Base updated: red check(s) [Lint & Static (ubuntu-latest, Node 22.x)] pass on current main — merged current main via update-branch; CI will re-run.

中文说明

🔀 已更新 base:红色检查 [Lint & Static (ubuntu-latest, Node 22.x)] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Addressed the latest review feedback (round 3/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 3/100 轮)。改动内容与我反驳保留之处如下:

Autofix round — PR #12513

Commit 93783707c9 (test(cli): pin grouped REST doc rows against bare SDK names and column shifts) addresses the two standing Suggestion findings on packages/cli/src/serve/rest-integration-docs-contract.test.ts. One file changed, 44 insertions, 18 deletions; no production code touched.

Feedback dispositions

R3-1 — grouped-row SDK check skipped bare method names (rc:4153835870, rc:4156815933) — resolved

Reproduced first: on the pre-round test, renaming the bare doc token `getSessionsLiveState` at docs/developers/daemon-rest-api-reference.md:128 to a nonexistent method left the suite green (the class-prefix-only regex never saw it), while renaming the qualified token in the same row reddened as designed.

Fix, combining both reviewer prescriptions: the SDK regex's class prefix is now optional ((?:(DaemonClient|WorkspaceDaemonClient)\.)?) and the prevailing class carries forward across bare continuation names, resetting at each explicit qualifier (the "Session organization" row switches classes mid-cell). Membership is checked against per-class name sets built by this file's own sdkMethods() helper, which was generalized to sdkMethods(root: object = DaemonClient.prototype) — its no-arg caller at :425 and its documented "including inherited accessors" semantics are unchanged, and the set-membership check also removes the latent getter-invocation side effect of the old typeof proto[method] probe.

R3-2 — hardcoded column indices could silently vacate both loops (rc:4153835887, rc:4156816262) — resolved

Reproduced first: replaying the reviewer's splice (an extra column before "Capability and scope") against the pre-round test left cells[3]/cells[4] pointing at the wrong cells with zero assertions and a green suite.

Fix: the two column indices are derived from the grouped table's own header row (/^\| Area\s*\|/, unique in the doc — verified) via indexOf('Capability and scope') / indexOf('TypeScript SDK'), each asserted found with a named message; per-table match counters are accumulated and asserted > 0, so a future shape change either fails loudly or keeps asserting — it cannot pass with zero assertions. The row selector is unchanged, so the neighbouring rows.length === operations.size invariant is unaffected, and expect(rows.length).toBe(operations.size) still forbids a first-cell-link row for the batch route.

R2-3 — integration lane never executed the envelope assertion (rc:4153835899) — no code change; thread left open with a reply

The finding itself states no source change is needed. Verified locally that a4bf0026c0f2 (the eslint.legacy-filenames.mjs change on main that tripped Check lint gate freshness) is already an ancestor of the branch head via the merge at 65f11e80d7, so this round's push re-validates that step under the current gate. Confirming Integration Tests (CLI, No Sandbox) runs green on the new head is only observable after the push (the lane is skipped by routing on previous heads and the file sits outside every npm workspace, so it cannot run locally); the thread is left open with that status posted as a reply. The cancelled macos-latest / Java 21 check is infrastructure cancellation, not a code signal.

The reviewer-deferred probe (multi-workspace-sessions.test.ts:7245 negative-half memory-only pin, recorded under qwen-review-deferred as "not requested in this round") was not actioned, per its own deferral.

Mutation probes (guard witnesses)

All probes mutated docs/developers/daemon-rest-api-reference.md only; the file was restored byte-identical after each (cmp verified), and the shipped tree contains no probe residue.

  1. Bare-token rename: `getSessionsLiveState` → `getSessionsLiveStateRenamed` → RED: "Persisted session catalog: getSessionsLiveStateRenamed is not an SDK method". (Survived on the pre-round test.)
  2. Header rename: Capability and scope → Capability → RED: grouped table must have a "Capability and scope" column: expected -1 to be greater than or equal to 0.
  3. Column insert (header + separator + all 13 rows) plus capability corruption `workspace_session_live_state_batch` → ..._batch_nope → RED: "...is not a registered serve capability". This exact combination was silent-green on the pre-round test.
  4. Column insert alone → GREEN, counters still satisfied — a legitimate table evolution neither false-positives nor goes vacuous.

Changed files

  • packages/cli/src/serve/rest-integration-docs-contract.test.ts (+44/−18): generalized sdkMethods(); header-derived column indices with found-assertions; carried-class SDK resolution over all backticked tokens; per-table assertion counters.

Verification

  • npm run build — passed (exit 0), both before and after the change (requires COREPACK_HOME=/tmp/corepack-home on this runner; the default corepack cache path is not writable in the sandbox)
  • npm run typecheck — passed (exit 0)
  • npm run lint — passed (exit 0)
  • npx vitest run src/serve/rest-integration-docs-contract.test.ts (packages/cli) — 14 passed, run pre-change (baseline), post-change, and post-commit
  • npx vitest run src/serve/multi-workspace-sessions.test.ts (packages/cli) — 185 passed
  • Mutation probes 1–4 above — all behaved as required (guarded cases red, legitimate evolution green)
  • Pre-commit hook (prettier + eslint --fix --max-warnings 0 on the staged file) — completed clean during the commit
  • Not run: integration tests — the change is a docs-contract unit test exercised directly by Vitest, not only through the bundled CLI; no settings source changed, so generate:settings-schema was not needed
中文说明

Autofix 本轮处理 — PR #12513

提交 93783707c9(test(cli): pin grouped REST doc rows against bare SDK names and column shifts)处理了 packages/cli/src/serve/rest-integration-docs-contract.test.ts 上两条仍然成立的 Suggestion 级发现。共改动一个文件,新增 44 行、删除 18 行;未改动任何生产代码。

发现处理结果

R3-1 —— 分组行的 SDK 校验跳过裸方法名(rc:4153835870、rc:4156815933)—— 已解决

先复现:在本轮之前的测试上,把 docs/developers/daemon-rest-api-reference.md:128 的裸写 token `getSessionsLiveState` 改名为一个不存在的方法,套件依然全绿(要求类名前缀的正则根本看不到它);而改名同一行里带前缀的 token 则如期变红。

修复(综合评审者两轮给出的方案):SDK 正则中的类名前缀改为可选((?:(DaemonClient|WorkspaceDaemonClient)\.)?),并让最近出现的类名向后传递给随后的裸写方法名,在每个显式限定名处重置("Session organization" 行在单元格中途切换了类)。成员判断改为对照按类预构建的名称集合,集合由本文件已有的 sdkMethods() 辅助函数生成——该函数泛化为 sdkMethods(root: object = DaemonClient.prototype)::425 处的无参调用方及其文档写明的「包括继承来的 accessor」语义保持不变;集合成员判断同时消除了旧写法 typeof proto[method] 潜在调用 getter 的副作用。

R3-2 —— 硬编码列下标可能让两个循环静默空转(rc:4153835887、rc:4156816262)—— 已解决

先复现:按评审者的做法重放(在 "Capability and scope" 之前插入一列),本轮之前的测试中 cells[3]/cells[4] 指向错误的单元格,断言次数为零,而套件依然全绿。

修复:两个列下标改为从分组表格自身的表头行(/^\| Area\s*\|/,已验证该表头在文档中唯一)通过 indexOf('Capability and scope') / indexOf('TypeScript SDK') 推导,并各自断言找到(带明确失败信息);同时在整张表上累计匹配次数并断言 > 0,因此将来表格形状变化要么显式失败、要么继续产生断言——不可能出现「零断言却通过」。行选择器保持不变,因此相邻测试的 rows.length === operations.size 不变量不受影响,:540 处的 expect(rows.length).toBe(operations.size) 仍然禁止为批量路由新增首单元格为链接的行。

R2-3 —— 集成通道从未执行过 envelope 断言(rc:4153835899)—— 无需改代码;线程保持开放并已回复

该发现本身写明不需要修改源码。已在本地验证 a4bf0026c0f2(main 上导致 Check lint gate freshness 失败的 eslint.legacy-filenames.mjs 变更)经由 65f11e80d7 的合入已是分支 head 的祖先提交,因此本轮推送会在当前门禁下重新验证该步骤为绿色。确认 Integration Tests (CLI, No Sandbox) 在新 head 上跑绿只能在推送之后观察(该通道在之前的 head 上因路由规则被跳过,且该文件不在任何 npm workspace 内,无法在本地运行);线程保持开放,并已将上述状态作为回复发布。macos-latest / Java 21 的取消属于基础设施取消,不是代码信号。

评审者自行延后的探针(multi-workspace-sessions.test.ts:7245 关于 memory-only 契约负半部分的钉住,记录在 qwen-review-deferred 下并注明「本轮不要求修改」)按其自身的延后决定未做处理。

变异探针(防护见证)

所有探针只改动 docs/developers/daemon-rest-api-reference.md;每次探针后该文件都按字节恢复原样(cmp 已验证),最终提交的树中没有任何探针残留。

  1. 裸名改名:`getSessionsLiveState` → `getSessionsLiveStateRenamed` → 变红:"Persisted session catalog: getSessionsLiveStateRenamed is not an SDK method"。(该变异在本轮之前的测试上存活。)
  2. 表头改名:Capability and scope → Capability → 变红:grouped table must have a "Capability and scope" column: expected -1 to be greater than or equal to 0。
  3. 插入一列(表头 + 分隔行 + 全部 13 个数据行)并同时损坏 capability `workspace_session_live_state_batch` → ..._batch_nope → 变红:「...is not a registered serve capability」。这一组合在本轮之前的测试上是静默全绿的。
  4. 仅插入一列 → 全绿,计数断言仍然满足——合理的表格演进既不会误报,也不会让检查变成空转。

改动文件

  • packages/cli/src/serve/rest-integration-docs-contract.test.ts(+44/−18):泛化 sdkMethods();列下标改为从表头推导并断言找到;SDK 校验改为沿用当前类名并覆盖全部反引号 token;新增全表断言计数。

验证

  • npm run build —— 通过(exit 0),改动前后各跑一次(本 runner 上需要 COREPACK_HOME=/tmp/corepack-home;沙箱内默认 corepack 缓存路径不可写)
  • npm run typecheck —— 通过(exit 0)
  • npm run lint —— 通过(exit 0)
  • npx vitest run src/serve/rest-integration-docs-contract.test.ts(packages/cli)—— 14 个通过,分别在改动前(基线)、改动后、提交后各跑一次
  • npx vitest run src/serve/multi-workspace-sessions.test.ts(packages/cli)—— 185 个通过
  • 上述变异探针 1–4 —— 行为全部符合要求(受防护的用例变红,合理演进保持绿色)
  • 提交前钩子(对暂存文件执行 prettier + eslint --fix --max-warnings 0)—— 在提交过程中干净通过
  • 未运行:集成测试——本次改动是文档契约单元测试,由 Vitest 直接覆盖,并非只能通过打包后的 CLI 触发;未改动 settings 源文件,因此无需运行 generate:settings-schema

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 2 selected review thread(s). · 已关闭全部选中的 2 条评审线程。

Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.7

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Addressed the latest review feedback (round 4/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 4/100 轮)。改动内容与我反驳保留之处如下:

Autofix round summary — PR #12513

Commit 1fd4eeb7a2 (test(cli): fail closed on grouped REST doc rows and ambiguous SDK qualifiers) addresses both round-5 inline findings. Both are on the same test, pins the grouped reference rows to registered capabilities and SDK methods, and both stem from one root cause the review body's convergence note anticipated: the pin was written to fail open. The fix rewrites that one test to fail closed at every granularity instead of patching instances.

rc:4159484825 (R5-1) — grouped-rows pin fails open at both granularities — IMPLEMENTED

packages/cli/src/serve/rest-integration-docs-contract.test.ts:

  • Rows now come from the table body under the | Area header (findIndex + take |-leading lines from headerIndex + 2 until the first non-table line), replacing the document-wide /^\| [A-Z]/ + [` shape heuristic. A row whose Area starts lowercase or with a digit, or whose Operations cell has no backticked link, is still selected and checked; no other table's rows are ever read at this table's column indices (the mirror failure the finding demonstrated is structurally impossible now).
  • Per-row assertions so one dropped row or token cannot hide behind table-wide totals: cell count must equal the header's; each row must contribute at least one capability check and at least one SDK check; and every backticked span in each asserted cell must be consumed by its matcher (span-count equality, counted by spans `[^`]*`, never by words — the prose spans at lines 122/128/131 of the reference doc are unaffected).
  • The qualifier alternation is built from Object.keys(sdkMethodSets), so a third client class cannot silently un-assert the methods named for it.
  • The document-wide vacuity counters (capabilityChecks/sdkChecks > 0) were removed; the per-row assertions are strictly stronger. Recorded in test-weakening.json.

rc:4159484840 (R3-1) — 'DaemonClient' seed certifies bare names against a class the document never named — IMPLEMENTED, with a measurement correction

  • sdkClass is now seeded as undefined: a cell's first SDK token must carry a class qualifier (must follow a class-qualified SDK method).
  • A bare token whose name exists on more than one SDK client must be qualified (exists on more than one SDK client and must be qualified), computed from Object.values(sdkMethodSets) so it tracks the same source as the token regex.
  • Measurement correction to the finding's premise: the names it cites as client-unique are not. DaemonClient exposes 279 members, WorkspaceDaemonClient 131, and 91 names exist on both prototypes — including workspaceMcp, workspaceSkills, workspaceProviders, workspaceEnv, workspacePreflight, createSessionGroup, updateSessionGroup, and deleteSessionGroup (verified against the built SDK with the test's own prototype walk). So the suggested guard could not "stay green on today's doc" unmodified: today's table carries 17 bare tokens that exist on both clients. Per the finding's own rule — an ambiguous bare name must be qualified — those 17 tokens are now qualified in docs/developers/daemon-rest-api-reference.md (6 rows: Workspace runtime status, Workspace-qualified history, Persisted session catalog, Session organization, Bulk persisted-session changes, Workspace configuration). The 12 remaining bare tokens are genuinely single-client (e.g. sessionTasks, setUserLanguage, exportArchivedSession) and keep the doc's continuation convention, which the finding endorsed. The demanded pin behaves as required: replacing the Session organization row's final `WorkspaceDaemonClient.updateSessionOrganization` with a bare `updateSessionOrganization` turns the test red; restoring the doc returns it to green.

The reference doc was re-formatted with prettier --write afterwards (the table re-padded to the new cell widths, 15 lines), restoring the file's prettier-clean state from HEAD.

Failed checks

  • web-shell E2E Smoke (ubuntu-latest Node 22.x) shows CANCELLED, not a test failure. This PR touches no web-shell files, and this round changes only a serve docs-contract test and the reference doc. No check logs are available without GitHub credentials; nothing in this change set can affect that job. The workflow's post-push CI remains the verification gate for it.

Review coverage disclosure

The review noted the one line this PR adds to integration-tests/cli/qwen-serve-routes.test.ts (the workspace_session_live_state_batch capability name) was reached only by typecheck because that suite was skipped in CI. The added name is asserted against SERVE_CAPABILITY_REGISTRY by the unit contract test above, which does run — the registry membership the integration line relies on is covered there.

Mutation probes (each guard witnessed red, then restored to green, on the final tree)

  1. Ambiguity guard: `WorkspaceDaemonClient.updateSessionOrganization` → bare `updateSessionOrganization` ⇒ red (exists on more than one SDK client and must be qualified); restore ⇒ green. (Both directions of the demanded R3-1 pin.)
  2. Undefined-seed guard: first token `DaemonClient.startDeviceFlow` → bare ⇒ red (must follow a class-qualified SDK method); restore ⇒ green. Probed with a single-client name so only the seed guard can fire.
  3. Row derivation: appended a 14th grouped row with a lowercase Area and a bogus capability ⇒ red (row selected, capability rejected); removed ⇒ green.
  4. Capability span equality: injected `session-info` into a capability cell ⇒ red (capability cell has backticked spans this test cannot read); restore ⇒ green.
  5. SDK span equality: injected `WorkspaceClient.remove` into an SDK cell ⇒ red (SDK cell has backticked spans this test cannot read); restore ⇒ green.
  6. Cell-count guard: dropped a row's last cell ⇒ red (grouped row must have 6 columns); restore ⇒ green.
  7. Per-row ≥1 SDK check: replaced a row's SDK cell with prose ⇒ red (row must name at least one backticked SDK method); restore ⇒ green.
  8. Per-row ≥1 capability check: replaced a row's capability cell with prose ⇒ red (row must name at least one backticked capability); restore ⇒ green.

Verification

  • npm run build — passed (exit 0). Required COREPACK_HOME=/tmp/corepack-cache: the sandbox denies writes to ~/.cache/node/corepack, which otherwise fails the build with EACCES before any package compiles.
  • npm run typecheck — passed (exit 0).
  • npm run lint — passed (exit 0).
  • npx vitest run src/serve --coverage.enabled=false (packages/cli) — 261 test files passed, 11698 passed | 54 skipped.
  • npx vitest run src/serve/rest-integration-docs-contract.test.ts --coverage.enabled=false — 14/14 passed, re-run after the pre-commit hook and again on the committed tree.
  • npx prettier --check on both touched files — clean (prettier --write applied to the doc first).
  • Pre-commit hook (lint-staged: prettier + eslint --fix --max-warnings 0) — passed on the staged files.
  • Mutation probes 1–8 above — each red with the expected message on the final tree, each restored to green.
  • Integration tests not run: this round's behavior (a Markdown-table contract test) is not exercised through the bundled CLI or integration harness, so the trigger condition does not apply.
中文说明

Autofix 本轮总结 — PR #12513

提交 1fd4eeb7a2(test(cli): fail closed on grouped REST doc rows and ambiguous SDK qualifiers)处理了第 5 轮的两条行内发现。两条都落在同一个测试 pins the grouped reference rows to registered capabilities and SDK methods 上,且都源于评审正文收敛提示所预言的同一根因:这个钉住测试写成了一路放行(fail-open)。本次修复把这一个测试改写为在每个粒度上都失败即拦截(fail-closed),而不是逐个修补实例。

rc:4159484825(R5-1)—— 分组行钉住测试在两个粒度上失败即放行 —— 已实现

packages/cli/src/serve/rest-integration-docs-contract.test.ts:

  • 行现在取自 | Area 表头下方的表体(findIndex 定位后,从 headerIndex + 2 起收集以 | 开头的行,直到第一行非表格行),替换了原先全文档的 /^\| [A-Z]/ + [` 形状启发式。Area 以小写字母或数字开头的行、或 Operations 单元格没有反引号链接的行,现在仍会被选中并校验;其他表格的行再也不会按本表的列索引被误读(发现中演示的反向误报在结构上已不可能发生)。
  • 按行断言,使少一行、少一个 token 都无法再藏在全表总量之下:单元格数必须与表头一致;每行必须至少贡献一次 capability 检查和一次 SDK 检查;每个被断言单元格里的每一个反引号片段都必须被对应匹配器消费(按反引号片段 `[^`]*` 计数的相等断言,绝不按词计数——参考文档第 122/128/131 行的散文片段不受影响)。
  • 限定符分支改由 Object.keys(sdkMethodSets) 生成,将来出现第三个客户端类时,为它写的方法名不会静默地不再被断言。
  • 删除了全文档空转计数器(capabilityChecks/sdkChecks > 0);按行断言严格更强。已记录在 test-weakening.json。

rc:4159484840(R3-1)—— 'DaemonClient' 初始值把裸名背书到文档从未指明的类 —— 已实现,并更正一处测量前提

  • sdkClass 现在以 undefined 为初始值:每个单元格的首个 SDK token 必须带类限定(must follow a class-qualified SDK method)。
  • 凡是存在于多个 SDK 客户端上的裸名都必须带限定(exists on more than one SDK client and must be qualified),该判断基于 Object.values(sdkMethodSets) 计算,与 token 正则保持同一来源。
  • 对发现前提的测量更正: 发现中 cited 为「客户端独有」的名字实际上不是。DaemonClient 暴露 279 个成员,WorkspaceDaemonClient 131 个,两者原型的交集有 91 个名字——包括 workspaceMcp、workspaceSkills、workspaceProviders、workspaceEnv、workspacePreflight、createSessionGroup、updateSessionGroup、deleteSessionGroup(用测试自身的原型链遍历对构建产物实测)。因此建议的守卫不可能「在今天的文档上保持绿色」:今天的表格里有 17 个裸 token 同时存在于两个客户端。按照发现自身的规则——有歧义的裸名必须带限定——这 17 个 token 已在 docs/developers/daemon-rest-api-reference.md 中补齐限定(涉及 6 行:Workspace runtime status、Workspace-qualified history、Persisted session catalog、Session organization、Bulk persisted-session changes、Workspace configuration)。其余 12 个裸 token 确实是单客户端独有(如 sessionTasks、setUserLanguage、exportArchivedSession),继续保留发现所认可的续写约定。要求钉住的变异行为符合预期:把 Session organization 行末尾的 `WorkspaceDaemonClient.updateSessionOrganization` 换成裸的 `updateSessionOrganization` 测试变红;恢复文档后重新变绿。

修改后对参考文档执行了 prettier --write(表格按新的单元格宽度重新对齐,共 15 行),恢复了该文件在 HEAD 上的 prettier 整洁状态。

失败的检查

  • web-shell E2E Smoke (ubuntu-latest Node 22.x) 显示为 CANCELLED(被取消),不是测试失败。本 PR 未触碰任何 web-shell 文件,本轮也只改了一个 serve 文档契约测试和参考文档。没有 GitHub 凭据无法获取该检查日志;本次变更集中没有任何内容能影响该任务。推送后工作流的 CI 仍是它的最终验证关口。

评审覆盖披露

评审指出本 PR 在 integration-tests/cli/qwen-serve-routes.test.ts 中新增的一行(workspace_session_live_state_batch capability 名)因 CI 跳过该套件而只经过了 typecheck。该名字在上面的单元契约测试中被断言必须存在于 SERVE_CAPABILITY_REGISTRY(该测试实际运行)——集成测试行所依赖的注册表成员资格已在该处覆盖。

变异探针(在最终代码树上,每个守卫先证红、再恢复为绿)

  1. 歧义守卫:`WorkspaceDaemonClient.updateSessionOrganization` → 裸 `updateSessionOrganization` ⇒ 红(exists on more than one SDK client and must be qualified);恢复 ⇒ 绿。(即 R3-1 要求的双向钉住。)
  2. 未定义初始值守卫:首个 token `DaemonClient.startDeviceFlow` → 裸 ⇒ 红(must follow a class-qualified SDK method);恢复 ⇒ 绿。特意选用单客户端独有的名字,确保只有初始值守卫能触发。
  3. 行来源:追加一行 Area 以小写字母开头、携带伪造 capability 的分组行 ⇒ 红(行被选中,capability 被拒绝);移除 ⇒ 绿。
  4. capability 片段相等:向 capability 单元格注入 `session-info` ⇒ 红(capability cell has backticked spans this test cannot read);恢复 ⇒ 绿。
  5. SDK 片段相等:向 SDK 单元格注入 `WorkspaceClient.remove` ⇒ 红(SDK cell has backticked spans this test cannot read);恢复 ⇒ 绿。
  6. 单元格数守卫:删除某行最后一个单元格 ⇒ 红(grouped row must have 6 columns);恢复 ⇒ 绿。
  7. 按行 ≥1 SDK 检查:把某行 SDK 单元格换成散文 ⇒ 红(row must name at least one backticked SDK method);恢复 ⇒ 绿。
  8. 按行 ≥1 capability 检查:把某行 capability 单元格换成散文 ⇒ 红(row must name at least one backticked capability);恢复 ⇒ 绿。

验证

  • npm run build —— 通过(退出码 0)。需要 COREPACK_HOME=/tmp/corepack-cache:沙箱禁止写入 ~/.cache/node/corepack,否则构建在任何包编译之前就以 EACCES 失败。
  • npm run typecheck —— 通过(退出码 0)。
  • npm run lint —— 通过(退出码 0)。
  • npx vitest run src/serve --coverage.enabled=false(packages/cli)—— 261 个测试文件通过,11698 通过 | 54 跳过。
  • npx vitest run src/serve/rest-integration-docs-contract.test.ts --coverage.enabled=false —— 14/14 通过,并在 pre-commit 钩子之后、提交后的代码树上再次运行确认。
  • npx prettier --check 检查两个被改文件 —— 整洁(文档先执行了 prettier --write)。
  • pre-commit 钩子(lint-staged:prettier + eslint --fix --max-warnings 0)—— 对暂存文件通过。
  • 上述变异探针 1–8 —— 在最终代码树上各自以预期报错信息变红,恢复后各自变绿。
  • 集成测试未运行:本轮改动的行为(Markdown 表格契约测试)并不只经由打包后的 CLI 或集成测试框架才能触达,触发条件不成立。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 2 selected review thread(s). · 已关闭全部选中的 2 条评审线程。

Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.7

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🔀 Base updated: red check(s) [Lint & Static (ubuntu-latest, Node 22.x)] pass on current main — merged current main via update-branch; CI will re-run.

中文说明

🔀 已更新 base:红色检查 [Lint & Static (ubuntu-latest, Node 22.x)] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下:

Autofix review round — no action required

No code change was made this round. Every item in this round's feedback was triaged and none is actionable, for the reasons below.

What arrived since the last evaluation

  1. Review rv:5386923384 (round 6, @qwen-code-ci-bot, COMMENTED) — downgraded from Approve to Comment because CI was failing at the pre-merge head (1fd4eeb): web-shell E2E Smoke (ubuntu-latest, Node 22.x) and Lint & Static (ubuntu-latest, Node 22.x).
    • Disposition: already resolved. The branch head is now the merge commit 9706556 ("Merge branch 'main' into codex/batch-workspace-live-state", pushed 2026-10-02T00:25:03Z). The CI run on that head shows both previously failing checks green: Lint & Static SUCCESS (completed 2026-10-02T00:37:53Z) and web-shell E2E Smoke SUCCESS (completed 2026-10-02T01:10:29Z). This round's feedback file confirms it: the Failed checks and Still-red checks sections are both empty. There is no failing check left to fix.
  2. Review rv:5388296230 (round 7, @qwen-code-ci-bot, COMMENTED, at the current head) — posted no new findings. It confirmed two Suggestion-level findings already reported on this PR (GET-route catalog-version cache-invalidation pin, spawned-daemon capability assertion collectability) and explicitly did not repeat them. Its disclosed review gaps (integration suite skipped in CI, reverse audit not converged, one chunk under-explored) are statements about review coverage, not code findings to act on.
  3. Deferred findings (rounds 6 and 7) — three items recorded under the convergence posture (rest-integration-docs-contract.test.ts:668, multi-workspace-sessions.test.ts:7321, docs/design/daemon-batch-session-live-state.md:13). Both reviews state these are "recorded, not requested in this round" — an audit record, not work — so they were intentionally left untouched.
  4. Inline comments and issue-level comments: none. The curated feedback lists zero in both sections, and the raw comment data confirms nothing was created after the last evaluation cutoff (2026-10-01T19:17:23Z).

Evidence checked

  • git log / git status — head is 97065561a7 (merge with main), working tree clean.
  • git diff --stat origin/main...HEAD — PR footprint reviewed for context (25 files, ~1299 insertions).
  • Workflow-captured checks.json for the current head — every check SUCCESS or SKIPPED; specifically the two checks flagged in round 6 are SUCCESS.
  • Raw rc.json / ic.json — zero review or issue comments newer than the last evaluation.
  • Raw rv.json — both review bodies verified complete and matching the curated feedback; no additional requested items.

No build/test commands were run because this round makes no code change and the authoritative CI at the exact current head is green; the workflow's independent CI remains the final gate.

中文说明

Autofix 审查轮次 —— 无需处理

本轮未做任何代码修改。本轮反馈中的每一项都已分类处理,均不可执行,原因如下。

自上次评估以来收到的内容

  1. 审查 rv:5386923384(第 6 轮,@qwen-code-ci-bot,COMMENTED) —— 由于在合并前的 head(1fd4eeb)上 CI 失败而从 Approve 降级为 Comment:web-shell E2E Smoke (ubuntu-latest, Node 22.x) 和 Lint & Static (ubuntu-latest, Node 22.x)。
    • 处置:已解决。 分支 head 现在是合并提交 9706556("Merge branch 'main' into codex/batch-workspace-live-state",推送于 2026-10-02T00:25:03Z)。在该 head 上的 CI 运行显示此前失败的两个检查均已通过:Lint & Static SUCCESS(完成于 2026-10-02T00:37:53Z),web-shell E2E Smoke SUCCESS(完成于 2026-10-02T01:10:29Z)。本轮反馈文件也证实了这一点:Failed checks 和 Still-red checks 两节均为空。没有遗留的失败检查需要修复。
  2. 审查 rv:5388296230(第 7 轮,@qwen-code-ci-bot,COMMENTED,针对当前 head) —— 未发布新发现。该审查确认了两条已在本 PR 上报告过的 Suggestion 级发现(GET 路由的 catalog 版本缓存失效 pin、spawned-daemon 能力断言的可收集性),并明确不再重复发布。其披露的审查缺口(CI 中跳过集成测试套件、反向审计未收敛、一个 chunk 未探索完全)是关于审查覆盖面的说明,而非可据以修改代码的发现。
  3. 延后的发现(第 6、7 轮) —— 在收敛姿态下记录的三条(rest-integration-docs-contract.test.ts:668、multi-workspace-sessions.test.ts:7321、docs/design/daemon-batch-session-live-state.md:13)。两份审查均明确说明这些条目**"已记录,本轮不要求修改"** —— 属于审计记录而非工作任务 —— 因此有意不做改动。
  4. 行内评论与 issue 级评论:无。 整理后的反馈在这两节中均为零条,原始评论数据也确认在最后评估截止时间(2026-10-01T19:17:23Z)之后没有任何新评论。

已核查的证据

  • git log / git status —— head 为 97065561a7(与 main 的合并),工作区干净。
  • git diff --stat origin/main...HEAD —— 为理解上下文查看了 PR 的改动范围(25 个文件,约 1299 行新增)。
  • 工作流为当前 head 抓取的 checks.json —— 所有检查均为 SUCCESS 或 SKIPPED;特别是第 6 轮标记的两个检查均为 SUCCESS。
  • 原始 rc.json / ic.json —— 最后评估之后没有任何新的审查评论或 issue 评论。
  • 原始 rv.json —— 两份审查正文已核实完整且与整理后的反馈一致;没有额外的请求项。

本轮未运行构建/测试命令,因为本轮不做任何代码修改,且针对当前确切 head 的权威 CI 已全部通过;工作流的独立 CI 仍是最终验证关口。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.7

@wenshao

wenshao commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

🔬 Local verification round (maintainer, sandboxed worktrees)

Verdict: merge-ready — 52/52 scripted assertions passed, 0 unexpected failures. Verified head 97065561a753b9431ea6c872c60c530e963bf20e (matches PR headRefOid). Local maintainer round on macOS arm64, Node 24.18.1.

中文摘要

结论:merge-ready —— 52/52 项脚本断言全部通过。

  • A/B 结论:head 守护进程上 POST /sessions/live-state 一次请求按输入顺序返回 2 个 workspace 快照 + 1 个 workspace_not_found 成员错误(23/23);同一场景在 base 构建上该路由为 404、capability 不公布,而单 workspace 路由与兄弟路由 /sessions/catalog 均正常(6/6)——改动是 load-bearing 的,且未误伤既有路由。见 01-ab-head-vs-base.png。
  • 信任门:开启 folder trust 并把 secondary 标记 DO_NOT_TRUST 后,批量请求中该成员返回 403 untrusted_workspace 且不携带 sessions,同一请求内受信 primary 成员照常成功;单 workspace 路由同样返回 403(5/5)。见 04-trust-gate.png。
  • SDK 实测:真实编译产物 DaemonClient.getSessionsLiveState 对真实 daemon 一次 POST 拿到类型化结果;abort 信号生效;timeoutMs: 0 的语义是「禁用超时」(SDK 文档化行为,非缺陷);timeoutMs: 100 能中止黑洞服务器上的请求(7/7)。见 02-sdk-wire.png。
  • 变异矩阵:4 个变异(中断 yield、响应前 abort 守卫、代际时效守卫、capability 公布)全部被 PR 自带测试击杀,阳性对照同样被击杀(线束有效)。唯一存活的子场景(replaced-during-read 在 M3 下存活)经核实是 readWorkspaceLiveState 内部 assertRuntimeOpen 的冗余防御——beginReplacement 会关闭代际 guard,内层断言抛错;beginDrain 不关闭 guard,因此时效守卫是 drain 路径的唯一防线。两层守卫各自 load-bearing。见 03-mutation-matrix.png。
  • 测试门:聚焦 serve 测试文件 6/7 绿;server.test.ts 在 head 两轮、base 一轮中失败集合互不相交且与 PR 改动面无关(附件/归档/转录等),证实为既有的负载敏感 flake(PR 作者也报告过)。SDK 520/520;integration qwen-serve-routes 42/42;与当前 main(ba09794545)试验性合并无冲突。
  • 发现:无阻塞项。两条非阻塞观察见下文 Findings。
  • 未覆盖:模型驱动的真实会话状态迁移(需模型凭据,单测以模拟 bridge 覆盖);Windows;真实 socket 上的 mid-batch 断连(单测覆盖,变异矩阵证实有效);512 KiB 超限成员的真实 E2E(单测覆盖);transitioning 成员的 E2E(单测覆盖);逐 commit 归因(10 个 commit,按聚合 diff 验证)。

Central claim

POST /sessions/live-state returns complete per-workspace live session snapshots for 1–20 explicitly selected registered workspaces in one request, members in input order, per-member success/error, capability workspace_session_live_state_batch advertised, no behavior change to the single-workspace route or the sibling catalog route.

A/B proof (real daemons, real loopback HTTP)

Identical harness (harness/live-state-e2e.mjs) spawned dist/cli.js serve --port 0 --no-web per arm with isolated HOME, registered a secondary workspace over POST /workspaces, then drove the batch route. Control purity: lockfiles untouched by the PR; readlink -f node_modules/@qwen-code/qwen-code-core inside each tree resolves into that tree (tmp/pr12513-head/... vs tmp/pr12513-base/...).

Cell Oracle base a7deb01 head 9706556
POST /sessions/live-state status 404 (route absent) ✔ 200 ✔
capability workspace_session_live_state_batch /capabilities features array absent ✔ present ✔
batch members (secondaryId, primaryCwd, unknown) order + shape n/a 2 snapshots + 1 workspace_not_found member, input order ✔
Cache-Control response header n/a no-store ✔
malformed bodies ×7 (empty/oversized/strict/non-string…) status + code n/a all 400 invalid_session_live_state_batch_request ✔
single route GET /workspaces/:ws/sessions/live-state status 200 ✔ 200 ✔, sessions payload identical to batch member ✔
sibling POST /sessions/catalog (refactored in this PR) 200 + success/not_found members 200 ✔ 200 ✔, both member kinds correct ✔

Totals: head 23/23, base 6/6 (base 404 cells are the expected-failure controls, encoded as assertions). Witness: evidence/01-ab-head-vs-base.png

A/B: batch live-state — head vs base, raw logs logs/head.txt, logs/base.txt.

Wire-oracle harnesses

SDK against the real daemon (harness/sdk-wire.mjs, head packages/sdk-typescript/dist/index.mjs → live head daemon): one POST returns the typed batch result; unknown member surfaces as typed error; pre-aborted signal aborts without a response; timeoutMs: 0 keeps the request alive (SDK-wide documented "disable" semantics — verified against fetchWithTimeout, not assumed); timeoutMs: 100 aborts a blackholed request in 100 ms. 7/7. Witness: evidence/02-sdk-wire.png

SDK wire oracle vs real daemon.

Trust gate (harness/trust-gate.mjs, folder trust enabled via isolated HOME, secondary DO_NOT_TRUST): batch returns 200 with the untrusted member as 403 untrusted_workspace without sessions, the trusted primary member succeeds in the same request, and the single-workspace route agrees with a 403. 5/5. Witness: evidence/04-trust-gate.png

Trust gate E2E.

Mutation matrix (vacuity check on the PR's new tests)

Ran against packages/cli/src/serve/multi-workspace-sessions.test.ts -t 'batch workspace session live-state route' (18 selected; baseline 18/18 green) plus the capabilities contract file. Each mutant applied to head source, suite run, tree restored (git status clean at the end).

Mutant Expected killer Result
M1 drop in-loop abort check + setImmediate yield disconnect-mid-batch tests killed (2 tests red)
M2 drop post-loop abort guard disconnect-during-last-read killed
M3 captureWorkspaceEntryCurrency → always current drain/replacement-during-read killed by the drain test (1 red)
M4 capability not advertised capabilities-docs-contract killed (1 red)
C1 control: rename the 400 code envelope test killed (harness live)

Adjudicated survivor: under M3, "fails closed if the selected runtime is replaced during its read" stays green. Verified cause, not assumed: beginReplacement closes the generation guard (workspace-registry.ts:419), so the inner assertRuntimeOpen() inside readWorkspaceLiveState throws and the catch maps it to the same 503 member — redundant defence, correct as written. beginDrain does not close the guard, so the currency check is the only pin on the drain path, which is exactly the test M3 killed. Both layers are load-bearing. Witness: evidence/03-mutation-matrix.png

Mutation matrix vs PR test suite, log logs/mutation-matrix.txt.

Gates

Gate Result
Focused serve files on head (multi-workspace-sessions, rate-limit, capabilities-docs-contract, rest-integration-docs-contract, telemetry, telemetry-catalog) green (in the combined run: 1677 passed / 1684, all 7 failures in server.test.ts)
server.test.ts attribution pre-existing flake, not PR-caused: three runs (head ×2, base ×1) produced three disjoint failure sets (archive-invalidation; attachments/status/artifacts/metadata/middleware; status-routes ×2) — none byte-identical across runs, none in the PR's surface; the PR touches this file only by adding one capability string to three lists, and the capability tests passed in every run. Matches the author's own note about this file.
SDK unit (DaemonClient.test.ts, daemon-public-surface.test.ts) 520/520
integration-tests/cli/qwen-serve-routes.test.ts (against the head bundle) 42/42
Trial merge into current main ba09794545 (git merge-tree --write-tree) conflict-free; main's delta since merge-base touches 4 files this PR also touches, all additive in unrelated areas (agent-collaboration field, memory-file content)

Findings

None blocking. Two non-blocking observations:

  1. server.test.ts is flaky under load on both base and head (disjoint failure sets). Pre-existing; worth a deflake pass independently of this PR.
  2. The batch route accepts duplicate selectors and answers each independently (verified: two success members). Harmless and arguably intended, but it spends one member slot of the 20 on a duplicate; the SDK does not dedupe either. Note only.

Not covered

  • Model-driven live session state transitions (running/waiting/idle) under a constant catalog version — needs model credentials; the PR's unit tests pin this with simulated bridges.
  • Windows; transitioning-member and 512 KiB oversize-member behavior over real sockets (unit-tested; oversize bound and transition semantics verified by source + killed mutants, not E2E).
  • Per-commit attribution: the PR has 10 commits; the aggregate baseRefOid..headRefOid diff was verified, not each commit.
  • Rate-limit tier classification (read tier for the new route) is covered only by rate-limit.test.ts, not by a live flood test.

Methodology

macOS arm64, Node 24.18.1. Two worktrees at the resolved baseRefOid/headRefOid, each installed via the repo's pinned pnpm (scripts/setup-worktree.js) and built with npm run build && npm run bundle; internal-link realpaths asserted per tree. All harnesses spawn the real bundled daemon on 127.0.0.1:0 with isolated HOME; no mocks anywhere in the E2E path. Harness sources and per-cell JSON/text logs live in this artifact directory (harness/, logs/). Mutations were applied to head source for vitest only; the bundle used by the E2E/SDK/trust cells was built from pristine head before any mutation, and the tree was verified clean after the matrix.

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下:

Autofix round — no change required

New feedback this round: one issue-level comment, the maintainer's sandboxed local verification round ([ic:5965045999], @wenshao). Verdict: merge-ready — 52/52 scripted assertions passed against head 9706556, including A/B proof against the base build, SDK wire-oracle, trust-gate E2E, and a mutation matrix confirming the PR's guards are load-bearing. The comment raises no blocking findings and requests no changes.

CI: all checks on the current head are green (SUCCESS or SKIPPED; no FAILURE). The "Failed checks" and "Still-red checks" sections are empty; the CI failures noted in the round-6 review were resolved by the merge with main in 9706556.

Disposition of the verification comment's two non-blocking observations:

  1. server.test.ts flake under load — verified real by the maintainer's evidence (three runs across base and head produced disjoint failure sets, none in this PR's surface) and explicitly scoped as "worth a deflake pass independently of this PR". Its fix lies outside this PR's footprint, so it is recorded in deferred-findings.json for the follow-up queue rather than addressed here.
  2. Duplicate selectors in the batch route — "harmless and arguably intended... note only" per the verifier; the SDK does not dedupe either. Declined: no change worth making, and altering member handling would change documented behavior the verifier confirmed as working as intended.

No code changes, no commits this round.

中文说明

Autofix 轮次 —— 无需改动

本轮新增反馈: 一条 issue 级评论,即维护者的沙箱本地验证轮([ic:5965045999],@wenshao)。结论:merge-ready(可合并) —— 针对 head 9706556 的 52/52 项脚本断言全部通过,包括与 base 构建的 A/B 对照、SDK 线路级实测、信任门 E2E,以及证实 PR 各守护均为 load-bearing 的变异矩阵。该评论没有阻塞性发现,也未要求任何改动。

CI: 当前 head 上的所有检查均为绿色(SUCCESS 或 SKIPPED;无 FAILURE)。"Failed checks" 与 "Still-red checks" 两节均为空;第 6 轮评审中提到的 CI 失败已由 9706556 中与 main 的合并解决。

验证评论中两条非阻塞观察的处理:

  1. server.test.ts 在负载下的 flake —— 维护者的证据证实其真实存在(base 与 head 上三次运行产生互不相交的失败集合,且均不在本 PR 的改动面内),并被明确界定「值得在本 PR 之外单独做一轮 deflake」。其修复位于本 PR 的范围之外,因此已记入 deferred-findings.json 进入跟进队列,而非在本轮处理。
  2. 批量路由接受重复 selector —— 按验证者说法「无害且可认定为有意为之……仅作记录」;SDK 同样不去重。不予改动:没有值得做的修改,而改变成员处理行为会改动经维护者确认符合预期的已文档化行为。

本轮无代码改动、无提交。

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.24.7

…-live-state

# Conflicts:
#	packages/cli/src/serve/capabilities-docs-contract.test.ts
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Addressed the latest review feedback (round 5/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 5/100 轮)。改动内容与我反驳保留之处如下:

Autofix round summary — PR #12513

Feedback triage

feedback.md contained no actionable items: no reviews, no inline comments, no issue-level comments, no failed checks, and no still-red checks newer than the last evaluation (2026-10-03T03:25:17Z). There was nothing to implement, decline, defer, or escalate.

Merge-conflict resolution (--conflict true)

The PR branch had diverged from main (merge-base a7deb01bcb; main advanced to f07c190cd6, ~50 commits). Merging origin/main produced exactly one textual conflict:

  • packages/cli/src/serve/capabilities-docs-contract.test.ts — both sides appended a new test after the same anchor (advertises batched session catalogs for single and multiple workspaces):

    • PR side: advertises batched workspace live-state snapshots (pins the PR's new workspace_session_live_state_batch capability).
    • Main side: derives the hosted persona tags from the registry, never a hand list (pins hostedPersonaServeFeatures derivation).

    Resolution: keep both tests, verbatim, in sequence. The auto-merged import block already carried both sides' imports (hostedPersonaServeFeatures, SERVE_CAPABILITY_REGISTRY from main). No test was deleted or weakened.

Semantic overlap was audited beyond the textual merge:

  • packages/cli/src/serve/capabilities.ts — merged cleanly; registry contains both workspace_session_live_state_batch (PR) and the hosted-persona entries (main).
  • packages/cli/src/serve/routes/session.ts — main's sendStandaloneActionUnsupported refactor and the side-task handler move are orthogonal to the PR's readWorkspaceLiveState extraction and new POST /sessions/live-state batch route.
  • packages/cli/src/serve/server/telemetry.ts — main's POST /workspace/trust/grant entries touch resolveDaemonTelemetryRoute, not legacySessionTelemetryRoutes; the PR's catalog addition keeps the drift-guard count at 80, which the merged test confirms.
  • packages/sdk-typescript/src/daemon/* — PR's batch live-state types/client method are additive and coexist with main's additions.
  • docs/developers/qwen-serve-protocol.md — merged doc carries both the PR's workspace_session_live_state_batch row/description and main's hosted-persona rows; the docs-contract test verifies the table against the registry.

Merge committed as 44ad0e9cb7 (merge of 97065561a7 + f07c190cd6), following the branch's existing merge-commit style.

No guards, branches, or behaviors were added by this round beyond the merge itself, so no mutation probe applies; no settings source changed, so the settings schema was not regenerated.

Environment note

The default corepack cache location (/home/github-runner/.cache/node/corepack) is not writable in this sandbox, so all package-manager commands ran with COREPACK_HOME=/tmp/corepack, one of the pre-provisioned caches containing pnpm 11.24.0. This is local environment setup only; no check was skipped.

Verification

All commands ran against the merged tree before the commit; the resolved file was additionally re-verified against the committed tree (working tree is clean and identical to HEAD).

  • npm run build — passed
  • npm run typecheck — passed
  • npm run lint — passed
  • vitest (packages/cli): capabilities-docs-contract.test.ts, multi-workspace-sessions.test.ts, rate-limit.test.ts, rest-integration-docs-contract.test.ts, server.test.ts, server/telemetry-catalog.test.ts, server/telemetry.test.ts — 7 files, 1695 tests passed
  • vitest (packages/sdk-typescript): test/unit/DaemonClient.test.ts, test/unit/daemon-public-surface.test.ts — 2 files, 524 tests passed
  • npm run bundle — passed
  • integration cli/qwen-serve-routes.test.ts (QWEN_SANDBOX=false) — 42 tests passed
  • Post-commit: prettier --check capabilities-docs-contract.test.ts — passed; re-run of capabilities-docs-contract.test.ts — 6 tests passed
中文说明

Autofix 本轮总结 — PR #12513

反馈分诊

feedback.md 中没有任何可处理的条目:自上次评估(2026-10-03T03:25:17Z)以来,没有新的评审、行内评论、issue 级评论、失败检查或持续失败的检查。没有需要实现、拒绝、延后或升级处理的内容。

合并冲突解决(--conflict true)

PR 分支已与 main 分叉(合并基点 a7deb01bcb;main 前进到 f07c190cd6,约 50 个提交)。合并 origin/main 只产生了一个文本冲突:

  • packages/cli/src/serve/capabilities-docs-contract.test.ts —— 双方在同一个锚点测试(advertises batched session catalogs for single and multiple workspaces)之后各自追加了一个新测试:

    • PR 侧:advertises batched workspace live-state snapshots(锁定本 PR 新增的 workspace_session_live_state_batch 能力)。
    • main 侧:derives the hosted persona tags from the registry, never a hand list(锁定 hostedPersonaServeFeatures 的派生逻辑)。

    解决方案:两个测试原样全部保留,按顺序排列。自动合并的 import 块已经同时携带了双方的导入(main 的 hostedPersonaServeFeatures、SERVE_CAPABILITY_REGISTRY)。没有删除或削弱任何测试。

除文本合并外,还审计了语义重叠:

  • packages/cli/src/serve/capabilities.ts —— 干净合并;注册表同时包含 workspace_session_live_state_batch(PR)和 hosted-persona 条目(main)。
  • packages/cli/src/serve/routes/session.ts —— main 的 sendStandaloneActionUnsupported 重构和 side-task 处理器迁移与本 PR 的 readWorkspaceLiveState 提取及新增 POST /sessions/live-state 批量路由互不干扰。
  • packages/cli/src/serve/server/telemetry.ts —— main 的 POST /workspace/trust/grant 条目作用于 resolveDaemonTelemetryRoute,而非 legacySessionTelemetryRoutes;本 PR 的目录新增使漂移防护计数保持为 80,合并后的测试已确认。
  • packages/sdk-typescript/src/daemon/* —— PR 的批量 live-state 类型和客户端方法均为纯增量,与 main 的新增内容共存。
  • docs/developers/qwen-serve-protocol.md —— 合并后的文档同时包含本 PR 的 workspace_session_live_state_batch 行/描述和 main 的 hosted-persona 行;docs-contract 测试会依据注册表校验该表。

合并已提交为 44ad0e9cb7(97065561a7 + f07c190cd6 的合并),沿用该分支既有的合并提交风格。

本轮除合并本身外没有新增任何守卫、分支或行为,因此不适用变异探针;未改动 settings 源文件,故未重新生成 settings schema。

环境说明

沙箱中默认的 corepack 缓存位置(/home/github-runner/.cache/node/corepack)不可写,因此所有包管理命令均使用 COREPACK_HOME=/tmp/corepack(预置缓存之一,内含 pnpm 11.24.0)。这只是本地环境设置,没有跳过任何检查。

验证

所有命令均在提交前针对合并后的工作树运行;冲突解决文件在提交后又针对已提交的树额外复验(工作树干净,与 HEAD 完全一致)。

  • npm run build —— 通过
  • npm run typecheck —— 通过
  • npm run lint —— 通过
  • vitest(packages/cli):capabilities-docs-contract.test.ts、multi-workspace-sessions.test.ts、rate-limit.test.ts、rest-integration-docs-contract.test.ts、server.test.ts、server/telemetry-catalog.test.ts、server/telemetry.test.ts —— 7 个文件,1695 个测试通过
  • vitest(packages/sdk-typescript):test/unit/DaemonClient.test.ts、test/unit/daemon-public-surface.test.ts —— 2 个文件,524 个测试通过
  • npm run bundle —— 通过
  • 集成测试 cli/qwen-serve-routes.test.ts(QWEN_SANDBOX=false)—— 42 个测试通过
  • 提交后:prettier --check capabilities-docs-contract.test.ts —— 通过;重跑 capabilities-docs-contract.test.ts —— 6 个测试通过

Base-conflict check · 基分支冲突检查: conflicted with main — resolved in this push. · 与 main 有冲突——已在本次推送中解决。

Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-review-bot

qwen-code-review-bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

🩺 serve daemon A/B

Built the PR base vs this PR head 62ecc1c, drove a fixed endpoint set against each, and diffed the JSON responses. Only fields that changed are shown.

capabilities

field PR base (before) this PR (after)
features[] — "workspace_session_live_state_batch"

health-deep-with-session

field PR base (before) this PR (after)
activeWorkStaleMs 5 6

— Qwen Code · serve A/B

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed. Suggestions are inline.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 1)": did not execute the new batch workspace session live-state route tests myself; pass/fail and flake status rest on the prior round's measured build/test result….

中文说明

已审查。 建议见行内评论。

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 1)":did not execute the new batch workspace session live-state route tests myself; pass/fail and flake status rest on the prior round's measured build/test result…。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment thread packages/cli/src/serve/multi-workspace-sessions.test.ts
Comment thread docs/developers/qwen-serve-protocol.md
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 AutoFix ran out of time before finishing (timeout (3600000ms)) (attempt 6/100) — it will retry on the next scan.

⚠️ This change was NOT pushed — any commit referenced below was made only in the runner workspace and has been discarded. What the agent reported:
Qwen failed during address-review: timeout (3600000ms).

See the Qwen Autofix agent step logs for model/tool output.

中文说明

🤖 AutoFix 在完成前耗尽了时间(timeout (3600000ms))(第 6/100 次尝试)—— 将在下次扫描时重试。

Run log: https://github.com/QwenLM/qwen-code/actions/runs/37609284075


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

qwen-code-autofix and others added 2 commits October 7, 2026 13:10
The autofix verification gate runs the packages/cli suite with the host's
real HOME on a machine shared with other jobs. A malformed
$HOME/.qwen/settings.json there (observed content: "X") makes every
operator-settings reader fail closed at once — readOperatorSettingsScopes
is designed to refuse a corrupted policy file — and one such file failed
408 tests across 15 files in the measured gate rejection.

Pin a private empty QWEN_HOME per test-file process in the suite setup
when the environment did not select one, following the file's existing
doctrine for operator-owned environment. Tests that exercise QWEN_HOME
set, stub, or delete it themselves. settings.test.ts keys every
user-settings and .env path off a mocked os.homedir, so it now deletes
an ambient QWEN_HOME explicitly; with any ambient value set, 13 of its
tests fail regardless of this change.
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Addressed the latest review feedback (round 7/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 7/100 轮)。改动内容与我反驳保留之处如下:

Autofix round summary — PR #12513 (verification repair)

Commit: 4862869926 test(cli): isolate the suite from the operator's settings home (#12513)

This was a Same-run verification repair round: the previous commit
(1730596b45, preserved untouched) was rejected by deterministic verification
because packages/cli tests failed — 15 files / 408 tests, every one with:

Cannot read operator sandbox policy from /home/github-runner/.qwen/settings.json:
SyntaxError: Unexpected token 'X', "X\n" is not valid JSON.

Root cause (evidence, not a guess)

  • The gate runs the suite with the host's real HOME on a machine the gate
    script itself documents as "shared with up to 20 other autofix jobs"
    (run-autofix-review-verification.sh:576). During that run, the shared
    $HOME/.qwen/settings.json contained X/X\n.
  • An exhaustive repo-wide search found no test or tool that writes X to a
    settings file — the pollution is external cross-job contamination of the
    shared runner home, not something this branch (or this repo's suite) writes.
  • readOperatorSettingsScopes fails closed on a malformed operator settings
    file by design (a corrupted confinement policy must never silently open the
    sandbox), so one polluted file took down every test that loads operator
    settings without stubbing QWEN_HOME.
  • Reproduced locally: with QWEN_HOME unset and HOME pointing at a
    directory whose .qwen/settings.json contains X\n,
    src/nonInteractive/session.test.ts fails exactly 42/42 tests with the
    gate's verbatim error (the gate's junit recorded 42 failures for that file).

Fix (two test-infrastructure files, both inside the PR's packages/cli footprint)

  1. packages/cli/test-setup.ts — pin a private empty QWEN_HOME
    (per-test-file-process temp dir, removed on exit) when the environment
    did not select one
    . This follows the file's existing doctrine for
    operator-owned environment (QWEN_RUNTIME_DIR, QWEN_REVIEW_SANDBOX, …).
    Tests that exercise QWEN_HOME semantics set, stub, or delete it
    themselves, exactly as before.
  2. packages/cli/src/config/settings.test.ts — delete an ambient QWEN_HOME
    before resolving its module-level user-settings path. This file mocks
    os.homedir and keys every user-settings/.env path off the mock; an
    ambient QWEN_HOME diverts those paths away from it. Measured: 13 of its
    tests fail with any ambient QWEN_HOME set, on the pristine tree too —
    a pre-existing latent fragility the pin made deterministic.

No production code changed; no assertion was removed or weakened (one
delete process.env statement was added). The fail-closed reader behavior is
untouched.

Mutation probe

  • Guard negated (pre-fix tree), polluted HOME: session.test.ts → 42
    failed / 42
    with the gate's exact error.
  • Guard restored, same polluted HOME: 42 passed / 42.
  • Breadth witness: all 15 suites the gate reported as failed (including
    run-qwen-serve, workspace-qualified-extensions, ControlDispatcher,
    systemController) run against the same polluted HOME with QWEN_HOME
    unset → 1440 passed (1440), 0 failures, zero occurrences of the
    operator-settings error in the log.
  • settings.test.ts with an ambient QWEN_HOME on the pristine tree → 13
    failures (proof the fragility predates this change); after the fix, with the
    pin active → 239 passed (239).

Feedback dispositions

  • rc:4205820796 (R1-1) and rc:4205820805 (R1-2) — both remain resolved
    in the code
    by the preserved commit 1730596b45 and were re-verified this
    round: accepts the documented maxima
    (multi-workspace-sessions.test.ts:7579), publishes the runtime-owned batch bounds in the protocol prose
    (rest-integration-docs-contract.test.ts:973), and the interpolated 413
    message (routes/session.ts:9772) all pass in the full-suite run. Listed in
    resolved-comments.txt.
  • ic:6031957776 (serve A/B) — informational; no action required.

Conflict notes

--conflict false; no merge performed. No defer-to-human items. No test
weakening (no test-weakening.json). No deferred findings. No unresolved
inline findings (no comment-replies.json). Self-review: off (not armed).

Verification

All run on this tree, in a gate-like environment (agent-specific
QWEN_*/SANDBOX variables scrubbed, QWEN_HOME unset, CI=true, writable
HOME; the gate's own env is an env -i allowlist with the host HOME):

  • npm run build — passed (exit 0)
  • npm run typecheck — passed (exit 0)
  • npm run lint — passed (exit 0)
  • npx eslint packages/cli/test-setup.ts + npx prettier --check on both
    changed files — clean
  • cd packages/cli && npx vitest run src/config/settings.test.ts (pin
    active) — 239 passed (239)
  • Pollution witness: 15 gate-failed suites against X\n-polluted
    $HOME/.qwen/settings.json, QWEN_HOME unset — 15 files / 1440 tests
    passed (witness exit 0)
  • Full packages/cli suite (npx vitest run, gate load clamps
    --maxWorkers=25% --testTimeout=60000 --hookTimeout=60000 --coverage.enabled=false) — 1214 files passed | 1 skipped, 39547 tests
    passed | 210 skipped, exit 0
  • Exact gate command (npm run test --workspace packages/cli --if-present -- --changed origin/main --passWithNoTests <clamps>) — same full-scope result
    (changing test-setup.ts places it in every test file's module graph, so
    --changed selects the entire suite): 1214 files / 39547 tests passed,
    exit 0
  • Pre-commit hook (prettier + eslint on staged files) — passed during commit

Note on scope: because this round touches test-setup.ts, the gate's
--changed origin/main selection expands to the whole packages/cli suite —
the full-suite result above is exactly what the gate will re-run.

中文说明

Autofix 轮次总结 — PR #12513(验证修复)

提交:4862869926 test(cli): isolate the suite from the operator's settings home (#12513)

本轮是同轮验证修复:上一个提交(1730596b45,原样保留)被确定性验证拒绝,原因是 packages/cli 测试失败——15 个文件 / 408 个用例,全部报:

Cannot read operator sandbox policy from /home/github-runner/.qwen/settings.json:
SyntaxError: Unexpected token 'X', "X\n" is not valid JSON.

根因(基于证据,而非猜测)

  • 验证门在宿主机上以真实 HOME 运行测试套件,而门脚本自己记录该机器「与多达 20 个其他 autofix 作业共享」(run-autofix-review-verification.sh:576)。运行期间,共享的 $HOME/.qwen/settings.json 内容为 X/X\n。
  • 对全仓库做了穷尽搜索,没有任何测试或工具会向 settings 文件写入 X——该污染是共享 runner 主目录上的跨作业外部污染,并非本分支(或本仓库测试套件)写入。
  • readOperatorSettingsScopes 对损坏的 operator settings 文件按设计 fail-closed(被损坏的隔离策略文件绝不能被静默放开),因此一个被污染的文件击垮了所有未 stub QWEN_HOME 而读取 operator settings 的测试。
  • 本地复现:在 QWEN_HOME 未设置、HOME 指向 .qwen/settings.json 内容为 X\n 的目录时,src/nonInteractive/session.test.ts 恰好 42/42 全部失败,报错与线上逐字一致(线上 junit 记录该文件 42 个失败)。

修复(两个测试基础设施文件,均在 PR 的 packages/cli 足迹内)

  1. packages/cli/test-setup.ts——当环境未指定时,为每个测试文件进程钉一个私有空 QWEN_HOME(临时目录,退出时清理)。这遵循该文件对 operator 环境变量的既有处理原则(QWEN_RUNTIME_DIR、QWEN_REVIEW_SANDBOX 等)。需要操作 QWEN_HOME 语义的测试照旧自行 set、stub 或 delete。
  2. packages/cli/src/config/settings.test.ts——在解析模块级 user-settings 路径前删除环境中的 QWEN_HOME。该文件 mock 了 os.homedir 并以它为基准推导所有 user-settings/.env 路径;环境中的 QWEN_HOME 会把这些路径引离 mock。实测:在原始树上只要设置了任意 QWEN_HOME,该文件就有 13 个用例失败——这是 pin 之前就存在的隐性脆弱,本次只是让它确定性地显现并修掉。

未改动任何生产代码;未删除或削弱任何断言(只新增了一条 delete process.env 语句)。fail-closed 的读取行为保持不变。

变异探测

  • 去掉防护(修复前的树)+ 被污染的 HOME:session.test.ts → 42/42 失败,报错与线上逐字一致。
  • 恢复防护 + 同样的污染 HOME:42/42 通过。
  • 广度见证:线上失败的全部 15 个套件(含 run-qwen-serve、workspace-qualified-extensions、ControlDispatcher、systemController)在同样的污染 HOME、未设 QWEN_HOME 下运行 → 1440 通过(1440),0 失败,日志中 operator-settings 错误出现次数为 0。
  • 在原始树上给 settings.test.ts 设置环境 QWEN_HOME → 13 个失败(证明该脆弱先于本次改动存在);修复后在 pin 生效下 → 239 通过(239)。

反馈处置

  • rc:4205820796(R1-1)与 rc:4205820805(R1-2)——两者在代码中仍由保留的提交 1730596b45 解决,本轮已复核:accepts the documented maxima(multi-workspace-sessions.test.ts:7579)、publishes the runtime-owned batch bounds in the protocol prose(rest-integration-docs-contract.test.ts:973)以及插值后的 413 消息(routes/session.ts:9772)在全量套件运行中均通过。已列入 resolved-comments.txt。
  • ic:6031957776(serve A/B)——信息性评论,无需处理。

冲突说明

--conflict false;未做合并。无需要人类决定的条目。未削弱任何测试(无需 test-weakening.json)。无延后发现。无未解决的行内发现(无需 comment-replies.json)。Self-review:off(未启用)。

验证

以下均在本树上、类线上环境中执行(清除 agent 特有的 QWEN_*/SANDBOX 变量,QWEN_HOME 未设,CI=true,HOME 可写;线上门本身是带宿主机 HOME 的 env -i 白名单环境):

  • npm run build — 通过(exit 0)
  • npm run typecheck — 通过(exit 0)
  • npm run lint — 通过(exit 0)
  • npx eslint packages/cli/test-setup.ts,并对两个改动文件执行 npx prettier --check — 全部通过
  • cd packages/cli && npx vitest run src/config/settings.test.ts(pin 生效)— 239 通过(239)
  • 污染见证:15 个线上失败的套件在 $HOME/.qwen/settings.json 内容为 X\n、QWEN_HOME 未设下运行 — 15 个文件 / 1440 个用例通过(exit 0)
  • packages/cli 全量套件(npx vitest run,线上负载钳制 --maxWorkers=25% --testTimeout=60000 --hookTimeout=60000 --coverage.enabled=false)— 1214 个文件通过 | 1 跳过,39547 个用例通过 | 210 跳过,exit 0
  • 逐字的线上门命令(npm run test --workspace packages/cli --if-present -- --changed origin/main --passWithNoTests <clamps>)— 同样的全量结果(改动 test-setup.ts 会进入每个测试文件的模块图,因此 --changed 选中整个套件):1214 个文件 / 39547 个用例通过,exit 0
  • 提交时运行的 pre-commit 钩子(对暂存文件执行 prettier + eslint)— 通过

范围说明:由于本轮改动了 test-setup.ts,线上门的 --changed origin/main 选择会扩展到整个 packages/cli 套件——上面的全量结果正是门将重跑的内容。

🦷 Gate advisory — this round's changed tests all pass on the pre-round tree (machine-measured, not agent-authored). Expected for a refactor or coverage addition; if this round was meant to FIX a defect, that defect did not reproduce. · 本轮改动的测试在轮前树上全部通过(门自动测量,非 agent 文本)。对重构或补充覆盖属正常;若本轮意在修复缺陷,则该缺陷未能复现。

Deferred non-Critical feedback

Critical-only mode is active: 5 change-producing rounds are complete. The workflow excluded the non-Critical feedback below from this round's actionable sections; the items remain open for human follow-up. Maintainer feedback is deferred only after its author has used 2 regular feedback batches in this window's Critical-only tail; authors at that budget, if any, are named below. (@qwen-code /retry starts a fresh counting window.)

中文说明

已进入仅处理 Critical 的模式:已完成 5 个产生改动的轮次。本轮可执行区域已排除下方非 Critical 反馈;这些条目保持开放,留待人工跟进。维护者反馈仅在其本人于本窗口 Critical-only 阶段已使用 2 批常规反馈预算后才会延后;达到预算的作者(如有)在下方点名。(评论 @qwen-code /retry 可开启新的计数窗口。)

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 2 selected review thread(s). · 已关闭全部选中的 2 条评审线程。

Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed. Suggestions are inline.

Not explored to full depth (tool budget reached): "agent 1c": full packages/cli suite run under the new QWEN_HOME pin — sampled the 21 highest-risk files (~1,070 tests) instead of all ~1,000 files.; "agent reverse-audit (round 1)": I did not empirically mutate-test the new guards (removing if (req.aborted || res.destroyed) abort() , the post-loop abort check, or the MAX_MEMBER_BYTES bra…; "agent 4": none — I stopped at ~16 tool calls with every candidate above resolved by inspection..

Deferred under the convergence posture (round 2, not a blocker) — recorded, not requested in this round:

  • packages/cli/src/serve/routes/session.ts:9785 — [review] The batch member's 500 path skips the shared bridge-error…
中文说明

已审查。 建议见行内评论。

未探索到全部深度(达到工具调用预算):"agent 1c":full packages/cli suite run under the new QWEN_HOME pin — sampled the 21 highest-risk files (~1,070 tests) instead of all ~1,000 files.;"agent reverse-audit (round 1)":I did not empirically mutate-test the new guards (removing if (req.aborted || res.destroyed) abort() , the post-loop abort check, or the MAX_MEMBER_BYTES bra…;"agent 4":none — I stopped at ~16 tool calls with every candidate above resolved by inspection.。

收敛姿态下延后(第 2 轮,非阻断)——已记录,本轮不要求修改:共 1 条(原文未翻译,列表见上方英文部分)。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment thread packages/cli/src/serve/multi-workspace-sessions.test.ts
Comment thread packages/cli/test-setup.ts Outdated
Comment thread packages/cli/test-setup.ts
Comment thread packages/cli/test-setup.ts Outdated
qwen-code-autofix added 2 commits October 7, 2026 19:55
- Pin on truthiness so an ambient empty QWEN_HOME no longer falls
  through to $HOME/.qwen, matching every production reader.
- Sweep the pinned home from afterAll instead of an 'exit' listener
  that vitest workers never fire (51 leaked dirs on src/config -> 0).
- Restore QWEN_DEBUG_LOG_FILE around loadCliConfig tests so the debug
  logger's fire-and-forget writes stop racing the sweep.
- Witness that the suite never resolves settings from the ambient
  $HOME/.qwen, and assert member shape in the 20-selector maxima half
  so twenty failed members cannot read as twenty snapshots.
@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Addressed the latest review feedback (round 8/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 8/100 轮)。改动内容与我反驳保留之处如下:

Autofix round — PR #12513 (verification repair)

This round was a same-run verification repair: the previous commit (07491c2f2f) was rejected by deterministic verification because the full packages/cli suite failed with one unhandled rejection — ENOENT: no such file or directory, open '/tmp/qwen-cli-test-pIOJVW/extension-store/lock' from ExtensionStore.prepareDirectories, originating in src/ui/AppContainer.test.tsx. The rejected commit is preserved; this round adds one verified follow-up commit: 62ecc1cdd4 test(cli): keep AppContainer store work off the pinned per-file QWEN_HOME (#12513) (1 file, +27/−6).

No review feedback was actionable this round (Critical-only mode; the deferred section was empty; all four inline Suggestions were resolved in the previous round and are re-verified holding below).

Root cause (probed, not guessed)

Reproduced the gate failure locally with the pin active (env -u QWEN_HOME): src/ui/AppContainer.test.tsx alone produced the unhandled rejection in 4 of 5 runs (extension_store_busy caused by ENOENT ... lstat '/tmp/.../qwen-cli-test-*' — the same class as the gate's raw ENOENT from prepareDirectories; both are the store's withLock chain finding its tree deleted mid-flight).

Instrumented probes (temporary, reverted) established the exact mechanism:

  • Every store op in the file is readConsistent (~200 per run), reached from the AppContainer mount effect's un-awaited config.initialize() → ExtensionManager.refreshCache() — one per test render.
  • Timestamped probe log showed a store op entering 14 ms before test-setup.ts's per-file afterAll rmSync of the pinned qwen-cli-test-* home (mid-chain when the sweep lands → ENOENT → unhandled rejection), and another op firing 3 ms after it (fires post-teardown — no in-file drain can prevent it).

Any in-worker deletion of the pinned tree is therefore inherently racy for this file: the writer is fire-and-forget production behavior that cannot be observed or awaited from the setup file, and vitest workers are terminated without an exit event (the worker's own SIGTERM → process.exit() handler is registered at bootstrap, before setup files, so no user handler can run at worker death either).

Fix

packages/cli/src/ui/AppContainer.test.tsx already solved this exact race for HOME: it redirects HOME to a scratch directory it deliberately never deletes, with a comment documenting the ENOENT mechanism. The suite's new QWEN_HOME pin reintroduced the race because QWEN_HOME outranks HOME for the store root (Storage.getGlobalQwenDir). The fix extends that existing redirect to QWEN_HOME (set and restored alongside HOME), so the file's in-flight store work targets the never-deleted scratch tree and the pin's afterAll sweep only ever deletes its own (empty) pinned dir.

Also added a witness test (keeps the extension store root on the never-deleted suite scratch home) asserting the redirect, so its accidental removal fails deterministically instead of flaking the gate.

Mutation probes (each reverted after measurement):

  • Mutant = the round's base itself (no redirect): 4/5 focused runs failed with the unhandled rejection. With the fix: 5/5 runs clean, 206/206 tests green, no Errors line.
  • Witness probe: removing the QWEN_HOME redirect line turns the new witness test red (1 failed | 205 skipped); restoring it returns the file to green.

Residual housekeeping: the file still leaves one qwen-appcontainer-home-* scratch dir per run (pre-existing, deliberate — deleting it is what caused this failure class) and an empty skills/ stub can reappear in the pinned dir when the exception-safe SkillManager.ensureUserSkillsDir (try/catch + debugLogger.warn) lands after the sweep — measured 3/3 pre-fix and 4/5 post-fix, i.e. pre-existing and unaffected by this change, and incapable of producing an unhandled rejection.

Prior findings re-verified holding

  • rc:4210466886 (R1-1): per-member toMatchObject({ workspaceId, v: 1, sessions }) assertion present in accepts the documented maxima (multi-workspace-sessions.test.ts:7586+). This round touches only AppContainer.test.tsx; the batch tests are unaffected.
  • rc:4210466909 (R2-1): truthiness guard if (!process.env['QWEN_HOME']) present (test-setup.ts:67).
  • rc:4210466921 (R2-2): test-setup.test.ts witness present and green (1 passed).
  • rc:4210466929 (R2-3): per-file afterAll cleanup present (test-setup.ts:73); focused AppContainer runs leave zero qwen-cli-test-* content behind except the pre-existing benign skills/ stub noted above.

Failed checks

  • Hosted process fault gates / MySQL 8.4 / Java 21 (SDK Java): not diagnosable from this environment — no GitHub credentials to read the job log, and the job's toolchain (JDK 21, Maven, MySQL 8.4 service container) does not exist on this runner, so no local reproduction or surrogate is possible. This round's change is test-only in packages/cli and cannot move that job in either direction. The workflow's post-push CI re-run remains the final gate.

Verification

  • npm run build — passed (exit 0)
  • npm run typecheck — passed (exit 0)
  • npm run lint — passed (exit 0)
  • npx prettier --check packages/cli/src/ui/AppContainer.test.tsx — passed
  • Focused: vitest run src/ui/AppContainer.test.tsx (pin active, private TMPDIR) ×5 — 206/206 passed, no unhandled errors in any run (base: 4/5 runs had the Errors 1 error rejection)
  • Mutation probe: redirect line removed → witness test red (1 failed | 205 skipped); restored → green
  • Pin witness: vitest run test-setup.test.ts — 1 passed
  • Gate mirror: vitest run --changed origin/main --passWithNoTests --maxWorkers=25% --testTimeout=60000 --hookTimeout=60000 --coverage.enabled=false (full packages/cli suite, 1216 files) — no unhandled errors (the rejection class this round repairs); 18 test failures in 7 files on this runner, all reproduced identically with the change stashed and traced to this agent shell's SANDBOX/QWEN_CODE_* env leaking into the run (docker/sandbox/browser environment detection): with a clean env the same 7 files drop to 2 failures, the pre-existing EACCES: mkdtemp under the read-only real $HOME in cdCommand.test.ts/directoryCommand.test.tsx documented last round (CI's Test (ubuntu-latest, Node 22.x) is green on this head, and the gate's own failing run showed 0 test failures)

No settings source changed, so generate:settings-schema was not needed; the change is test-only, so no integration run applies. No test was deleted, skipped, or weakened (one test added), so no test-weakening.json; no unresolved inline findings, so no comment-replies.json.

中文说明

Autofix 本轮处理 — PR #12513(验证修复)

本轮是一次同轮验证修复:上一个提交(07491c2f2f)被确定性验证拒绝,原因是 packages/cli 完整套件因一个未处理的 rejection 失败——ENOENT: no such file or directory, open '/tmp/qwen-cli-test-pIOJVW/extension-store/lock',来自 ExtensionStore.prepareDirectories,归属于 src/ui/AppContainer.test.tsx。被拒绝的提交保留不动;本轮在其上追加了一个经过验证的修复提交:62ecc1cdd4 test(cli): keep AppContainer store work off the pinned per-file QWEN_HOME (#12513)(1 个文件,+27/−6)。

本轮没有可执行的评审反馈(仅 Critical 模式;延后区段为空;4 条行内 Suggestion 已在上一轮解决,并在下方重新验证仍然成立)。

根因(探针实测,非猜测)

在 pin 生效的条件下(env -u QWEN_HOME)本地复现了 gate 失败:仅运行 src/ui/AppContainer.test.tsx,5 次中有 4 次出现该未处理 rejection(extension_store_busy,其 cause 为 ENOENT ... lstat '/tmp/.../qwen-cli-test-*'——与 gate 上 prepareDirectories 抛出的原始 ENOENT 同属一类:两者都是 store 的 withLock 链在中途发现自己的目录树被删除)。

通过(临时、已还原的)埋点探针确认了确切机制:

  • 该文件中的每个 store 操作都是 readConsistent(每次运行约 200 次),来自 AppContainer 挂载副作用中未 await 的 config.initialize() → ExtensionManager.refreshCache()——每次测试渲染一次。
  • 带时间戳的探针日志显示:一个 store 操作在 test-setup.ts 按文件注册的 afterAll 对被 pin 的 qwen-cli-test-* 目录执行 rmSync 之前 14 毫秒进入(清扫落在链条中途 → ENOENT → 未处理 rejection),另一个操作在其之后 3 毫秒才触发(在 teardown 之后才发起——任何文件内的 drain 都无法阻止它)。

因此,对该文件而言,worker 内对被 pin 目录树的任何删除本质上都是竞态:写入方是 fire-and-forget 的生产行为,setup 文件无法观察也无法等待它;而且 vitest worker 被终止时不会触发 exit 事件(worker 自己的 SIGTERM → process.exit() 处理器在引导阶段注册,早于所有 setup 文件,因此用户也不可能在 worker 死亡时运行清理处理器)。

修复

packages/cli/src/ui/AppContainer.test.tsx 其实早已为 HOME 解决过完全相同的竞态:它把 HOME 重定向到一个刻意从不删除的临时目录,并有注释记录了 ENOENT 机制。套件新增的 QWEN_HOME pin 重新引入了这个竞态,因为对于 store 根目录,QWEN_HOME 的优先级高于 HOME(Storage.getGlobalQwenDir)。修复方式是把这个既有的重定向扩展到 QWEN_HOME(与 HOME 一起设置、一起还原),使该文件在途的 store 操作指向那棵永不删除的临时树,而 pin 的 afterAll 清扫只删除它自己的(空)pin 目录。

同时新增了一个见证测试(keeps the extension store root on the never-deleted suite scratch home)断言该重定向存在,这样如果重定向被意外移除,会以确定性失败暴露,而不是让 gate 偶发失败。

变异探针(每个均在测量后还原):

  • 变异体即本轮基线本身(无重定向):5 次聚焦运行中 4 次因未处理 rejection 失败。修复后:5/5 次运行干净,206/206 测试通过,无 Errors 行。
  • 见证探针:删除 QWEN_HOME 重定向行后新见证测试变红(1 failed | 205 skipped);还原后恢复绿色。

残留 housekeeping:该文件每次运行仍会留下一个 qwen-appcontainer-home-* 临时目录(既有行为、刻意保留——删除它正是这次失败类别的成因);当被 try/catch + debugLogger.warn 保护、异常安全的 SkillManager.ensureUserSkillsDir 在清扫之后才落盘时,pin 目录里可能重新出现一个空的 skills/ 残壳——实测修复前 3/3、修复后 4/5,即既有现象、与本改动无关,且不可能产生未处理 rejection。

既有发现重新验证仍然成立

  • rc:4210466886(R1-1):accepts the documented maxima 中按成员的 toMatchObject({ workspaceId, v: 1, sessions }) 断言存在(multi-workspace-sessions.test.ts:7586+)。本轮只改动 AppContainer.test.tsx,批量路由测试不受影响。
  • rc:4210466909(R2-1):真值守卫 if (!process.env['QWEN_HOME']) 存在(test-setup.ts:67)。
  • rc:4210466921(R2-2):test-setup.test.ts 见证存在且为绿(1 passed)。
  • rc:4210466929(R2-3):按文件的 afterAll 清理存在(test-setup.ts:73);聚焦运行后 qwen-cli-test-* 无内容残留,仅有上述既有的良性 skills/ 残壳。

失败的检查

  • Hosted process fault gates / MySQL 8.4 / Java 21(SDK Java):本环境无法诊断——没有 GitHub 凭据读取该 job 日志,且本 runner 上没有该 job 的工具链(JDK 21、Maven、MySQL 8.4 服务容器),无法在本地复现或构造替代验证。本轮改动仅涉及 packages/cli 的测试,不可能从任何方向影响该 job。推送后的 CI 重跑仍是最终验收门禁。

验证

  • npm run build — 通过(exit 0)
  • npm run typecheck — 通过(exit 0)
  • npm run lint — 通过(exit 0)
  • npx prettier --check packages/cli/src/ui/AppContainer.test.tsx — 通过
  • 聚焦:vitest run src/ui/AppContainer.test.tsx(pin 生效、私有 TMPDIR)×5 — 206/206 通过,任何一次运行都无未处理错误(基线:5 次中 4 次出现 Errors 1 error 的 rejection)
  • 变异探针:删除重定向行 → 见证测试变红(1 failed | 205 skipped);还原 → 变绿
  • pin 见证:vitest run test-setup.test.ts — 1 通过
  • gate 镜像:vitest run --changed origin/main --passWithNoTests --maxWorkers=25% --testTimeout=60000 --hookTimeout=60000 --coverage.enabled=false(packages/cli 完整套件,1216 个文件)— 无未处理错误(即本轮修复的 rejection 类别);本 runner 上有 7 个文件、18 个测试失败,全部在改动被 stash 后以完全相同的方式复现,并追溯到本 agent shell 的 SANDBOX/QWEN_CODE_* 环境变量泄漏进测试运行(docker/sandbox/浏览器环境检测):用干净环境重跑这 7 个文件后只剩 2 个失败,即上一轮已记录的、在只读真实 $HOME 下 EACCES: mkdtemp 的 cdCommand.test.ts/directoryCommand.test.tsx 既有问题(CI 的 Test (ubuntu-latest, Node 22.x) 在该 head 上为绿,且 gate 自己那次失败运行的测试失败数为 0)

未改动任何 settings 源文件,故无需运行 generate:settings-schema;改动仅涉及测试,无需集成测试运行。没有删除、跳过或削弱任何测试(新增了一个),故无 test-weakening.json;没有未解决的行内发现,故无 comment-replies.json。

🦷 Gate advisory — this round's changed tests all pass on the pre-round tree (machine-measured, not agent-authored). Expected for a refactor or coverage addition; if this round was meant to FIX a defect, that defect did not reproduce. · 本轮改动的测试在轮前树上全部通过(门自动测量,非 agent 文本)。对重构或补充覆盖属正常;若本轮意在修复缺陷,则该缺陷未能复现。

Deferred non-Critical feedback

Critical-only mode is active: 5 change-producing rounds are complete. The workflow excluded the non-Critical feedback below from this round's actionable sections; the items remain open for human follow-up. Maintainer feedback is deferred only after its author has used 2 regular feedback batches in this window's Critical-only tail; authors at that budget, if any, are named below. (@qwen-code /retry starts a fresh counting window.)

中文说明

已进入仅处理 Critical 的模式:已完成 5 个产生改动的轮次。本轮可执行区域已排除下方非 Critical 反馈;这些条目保持开放,留待人工跟进。维护者反馈仅在其本人于本窗口 Critical-only 阶段已使用 2 批常规反馈预算后才会延后;达到预算的作者(如有)在下方点名。(评论 @qwen-code /retry 可开启新的计数窗口。)

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。

🧵 Resolved all 4 selected review thread(s). · 已关闭全部选中的 4 条评审线程。

🩸 Regression charged to round 7: that round pushed onto a head whose checks were all green and left them red. It no longer counts as progress for the consecutive-failure brake. · 已将回归记在第 7 轮:该轮在检查全绿的 head 上推送后检查转红,因此不再计入连续失败熔断的有进展判定。

Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed. Suggestions are inline.

Not explored to full depth (tool budget reached): "agent 1a": full packages/cli suite run to measure the new QWEN_HOME pin's blast radius beyond the ten files I targeted individually (a single whole-package run exceeds…; "agent 5": I did not execute packages/cli tests or any mutation (no run performed) — all mutation verdicts above are read-based..

Deferred under the convergence posture (round 3, not a blocker) — recorded, not requested in this round:

  • packages/cli/src/serve/multi-workspace-sessions.test.ts:7566 — [review] the invalid-envelope enumeration omits the empty-selector case, leaving the schema's element-level min(1) unpinned
中文说明

已审查。 建议见行内评论。

未探索到全部深度(达到工具调用预算):"agent 1a":full packages/cli suite run to measure the new QWEN_HOME pin's blast radius beyond the ten files I targeted individually (a single whole-package run exceeds…;"agent 5":I did not execute packages/cli tests or any mutation (no run performed) — all mutation verdicts above are read-based.。

收敛姿态下延后(第 3 轮,非阻断)——已记录,本轮不要求修改:共 1 条(原文未翻译,列表见上方英文部分)。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment on lines +19 to +20
it('keeps the resolved user settings dir off the ambient $HOME/.qwen', () => {
expect(getUserSettingsDir()).not.toBe(path.join(homedir(), '.qwen'));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R2-2: (fix-induced) The witness this round added to answer R2-2 asserts more than the guard it witnesses. test-setup.ts pins a private home only when QWEN_HOME is falsy, but this assertion demands unconditionally that the resolved settings directory differ from path.join(homedir(), '.qwen'). When an environment legitimately selects a QWEN_HOME that resolves to that same default — ~/.qwen, $HOME/.qwen, or a trailing-slash spelling, all of which Storage.resolvePath normalizes to the identical string — the pin correctly stands aside, both sides compare equal, and the whole packages/cli suite goes red with a message that reads as an isolation breach while the guard is behaving exactly as designed. That is precisely the case the header comment above this test promises is safe: it claims the assertion holds whether the pin applied or the environment legitimately selected its own QWEN_HOME. The in-tree carrier today is .github/workflows/qwen-triage.yml:4416-4418, which runs the verify agent with HOME=$AGENT_HOME and QWEN_HOME=$AGENT_HOME/.qwen; no CI job that runs this suite carries the coincidence yet (ci.yml:2119-2123 does pair them, but only for typecheck:integration and test:integration:no-ak, and integration-tests has no setupFiles entry pointing at packages/cli/test-setup.ts), so this is a false alarm waiting for an operator shell profile or a new job rather than a live red. The same assertion is also vacuous in the other direction: with an empty HOME, os.homedir() returns '', the left side resolves to <tmpdir>/.qwen while the right side is '.qwen', so it passes whether or not the pin applied.

Witness:

ARM 1  HOME=/tmp/hh, QWEN_HOME unset:
  Test Files  1 passed (1)
       Tests  1 passed (1)
ARM 2  HOME=/tmp/hh QWEN_HOME=/tmp/hh/.qwen:
  x test-setup QWEN_HOME pin > keeps the resolved user settings dir off the ambient $HOME/.qwen
  AssertionError: expected '/tmp/hh/.qwen' not to be '/tmp/hh/.qwen' // Object.is equality
  Test Files  1 failed (1)
       Tests  1 failed (1)      known-fail rc=1
Same environment, sibling suites unaffected:
  src/config/path-freshness.test.ts (3 tests) and src/config/settings.test.ts (239 tests) pass
  -> Tests 1 failed | 242 passed (243)
Empty-homedir arm (pin absent, real resolver):
  homedir()               = ""
  LHS getGlobalQwenDir()  = "<tmpdir>/.qwen"
  RHS join(homedir,.qwen) = ".qwen"
  witness .not.toBe(rhs) would PASS (vacuous)
  control, HOME=/tmp/hh:  LHS = "/tmp/hh/.qwen"  RHS = "/tmp/hh/.qwen" -> FAIL (witness bites)
Suggested change
it('keeps the resolved user settings dir off the ambient $HOME/.qwen', () => {
expect(getUserSettingsDir()).not.toBe(path.join(homedir(), '.qwen'));
it('keeps the resolved user settings dir off the ambient $HOME/.qwen', () => {
const resolved = getUserSettingsDir();
if (resolved === path.join(homedir(), '.qwen')) {
// Reaching the ambient default is legitimate only when the environment
// selected it itself; the pin stands aside for a truthy QWEN_HOME.
expect(process.env['QWEN_HOME']).toBeTruthy();
}

Any guard here has to compare the resolved value using the same truthiness test the pin uses, never !== undefined: packages/core/src/config/storage.ts:194-196 is const envDir = process.env['QWEN_HOME']; if (envDir) { return Storage.resolvePath(envDir); }, and resolvePath expands a leading ~. packages/cli/src/config/path-freshness.test.ts:26-30 also deletes QWEN_HOME in beforeEach and asserts the default at :49, so the pin must stay a single setup-file assignment that a test can delete rather than something re-applied from a hook. Exporting the pinned directory from test-setup.ts and asserting against it would additionally close the empty-HOME vacuity and keep a positive assertion in the common case, at the cost of a little module surface.

Please confirm the mutation still bites after the change: with QWEN_HOME unset, delete the if (!process.env['QWEN_HOME']) { … } block from test-setup.ts:67-76 and check this test goes red, and check the QWEN_HOME=$HOME/.qwen run goes green.

中文说明

[Suggestion] R2-2:(由修复引入)本轮为回应 R2-2 而新增的见证测试,断言的范围超出了它所见证的那个守卫。test-setup.ts 只在 QWEN_HOME 为假值时才固定一个私有 home,但这条断言无条件要求解析出的 settings 目录不等于 path.join(homedir(), '.qwen')。当环境合法地把 QWEN_HOME 选为这个默认位置时——~/.qwen、$HOME/.qwen,或带尾斜杠的写法,Storage.resolvePath 会把它们全部规范化为同一个字符串——pin 会正确地不介入,两侧比较相等,于是整个 packages/cli 测试套件变红,报错信息读起来像是隔离被破坏,而守卫其实完全按设计工作。这正是本文件上方注释承诺安全的场景:注释声称该断言「无论 pin 是否生效、或环境是否合法地选择了自己的 QWEN_HOME,都成立」。目前仓库内的触发点是 .github/workflows/qwen-triage.yml:4416-4418,它以 HOME=$AGENT_HOME 和 QWEN_HOME=$AGENT_HOME/.qwen 运行 verify agent;但运行本套件的 CI 作业目前都不带这个巧合(ci.yml:2119-2123 确实成对设置了二者,但只用于 typecheck:integration 与 test:integration:no-ak,而 integration-tests 的 setupFiles 并不指向 packages/cli/test-setup.ts),所以这是一枚等待某个操作者 shell 配置或新作业触发的误报,而不是当前的红。同一条断言在另一个方向上也是空的:HOME 为空时 os.homedir() 返回 '',左侧解析为 <tmpdir>/.qwen,右侧是 '.qwen',因此无论 pin 是否生效都会通过。

修复必须比较解析后的值,并使用与 pin 相同的真值判断,而不是 !== undefined:packages/core/src/config/storage.ts:194-196 是 const envDir = process.env['QWEN_HOME']; if (envDir) { return Storage.resolvePath(envDir); },且 resolvePath 会展开开头的 ~。packages/cli/src/config/path-freshness.test.ts:26-30 也会在 beforeEach 中删除 QWEN_HOME 并在 :49 断言默认值,所以 pin 必须保持为「setup 文件中一次性赋值、测试可以删除」的形式,而不是从钩子中反复施加。若从 test-setup.ts 导出被固定的目录并对其断言,还能同时消除空 HOME 下的空断言问题,并在常规场景保留一条正向断言,代价是增加一点模块导出面。

请在改动后确认变异仍然能被捕获:在不设置 QWEN_HOME 的情况下,删除 test-setup.ts:67-76 的 if (!process.env['QWEN_HOME']) { … } 代码块,本测试应变红;同时 QWEN_HOME=$HOME/.qwen 的运行应变绿。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment on lines +1464 to +1468
if (originalDebugLogFile === undefined) {
delete process.env['QWEN_DEBUG_LOG_FILE'];
} else {
process.env['QWEN_DEBUG_LOG_FILE'] = originalDebugLogFile;
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R3-1: This restore is the only guard the round added that no test can detect being removed. Every assertion on QWEN_DEBUG_LOG_FILE in this file establishes its own precondition first — :1851 deletes the variable and then asserts toBe('1'), :1921 assigns '0' and then asserts toBe('0'), :1931 deletes and then asserts toBeUndefined() — and a grep of the file turns up no other read of it, so deleting these five lines leaves all 509 tests green while the leak the comment above them describes comes straight back. config.ts:1721-1722 assigns '1' whenever debug mode is on and the variable is undefined, and a bare delete plus a bare assignment are both invisible to vi.unstubAllEnvs(), so the '1' written by the --debug test at :1850 survives for the roughly 330 tests that follow it. The leak is measurable rather than theoretical: with a fixed QWEN_HOME, removing the guard turns 0 session log files into 26 files totalling 4066 B under $QWEN_HOME/debug, and with the pin engaged it resurrects a pin directory that test-setup.ts's new afterAll sweep had already deleted, undoing the sibling cleanup added in the same round. It cannot fail the run, which is why this is a Suggestion and not a Critical: every debugLogger write path swallows its own rejection, and the mutated tree exited 0 on Linux where dangerouslyIgnoreUnhandledErrors is false.

Witness:

Eight-arm matrix, `vitest run src/config/config.test.ts --reporter=default`,
fresh TMPDIR and QWEN_HOME per row, on an out-of-tree copy of packages/cli.

ARM 0   intact, no witness            -> Tests 509 passed (509), exit 0
ARM 0b  intact, fixed QWEN_HOME       -> debug/ + dangling `latest`, 0 log files, 40 B
ARM 1   MUTANT (1464-1468 deleted)    -> Tests 509 passed (509), exit 0   <-- guard has no witness
ARM 1   same, fixed QWEN_HOME         -> 26 session .txt files, 4066 B
ARM 1-pin MUTANT, pin engaged         -> exit 0 on Linux (dangerouslyIgnoreUnhandledErrors:false);
                                         1 resurrected pin dir <TMPDIR>/qwen-cli-test-97DR32/debug
ARM 2   MUTANT + proposed witness     -> 1 failed | 509 passed (510)
                                         x loadCliConfig > restores QWEN_DEBUG_LOG_FILE after a --debug run
                                         AssertionError: expected '1' to be '0' // Object.is equality
ARM 3   intact + proposed witness     -> Tests 510 passed (510); pin dirs left: 0
ARM 4   intact + witness, ambient QWEN_DEBUG_LOG_FILE=1 -> 1 failed | 509 passed (spurious)
ARM 6   MUTANT + toBe(originalDebugLogFile) -> 510 passed -- NO FLIP (tautological)
ARM 7   MUTANT + collection-time seeded const -> 1 failed | 509 passed, expected '1' to be '0'
ARM 8/8b intact + seeded const, ambient '1' / unset -> 510 passed both

A witness placed right after the --debug test at :1858, reading the variable rather than touching it:

// at describe('loadCliConfig') scope, evaluated at collection time
const seededDebugLogFile = process.env['QWEN_DEBUG_LOG_FILE'];

// immediately after the --debug test at :1858
it('restores QWEN_DEBUG_LOG_FILE after a --debug run', () => {
  expect(process.env['QWEN_DEBUG_LOG_FILE']).toBe(seededDebugLogFile);
});

test-setup.ts:19-21 seeds the variable to '0' only when it is undefined, so the baseline is '0' on any run through the CLI setup file and never undefined; and because config.ts:1721 fires only on undefined, the new test must read the variable and never delete it — deleting would reproduce the :1851 setup and make the assertion self-fulfilling. Do not write the obvious toBe(originalDebugLogFile) either: the witness's own beforeEach re-captures the leaked '1' and compares it to itself, which measured green against the mutant (arm 6). A plain toBe('0') also flips correctly and is fine on CI, but it fails spuriously on a machine that exports QWEN_DEBUG_LOG_FILE=1 — an opt-in test-setup.ts:17 explicitly invites — which is why the collection-time constant is the form that is correct in both.

Please confirm the mutation: remove these five lines and run npx vitest run src/config/config.test.ts — the new test must fail on expected '1' to be '0' while the other 509 stay green.

中文说明

[Suggestion] R3-1:本轮新增的守卫中,只有这一处没有任何测试能发现它被删除。文件中所有针对 QWEN_DEBUG_LOG_FILE 的断言都先自行设置前提——:1851 先删除该变量再断言 toBe('1'),:1921 先赋值 '0' 再断言 toBe('0'),:1931 先删除再断言 toBeUndefined()——而在整个文件中搜索也找不到其它读取点,因此删除这五行后 509 个测试依然全绿,而它们上方注释所描述的泄漏会原样回来。config.ts:1721-1722 在 debug 模式开启且变量为 undefined 时赋 '1',而裸 delete 与裸赋值对 vi.unstubAllEnvs() 都不可见,所以 :1850 那个 --debug 测试写入的 '1' 会在其后约 330 个测试中一直存活。这个泄漏是可测量的,不是理论上的:在固定 QWEN_HOME 的情况下,去掉守卫会让 $QWEN_HOME/debug 下的会话日志从 0 个变成 26 个、共 4066 B;在 pin 生效时,它还会重建一个 test-setup.ts 新增的 afterAll 清扫已经删除的 pin 目录,等于抵消同一轮加入的姊妹清理。它不会让运行失败,这也是本条定为 Suggestion 而非 Critical 的原因:debugLogger 的每条写入路径都自行吞掉 rejection,且在 dangerouslyIgnoreUnhandledErrors 为 false 的 Linux 上,变异后的树仍以 0 退出。

test-setup.ts:19-21 只在变量为 undefined 时把它播种为 '0',因此经过 CLI setup 文件的运行中基线一定是 '0' 而不会是 undefined;又因为 config.ts:1721 只在 undefined 时触发,新测试必须读取该变量而绝不能删除它——删除会重演 :1851 的前提,使断言自我实现。也不要写成显而易见的 toBe(originalDebugLogFile):见证测试自己的 beforeEach 会重新捕获泄漏的 '1' 并与自身比较,实测在变异体上仍为绿(arm 6)。直接写 toBe('0') 同样能翻转,在 CI 上也没问题,但在导出了 QWEN_DEBUG_LOG_FILE=1 的机器上会误报——而 test-setup.ts:17 明确邀请这种显式覆盖——因此采集期常量是两种场景都正确的写法。

请确认变异效果:删除这五行后运行 npx vitest run src/config/config.test.ts,新测试必须以 expected '1' to be '0' 失败,而其余 509 个测试保持绿色。

— qwen3.8-max via Qwen Code /review (v0.25.0)

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下:

Autofix round: no action required

No code changes were made in this round.

Why

  • No actionable feedback. Critical-only mode is active for this PR (5 change-producing rounds complete). The only open feedback — one review body and two inline comments from @qwen-code-review-bot (rc:4213759653 on packages/cli/test-setup.test.ts:20, rc:4213759676 on packages/cli/src/config/config.test.ts:1468) — is non-Critical and was excluded from this round by the workflow's deterministic brake. Per the Critical-only rules, those items are an audit record for human follow-up: they were not implemented, and their threads were neither resolved nor replied to.
  • No failed checks. All 59 reported checks are SUCCESS, SKIPPED, or NEUTRAL; the "Failed checks" and "Still-red checks" sections are empty.
  • No base conflict. The round was invoked with --conflict false, and no merge was requested or needed.
  • No growth audit. The "Diff growth this window" section reports source 26 / test 103 net lines against budgets of 400/400 with zero prior over-budget rounds, and no Growth audit required section is present, so no audit verdict was owed.

The PR is left exactly as it was at the start of this round. The deferred non-Critical items remain open for maintainer follow-up (start a comment with [Critical], submit a Request changes review, or comment @qwen-code /retry for a fresh window).

Verification

  • git status --short — clean; no files modified or created in the repository this round.
  • checks.json inspection — 59 checks, none failing (the single non-CheckRun entry is a StatusContext with state SUCCESS).
中文说明

Autofix 本轮:无需处理

本轮未做任何代码改动。

原因

  • 没有可执行的反馈。 该 PR 已进入仅处理 Critical 的模式(已完成 5 个产生改动的轮次)。仅剩的开放反馈——来自 @qwen-code-review-bot 的一条 review 正文和两条行内评论(packages/cli/test-setup.test.ts:20 上的 rc:4213759653、packages/cli/src/config/config.test.ts:1468 上的 rc:4213759676)——均为非 Critical,已被工作流的确定性刹车机制排除在本轮之外。按照 Critical-only 规则,这些条目是留待人工跟进的审计记录:本轮未实现它们,也未对其讨论串进行解决或回复操作。
  • 没有失败的检查。 全部 59 项报告的检查均为 SUCCESS、SKIPPED 或 NEUTRAL;"Failed checks" 和 "Still-red checks" 两个部分均为空。
  • 没有与基线的冲突。 本轮以 --conflict false 调用,未被请求也不需要合并基线。
  • 无需增长审计。 "Diff growth this window" 部分显示本窗口净增长为源码 26 行 / 测试 103 行,预算均为 400,且此前没有超预算轮次;反馈中也没有 Growth audit required 部分,因此本轮不需要提交审计结论。

PR 保持本轮开始时的原样。被延后的非 Critical 条目保持开放,等待维护者跟进(以 [Critical] 开头评论、提交 Request changes 审查,或评论 @qwen-code /retry 开启新的计数窗口)。

验证

  • git status --short —— 干净;本轮未在仓库中修改或新建任何文件。
  • 检查 checks.json —— 59 项检查,无失败(唯一非 CheckRun 的条目是状态为 SUCCESS 的 StatusContext)。

Deferred non-Critical feedback

Critical-only mode is active: 5 change-producing rounds are complete. The workflow excluded the non-Critical feedback below from this round's actionable sections; the items remain open for human follow-up. Maintainer feedback is deferred only after its author has used 2 regular feedback batches in this window's Critical-only tail; authors at that budget, if any, are named below. (@qwen-code /retry starts a fresh counting window.)

中文说明

已进入仅处理 Critical 的模式:已完成 5 个产生改动的轮次。本轮可执行区域已排除下方非 Critical 反馈;这些条目保持开放,留待人工跟进。维护者反馈仅在其本人于本窗口 Critical-only 阶段已使用 2 批常规反馈预算后才会延后;达到预算的作者(如有)在下方点名。(评论 @qwen-code /retry 可开启新的计数窗口。)

Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。


🧠 Handled by Qwen Code · model/模型 kimi-k3 · CLI 0.25.0

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

autofix/takeover Summon the autofix loop to manage this PR (remove to release; needs triage+) review/self-reported The linked issue was opened by the PR author (self-reported)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(serve): batch workspace session live-state snapshots

6 participants