Skip to content

feat(managed-agent): enable G0 public Workspace file turns - #12955

Merged
wenshao merged 8 commits into
mainfrom
codex/g0-workspace-admission
Sep 29, 2026
Merged

wenshao merged 8 commits into
mainfrom
codex/g0-workspace-admission

Conversation

@wenshao

@wenshao wenshao commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Allows an authorized Workspace Session to include one initial Read/Write/Edit Turn when the deployment explicitly enables QWEN_MANAGED_AGENT_WORKSPACE_FILES_ENABLED. Public REST and WebShell creation share the same admission, use the persisted Workspace identity, and run through the production Hosted Harness, Broker and worker. The server selects a fixed file-tool profile and rechecks the binding before attachment or cache reuse. The opt-in defaults to false. A permanent Workspace authorization refusal before Turn submission fails immediately with its original error code; retryable failures and uncertain submission outcomes retain recovery retries.

Why it's needed

Hosted file tools already work through private integration paths, but public Workspace creation rejects initial input. This connects the existing public admission and durable Session binding to that execution path for the minimal G0 slice.

Reviewer Test Plan

How to verify

  1. Enable the opt-in in a supported local-process deployment with the Hosted Harness, HTTP Session Store, preapproved execution, canonical Workspace mounts and a trusted actor with read/create grants. Create a Session with initial input through REST, then through WebShell in a different Workspace. Each Turn should write, edit and read under the selected relative cwd, leave the Harness cwd untouched, and produce durable results and one public terminal event.
  2. Repeat a completed creation with the same key and payload. It should return the original Session/Turn without additional model or tool effects. Changed input under the same key should conflict.
  3. Verify that inaccessible, draining, unmounted or unsupported Workspaces are refused, read-only actors cannot create, and public metadata cannot select a tool profile. With the opt-in disabled, creation with input remains refused while empty Workspace creation still works.
  4. Revoke Workspace authority before the initial Turn is submitted. It should fail as workspace_unavailable without repeated dispatch retries or a misleading Harness-unavailable error. A Turn that may already have been submitted must retain recovery behavior.
  5. Verify that later Workspace submit/cancel/lifecycle operations remain gated and existing unbound no-tool Sessions retain their behavior.

Evidence (Before & After)

Before: initial Workspace input is refused by the existing admission gates. The baseline admission suite passes 17 tests confirming the shared creation gate and preserved empty creation/read behavior.

After: the packaged real-process test completes initial Write/Edit/Read Turns in two Workspaces through REST and WebShell, with three tool executions and one persisted terminal event per Session. Completed creation replay adds no model or tool calls. All 71 targeted server tests and the original 18 SDK tests pass; build, typecheck, bundle, Checkstyle and formatting checks pass. Regression tests fail on the unfixed authorization-denial path and when the uncertain-submission guard is removed. An independent dispatch probe passes 12 denial and recovery-boundary scenarios. Repeated full-diff undirected and reverse audits ended with two consecutive clean passes, followed by an independent minimal review. Only Critical fixes were eligible after round five; remaining Suggestions are recorded in the review threads.

Tested on

OS Status
🍏 macOS ✅ tested
🪟 Windows ⚠️ not tested
🐧 Linux ⚠️ not tested

Environment (optional)

macOS arm64, Node.js 22.22.2, Zulu JDK 21.0.11, Maven 3.8.4, isolated H2, deterministic local model and a test trusted-principal adapter. The coordinator, SQL persistence, Broker and worker use production wiring.

Risk & Scope

  • Main risk or tradeoff: this is an opt-in, same-host, preapproved file-tool path. Workspace mounts are trusted deployment data, not a filesystem sandbox. Disabling the opt-in also refuses input-bearing creation replays.
  • Not validated / out of scope: local MySQL, real model providers, production ingress authentication, concurrent or restart replay and SSE reconnect were not exercised. The existing Hosted MySQL CI suite discovers the new integration test. Later Turns, Shell, lifecycle/cwd enablement and G1–G3 recovery remain separate.
  • Breaking changes / migration notes: no schema migration or public request field is added. Existing deployments retain the closed gate until explicitly enabled.

Design: English · 简体中文. Both versions are complete and synchronized.

Linked Issues

Refs #12952, #12380. This PR implements the minimal G0 slice and does not close either tracker.

中文说明

本次改动

部署显式启用 QWEN_MANAGED_AGENT_WORKSPACE_FILES_ENABLED 后,获授权的 Workspace 会话可以在创建时携带一次初始 Read/Write/Edit 轮次。公开 REST 和 WebShell 创建共享准入,使用持久 Workspace 身份,经生产 Hosted Harness、Broker 和 worker 执行。服务端选择固定的文件工具 profile,并在连接或复用附件缓存前重新检查绑定。开关默认关闭。Turn 提交前的永久 Workspace 授权拒绝会立即按原错误码失败;可重试错误与提交结果不确定的情况继续保留恢复重试。

动机

Hosted 文件工具已能通过私有集成路径执行,但公开 Workspace 创建仍拒绝初始输入。本次把既有公开准入和持久会话绑定接到该执行路径,完成最小 G0 切片。

审阅者测试计划

验证方法

  1. 在受支持的 local-process 部署中启用开关,配置 Hosted Harness、HTTP Session Store、预授权执行、规范路径的 Workspace 挂载,以及具有读取/创建权限的可信 actor。分别通过 REST 和 WebShell 在不同 Workspace 创建带初始输入的会话。每轮应在选定相对 cwd 下写入、编辑并读取文件,不触碰 Harness cwd,且产生持久结果和一个公开终态事件。
  2. 在创建完成后用相同幂等键和载荷重试。应返回原 Session/Turn,不增加模型或工具副作用;同键改变输入应冲突。
  3. 验证不可访问、draining、未挂载或配置不受支持的 Workspace 被拒绝,只读 actor 无法创建,公开 metadata 不能选择工具 profile。关闭开关后仍拒绝带输入的创建,空输入 Workspace 创建继续可用。
  4. 在初始 Turn 提交前撤销 Workspace 权限,应立即按 workspace_unavailable 失败,不反复调度重试,也不误报 Harness 不可用。可能已提交的 Turn 应保留恢复行为。
  5. 验证后续 Workspace submit/cancel/生命周期操作继续受门禁限制,原有无绑定的无工具会话行为不变。

前后对照证据

改动前:初始 Workspace 输入被既有准入门禁拒绝。基线准入套件的 17 项测试通过,确认共享创建门禁以及保留的空输入创建/读取行为。

改动后:打包 CLI 的真实进程测试分别通过 REST 和 WebShell 在两个 Workspace 完成初始 Write/Edit/Read,每个会话有三次工具执行和一个持久终态事件。完成后的创建重试不增加模型或工具调用。71 项定向服务端测试和原有 18 项 SDK 测试全部通过;build、typecheck、bundle、Checkstyle 和格式检查通过。授权拒绝修复前及移除不确定提交保护时,回归测试均能报错。独立派发探针的 12 项拒绝与恢复边界场景通过。多轮完整差异的无方向与反向审计以连续两轮干净结束,随后完成最小独立审查。第五轮后仅接受 Critical 修正,其余 Suggestion 已在审阅线程记录。

测试平台

OS 状态
🍏 macOS ✅ 已测试
🪟 Windows ⚠️ 未测试
🐧 Linux ⚠️ 未测试

环境

macOS arm64、Node.js 22.22.2、Zulu JDK 21.0.11、Maven 3.8.4、隔离 H2、确定性本地模型和测试可信 principal 适配器。coordinator、SQL 持久化、Broker 与 worker 使用生产接线。

风险与范围

  • 主要风险或取舍:这是显式启用、同机、预授权的文件工具路径。Workspace 挂载属于可信部署数据,不是文件系统沙箱。关闭开关也拒绝携带输入的创建重试。
  • 未验证或范围外:未在本地测试 MySQL、真实模型服务、生产入口认证、并发或重启重试及 SSE 重连。现有 Hosted MySQL CI 套件会发现新增集成测试。后续轮次、Shell、生命周期/cwd 开放与 G1–G3 恢复继续另行推进。
  • 兼容性与迁移:没有新增数据库迁移或公开请求字段。既有部署在显式启用之前仍保持门禁关闭。

设计:English · 简体中文。两种语言版本完整并保持同步。

关联 Issue

关联 #12952、#12380。本 PR 实现最小 G0 切片,不关闭这两个跟踪 Issue。

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

G0 E2E verification report

Verified commit: 3c53e186085eb48dd081c045de8a7a5fe5a4d27f.

Result: passed on macOS arm64, using Node.js 22.22.2, Zulu JDK 21.0.11 and Maven 3.8.4. The packaged CLI ran against isolated H2 with a deterministic local model and a test trusted-principal adapter. The Spring coordinator, SQL persistence, Hosted Harness, production Broker and separate worker were real components; Sessions were admitted through public HTTP, without direct store creation or a replacement Broker.

Verification Observed result
Public REST creation in Workspace A and WebShell creation in Workspace B Both initial Turns completed Write/Edit/Read under the selected child cwd; each resulting file contained after; the Harness decoy directory remained untouched.
Tool profile and persistence The model received exactly read_file, write_file and edit; durable journal heads used the selected Workspace and persistent messages contained the tool results.
Completion and replay Each Session had one Turn, three durable tool executions and one persisted public terminal event. Public event responses contained the final output. Completed creation replay returned the original Session and added no model/tool calls; changed input conflicted. Eight model calls total for the two Sessions.
Rejection paths Missing or unauthorized actor, read-only actor, unknown/draining/unmounted Workspace, unsupported Agent/configuration and profile-override metadata were rejected. A tenant header inconsistent with the trusted principal was rejected on listing. Later submit remained gated.
Default-off and compatibility Existing admission, API, coordinator and unbound no-tool regressions remained green.
Reverse check Deliberately routing a bound Session to the global Workspace made the connector assertion and real-process completion test fail. Original source was restored byte-for-byte; the corrected test teardown also exited normally on this failure path.

Final verification after restoration:

  • Targeted server suite: 66 tests, 0 failures, 0 errors, 0 skipped; the production-path test took 3.757 seconds in the final run.
  • SDK suite: 18 tests, 0 failures, 0 errors, 0 skipped.
  • npm run build, npm run typecheck, npm run bundle, Checkstyle, Prettier checks and git diff --check: passed. Checkstyle reported zero violations.
  • Five complete undirected/reverse audit rounds ended with two consecutive clean passes. An additional fresh-context minimal review read all 21 changed files and found no reportable issue. After round five, only Critical fixes were eligible; no further source change was needed.

Baseline: the globally installed CLI could not start the Hosted profile, so the baseline used the existing isolated admission suite (17 passing tests). That establishes the original creation gate, not a passing tool-execution baseline.

Limits: local MySQL, real provider behavior, production authentication, concurrent/restart replay and SSE client reconnect were not tested. The terminal evidence is a persisted public-event count plus public event content. The existing Hosted MySQL CI profile includes the new integration test. G1–G3 recovery and later Workspace operations remain outside this change.

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification — G0 public Workspace file turns @ 3c53e186

Verdict: no blocker. The G0 claims hold on a real MySQL stack, on both creation surfaces and with a real model. I suggest landing it together with the two-commit candidate patch (+154/−0; it applies to the head and to head ⊕ main): F1 is a 10-line coordinator fix and F2 is one test. F3 needs a scope sentence in Risk, not a code change in this PR.

Setup. The stack was the production Spring jar with its embedded Runtime Broker, a separate managed-runtime-worker, the packaged Hosted Harness (dist/cli.js serve --profile hosted-harness) and MySQL 8.4.7. Only two parts were stand-ins: a deterministic OpenAI-compatible model and a trusted-principal filter of the same shape as the IT's. Every Session was created through public HTTP, and a wire tap recorded Spring → Harness traffic. I ran four arms:

  • the PR head, 3c53e186;
  • the base, bc8879ee;
  • PR ⊕ main 9f6138ae (a clean merge whose jar classes are CRC-identical to the PR's);
  • the candidate.

The author's E2E ran on H2; this run adds local MySQL.

Web Shell UI rendering a G0 Session that a real model created through the WebShell adapter

What holds

Claim Evidence
REST and the WebShell adapter run one initial Write/Edit/Read Turn in the selected Workspace and relative cwd ws-a / ws-b child/proof.txt = "after"; the Harness decoy cwd is untouched; the journal head workspace_id is the selected Workspace; each Session has 3 executions, 1 terminal event and 4 model calls
The profile is server-owned and the store identity is per Workspace On the wire, bound POST /session carries toolProfile: hosted-workspace-files/1 and managedSessionStore.workspaceId = ws-a / ws-b. All 61 deterministic-model G0 requests offered exactly read_file, write_file and edit. metadata.toolProfile returns 400 unsupported_feature
Replay and conflict The same key and payload return the same Session with 1 Turn and +0 model calls and executions, on both surfaces. A changed input returns 409 idempotency_conflict
Real model (qwen3.8-max) WebShell adapter: the haiku was written, edit applied and the file read back (4 executions, 25.4 s). REST: todo.txt written and read (2 executions). The decoy was untouched in both
Refusals 9-case matrix (image 03): draining, unmounted, unsupported config, read-only actor, other agent, actor without a grant, no actor, unknown Workspace, profile override. Empty creation behaves as before
Later operations stay gated submit, cancel, rename, close, archive and delete all return 409 workspace_unavailable; so does WebShell /turns/submit
Default off and compatibility With the opt-in off on the same DB, input returns 409 on both surfaces, empty creation returns 202, and completed G0 Sessions stay readable. Base bc8879ee behaves the same. An unbound Session sends no profile, uses global-ws, is offered no tools, and completes
Startup validation The real jar boot, with the opt-in on and each missing prerequisite (Harness, Store, Broker, approval-mode=default, isolation-class=workspace, no mounts), refuses to start with the documented message. The opt-out variants start
Integration tests HostedPublicWorkspaceIT passes on PR and on PR ⊕ main, each on H2 and MySQL. All Hosted*IT classes pass 9/9 on PR ⊕ main with MySQL. The CI MySQL job ran it (11.3 s)

G0 happy path, wire tap, real model and ITs

Gates, later operations, opt-out, base and startup validation

Findings

F1 — An authorization refusal at dispatch is retried and reported as "Harness unavailable" (minor; candidate fix).

  • Cause: QwenHostedHarnessConnector.createOrLoad rechecks the binding through WorkspaceExecutionStore.authorize, which throws RuntimeBrokerException(409, workspace_unavailable, retryable=false). HarnessCoordinator.coordinate catches that as a generic RuntimeException, so transientFailure keeps retrying it until the pre-admission retry budget (5) is spent. The Turn then fails with hosted_harness_unavailable and the message "Hosted Harness remained unavailable before Turn admission."
  • Reproduction: drain the Workspace, or revoke the creator's create grant, between admission and the first successful dispatch. The Turn fails after 32.9 s with the wrong code.
  • Candidate fix: fail with workspace_unavailable when this refusal happens before submission. Refusals after submission, and all other Broker errors, keep the retry path.
  • A/B on the real stack: the candidate fails in 1.8 s with the correct code, and the control Session still completes. The new unit test fails without the change.

F1 A/B

F2 — No test pins the default-off gate for an otherwise admissible Workspace (test gap; candidate test).

  • Removing both opt-in checks (!workspaceFilesEnabled in the store and !harness.isWorkspaceFilesAvailable() in the service) survives all 180 server unit tests and the IT.
  • The masking comes from the fixture. ManagedWorkspaceAdmissionTest registers config-ws-a and configures no mounts, so the new profile/mount check refuses first and the opt-in gate is never reached.
  • The candidate ManagedWorkspaceFilesOptInTest uses the fixed profile, a matching mount and the opt-in off:
    • it passes on the PR;
    • it passes with either gate removed on its own;
    • it fails with both gates removed (202 instead of 409).
  • One more survivor, not covered by the patch: mount matching that ignores tenantId also survives, because the IT uses a single tenant.

F3 — One active G0 Turn per Workspace, and a crashed Turn keeps the Workspace (scope note; not introduced here).

  • Concurrent creations: two creations in the same Workspace at once make the second Turn fail in 2 s. The Broker returns 409 workspace_busy, but the public event says only hosted_turn_failed. The Session cannot retry, because later submit is gated.
  • Harness crash: SIGKILL of the Harness mid-Turn (after write_file ran) fails the Turn, and public cancel returns 409.
  • Lease stays held: the crashed Turn's qwen_runtime_binding stays READY. Every new G0 Session in that Workspace then fails with hosted_turn_failed (workspace_busy). This still held about 10 minutes and two Spring restarts later.
  • Scope: other Workspaces are unaffected.

This is the same class as #12904 (a held Workspace lease has no supported recovery) and falls under the PR's stated G1–G3 exclusion, but G0 makes it reachable from public creation. I suggest two things:

  1. Surface workspace_busy in the terminal event data instead of the generic hosted_turn_failed.
  2. Add one sentence to the README/design Risk section naming both limits.

F3

Observations (no merge action needed)

Mutation testing

I ran 20 single mutants plus 1 combination on the Java diff:

  • 14 are killed by the unit tests.
  • M6–M8 (the mount, config/policy and agent admission checks) are killed only by the IT.
  • M9 and M10 are equivalent when applied alone.
  • M9+M10 and M16 survive (see F2).

With the candidate applied, the server suite passes 182/182 with Checkstyle.

Mutation

Not covered: Windows and Linux hosts, SSE reconnect, concurrent replay of one key, and production ingress authentication.

Evidence: rig, probes, results and candidate patch. harness/candidate-f1-f2.patch can be applied with git am. Before the reported runs I ruled out three rig artifacts:

  • The tap's Host header caused 403s.
  • The tap did not propagate an upstream abort, which made the crash look like an endlessly RUNNING Turn.
  • The first S3 variant restarted the Harness and was confounded by the generation issue above. The report uses S3b, the clean variant.
中文版

真实环境验证 —— G0 公开 Workspace 文件轮次 @ 3c53e186

结论:没有阻断项。G0 的主张在真实 MySQL 栈上、两个创建入口上、以及真实模型下都成立。 建议连同两个 commit 的候选补丁(+154/−0,可同时应用到 head 和 head ⊕ main)一起合入:F1 是 10 行 coordinator 修复,F2 是一个测试。F3 需要在 Risk 里补一句范围说明,不需要本 PR 改代码。

环境。 栈的组成:生产 Spring jar(内嵌 Runtime Broker)、独立的 managed-runtime-worker、打包后的 Hosted Harness(dist/cli.js serve --profile hosted-harness)、MySQL 8.4.7。替身只有两个:确定性的 OpenAI 兼容模型,以及与 IT 同形态的可信 principal filter。所有会话都通过公开 HTTP 创建,并在 Spring → Harness 之间放了抓包代理。共跑四个对照臂:

  • PR head 3c53e186;
  • base bc8879ee;
  • PR ⊕ main 9f6138ae(干净合并,jar 内 Java 类与 PR 的 CRC 一致);
  • 候选补丁。

作者的端到端验证跑在 H2 上,本次补上了本地 MySQL。

成立的部分

主张 证据
REST 与 WebShell adapter 都能在选定 Workspace 和相对 cwd 下完成一次初始 Write/Edit/Read ws-a / ws-b 的 child/proof.txt = "after";Harness 的 decoy cwd 未被触碰;journal head 的 workspace_id 为选定 Workspace;每个会话 3 次执行、1 个终态事件、4 次模型调用
profile 由服务端决定,store 身份按 Workspace 区分 抓包显示:绑定会话的 POST /session 携带 toolProfile: hosted-workspace-files/1,managedSessionStore.workspaceId 分别为 ws-a / ws-b。61 次确定性模型的 G0 请求都恰好只提供 read_file、write_file、edit。metadata.toolProfile 返回 400 unsupported_feature
重试与冲突 相同幂等键和载荷返回同一会话,仍是 1 个 Turn,模型调用和执行次数都不增加(两个入口相同)。改变输入返回 409 idempotency_conflict
真实模型(qwen3.8-max) WebShell adapter:写入俳句、edit 生效、读回(4 次执行,25.4 s)。REST:写入并读回 todo.txt(2 次执行)。两者都未触碰 decoy
拒绝路径 9 项矩阵(图 03):draining、未挂载、不支持的 config、只读 actor、其他 agent、无授权 actor、无 actor、未知 Workspace、profile 覆盖。空输入创建的行为保持不变
后续操作仍受门禁 submit、cancel、rename、close、archive、delete 全部返回 409 workspace_unavailable,WebShell /turns/submit 同样如此
默认关闭与兼容性 同一数据库关闭开关后:两个入口带输入都返回 409,空输入创建返回 202,已完成的 G0 会话仍可读。base bc8879ee 表现相同。无绑定会话不带 profile、使用 global-ws、没有工具,正常完成
启动校验 真实 jar 启动时,开关打开且逐一缺少前置条件(Harness、Store、Broker、approval-mode=default、isolation-class=workspace、无挂载),都会拒绝启动并打印文档中的提示;关闭开关的变体都能启动
集成测试 HostedPublicWorkspaceIT 在 PR 与 PR ⊕ main 上、H2 与 MySQL 下都通过。PR ⊕ main + MySQL 上全部 Hosted*IT 类 9/9 通过。CI 的 MySQL 作业确实运行了该测试(11.3 s)

发现

F1 —— 派发时的授权拒绝被当作瞬时故障重试,最终报成"Harness 不可用"(次要;附候选修复)。

  • 原因: QwenHostedHarnessConnector.createOrLoad 通过 WorkspaceExecutionStore.authorize 复核绑定,拒绝时抛出 RuntimeBrokerException(409, workspace_unavailable, retryable=false)。HarnessCoordinator.coordinate 把它当作普通 RuntimeException 捕获,于是 transientFailure 一直重试到准入前重试预算(5 次)耗尽,最后以 hosted_harness_unavailable("Hosted Harness remained unavailable before Turn admission.")结束 Turn。
  • 复现: 在准入之后、首次成功派发之前,把 Workspace 置为 DRAINING,或撤销创建者的 create 授权。Turn 在 32.9 s 后以错误的 code 失败。
  • 候选修复: 当这种拒绝发生在 submit 之前时,直接以 workspace_unavailable 失败。submit 之后的拒绝以及其他 Broker 错误仍走原来的重试路径。
  • 真实栈 A/B: 候选版本 1.8 s 失败且 code 正确,对照会话照常完成。新增的单测在去掉修复后会失败。

F2 —— 对于其他条件都可准入的 Workspace,没有测试钉住默认关闭的门禁(测试缺口;附候选测试)。

  • 同时去掉两处开关检查(store 的 !workspaceFilesEnabled 和 service 的 !harness.isWorkspaceFilesAvailable()),180 个服务端单测和 IT 仍全部通过。
  • 被掩盖的原因在夹具:ManagedWorkspaceAdmissionTest 注册的是 config-ws-a 且没有配置挂载,新增的 profile/挂载检查会先拒绝,开关门禁根本走不到。
  • 候选测试 ManagedWorkspaceFilesOptInTest 使用固定 profile、匹配的挂载,并关闭开关:
    • 在 PR 上通过;
    • 只去掉任意一处门禁时仍通过;
    • 两处都去掉时失败(返回 202 而非 409)。
  • 另有一个存活变异体未纳入补丁:挂载匹配忽略 tenantId 也能存活,因为 IT 只有单个租户。

F3 —— 每个 Workspace 同时只能有一个 G0 Turn,且崩溃的 Turn 会一直占住该 Workspace(范围说明;并非本 PR 引入)。

  • 并发创建: 同一 Workspace 同时创建两个会话,第二个 Turn 在 2 s 内失败。Broker 返回的是 409 workspace_busy,但公开事件只显示 hosted_turn_failed。由于后续 submit 受门禁限制,该会话也无法重试。
  • Harness 崩溃: 在 Turn 中途(write_file 已执行)对 Harness 发 SIGKILL,Turn 失败,公开 cancel 返回 409。
  • 租约不释放: 崩溃 Turn 的 qwen_runtime_binding 一直保持 READY。之后在该 Workspace 新建的每个 G0 会话都以 hosted_turn_failed(workspace_busy)失败,约 10 分钟、两次 Spring 重启后依然如此。
  • 影响范围: 其他 Workspace 不受影响。

这与 #12904 属同一类问题(被占住的 Workspace 租约没有受支持的恢复路径),也在 PR 声明的 G1–G3 排除范围内,但 G0 让它可以从公开创建入口触发。建议两点:

  1. 在终态事件数据里给出 workspace_busy,而不是笼统的 hosted_turn_failed。
  2. 在 README/设计文档的 Risk 部分用一句话写明这两个限制。

观察(不影响合并)

变异测试

对 Java 改动做了 20 个单变异体和 1 个组合变异体:

  • 14 个被单测杀死。
  • M6–M8(挂载、config/policy、agent 准入检查)只有 IT 能杀死。
  • M9、M10 单独应用时是等价变异体。
  • M9+M10 与 M16 存活(见 F2)。

应用候选补丁后,服务端测试套件含 Checkstyle 182/182 通过。

未覆盖: Windows 与 Linux 宿主、SSE 重连、同一幂等键的并发重试、生产入口认证。

证据: 装置、探针、结果与候选补丁都在这里,harness/candidate-f1-f2.patch 可直接 git am。正式运行前排除了三个装置伪影:

  • 抓包代理转发的 Host 头导致 403。
  • 代理没有传递上游中断,使崩溃看起来像一个永远 RUNNING 的 Turn。
  • 第一版 S3 重启了 Harness,结果被上面的代际问题干扰;报告采用的是干净的 S3b。

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Downgraded from Request changes to Comment: self-PR; CI still running. Reviewed. Suggestions are inline.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 2)": compiling and running ManagedAgentPropertiesTest / QwenHostedHarnessConnectorTest under Maven to demonstrate the two surviving mutations empirically — both …; chunk 5: none — I did not compile or execute the Java tests, so "the new tests pass" rests on the source-level trace above rather than a run.; "agent reverse-audit (round 3)": did not execute a forced-timeout probe ( @Timeout(1) arm) to observe the orphaned Tomcat/Node process directly; the mechanism is read from JUnit's documented ….

中文说明

⚠️ 已从请求修改降级为评论:self-PR; CI still running。 已审查。 建议见行内评论。

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 2)":compiling and running ManagedAgentPropertiesTest / QwenHostedHarnessConnectorTest under Maven to demonstrate the two surviving mutations empirically — both …;chunk 5:none — I did not compile or execute the Java tests, so "the new tests pass" rests on the source-level trace above rather than a run.;"agent reverse-audit (round 3)":did not execute a forced-timeout probe ( @Timeout(1) arm) to observe the orphaned Tomcat/Node process directly; the mechanism is read from JUnit's documented …。

— qwen3.8-max via Qwen Code /review (v0.24.6)

Comment thread packages/sdk-java/managed-agent-server/README.md
Comment thread packages/sdk-java/managed-agent-server/README.md
through `POST /v1/agents/sessions`. The server chooses the fixed
`hosted-workspace-files/1` private profile and uses the persisted Workspace ID
for the Session Store. Public callers cannot choose the profile. Configure the
Harness's deployment-owned `--managed-runtime-broker-url` and

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-3: This sentence makes --managed-runtime-broker-url / --managed-runtime-broker-token a required documented deployment step, while the options' own --help text (packages/cli/src/commands/serve.ts:654/:660, unchanged) still declares them "Reserved … not implemented and rejects startup." An operator following this README who runs qwen serve --help to confirm the flags before enabling G0 reads that they are fatal and concludes the documented deployment cannot work — or files a bug against flags that do work (serve.ts:1013-1020 consumes both; this PR's IT passes both and completes three real tool executions). The staleness predates this PR, but this new instruction is what makes it load-bearing — before it, no README procedure told anyone to set these options. Fix: correct the two one-line description strings in serve.ts in this PR (e.g. "Private Broker URL for --profile hosted-harness; required for Workspace tool turns."), or, if touching packages/cli is out of scope here, drop the flag names from this sentence and point at the "Broker deployment" section that already documents them.

Witness:

$ node dist/cli.js serve --help
 --managed-runtime-broker-url   Reserved Broker URL for --profile hosted-harness; not implemented and rejects startup. [string]
 --managed-runtime-broker-token Reserved Broker credential for --profile hosted-harness; not implemented and rejects startup. [string]
vs. HostedPublicWorkspaceIT.java:261-262 passing both flags — Tests run: 1, Failures: 0
中文说明

这句话把 --managed-runtime-broker-url / --managed-runtime-broker-token 变成了文档化的必需部署步骤,而这两个选项自己的 --help 文本(packages/cli/src/commands/serve.ts:654/:660,未改动)仍声明“Reserved … not implemented and rejects startup.”。按本 README 操作、在启用 G0 前先用 qwen serve --help 确认参数的运维人员,会读到它们是致命项,从而断定文档化的部署无法工作——或者对一个可用的参数提 bug(serve.ts:1013-1020 会消费两者;本 PR 的 IT 传入两者并完成三次真实工具执行)。过时早于本 PR,但正是这条新增说明使其变得承重——此前没有任何 README 流程要求设置这两个选项。修复:在本 PR 中修正 serve.ts 的两行 description(例如 "Private Broker URL for --profile hosted-harness; required for Workspace tool turns.");若不便改 packages/cli,则从这句话中去掉参数名,指向已有文档的 "Broker deployment" 一节。

见证:serve --help 实际输出与 IT 绿色运行的对照(见上)。

— qwen3.8-max via Qwen Code /review (v0.24.6)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Acknowledged as a Suggestion and deferred under the requested Critical-only policy after round five. The stale Broker flag help is real; a follow-up should correct both help descriptions. The runtime options are consumed and the G0 integration test exercises them.

按第五轮后仅处理 Critical 的约定延后;后续同步修正两项参数帮助文案。

Tracked in the follow-up/deferral report; leaving this thread open.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tracked in docs(cli): correct Hosted Runtime Broker option help. The issue preserves the original suggestion and acceptance criteria. This is follow-up tracking, not an implemented fix; leaving the conversation open.

已转为独立 Issue:https://github.com/QwenLM/qwen-code/issues/13044,保留验收条件;尚未实施修复。

for (String name : List.of("PATH", "SystemRoot", "WINDIR", "COMSPEC", "PATHEXT")) {
if (System.getenv(name) != null) environment.put(name, System.getenv(name));
}
environment.putAll(Map.of("HOME", home.toString(), "USERPROFILE", home.toString(), "QWEN_HOME", config.toString(),

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-5: startHarness clears the child environment (line 265) but never pins TMPDIR/TMP/TEMP into the test's @TempDir — the sibling HostedHarnessMySqlIT.java:384-386 pins all three after the same clear(). The spawned Node harness therefore resolves os.tmpdir() to the runner's shared /tmp. Measured in the current flow the escape does not materialise — a green run leaves exactly one orphaned tomcat.* work dir (from the test JVM itself), no qwen-* entries — so this stands on the divergence from the sibling's declared contract and its latent cost: the harness paths that do write to os.tmpdir() (TLS pem-block oracle, sandbox, extension upload, skill checkout) are one flag away from being exercised by this IT, at which point writes escape CleanupMode.ON_SUCCESS (which only deletes the @TempDir) into a directory shared with concurrent failsafe forks and other agents' sessions. Fix — after the putAll:

environment.put("TMPDIR", temporary.toString());
environment.put("TMP", temporary.toString());
environment.put("TEMP", temporary.toString());

The fix must not violate: the sibling's contract at HostedHarnessMySqlIT.java:384-386 (temporary.toString() values), applied AFTER environment.clear() — lines 265-268 wipe the inherited parent env, so pins placed before it are erased.

Witness:

ls -1a /tmp: 4023 entries before; after a green run (Tests run: 1, Failures: 0 — 11.01 s):
> tomcat.46511.7473523348564221804    ← exactly one new entry, no qwen-* entry
sibling contract: HostedHarnessMySqlIT.java:384-386
environment.put("TMPDIR", temporary.toString()); … TMP … TEMP … after the same clear()
中文说明

startHarness 清空了子进程环境(265 行),但没有像同级 HostedHarnessMySqlIT.java:384-386 在同样 clear() 之后那样把 TMPDIR/TMP/TEMP 钉进测试自己的 @TempDir。因此派生的 Node harness 会把 os.tmpdir() 解析到运行器共享的 /tmp。实测当前流程中逃逸并未发生——绿色运行只留下一个孤儿 tomcat.* 工作目录(来自测试 JVM 本身),没有任何 qwen-* 条目——所以本条立足于与同级契约的偏离及其潜在成本:真正会写 os.tmpdir() 的 harness 路径(TLS pem 预言机、sandbox、扩展上传、skill checkout)距离被本 IT 触发只差一个开关,一旦触发,写入会逃过 CleanupMode.ON_SUCCESS(只删除 @TempDir),落到与并发 failsafe fork 及其他代理会话共享的目录里。修复:在 putAll 之后补上三个 environment.put(代码见上)。

约束:取值遵循同级契约(HostedHarnessMySqlIT.java:384-386 的 temporary.toString()),且必须放在 environment.clear() 之后——265-268 行会清掉继承的环境变量,放在之前的钉值会被抹掉。

见证:绿色运行前后 /tmp 目录对比(仅新增一个 tomcat.* 条目)与同级契约引文。

— qwen3.8-max via Qwen Code /review (v0.24.6)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deferred as test-fixture hardening under the Critical-only policy after round five. A follow-up should pin all three temporary-directory variables after clearing the child environment. The report found no current Qwen temporary-file escape in this test flow.

作为测试隔离加固记录延后;后续在清空环境后设置三个临时目录变量。

Tracked in the follow-up/deferral report; leaving this thread open.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tracked in test(managed-agent): isolate G0 Harness temporary directories. The issue preserves the original suggestion and acceptance criteria. This is follow-up tracking, not an implemented fix; leaving the conversation open.

已转为独立 Issue:https://github.com/QwenLM/qwen-code/issues/13045,保留验收条件;尚未实施修复。

if (!input.isEmpty()
&& (!"qwen-code".equals(agentId)
|| !WorkspaceExecutionProfile.CONFIG_REF.equals(workspace.configRef())
|| !WorkspaceExecutionProfile.POLICY_REF.equals(workspace.policyRef())

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-7: Two of the four legs of this new creation-admission predicate — the frozen policy_ref equality here and the mount's tenantId equality at :247-248 — have no test that can distinguish them from being absent. HostedPublicWorkspaceIT is the only flag-enabled test and its fixtures cannot reach those legs (every registry row comes from register() with the frozen POLICY_REF; every configured mount carries the one fixture tenant; it varies only config_ref='unsupported' and an unmounted storage-id), and no unit test enables the flag at all. Executed: deleting either leg leaves everything the repo runs green, while deleting the sibling CONFIG_REF leg reddens the IT — an unowned contract. The cost is concrete: a Workspace registered with the frozen config but a drifted policy_ref, or a second tenant whose storage_id collides with an existing mount's, gets a 202 creation and an admitted initial Turn that dies later at WorkspaceExecutionStore.authorize / WorkspaceRuntimeResolver — a clean creation-time 409 converted into a failed Turn after warmRuntime already appended environment.provisioning. (No cross-tenant file access is possible — the resolver's mount key is tenant-scoped — which is why this is a Suggestion, not a blocker.) Fix: add a flag-enabled fast store case beside the existing admission tests: seed a registry row with CONFIG_REF + policy_ref="preapproved-workspace-tools/2" and assert non-empty-input insertWorkspaceSessionCommand throws workspace_unavailable; then configure a mount for a different tenant carrying the same storage-id and assert the same refusal.

Witness (mutation runs, every arm recompiled):

CONTROL (CONFIG_REF leg removed): [ERROR] HostedPublicWorkspaceIT — expected: 409 but was: 202
MUTATION B (POLICY_REF leg removed):   Tests run: 180, Failures: 0 + IT 1/1 green — BUILD SUCCESS
MUTATION C (mount tenantId leg removed): Tests run: 180, Failures: 0 + IT 1/1 green — BUILD SUCCESS

The new case must go red when either leg is removed from the predicate at :245 or :247-248 — please prove that mutation. The fix must not violate: .github/workflows/sdk-java.yml:307 gives the hosted step timeout-minutes: 12 and it already took 14m42s on this PR, so the cases belong in a fast store test, not in HostedPublicWorkspaceIT.

中文说明

这个新创建准入谓词的四条分支中,有两条——此处的冻结 policy_ref 相等与 :247-248 的挂载 tenantId 相等——没有任何测试能把它们与“不存在”区分开。HostedPublicWorkspaceIT 是唯一开启开关的测试,其 fixture 触达不到这两条分支(所有 registry 行都由 register() 以冻结 POLICY_REF 写入;所有配置挂载都携带同一个 fixture 租户;它只变化 config_ref='unsupported' 与未挂载的 storage-id),而单元测试完全不开开关。实测:删除任一条分支,仓库运行的所有测试仍绿;删除同级 CONFIG_REF 分支则 IT 变红——这是一份无人认领的契约。成本是具体的:以冻结 config 但漂移 policy_ref 注册的 Workspace,或 storage_id 与既有挂载碰撞的第二个租户,会得到 202 创建与被准入的初始 Turn,随后才在 WorkspaceExecutionStore.authorize / WorkspaceRuntimeResolver 处失败——干净的创建期 409 变成了 warmRuntime 已追加 environment.provisioning 之后的 Turn 失败。(跨租户文件访问不可能——resolver 的挂载键含租户——因此这是 Suggestion 而非阻断项。)修复:在既有准入测试旁增加开启开关的快速 store 用例:以 CONFIG_REF + policy_ref="preapproved-workspace-tools/2" 写入 registry 行,断言非空 input 的 insertWorkspaceSessionCommand 抛 workspace_unavailable;再为另一个租户配置相同 storage-id 的挂载,断言同样拒绝。

见证(变异运行,每臂均重新编译):删除 CONFIG_REF 分支的正控使 IT 变红(expected 409 but was 202);删除 POLICY_REF 或 tenantId 分支后 180/180 快速套件与 IT 全绿。

新用例必须在 :245 或 :247-248 任一分支被移除时变红——请做该变异验证。约束:.github/workflows/sdk-java.yml:307 给 hosted 步骤的 timeout-minutes: 12 在本 PR 上已用掉 14m42s,新用例应放进快速 store 测试而非 HostedPublicWorkspaceIT。

— qwen3.8-max via Qwen Code /review (v0.24.6)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deferred as missing regression coverage under the Critical-only policy after round five. Both production predicate legs are present. A follow-up should add fast, flag-enabled cases for a drifted policy reference and a different tenant sharing a storage ID, including the proposed mutation checks.

生产检查仍在;针对 policy 和跨租户 storage 碰撞的测试及变异验证记录为后续工作。

Tracked in the follow-up/deferral report; leaving this thread open.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tracked in test(managed-agent): cover G0 policy and mount-tenant admission guards. The issue preserves the original suggestion and acceptance criteria. This is follow-up tracking, not an implemented fix; leaving the conversation open.

已转为独立 Issue:https://github.com/QwenLM/qwen-code/issues/13046,保留验收条件;尚未实施修复。

}

@PostConstruct
void validateWorkspaceFiles() {

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-8: The guard's logic is well covered — ManagedAgentPropertiesTest walks seven single-condition mutations plus a valid baseline — but its activation is not: the only test calls validateWorkspaceFiles() as a plain method, and the only Spring boot in the suite (HostedPublicWorkspaceIT.startSpring) passes a fully valid combination. Removing @PostConstruct (or relocating the check into a helper the bean no longer invokes) leaves every test green, while a deployment setting QWEN_MANAGED_AGENT_WORKSPACE_FILES_ENABLED=true next to e.g. runtime-broker.isolation-class=workspace or harness.approval-mode=default starts and serves traffic: the store gate keys only off workspaceFilesEnabled (ManagedAgentStore.java:229) and the service gate only off harness.isWorkspaceFilesAvailable() (ManagedAgentService.java:145), so public input-bearing creation is admitted into precisely the configuration this boot check exists to refuse. Fix: add a context-startup test — ApplicationContextRunner or @SpringBootTest — asserting IllegalStateException / hasFailed() with an explicitly invalid combination.

Witness (probe, both arms):

ARM 1 (HEAD, annotation present): PROBE-POSTCONSTRUCT: context REFUSED with
  java.lang.IllegalStateException: Hosted Workspace files require a preapproved Harness, Session Store and Session-isolated local-process Broker with Workspace mounts
ARM 2 (@PostConstruct removed):   PROBE-POSTCONSTRUCT: context STARTED with workspace-files-enabled=true + isolation-class=workspace
  [INFO] Tests run: 181, Failures: 0, Errors: 0 — BUILD SUCCESS

The new startup test must go red if @PostConstruct is removed from validateWorkspaceFiles() — no test in the current suite does; please prove that mutation. The fix must not violate: the defaults are already valid (approvalMode="yolo" at :68, isolationClass="session" at :309), so the test must set an explicitly invalid value — relying on defaults asserts nothing.

中文说明

该守卫的逻辑覆盖良好——ManagedAgentPropertiesTest 遍历了七个单条件变异加一个有效基线——但其“激活”没有覆盖:唯一的测试把 validateWorkspaceFiles() 当普通方法调用,套件中唯一的 Spring 启动(HostedPublicWorkspaceIT.startSpring)传入的是完全有效的组合。移除 @PostConstruct(或把检查挪到 bean 不再调用的辅助方法)后所有测试仍绿;而此时把 QWEN_MANAGED_AGENT_WORKSPACE_FILES_ENABLED=true 与例如 runtime-broker.isolation-class=workspace 或 harness.approval-mode=default 放在一起的部署会正常启动并服务流量:store 门禁只看 workspaceFilesEnabled(ManagedAgentStore.java:229),service 门禁只看 harness.isWorkspaceFilesAvailable()(ManagedAgentService.java:145),于是带输入的公开创建恰好被放行进这个启动检查本应拒绝的配置。修复:新增上下文启动测试(ApplicationContextRunner 或 @SpringBootTest),用显式无效组合断言 IllegalStateException / hasFailed()。

见证(双臂探针):HEAD 下无效组合的上下文被拒绝(IllegalStateException);移除注解后同样的上下文成功启动,且全套件 181/181 绿。

新启动测试必须在 @PostConstruct 被移除时变红——当前套件没有任何测试会;请做该变异验证。约束:默认值本就有效(:68 的 approvalMode="yolo"、:309 的 isolationClass="session"),测试必须显式设置无效值——依赖默认值什么都断言不了。

— qwen3.8-max via Qwen Code /review (v0.24.6)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deferred as startup regression coverage under the Critical-only policy after round five. The validation callback is currently annotated and active. A follow-up should boot a Spring context with an explicitly invalid opt-in configuration and demonstrate failure if callback activation is removed.

当前启动校验已启用;无效 Spring 配置与回调移除的回归验证记录延后。

Tracked in the follow-up/deferral report; leaving this thread open.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tracked in test(managed-agent): verify G0 deployment validation runs at Spring startup. The issue preserves the original suggestion and acceptance criteria. This is follow-up tracking, not an implemented fix; leaving the conversation open.

已转为独立 Issue:https://github.com/QwenLM/qwen-code/issues/13047,保留验收条件;尚未实施修复。

.contains("toolProfile=hosted-workspace-files/1", "workspaceId=selected-workspace", "tenantId=tenant-a")
.doesNotContain("workspaceId=workspace-a");
}
doThrow(new IllegalStateException("grant revoked")).when(execution).authorize(session);

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-12: This stub injects an IllegalStateException("grant revoked") — a type the real WorkspaceExecutionStore.authorize never throws (it fails only with RuntimeBrokerException via unavailable(), 409 workspace_unavailable, retryable=false; busy() comes from claim/assertHeld/release). The test therefore pins message pass-through, not the denial contract. Executed: wrapping the recheck in catch (RuntimeBrokerException ignored) { } inside QwenHostedHarnessConnector.createOrLoad keeps the fast suite green (180/180) AND the CI-mandatory HostedPublicWorkspaceIT green (1/1 — its denial cases all fail earlier, at creation; no test drives a grant withdrawn between creation and attach), while in production a revoked-grant Workspace would attach the Session and run hosted-workspace-files/1 file tools against it. It also leaves R1-1's isRetryable() contract without a witness at the boundary where the exception is produced. Fix — stub the production factory and assert the propagated contract:

doThrow(WorkspaceExecutionStore.unavailable()).when(execution).authorize(session);
assertThatThrownBy(() -> connector.createOrLoad("tenant-a", SESSION_ID, true))
        .isInstanceOfSatisfying(RuntimeBrokerException.class, error -> {
            assertThat(error.getCode()).isEqualTo("workspace_unavailable");
            assertThat(error.getStatusCode()).isEqualTo(409);
            assertThat(error.isRetryable()).isFalse();
        });

(adds a RuntimeBrokerException import)

With the corrected stub, adding a catch-and-ignore wrapper around the recheck in createOrLoad must redden boundCreateConflictLoadsOriginalWorkspaceAndProfileAndRechecksCachedGrant — please prove that mutation. The fix must not violate: unavailable() is public static; busy() is private static, so the retryable leg must be built through the public ctor new RuntimeBrokerException(409, "workspace_busy", …, true) — which throws IllegalArgumentException for a statusCode outside 400-599 (RuntimeBrokerException.java:26-29).

Witness:

swallow mutant + today's stub:   fast suite Tests run: 180, Failures: 0; IT Tests run: 1, Failures: 0 (10.77 s)
corrected stub + swallow mutant: EXIT=1, Failures: 1 — "Expecting code to raise a throwable."
corrected stub + intact code:    EXIT=0 — BUILD SUCCESS
中文说明

这个桩注入的是 IllegalStateException("grant revoked")——真实的 WorkspaceExecutionStore.authorize 从不抛这个类型(它只经 unavailable() 抛 RuntimeBrokerException,409 workspace_unavailable、retryable=false;busy() 来自 claim/assertHeld/release)。因此测试钉住的是消息透传,而不是拒绝契约。实测:在 QwenHostedHarnessConnector.createOrLoad 里把复检包进 catch (RuntimeBrokerException ignored) { },快速套件全绿(180/180),CI 强制运行的 HostedPublicWorkspaceIT 也全绿(1/1——其拒绝用例都在创建期更早失败;没有测试驱动“创建与 attach 之间授权被撤销”),而生产中被撤销授权的 Workspace 会 attach 该 Session 并对其执行 hosted-workspace-files/1 文件工具。这也使 R1-1 的 isRetryable() 契约在异常产生边界上没有见证。修复——用生产工厂做桩并断言传播的契约(代码见上,需新增 RuntimeBrokerException import)。

使用修正后的桩时,在 createOrLoad 的复检外加 catch-and-ignore 包装必须使 boundCreateConflictLoadsOriginalWorkspaceAndProfileAndRechecksCachedGrant 变红——请做该变异验证。约束:unavailable() 是 public static;busy() 是 private static,可重试分支须经公共构造器 new RuntimeBrokerException(409, "workspace_busy", …, true) 构造——该构造器对 400-599 之外的 statusCode 抛 IllegalArgumentException(RuntimeBrokerException.java:26-29)。

见证:吞异常变异 + 现有桩 → 快速套件与 IT 全绿;修正桩 + 变异 → 红("Expecting code to raise a throwable.");修正桩 + 原样代码 → 绿。

— qwen3.8-max via Qwen Code /review (v0.24.6)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new R1-1 coordinator regression uses the real WorkspaceExecutionStore.unavailable() contract and proves the typed handling. Strengthening this separate connector test and its catch-and-ignore mutation remains a deferred Suggestion under the Critical-only policy after round five; it is not being marked fixed by the coordinator test.

R1-1 的新增协调器测试已使用真实异常契约;这条连接器独立测试建议仍明确延后,不以协调器测试冒充完成。

Tracked in the follow-up/deferral report; leaving this thread open.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tracked in test(managed-agent): pin the connector Workspace refusal contract. The issue preserves the original suggestion and acceptance criteria. This is follow-up tracking, not an implemented fix; leaving the conversation open.

已转为独立 Issue:https://github.com/QwenLM/qwen-code/issues/13048,保留验收条件;尚未实施修复。

await().atMost(Duration.ofSeconds(35)).failFast(() -> {
String status = jdbc.queryForObject("SELECT status FROM managed_agent_turn WHERE session_id = ?",
String.class, session);
assertThat(status).as("Turn: %s; events: %s; model requests: %s; Harness: %s",

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-14: This failFast diagnostic is evaluated on every Awaitility poll (measured mean interval ~103 ms; up to ~340 evaluations per workspace against atMost(35s), ~680 per test method across the two-root loop): two SELECTs — one materializing the Turn's whole event history with its growing data_json — plus a full Files.readString of the continuously appended harness.log, all as eagerly evaluated as(...) arguments for a description discarded on the ~99% of polls where status != FAILED. (AssertJ does defer the formatting itself — the per-poll cost is the argument expressions.) Measured in a green run: 42 polls, 84 extra SELECTs, ~88 KB of transient strings; the quadratic shape (polls × a growing log) materializes only in a long stuck-but-not-FAILED wait — exactly the case that eats the @Timeout(150) headroom (internal awaits already reserve ~100 s) against the CI step's timeout-minutes:12. The same shape repeats at line 276 (assertThat(harness.isAlive()).as(Files.readString(log)) inside the 30 s capability poll, ~300 full log reads per boot). Fix — build the dump only on the FAILED branch, preserving the identical message when it fires:

}).failFast(() -> {
    String status = jdbc.queryForObject("SELECT status FROM managed_agent_turn WHERE session_id = ?",
            String.class, session);
    if ("FAILED".equals(status)) {
        assertThat(status).as("Turn: %s; events: %s; model requests: %s; Harness: %s",
                jdbc.queryForList("SELECT status, error_code FROM managed_agent_turn WHERE session_id = ?", session),
                jdbc.queryForList("SELECT event_type, data_json FROM managed_agent_event WHERE session_id = ?", session),
                modelRequests.size(), Files.readString(temporary.resolve("harness.log"))).isNotEqualTo("FAILED");
    }
})

The fix must not violate: the guarded version must still THROW when status is FAILED — failFast terminates the wait only because the runnable raises; a restructure that skipped the assertion instead of running it on the FAILED branch would leave untilAsserted polling the full 35 s per workspace, doubling the worst-case wait against @Timeout(150) (:64).

Witness:

BASE (instrumented counters, production code at HEAD): PROBE-IT failFastPolls=42 totalHarnessLogCharsRead=65462 totalSqlDumpChars=22418
FIXED (guarded form): PROBE-IT failFastPolls=40 totalHarnessLogCharsRead=0 totalSqlDumpChars=0 — Tests run: 1, Failures: 0 (10.97 s)
failed-arm check: threw=yes type=TerminalFailureException messageHasTurnDump=true messageHasEventDump=true messageHasLogDump=true
poll rate: 21 failFast evaluations in 2175 ms (mean 103 ms); AssertJ passing-arm descriptionArgumentEvaluations=0
中文说明

这个 failFast 诊断在 Awaitility 的每次轮询都会被求值(实测平均间隔约 103ms;对 atMost(35s) 每个 workspace 最多约 340 次,双 root 循环下每个测试方法约 680 次):两条 SELECT——其中一条物化 Turn 的全部事件历史及其不断增长的 data_json——加上对持续追加的 harness.log 的整文件 Files.readString,全部作为急切求值的 as(...) 实参,而这个描述在约 99% 状态非 FAILED 的轮询里被丢弃。(AssertJ 确实会推迟格式化本身——每轮成本在于实参表达式。)绿色运行实测:42 次轮询、84 次多余 SELECT、约 88KB 临时字符串;二次方形态(轮询数 × 增长的日志)只在“卡住但未 FAILED”的长等待中出现——而那正是吞噬 @Timeout(150) 余量(内部等待已预留约 100s)与 CI 步骤 timeout-minutes:12 的场景。同样的形态在 276 行重复(30s capability 轮询里的 assertThat(harness.isAlive()).as(Files.readString(log)),每次启动约 300 次整文件读取)。修复——只在 FAILED 分支构建 dump,触发时保持完全相同的消息(代码见上)。

约束:守卫版本在状态为 FAILED 时仍必须抛出——failFast 终止等待正是因为 runnable 抛异常;若改成在 FAILED 分支跳过断言而非执行断言,untilAsserted 会把每个 workspace 轮询满 35s,使最坏等待对 @Timeout(150)(:64)翻倍。

见证:基线(探针计数器,生产代码为 HEAD)42 次轮询读日志 65462 字符、SQL dump 22418 字符;守卫形式两项均为 0 且测试绿(10.97s);失败臂确认仍抛出并携带完整 dump;轮询速率实测均值 103ms,AssertJ 通过臂描述实参求值为 0。

— qwen3.8-max via Qwen Code /review (v0.24.6)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deferred as test-diagnostic efficiency work under the Critical-only policy after round five. A follow-up should construct SQL/log diagnostics only on the failure branch while retaining the throwing assertion that makes failFast terminate immediately.

诊断性能优化记录延后;后续仅在失败时构造信息,并保留触发 failFast 的抛错断言。

Tracked in the follow-up/deferral report; leaving this thread open.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tracked in test(managed-agent): build G0 failure diagnostics only on failure. The issue preserves the original suggestion and acceptance criteria. This is follow-up tracking, not an implemented fix; leaving the conversation open.

已转为独立 Issue:https://github.com/QwenLM/qwen-code/issues/13049,保留验收条件;尚未实施修复。

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Review follow-up pushed in 7528ba6.

All three Critical findings are fixed:

  • R1-1 / real-stack F1: permanent Workspace authorization refusal before submission now fails immediately with its original code/message. Retryable failures and Turns that were already or may have been submitted keep recovery retries.
  • R1-2: README execution guidance now describes the opt-in initial file Turn and the remaining later-operation/Shell gates; stale Hosted-startup rejection claims were removed while the copied scripts' missing worker-artifact limitation remains explicit.
  • R1-9: Store scope documentation now distinguishes persisted Workspace identity for bound Sessions from the configured global identity for unbound Sessions, retaining the global-ID startup requirements.

Validation: 71 targeted server tests passed, no failures/errors/skips, including the real Hosted Harness REST/WebShell file-tool integration test. Build, typecheck, bundle, Checkstyle and formatting checks passed. Before the repair, two coordinator cases failed for the reported retry/misclassification; removing the submission guard also causes two failures. After restoration, an independent probe passed 12 scenarios, including persisted and in-call submission markers beyond the pre-admission retry limit. Two complete undirected/reverse audit passes over this follow-up diff were clean, followed by an independent minimal review with no findings.

The independent dispatch probe uses mocked persistence/connector/warmer around the real coordinator; it is not a claim of a new real-DB revocation or crash test. The existing reviewer-supplied real-stack evidence remains separate.

Per the requested Critical-only policy after audit round five, these Suggestions are deferred and their threads remain open:

Finding Follow-up scope
R1-3 Update stale Broker flag help text.
R1-4 Clarify capability descriptions and regenerate client types together.
R1-5 Pin the test Harness temporary-directory environment.
R1-7 Add frozen-policy and cross-tenant mount predicate coverage.
R1-8 Pin validation callback activation with an invalid Spring startup case.
R1-12 Strengthen the connector's own exception-contract regression; the new coordinator regression already uses the real refusal factory.
R1-14 Defer expensive integration diagnostics until failure.

Additional review notes are recorded, without widening G0: real-stack F2 (otherwise-admissible default-off regression fixture) is deferred; F3 (same-Workspace contention and held lease after a crash) remains in #12904 / G1–G3, with no timeout takeover introduced. The current README already describes retained holders. H1 #12946 is still open; the second landing must reconcile the API contract version and generated output. UI projection/banner, production ingress and existing worker/restart behavior remain outside this change.

中文说明

已修复全部三条 Critical:提交前的永久 Workspace 拒绝立即保留原错误,可能已提交的轮次继续恢复重试;同时更正 README 的执行边界和 Store 身份说明。71 项定向测试、12 项独立探针场景及构建/类型/格式检查通过;本次完整差异经两轮无方向及反向审计、一次最小独立审查。

按第五轮后仅处理 Critical 的约定,上表七条 Suggestion 已记录延后,线程保持开放。真实环境报告的默认关闭测试补强同样延后;同 Workspace 争用、崩溃后租约保留等归 #12904 / G1–G3。独立探针不等同于实际数据库撤权或重启端到端验证;H1 后合入的一方需协调契约版本。

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Merged current main (d5a157c45e) into this branch in a2a78369c5; GitHub now reports MERGEABLE (no conflicts).

The conflict was the shared API version/changelog. Main's v1.21 task-route 403 contract is preserved, and G0 is now v1.22 (1.22.0). Regenerated client types preserve both changes. No new feature or deferred Suggestion was added.

Merged-tree verification: 85 server tests + 1 real-process G0 integration test + 2 generated-type tests passed, with no failures or skips. The G0 test uses the rebuilt CLI and completes file-tool Turns through REST and WebShell. Build, typecheck, bundle, Checkstyle, formatting and commit-hook lint passed. Two resolution audit passes checked the full manual resolution, all five overlapping files and preservation of the 16 unchanged G0 files. The committed tree exactly matches the tested and audited tree.

已解决与主分支的冲突并推送:保留 v1.21 的任务接口 403 契约,G0 顺延至 v1.22。上述 88 项测试及构建、类型与格式检查通过;两轮解决冲突审计无新增问题。GitHub 已确认无冲突,合并仍以 CI 和所需审批为准。

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Fixed the CI failure and the current Critical capability-description mismatch in 312d96b.

The Ubuntu/Node 22 job failed seven positive skill-listing cases because the merged test fixture supplied a SkillManager but reported no registered Skill tool. The fix registers Skill only in that matrix's existing mock; all original assertions remain. The original file went from 7 failed / 74 passed to 81 passed. Production resume and tool-policy gates are unchanged.

Validation:

  • 174 focused core tests and 2 generated-contract tests pass.
  • Reverse verification: bypassing the per-agent policy makes all four denied/unreachable cases fail; the original production code was restored.
  • Root build, typecheck and bundle pass, as do changed-file lint/format checks.
  • Independent test-engineer verification passes the original 81-case file and the packaged CLI skill-discovery smoke. That smoke starts a fresh session; it is not a real paused-background-agent recovery test.
  • Audit rounds 10 and 11 are consecutive clean full-diff passes; an additional independent minimal review found no actionable defect.

The three renewed Critical reports against the old 3c53e18 implementation were checked against the current code; their 7528ba6 fixes remain present, and those duplicate conversations are resolved. R1-4's current contract text and generated client are corrected, and both conversations for it are resolved. The six remaining Suggestions (R1-3, R1-5, R1-7, R1-8, R1-12 and R1-14) remain deferred and open under the requested Critical-only policy after round five. The new GitHub CI run is pending; local results do not claim a green Ubuntu run.


已修复 CI 的七个失败用例及 R1-4 的能力契约说明。176 项定向测试、build/typecheck/bundle、两轮完整自查和独立验证通过;已修复问题的评论线程均已 Resolve。其余六条 Suggestion 按第五轮后的规则继续延期,新提交的 GitHub CI 正在等待结果。

@wenshao

wenshao commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator Author

Follow-up to the previous CI report: the remote rerun exposed an integration-tree difference that the previous local checks missed. The workflow checked out a synthetic merge containing the newer takeover reconciliation change from main. Acquisition now scans the original worker journal and settles the execution before returning, so the old UNKNOWN / RESOLVED expectations were no longer valid. The earlier held-response fix addressed a separate reproducible fixture race, but was insufficient for this merged contract.

Commit 97081c21f3 synchronizes main and keeps the deterministic crash window. The test requires no committed result before the first Broker is killed, requires SETTLED immediately after the replacement Broker acquires the original session, and requires same-execution replay followed by ALREADY_SETTLED. Original worker identity, binding generation, endpoint, marker contents, one initial execute and zero replayed executes remain checked. No new production behavior is introduced by the CI repair.

Verification on the integrated tree: root build, typecheck and bundle passed; 174 core regression tests, 149 Broker/takeover/repository tests, all 39 process fault gates, 39 focused Managed Agent/G0 tests, and the real-CLI HostedPublicWorkspaceIT passed without skips. Checkstyle passed. An independent test engineer verified both placements with an additional 3.5-second delay, and confirmed that deliberately delivering the result before crash or preventing worker execution makes the corresponding assertions fail. These local process checks use macOS/H2; GitHub provides the Linux/MySQL confirmation.

Two further complete self-audit passes, including reverse checks of the test evidence, and an independent scoped static review found no actionable issue. The committed tree exactly matches the tested and audited tree; commit-hook formatting and lint passed. The six previously recorded Suggestions remain deferred under the requested Critical-only policy; the eight fixed conversations are resolved.

Remote confirmation: the Java SDK workflow for 97081c2 is now fully successful. The previously failing Hosted process fault-gate job checked out synthetic merge 8290747, whose complete tree matches the locally tested commit. Its logs confirm 190 Managed Agent unit tests, all 10 Hosted integration cases (including G0), and all 39 Runtime Broker fault gates passed with zero failures, errors or skips; the durable-worker class passed all 10 cases. The platform matrix, MariaDB integration and real daemon E2E jobs also passed. This report supersedes the earlier local-only diagnosis.

At this update, unrelated Node tests, visual capture and automatic review checks are still running; this is confirmation of the repaired Java workflow, not a claim that every PR check has finished.

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Merged current main (fa4a4c92ce) in defcd631ad and resolved the single conflict with its parallel durable-worker fault-gate repair.

The resolution preserves this branch's deterministic held-response crash window and strict takeover-settlement assertions, which already passed the prior GitHub run. It retains original worker identity, successful result, and no-replay checks. All 23 existing authored paths are unchanged; the other 12 incoming paths match main exactly. No G0 production behavior or deferred Suggestion was added.

Merged-tree verification: build, typecheck and bundle passed; 65 trajectory UI tests and all 10 durable-worker fault-gate cases passed with no failures or skips. Checkstyle passed. Two clean resolution-audit passes, including a reverse check of the test evidence and exact path preservation, are complete.

The committed tree exactly matches the verified tree; commit-hook lint and formatting passed. GitHub now reports MERGEABLE (no conflicts). The new Java CI run is queued; this does not claim the new commit has completed CI.

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Merged main a903de1112 (W0e-3 trusted local reboot recovery and exact-integer validation) in fcb77a33f1. The only manual conflict was the deployment-boundary paragraph: it now preserves the new trusted-host reboot recovery scope and the existing G0 opt-in initial file Turn, while keeping later public operations and public Shell gated.

The two configuration overlaps preserve both independent, default-off options and their existing validations. No production source was manually changed for the conflict. All 20 unaffected authored paths match the previous PR head; all 46 non-overlapping incoming paths match main.

Merged-tree verification passed: root build, typecheck and bundle; 276 Broker tests, 65 focused Managed Agent tests, two real-worker/CLI integration tests and all 41 fault gates (384 tests total, no failures or skips). Checkstyle and two consecutive resolution audits, including reverse verification of the assertions and path preservation, passed. The integration tests use H2 and synthetic boot identity; they are not a new physical Linux reboot certification.

The committed tree exactly matches the tested/audited tree, and the working tree is clean. GitHub confirms MERGEABLE (no conflicts). The new Java CI run is queued; its result is still pending.

@doudouOUC doudouOUC left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed current head fcb77a3 across public and WebShell admission, actor and Workspace checks, Harness attachment and recovery, Broker routing, the fixed file-tool profile, API documentation, and tests. The previously reported Critical findings are fixed in the current code. I found no new blocking issue.

Local verification: 21 focused Java tests passed (configuration, connector, and coordinator); git diff --check passed. The PR's Java, Hosted MySQL, Node, lint, and WebShell CI checks are green. The remaining Suggestion threads are already documented for follow-up.

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head fcb77a33f1dff4351584600c25587adb1c9270e0

Verdict: COMMENT, on budget rather than on a defect. All four Criticals ever filed against this PR are fixed at this head, and I verified each one in the current code rather than accepting the round-3 ledger or the existing approval. My own Critical-only scan of the current diff did not reach the whole production surface before this run's time budget expired, so under this channel's rules I cannot issue an Approve. No new Critical is filed.

Historical blockers — all four confirmed FIXED at this head

R1-1 (HarnessCoordinator swallows the non-retryable workspace denial): FIXED. The prescribed catch arm exists in service/HarnessCoordinator.java at lines 194–197:

} catch (RuntimeBrokerException error) {
    terminal = !submissionAttempted.get() && !error.isRetryable()
            ? fail(claimed, error.getCode(), error.getMessage())
            : transientFailure(claimed, submissionAttempted.get(), error);
} catch (RuntimeException error) {

Every constraint the finding attached to the fix holds. It sits before catch (RuntimeException) (line 198), which is required because RuntimeBrokerException is final; it sits after catch (DaemonHttpException) (line 184), whose 4xx-except-409 ordering is byte-for-byte unchanged; and it classifies on error.isRetryable() rather than on the exception type, so WorkspaceExecutionStore.busy()'s retryable 409 still routes to transientFailure and keeps retrying. The failure carries the exception's own code, so a withdrawn Workspace authorization now terminates as workspace_unavailable on the first attempt instead of burning maxPreAdmissionRetries and certifying hosted_harness_unavailable. The added !submissionAttempted.get() conjunct is a sound tightening, not a weakening: once a submission has been attempted the Turn keeps the transient path rather than being failed twice.

R1-2 (README.md W0c-3 section certifies the public bound-Turn path as closed): FIXED. The false present-tense sentence is gone — remain closed and full Hosted tool loop no longer appear anywhere in the file. Line 425 now reads "Public bound Turn admission is limited to the opt-in initial file Turn described…", which is the narrowing the finding asked for rather than a deletion: the ### Private Workspace tool execution (W0c-3) heading (line 393) and the mount safety warning "provider and file tools do not confine access to the mount root" (line 420) are both still present, so an operator reading that section as the authoritative boundary now gets the correct one.

R1-9 (README.md carries two contradictory Store-scope rules): FIXED. Lines 210–213 now state the two cases the diff actually implements — "Workspace-bound Sessions use their persisted Workspace ID for the Store scope; unbound Sessions use QWEN_MANAGED_AGENT_WORKSPACE_ID" — matching QwenHostedHarnessConnector.managedSessionStore's session.workspace() == null ? workspaceId : session.workspace().getWorkspaceId(). The operationally load-bearing instruction that rides on that sentence ("must still be set explicitly when enabling the Session Store; if the Runtime Broker also has an explicit ID, startup rejects a mismatch") is preserved, so journal inspection and retention/GC scoping now point at the right ID for bound Sessions.

R1-4 (public contract certifies "does not enable execution"): FIXED. Both live capability-field descriptions were rewritten. PublicWorkspaceList.capabilities.workspace_binding in managed-agent-public-api.openapi.json and WebShellWorkspacePage.capabilities.workspaceBinding now read "Supports authorized Workspace discovery, Session creation, and saved binding read-back. This capability does not advertise execution readiness. Deployments may separately opt in to an initial Workspace Read/Write/Edit Turn at creation; later submit, cancel and lifecycle operations remain gated." The phrases empty Session creation and does not enable execution no longer occur as descriptions anywhere in the contract, and the regenerated packages/web-shell/client/components/managed/generated/managed-agent-api.ts carries the corrected text on both the capability field and the WebShell create route, so no shipped client embeds the old certification. One residue is deliberate and non-blocking: without enabling workspace execution survives once, inside the dated v1.14 clause of info.description (v1.14 integrates W0d workspace discovery … as partial; workspace_binding/workspaceBinding advertises this narrow flow without enabling workspace execution). That is a version-history entry scoped to v1.14 and rewriting it would falsify the changelog; it is not a machine-readable certification and does not reach a generated client.

Round 3 at this head posted seven findings, all severity S, and @doudouOUC approved it. Neither is the basis for the four verdicts above — each was re-read in the current code.

Gate not completed: my scan of the current diff

What I did read of the new production code raised no Critical. The opt-in is fail-closed in both directions I checked: ManagedAgentProperties.validateWorkspaceFiles() is a @PostConstruct guard that refuses startup unless harness, Session Store and Runtime Broker are all enabled with local-process provisioner, session isolation class, non-empty workspace-mounts and yolo approval mode, and workspaceFilesEnabled defaults to false; QwenHostedHarnessConnector.createOrLoad now resolves the SessionRecord first, throws for a bound Session when files are disabled, and calls workspaceExecution.authorize(session) — the recheck R1-1 depends on — while HarnessCoordinator.runClaimed independently fails a bound Turn with workspace_unavailable when isWorkspaceFilesAvailable() is false.

Not read within budget, and therefore unconfirmed rather than clear:

  • The remainder of QwenHostedHarnessConnector.java (+43/−17) — in particular toolProfile(session), which is the control that decides the tool surface a public bound Turn actually gets, and the managedSessionStore scope selection.
  • ManagedAgentStore.java (+15/−1), ManagedAgentService.java (+1/−1) — the public admission path that decides a 202-with-input is acceptable.
  • CreateHarnessSession.java (+11) and LoadHarnessSession.java (+11) — the wire fields carrying the profile.
  • HarnessConnector.java (+4) — the new interface member's default for non-Hosted connectors.
  • application.yml (+1) — the shipped default for the new flag.
  • HostedPublicWorkspaceIT.java (+355, new) and the design docs — unread, so I cannot state whether the R1-1 regression test the finding asked for (a bound Session with createOrLoad stubbed to WorkspaceExecutionStore.unavailable(), asserting failTurn(…, "workspace_unavailable", …) and never()).scheduleTurnRetry(…)) was added.

Because this PR's whole subject is opening a public file-tool Turn behind an opt-in, the unread toolProfile and admission path are exactly where a Critical would live. That is why I am not approving on the strength of the parts I did clear.

CI

Green at this head. Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), web-shell E2E Smoke, Hosted process fault gates / MySQL 8.4 / Java 21, Runtime Broker and Managed Agent MariaDB / Java 21, Real daemon E2E / Java 11 and the whole Java matrix all pass. review-pr is still pending, which this channel does not treat as a gate.

Next step: nothing further is needed on R1-1, R1-2, R1-9 or R1-4 — all four are cleared at this head. What remains is one scan pass over the six unread production paths above, at the head that will actually merge; toolProfile(session) and the ManagedAgentService/ManagedAgentStore admission change are the two worth reading first.

wenshao added a commit to wenshao/qwen-code that referenced this pull request Sep 29, 2026
@wenshao
wenshao added this pull request to the merge queue Sep 29, 2026
Merged via the queue into main with commit a631573 Sep 29, 2026
185 checks passed
@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

PR 12955 — local real-environment verification (follow-up round)

Verdict: findings — 15 passed · 1 failed · 16 total scripted assertions. The single failure is pre-existing on this host and proven not PR-caused (A/A control on base fails identically; see Finding 1). The PR's own claims all held: the central G0 claim passed its A/B on both H2 and a real local MySQL 8.0.45, and the only PR-authored production change since the last verified head is pinned by its new tests.

Verified head: fcb77a33f1dff4351584600c25587adb1c9270e0 (gh pr view headRefOid, fetched and checked out). Base: a903de1112a69259b30b914dbba84eb554802ed1 (baseRefOid). Host: Linux aarch64 (Orange Pi 5, Armbian/Ubuntu 22.04), Node v24.13.0, Oracle JDK 21+35, Maven 3.9.0, MySQL 8.0.45 — a platform cell no earlier round covered (author: macOS arm64; CI: Linux x86_64).

中文摘要

结论:findings —— 16 项脚本断言 15 通过 1 失败。唯一的失败经 A/A 对照证明是本机(Linux aarch64)既有问题,与本 PR 无关(见发现 1):base 在同一断言上以相同方式失败,失败方法未被 PR 修改,broker 生产代码在 base 与 head 之间逐字节一致,且 GitHub CI 的 fault-gate 作业在此 head 上是绿的。PR 自身的主张全部成立。

A/B 结论:中央主张在新 head 上重新测量(下表)。head 臂在 H2 与真实本地 MySQL 8.0.45 上均 1/1 通过(REST 与 WebShell 各 3 次工具执行、proof.txt=after、decoy 未被触碰、8 次模型请求);base 臂在两种数据库上都以预期的 409 workspace_unavailable 拒绝(IT:96,expected 202 but was 409)。真实 MySQL 臂补齐了上一轮 CI 验证的"未覆盖"项。

增量验证(相对上轮已验证的 3c53e18608):此后唯一的 PR 侧生产改动是 7528ba60a2(coordinator 拒绝分类,+5 行)。反向应用该补丁后,恰好两个新增分类用例变红(classifiesWorkspaceRefusalBeforeSubmission[1],[2],断言点为 failTurn(..., "workspace_unavailable", ...)),恢复后 17/17 全绿——修复被其新测试钉住。合并完整性:API 版本 head=1.22.0、main=1.21.0 且 v1.21 文本逐字保留;与最新 main(c6f52699c5)试合并干净;git diff base..head 足迹 23 文件 +875/−70 与 GitHub 元数据逐字一致。

门禁:server 模块 194/194 + Checkstyle 0 违规;qwencode SDK 165 项 0 失败 9 跳过(HostedHarnessClientTest 18/18);core 的 background-agent-resume.test.ts 81/81;fault-gates 41 项中 40 绿。

遗留状态:上轮建议(store 准入口径只有 IT 钉住)仍然成立;F2 按作者声明延期;F3 属 #12904/G1–G3。均未恶化。

未覆盖:Windows、真实模型、生产入口认证、并发/重启重试、SSE 重连、其余 Hosted*IT 类、F3 重测、取消转发竞态在 main 代码中的根因定位(PR 范围外,归因已由 A/A 完成)。

Previous-finding status (carried forward, re-measured at fcb77a33f1)

# Finding Severity Status at new head
F1 / R1-1 Pre-submission Workspace authorization refusal misclassified as transient → retried ~33 s, reported as hosted_harness_unavailable Critical (review) Fixed in 7528ba60a2. Re-measured: reverting the hunk turns exactly the two new pre-submission cases red (classifiesWorkspaceRefusalBeforeSubmission[1],[2] at the failTurn(..., "workspace_unavailable", ...) verification, IT of intended assertion — not a compile break); restored tree is 17/17 green (image 08).
R1-2 README stale "Hosted Harness rejects startup" guidance Critical (review) Fixed — scripted text check: stale claim absent, opt-in initial file Turn documented, later ops gate stated (C2).
R1-9 Store scope doc (bound vs unbound identity) Critical (review) Fixed — README now distinguishes persisted Workspace ID (bound) from QWEN_MANAGED_AGENT_WORKSPACE_ID (unbound) (C2).
R1-4 Capability description mismatch Critical (review) Fixed in 312d96b8e1 — new contract text present in openapi.json; re-running the generator reproduces managed-agent-api.ts byte for byte (C4).
CI-round Suggestion (M3) Store fixed-profile/mount admission block pinned only by the IT Suggestion Stands. No store-level unit test with an opt-in-enabled fixture was added (ManagedWorkspaceAdmissionTest still configures no mounts/opt-in-off). The guarantee still holds and is still pinned by the IT — re-confirmed green here on two databases (A1/A3). Coverage-distribution note, not a defect.
F2 No test pins the default-off gate for an otherwise admissible Workspace Suggestion Declined-with-rationale — author deferred under the Critical-only-after-round-five policy; ManagedWorkspaceFilesOptInTest was not added. I agree this is a test gap, not a defect; the deferral is recorded in the PR thread.
F3 Same-Workspace contention; crashed Turn holds the Workspace lease Scope note Deferred to #12904 / G1–G3 as declared; not re-measured here (out of scope, see Not covered).
Merge-order vs #12946 Both PRs had set the contract to 1.21.0 Observation Resolved — head is 1.22.0 with the v1.21 403-contract sentence preserved verbatim (C1); trial merge into today's main c6f52699c5 is conflict-free (C3); diff footprint 23 files, +875/−70 matches GitHub metadata exactly (C5).

Central claim and A/B (re-measured)

Central claim. With the opt-in enabled on a validated deployment shape, an authorized public Workspace Session creation (REST and WebShell adapter) runs one initial Read/Write/Edit Turn through the production Hosted Harness → Broker → worker under the selected Workspace cwd; at base, the same creation is refused.

Same IT blob on both arms (HostedPublicWorkspaceIT.java, sha256 559c960a…, cp-verified; byte-identical to round 1's), same bundled CLI on both arms (dist/cli.js built from head — a valid constant: base's daemon already carries toolProfile support, 11 references in hosted-harness-session.ts at a903de1112, verified by git grep). Base Java artifacts (qwencode, runtime-broker, managed-agent-server) were built from the base worktree into an isolated Maven repository (-Dmaven.repo.local=/root/.m2-pr12955-base) so no head artifact could leak into the control.

arm DB oracle result
head fcb77a33f1 H2 (MODE=MySQL) failsafe verdict PASS 1/1 — REST + WebShell, 3 tool executions per Session, proof.txt == "after", decoy untouched, 8 model requests, replay adds +0 calls, 409 on changed input, 400 on profile override, 403 cross-tenant
head fcb77a33f1 real MySQL 8.0.45 (local instance, dedicated hosted_harness_test schema) same PASS 1/1 — closes round 1's Not covered: real MySQL
base a903de1112 H2 same predicted red — POST /v1/agents/sessions → 409 workspace_unavailable, expected: 202 but was: 409 at IT:96
base a903de1112 real MySQL 8.0.45 same predicted red — identical 409 shape (refusal is not H2-dialect-specific)

Both base reds fail the intended assertion with named expected/actual values and are scored as passes. Witness: 07-ab-it-cells-h2-mysql.png.

Delta since the last verified head (3c53e18608 → fcb77a33f1)

The PR-authored delta is four commits; everything else is inbound main via three merges (verified under Previous-finding status / C1/C3/C5).

  • 7528ba60a2 — the only production change (+5 lines in HarnessCoordinator.coordinate: a dedicated catch (RuntimeBrokerException) that fails fast with the original code/message when the refusal is non-retryable and submission was not attempted; recovery retries otherwise). Load-bearing proof: reverting the hunk in a scratch checkout turns exactly classifiesWorkspaceRefusalBeforeSubmission[1] and [2] red at the intended failTurn(..., "workspace_unavailable", ...) verification (17 run / 2 failures); restored, the class is 17/17 green. The workspaceRefusal=true variant of doesNotExhaustAfterSubmissionMayHaveBeenAdmitted correctly survives the revert — it pins the unchanged post-submission recovery path. Witness: 08-mutation-coordinator-pins-fix.png.
  • 312d96b8e1 — capability contract text + resume-listing fixture. Verified: regenerated client types byte-identical (C4); packages/core background-agent-resume.test.ts 81/81 (D4).
  • 147740556b / 97081c21f3 — durable-worker crash-window pin + fault-gate alignment (test-only). The fault-gate suite ran here with one pre-existing failure — see Finding 1. The pinned method itself (brokerCrashAdoptsOriginalWorkerWithoutReplaying, the method the PR modifies) passed in every run (class-level 9–10/10 each time, only adoptedWorkerCanCancelItsOriginalActiveCall failed).

Targeted gates (head)

gate result
managed-agent-server surefire + checkstyle 194/194, 0 violations (round 1: 180; growth is inbound main + delta tests)
qwencode SDK surefire 165 run, 0 failures, 9 skipped (pre-existing conditional skips); HostedHarnessClientTest 18/18 — matches the PR body's "18 SDK tests"
runtime-broker -Pfault-gates (CI's exact invocation) 40/41 — Finding 1
packages/core resume fixture file 81/81

Findings

1 — Pre-existing, not PR-caused: adoptedWorkerCanCancelItsOriginalActiveCall deterministically fails on this Linux aarch64 host. (Does not block this PR; likely wants a tracking issue.)

  • Symptom. runtime-broker -Pfault-gates: 40/41 green; the failing case asserts proxy.count("cancel") == 1 ("cancellation must reach the original active worker") and gets 0 — the second Broker answers cancel with OK but never forwards it to the adopted original worker. The test's downstream assertions (which this run never reached) imply the worker then runs to completion — a quiet wrong-answer shape, not a loud crash.
  • Attribution (A/A control, measured). The base tree fails the same assertion the same way (1/1 class run); on head the method failed 4/4 single-method runs (both placements) and 2/2 full-class runs; which placement loses varies between runs, marking a timing race widened on this slow board. The failing method is not modified by the PR (the PR touches only brokerCrashAdoptsOriginalWorkerWithoutReplaying in the same file); the PR diff contains zero runtime-broker production files (base already carries feat(runtime-broker): reconcile executions on session takeover #12964 takeover reconciliation); GitHub CI's fault-gate job is green at this head (run 36559630341, MySQL 8.4 / x86_64). Conclusion: host/platform-specific pre-existing race in main's takeover cancel path, not a regression this PR introduces.
  • Suggested handling. Not a merge condition for feat(managed-agent): enable G0 public Workspace file turns #12955. Worth a repo tracking issue for Linux aarch64 (or slow-host) fault-gate coverage: on this platform a cancel acknowledged-but-never-forwarded is deterministic, and CI currently has no lane that would see it. Witness: 09-faultgate-preexisting-aa.png.

No other findings. Round-1's carried Suggestion (store admission block pinned only by the IT) still stands and is recorded in the status table above; it remains completeness reporting, not a merge condition.

Not covered

  • Windows and local x86_64 Linux (CI covers x86_64; this run adds aarch64).
  • Real model providers and production ingress authentication — the IT supplies a deterministic local model and a trusted-principal filter, same as round 1 and the author's rig; triage's ingress point stands.
  • Concurrent/restart replay and SSE reconnect — author-declared out of scope; not re-attempted.
  • The rest of the Hosted*IT family — invocation narrowed to -Dit.test=HostedPublicWorkspaceIT (the CI report gate would reject this narrowing; locally it is deliberate scoping, not a claim about the family).
  • F3 re-measurement (same-Workspace contention / held lease) — deferred to Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 / G1–G3 by declaration; no new evidence here.
  • Root-cause of Finding 1 in main's broker code — attribution is settled by the A/A control; the mechanism (takeover re-attach losing the race to cancel on slow hosts) is a hypothesis from the failure shape, not a proven root cause.
  • Web Shell UI — the IT exercises the WebShell adapter over HTTP, not the browser UI; the round-1 observation about the stale UI banner was not re-checked.
  • Repo-wide JS gates (lint, typecheck, full vitest suites) — not re-run; the PR's JS footprint is one generated-file doc comment + one test fixture, both directly exercised above.
  • The 9 skipped SDK tests — pre-existing conditional skips, uninvestigated.

Methodology

Local maintainer-driven round on an Orange Pi 5 (aarch64, 8 cores), worktrees at tmp/pr12955-head (fcb77a33f1) and tmp/pr12955-base (a903de1112). Head JS side installed with corepack pnpm install --frozen-lockfile (the repo retired package-lock.json in #11859; npm ci fails on these trees — harness fault #1, corrected) then npm run build && npm run bundle; dist/cli.js smoke-checked (--version → 0.24.6). Java side built with Oracle JDK 21 + Maven 3.9.0 with -Dgpg.skip=true (the qwencode pom signs by default and this host has no secret key — harness fault #2, corrected); head artifacts in ~/.m2, base artifacts in an isolated /root/.m2-pr12955-base seeded by a full copy. The A/B drove the same IT blob + same dist/cli.js against head and base production code; the MySQL arms used a dedicated hosted_harness_test schema + hosted user on the host's existing MySQL 8.0.45 (created for this round; no pre-existing schemas touched). The mutation probe used git apply -R of the exact 7528ba60a2 hunk with a trap-guaranteed restore; tree verified clean afterwards (git status --porcelain empty). Every assertion above maps to a scripted check whose raw log is in this directory: logs-it-{head,base}-{h2,mysql}.txt, logs-D1-server-tests.txt, logs-D2-sdk-tests.txt, logs-D3-fault-gates.txt, logs-D3b-durable-head-rerun.txt, logs-D3c-durable-base.txt, logs-D4-resume-test.txt, logs-mut-B2.txt, logs-C4-generate.txt, build logs. Assertion ledger: record.mjs → assertions.json / assertions-detail.json. Captures made with scripts/verify-capture.mjs.

Evidence images

07-ab-it-cells-h2-mysql

08-mutation-coordinator-pins-fix

09-faultgate-preexisting-aa

— Local maintainer verification round (verify-pr skill, local publish path)

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

All 13 unresolved Suggestions from this PR now have separate follow-up issues: the six originally deferred items and seven from the later review. Each issue links its source discussion and records the requested work and acceptance criteria.

Suggestion Follow-up issue
R1-3 #13044 — docs(cli): correct Hosted Runtime Broker option help
R1-5 #13045 — test(managed-agent): isolate G0 Harness temporary directories
R1-7 #13046 — test(managed-agent): cover G0 policy and mount-tenant admission guards
R1-8 #13047 — test(managed-agent): verify G0 deployment validation runs at Spring startup
R1-12 #13048 — test(managed-agent): pin the connector Workspace refusal contract
R1-14 #13049 — test(managed-agent): build G0 failure diagnostics only on failure
R3-1 #13050 — docs(managed-agent): align G0 design scope with the merged change
R3-2 #13051 — test(core): remove the redundant Skill registry override in the resume matrix
R3-3 #13052 — docs(managed-agent): correct the real-model and failover script prerequisites
R3-4 #13053 — test(managed-agent): assert G0 Workspace admission error codes
R3-5 #13054 — feat(managed-agent): define recovery visibility for post-submission Workspace refusal
R3-6 #13055 — test(managed-agent): clarify retryable Workspace refusal coverage
R3-7 #13056 — test(managed-agent): clarify the post-submission refusal regression witness

The original conversations have issue links and remain open until their fixes are implemented. The post-submission refusal item tracks a recovery-policy decision; it does not approve failing an uncertain submission based only on a missing local watermark. The two related coordinator-test items link to that policy issue.

已把全部 13 条未解决 Suggestion 分别转为 Issue(此前 6 条及后续 7 条),并回填原讨论;创建 Issue 不代表已经修复。

pull Bot pushed a commit to mcx/qwen-code that referenced this pull request Sep 30, 2026
…wenLM#13099)

* test(managed-agent): state which coordinator refusal paths are real

Follow-ups from the QwenLM#12955 review (R3-6, R3-7), test-only:

- Split the pre-submission refusal matrix so each case names what it
  covers. A Workspace authorization refusal before submission uses the
  real WorkspaceExecutionStore.unavailable() and fails the Turn; the
  same refusal on a claim that already recorded a submission is retried.
  The retryable RuntimeBrokerException case is kept as explicitly
  defensive coverage with a neutral fixture, since no current producer
  raises one before submission and Broker lease contention does not
  reach that arm (QwenLM#13055).
- Cover the path lease contention actually takes: a 409 reported by the
  Harness as DaemonHttpException is retried (QwenLM#13055).
- Rename the post-submission test to the invariant it protects, keep
  both rows including the lost-response case, and record that weakening
  the RuntimeBrokerException arm's guard, not deleting the arm, is the
  negative control for the refusal row (QwenLM#13056).

No production change. Not built or run locally.

* test(managed-agent): correct the 409 attribution and pin the 4xx/5xx cells

Review follow-up on this PR:

- The comments claimed Broker lease contention (workspace_busy) reaches
  the coordinator as a DaemonHttpException 409. It does not: the daemon
  ends the turn with a turn_error event before any coordinator call can
  see it, and a create-time 409 is answered by loading. Describe the
  409 test as pinning session-level Harness conflicts instead.
- Pin the other two cells of the DaemonHttpException classification
  after a recorded submission: a 400 fails the Turn as
  hosted_harness_rejected, and a 503 is retried.

---------

Co-authored-by: yiliang114 <[email protected]>
he-yufeng pushed a commit to he-yufeng/qwen-code that referenced this pull request Sep 30, 2026
* test(managed-agent): close the G0 admission follow-ups

Post-merge follow-ups from the QwenLM#12955 review, all test or documentation
only:

- Pin TMPDIR, TMP and TEMP to the test directory after the Hosted
  Harness environment is cleared, as the sibling MySQL fixture does
  (QwenLM#13045).
- Add a fast admission case with the file flag and mounts enabled: a
  frozen configuration with a drifted policy and a mount for another
  tenant on the same storage ID are both refused with 409
  workspace_unavailable and persist nothing, while the same store
  admits the Workspace that passes every guard (QwenLM#13046).
- Build the Turn and Harness-startup diagnostics only when the Turn has
  failed or the Harness has exited, instead of on every poll (QwenLM#13049).
- Assert error.code alongside the 409 for the unsupported-agent,
  draining, unmounted and unsupported-configuration refusals (QwenLM#13053).
- Separate G0's production scope from the generated WebShell types, the
  Broker fault-gate change and the core resume-test adjustment in both
  language versions of the design (QwenLM#13050).

Not built or run locally; CI is expected to confirm.

* test(managed-agent): run the gated admission case inside a transaction

The test constructs its own ManagedAgentStore to enable Workspace files,
so calls miss the Spring bean's @transactional proxy, and
ManagedWorkspaceRegistry.resolveForCreation refuses outside a creation
transaction. CI failed on the ws-policy case with that
IllegalStateException instead of the expected refusal. Wrap each call
in a TransactionTemplate, as ManagedSessionLifecycleTest does.

* test(managed-agent): name the refused Workspace and pin the store guard

Review follow-up on this PR:

- Capture the refusal with catchThrowable before asserting, so the
  ws-policy / ws-tenant label is kept when nothing is thrown.
- Assert the store's message, which the registry's own refusal does not
  share, so a refusal moved to the registry layer no longer passes.
- Fold registerFrozen into a five-argument register so the registry
  INSERT exists once in this file.

---------

Co-authored-by: yiliang114 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants