Skip to content

feat(serve): implement generic Broker provider controls - #12868

Merged
wenshao merged 19 commits into
mainfrom
codex/broker-provider-control-12765
Sep 29, 2026
Merged

wenshao merged 19 commits into
mainfrom
codex/broker-provider-control-12765

Conversation

@wenshao

@wenshao wenshao commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Connects the Broker provider to a versioned worker contract for manifest, turn preparation, tool preparation, approval, preflight, and file history. Provider invocations reserve a durable seven-field reference and start against the original prepared worker state, with no tool arguments stored in the execution reference. Session release requires the worker to close admission before Workspace activation and storage ownership are released. Invalid tool inputs and unsupported provider profiles retain specific, bounded diagnostic reasons across the worker, Java Broker and TypeScript client.

Refused provider acquisition releases its provisional admission claim so otherwise valid raw-tool work remains available. Status and cancellation report unknown for forgotten invocations, while changed references to retained invocations and attempts to replay forgotten execution remain rejected.

The raw four-field reference plus payloadJson reserve/start path from #12831 remains available. The stored reference fixes which execution contract applies; retries cannot switch protocols. The fault gates now exercise production Session release and reserve/start with separate tool payloads, removing both workarounds described in #12765.

Why it's needed

The existing provider exposed controls that the production HTTP transport rejected, and the original fault gates supplied local Session acknowledgements and stored tool arguments inside references. This completes the private provider control path while preserving the separately gated Hosted Workspace tool loop.

Reviewer Test Plan

How to verify

  • Acquire a provider Session, bind its file history, read its manifest, begin a turn, and prepare a write. Preparation and durable reservation must leave the file untouched; approval, preflight, and start should execute the original invocation once.
  • Repeat reservation/start and inspect the stored reference. Provider references must contain exactly seven identity fields, without tool names or arguments. Raw reservations must still require the matching payloadJson; mixing the two start contracts must fail before dispatch.
  • Cancel a prepared invocation, then release its Session. Release must wait for cleanup, reject active work, preserve existing evidence, and refuse new work or a conflicting owner. A lost execution response must remain UNKNOWN without replay.
  • Submit a relative write path, an unknown tool and an argument above 256 KiB. The caller should receive the specific validation reason; corrected input in the same Session should prepare and execute successfully. Unsupported provider profiles should fail explicitly. An execution refusal must remain UNKNOWN without another dispatch.
  • Refuse provider acquisition before Workspace activation or under an unsupported provider profile, then execute through the raw protocol after valid setup. Raw work should remain available, while successful provider acquisition and explicit release should retain their fences. Advance a settled invocation to the next turn: status/cancel should report exactly unknown, and execute must refuse replay.
  • Exercise the existing Hosted Workspace read/write/edit loop across two Workspaces. History reload, lost start acknowledgement, Shell refusal, and unknown-result blocking must retain their existing behavior.

Evidence (Before & After)

N/A for TUI/screenshots; this is a private protocol change. The baseline production Session path returned 501 and the global CLI lacked the provider control route. The real TS provider → Java Broker → bundled worker path now passes 40 provider E2E groups and 8 error/recovery groups (including four readiness groups in total). The packaged Hosted Workspace tool loop was rerun successfully after merging main.

Latest review fixes: build/typecheck/bundle, ESLint/Prettier, 136 core tests and 859 CLI tests passed. Three isolated rollback/lookup mutations were detected by the regression tests. Real-worker verification and the 40-group provider plus 8-group error/recovery chains passed. Three full open-ended/reverse audit rounds completed: round 1 corrected an old test expectation; rounds 2 and 3 were clean.

Validation of the merge with main at d66fdadd2794: full build, typecheck and bundle; changed-file ESLint/Prettier; 115 CLI tests; all 402 Broker tests with zero skips; 17 Workspace/schema tests; and the packaged Hosted Workspace tool-loop integration. Java Checkstyle passed. Real-worker probes confirm that released terminal receipt reads, cancellation and identical reservation retries add no worker requests and leave the released Session unchanged. The earlier feedback revision also passed 136 core and 892 CLI tests and four HTTP guard mutation checks; those broader TypeScript suites were not rerun for this merge. Before committing, two consecutive open-ended and reverse audit rounds found no new Critical issues in the conflict resolutions and their affected consumers.

Tested on

OS Status
🍏 macOS ✅ tested
🪟 Windows ⚠️ not tested
🐧 Linux ⚠️ not tested

Environment (optional)

Node.js 22.22.2, JDK 21, real local worker processes and JDBC execution storage. The Hosted integration uses a deterministic local model endpoint with the packaged Harness, Spring Broker, workers, and SQL store.

Risk & Scope

  • Main risk or tradeoff: provider lifecycle and raw-tool dispatch now share the Broker journal while retaining distinct wire contracts. Admission, authority, cancellation, and uncertain outcomes fail closed.
  • Not validated / out of scope: live-model behavior, arbitrary configuration loading, public Hosted admission expansion, prepared-state recovery after worker restart, and cross-platform execution. Worker observations retain current-turn scope. Broker HTTP can return owned persisted terminal execution receipts without a READY Session; live observation and history controls still require READY. Runtime-loss recovery can fence uncertainty as ABANDONED, which remains permanently unknown and cannot be replayed. Released worker Sessions currently retain heavyweight runtime state until worker shutdown, and an execution refusal conservatively becoming UNKNOWN can retain Workspace ownership; both remain documented follow-up work.
  • Breaking changes / migration notes: the raw reserve/start API remains compatible and no database migration is required. The new provider controls require a worker that implements the matching protocol. Boot v1 uses its fixed tool configuration and leaves confirmation enforcement to the caller; boot v2 requires the existing installed and activated Workspace profile.

Design: English · 简体中文. Both versions describe the same contract, boundaries, and acceptance criteria.

Linked Issues

Fixes #12765. Related to #12380; preserves the separately gated behavior introduced by #12831.

中文说明

What this PR does

通过版本化 worker 契约接通 Broker provider 的 manifest、回合准备、工具准备、审批、preflight 与文件历史控制。Provider 调用持久化预留七字段 reference,再针对原 worker 已准备状态启动;执行 reference 不保存工具参数。释放 Session 时,必须先由 worker 关闭准入,再停用 Workspace 激活并释放存储持有。 无效工具参数与不支持的 provider profile 会跨 worker、Java Broker 和 TypeScript 客户端保留明确且限长的错误原因。

provider 获取被拒时撤销临时准入登记,保留原本合法的原始工具调用。已遗忘调用的状态和取消返回 unknown;仍保留调用的引用被改动,以及重放已遗忘执行的尝试,继续被拒绝。

保留 #12831 的原始四字段 reference 加 payloadJson reserve/start 路径。保存的 reference 固定执行契约,重试不能切换协议。故障门禁现在使用生产 Session release 和参数分离的 reserve/start,移除了 #12765 描述的两处绕行。

Why it's needed

既有 provider 暴露的控制被生产 HTTP transport 拒绝,原故障门禁还在本地应答 Session 生命周期,并将工具参数存进 reference。本次补全私有 provider 控制链路,同时保留独立门控的 Hosted Workspace 工具循环。

Reviewer Test Plan

How to verify

  • 获取 provider Session、绑定文件历史、读取 manifest、开始回合并准备写入。准备和持久预留都不能修改文件;审批、preflight 与 start 后,原调用应只执行一次。
  • 重复预留和启动并检查保存的 reference。Provider reference 必须恰有七个身份字段,不含工具名或参数。原始预留仍必须提供匹配的 payloadJson;混用两种 start 契约必须在派发前失败。
  • 取消已准备调用,再释放 Session。释放必须等待清理、拒绝活跃工作、保留已有证据,并拒绝新工作或冲突 owner。执行响应丢失必须保持 UNKNOWN,不得重放。
  • 提交相对写入路径、未知工具与超过 256 KiB 的参数。调用方应收到具体校验原因;同一 Session 中修正参数后应能准备并执行。不支持的 provider profile 应明确拒绝。执行阶段拒绝仍应保持 UNKNOWN,不能再次派发。
  • 在 Workspace 激活前或不支持的 provider profile 下尝试获取并被拒,再完成合法配置后调用原始工具,原始调用应可用;成功的 provider 获取与显式释放仍应保持各自准入限制。将已结算调用推进到下一回合:旧引用的状态/取消应只返回 unknown,execute 必须拒绝重放。
  • 在两个 Workspace 中验证已有 Hosted 读写编辑循环。历史重载、start 回执丢失、Shell 拒绝和未知结果阻塞必须保持原行为。

Evidence (Before & After)

TUI/截图不适用;这是私有协议变更。基线生产 Session 路径返回 501,全局 CLI 不提供 provider 控制路由。真实 TS provider → Java Broker → bundled worker 链路现通过 40 组 provider E2E 和 8 组错误/恢复检查(合计含四组启动检查)。合并 main 后已重新验证打包后的 Hosted Workspace 工具循环。

最新评审修复通过 build、typecheck、bundle、ESLint/Prettier、136 个 core 测试和 859 个 CLI 测试。三项独立的清理/观察逻辑变异均被回归测试检出。真实 worker 验证以及 40 组 provider、8 组错误/恢复链路检查通过。完成三轮完整无方向审计与反向审计:第一轮修正一条旧测试预期,第二、三轮连续 clean。

合并 main d66fdadd2794 后,通过完整 build、typecheck、bundle、变更文件 ESLint/Prettier、115 个 CLI 测试、全部 402 个 Broker 测试(无跳过)、17 个 Workspace/数据库结构测试,以及打包后的 Hosted Workspace 工具循环集成。Java Checkstyle 通过。真实 worker 探针确认释放后读取终态回执、取消与同键预留重试均不增加 worker 请求,也不改变已释放的 Session。此前反馈修订还通过了 136 个 core 和 892 个 CLI 测试及四项 HTTP 校验变异检查;本次合并未重跑这些较广的 TypeScript 套件。提交前连续两轮无方向审计与反向审计未在冲突解决及受影响调用链中发现新增 Critical。

Tested on

OS Status
🍏 macOS ✅ 已测试
🪟 Windows ⚠️ 未测试
🐧 Linux ⚠️ 未测试

Environment (optional)

Node.js 22.22.2、JDK 21、真实本地 worker 进程及 JDBC 执行存储。Hosted 集成使用本地确定性模型端点,Harness、Spring Broker、worker 和 SQL 存储均为打包后的真实实现。

Risk & Scope

  • 主要风险或权衡:provider 生命周期和原始工具派发共用 Broker 日志,但保留不同的线上契约。准入、授权、取消与不确定结果均失败关闭。
  • 未验证或范围外:真实模型行为、任意配置加载、公开 Hosted 准入扩展、worker 重启后的准备状态恢复,以及跨平台执行。Worker 观察仍限当前回合。Broker HTTP 可在 Session 非 READY 时返回归属已验证的持久终态回执;实时观察与历史控制仍要求 READY。Runtime 丢失恢复可将不确定执行封存为 ABANDONED,结果永久未知,不得重放。已释放的 worker Session 当前仍保留重型 runtime 状态直到 worker 关闭,执行拒绝保守转为 UNKNOWN 时也可能继续持有 Workspace;两项均已记录为后续工作。
  • 破坏性变更或迁移:原始 reserve/start API 保持兼容,无需数据库迁移。新 provider 控制需要实现匹配协议的 worker。Boot v1 使用固定工具配置,确认决策由调用方落实;boot v2 要求既有 Workspace profile 已安装并激活。

设计:English · 简体中文。两版契约、边界及验收标准完整同步。

Linked Issues

Fixes #12765。关联 #12380;保留 #12831 引入的独立门控行为。

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

E2E verification on 233d49b5cc plus this feature, using a fresh build/typecheck/bundle and JDK 21:

  • Provider contract: 39/39 check groups passed, no skips. The built TypeScript provider calls the production Java Broker/HTTP transport and the bundled boot-v1 worker, with H2 JDBC execution persistence. Checks cover all nine controls, exact seven-field persisted references, no effects before start, one dispatch across retries, approval/cancellation, active-release rejection, unreserved cleanup, identity/fencing, and lost-response UNKNOWN without replay.
  • HostedWorkspaceToolTurnIT: 1/1 passed. The packaged Hosted Harness, Spring Broker, boot-v2 Workspace workers and SQL store preserve feat(serve): add gated Hosted Workspace file tool turns #12831's raw payloadJson path: two Workspaces, read/write/edit, history reload, lost start acknowledgement without duplicate dispatch, Shell refusal, and blocking on unknown outcomes.
  • Fault gates: 28/28 passed (context 12, lost response 7, process crash 6, concurrency/storage 2, control 1). After correcting the managed test adapter to use real provider release before Workspace deactivation, the affected context group passed again, including the observed control → activation request order.
  • Other validation: 966 relevant TypeScript tests; 384 Broker unit tests with zero failures/errors and one existing optional real-worker test skipped; 12 Workspace tests; build/typecheck/bundle, changed-file ESLint/Prettier, and Checkstyle for both affected Java modules.

The provider E2E uses a static pre-attested lease and deterministic Session/binding fixtures; the execution transport, worker, tools and H2 execution repository are real. The Hosted test uses a deterministic local model endpoint. These runs do not claim live-model, MySQL, cross-platform or worker-restart recovery coverage. Worker observation remains current-turn scoped; Broker HTTP observation requires a READY Session, while durable execution evidence survives release.

提交前验证:provider 39/39 组、Hosted 集成 1/1、故障门禁 28/28 均通过;审计修正后受影响的 12 项故障测试再次通过,并确认真实 release 在 Workspace 停用前完成。另有 966 个 TypeScript 测试、384 个 Broker 单测(既有可选跳过 1 个)和 12 个 Workspace 测试通过,构建、类型、格式、lint 及两个 Java 模块的 Checkstyle 均通过。覆盖边界如上,未把确定性模型或本地 H2 测试宣称为线上、MySQL 或重启恢复验证。

Three complete open-ended and reverse audit rounds preceded submission. The first found and corrected the managed fault-gate release adapter; rounds 2 and 3 independently reviewed the full updated diff and found no remaining actionable findings. The final staged tree matched the reviewed snapshot, and the actual E2E runtime artifact hashes remained unchanged. This records manual audits; the earlier model-CLI review timed out and is not counted as a passing check.

提交前完成三轮完整无方向审计与反向审计。第一轮修正 managed 故障测试的 release 适配层,第二、三轮独立全量复核均无剩余可操作问题。最终暂存内容与审计快照完全一致,E2E 运行产物哈希保持匹配。此处记录人工式独立代码审计;此前模型 CLI review 超时,未计入通过项。

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification of f67bde6b41

Verdict: ready to merge from the real-stack side. No blocking finding. Every item of the reviewer test plan held on a stack the PR's own evidence did not cover: MySQL 8.4.7 instead of H2, workers provisioned by the production local-process provisioner instead of a static lease, and a live model on the Hosted Workspace loop. Three follow-up observations (F1–F3) and two notes are listed below; none of them changes an outcome the PR claims.

What was run Result
A/B walk of the provider path, merge-base 233d49b5cc vs this head base: 501 runtime_session_verb_unsupported on every control · head: all steps ok, 19 provider requests reach the worker
Reviewer test plan on the real stack, boot v1 and boot v2 76/76 checks
Fault injection: lost execute reply, worker SIGKILL, races, refusals 15/15 checks
Hosted Workspace loop (#12831 raw path) with a live model, 3 turns 5 tool executions SETTLED/success, never recovery-blocked
The PR's own Hosted driver, unchanged, on MySQL 8.4.7 HOSTED_WORKSPACE_TOOLS_OK, 7 executions SETTLED/success
Repository suites on this head (JDK 21, MySQL 8.4.7) all green, see table below
Mutation matrix, 40 single-line mutants of the new production code 33 killed by the repository's unit suites, 7 survive

Setup

Everything on the request path is the shipped artifact: the built TypeScript BrokerManagedRuntimeProvider (packages/cli/dist), the Spring managed-agent-server jar with its embedded Runtime Broker and production HttpRuntimeTransport / WorkspaceRuntimeTransport, Flyway-migrated MySQL 8.4.7, and bundled workers (dist/cli.js managed-runtime-worker) started by the local-process provisioner. Sessions are created through the public /v1/agents/sessions API: a plain Session gives boot v1, a Workspace Session gives boot v2 with installed context, activation and storage ownership.

Two things are rig-only: a small auth adapter that names the actor for the public API, and an HTTP forward proxy between Broker and worker (JVM -Dhttp.proxyHost, empty nonProxyHosts). The proxy keeps a ledger of every Broker → worker request and injects the faults. Effects are counted with a shell command that appends a line; records are read from MySQL directly.

1. Before and after

A/B provider walk

2. The reviewer test plan

Contract checks

Test plan item Measured
Preparation and durable reservation leave the file untouched; approval, preflight and start execute once A.3–A.9: one PREPARED row at dispatch generation 0 and no file; after start one appended line. Two more provider starts, two raw :start posts and a restarted provider return the same execution: 1 row, generation 1, 1 execute on the wire, 1 line. 24 concurrent starts: still 1 (H.1)
Stored reference has exactly seven identity fields, no tool name or arguments A.5/A.6 and the row dump below. The argument text appears in one request only (prepare) and nowhere in MySQL
Raw reservations still need the matching payloadJson; mixing contracts fails before dispatch C.1 provider reservation + payload → 409 runtime_execution_conflict, 0 worker requests · C.5 raw reservation without payload → 400 runtime_payload_invalid, 0 worker requests · C.4 immediate route + seven-field reference → 400
Cancel, then release: waits for cleanup, rejects active work, keeps evidence, refuses new work and a conflicting owner D.1–D.14. Release is refused with 409 runtime_session_busy while a reservation is pending or a command runs; cancelling a running command settles cancelled within 60 ms and leaves no late effect; the released Session refuses acquire and control; another Harness Session or turn kind gets 409 runtime_session_conflict
Lost execution response stays UNKNOWN without replay F.1/F.2: the command ran once, the record is UNKNOWN, and retries, a restarted provider, inspect and reconcile never produce a second execute
Hosted Workspace loop keeps its behaviour: history reload, lost start acknowledgement, Shell refusal, unknown-result blocking Section 4: the PR's own driver on MySQL, plus a live model

Stored reference and wire

3. Fault injection

Faults

After the worker is killed, start for the reserved invocation fails closed with no worker request at all, and nothing is rebuilt from the stored reference. That matches the open boundary the design document states. On boot v1 the next Runtime Session gets a replacement worker and runs normally.

4. Live model on the Hosted Workspace loop

The PR changes the release path of raw-only Sessions too: the worker now acknowledges release before the Workspace is deactivated. Two runs cover it.

HostedWorkspaceToolTurnIT pins an in-memory H2 database, so I pointed the PR's own driver, unchanged, at the Spring server on MySQL 8.4.7. It prints HOSTED_WORKSPACE_TOOLS_OK (two Workspaces, two tool rounds, lost start acknowledgement, replay, Shell refusal, unknown blocking), with 7 executions SETTLED/success and both worker release acknowledgements followed by the Workspace deactivation.

Then the same loop with qwen3.8-max:

Live model

5. Repository suites on this head

Suite Result
runtime-broker unit 384 run, 0 failed, 1 skipped (the optional real-worker test)
runtime-broker fault gates (-Pfault-gates) 28/28: context 12, lost response 7, process crash 6, concurrency/storage 2, control 1
runtime-broker MySQL integration, MySQL 8.4.7 2/2
managed-agent-server unit + MySQL integration + Checkstyle 138/138 and 10/10, 0 violations
HostedHarnessMySqlIT on MySQL 8.4.7, HostedWorkspaceToolTurnIT on its built-in H2 1/1 and 1/1
runtime-broker Checkstyle pass
cli src/serve (245 files) 11,054 passed; 37 failed in 5 files this PR does not touch
core managed-tool* 189/189

The 5 failing cli files are local to my machine: ssh-workspace, session-pr-backfill, session-pr-refresh and aone-mrs fail with the same 36 tests on the merge-base build, and web-shell-pairing passed on rerun. CI is green on this head (23 checks passed), and the bot review approved this commit.

6. Mutation matrix

Mutation matrix

33 of 40 single-edit mutants are killed by the repository's unit suites. The 7 survivors are all second guards; the first guard for the same input sits on the other side of the wire.

  • J4, J5, J13, J14: the closed field sets of the reserve and start bodies, unknown fields in a control operation, and the file-history owner check in the Broker. I compiled these four into one server jar and reran the probes. B.7, B.9, C.2 and C.3 fail on that jar, so the real stack notices what the unit suites do not. Four small HTTP-level tests would pin them.
  • T2, T8, T18: closeSessionAdmission on provider release (the Session is already fenced by providerSessions), the worker's refusal to release with pending controls (the Broker refuses first), and the worker's 413 for an oversized response (the transport has its own bound).

J2 mutates a guard that predates this PR; it is listed because the new mixing rule depends on it. Four failures in hosted-harness-session.test.ts appeared during the full runs under load. They are not kills: with the same mutants applied, that file passes 16/16 on its own.

Points from the bot review, measured

The bot review read the code and named the missing end-to-end run as its one gap. This report is that run. Its other points, against measurements:

Bot point Measured
1. Worker error codes collapse distinct failures Confirmed and wider than stated: 10 different reasons arrived under 409 managed_runtime_identity_conflict, and the Broker drops the reason text (F2)
2. observePreparedCancellation polls every 25 ms without backoff In the normal case one cancel and one status reach the worker (D.4). A worker that never settles was not tested
3. Released provider Sessions live for the worker generation Measured at about 80 KiB per released Session (F3). The review expects the state to serve post-release observation. It cannot: after release, history, status and cancel answer 404 runtime_session_not_found at the Broker, the provider refuses locally, and 0 requests reach the worker
4. A cancelled shell result now reports cancelled Seen on the real worker: a cancelled running command settles cancelled (D.8)
Open question: does JsonCodec writing nulls change persisted bytes? No, for the local-process provisioner. resource_handle_json is {"provider":"local-process"} in the 4 rows the merge-base build wrote and in the 44 rows this build wrote. The provision seed has six non-null fields. I also started this build on the database the merge-base build had written: Flyway stays at rank 15, the provider walk passes on both boots, and the server log has no ERROR line. Static and Kubernetes provisioners were not tested

Observations (non-blocking)

Observations

F1. A definite worker refusal is recorded as UNKNOWN, and the Session can then never be released. When the worker answers execute with 409 and a reason, the Broker stores UNKNOWN, exactly as it does for a lost reply. UNKNOWN counts as active work, :resolve answers 501, and the HTTP face has no reconcile route, so release answers 409 runtime_session_busy from then on. On boot v2 the storage claim stays: the next Runtime Session and a sibling Harness Session of the same Workspace both get 409 workspace_busy.

This is inherited, not introduced. The merge-base build does the same on the raw path (Managed Runtime does not admit this tool. → UNKNOWN → release refused, storage held). What this PR adds is more refusals that land there. I saw three new ones: preflight has not permitted execution, invocation cancelled and Managed Runtime protocol conflicts (a raw call into a provider-owned Session, C.6). So "mixed admission is rejected" holds without effect, but it costs the Session. A possible direction: let the worker mark refusals that happen before dispatch with their own code, and let the transport settle those as not_started, the way WorkspaceRuntimeTransport.execute already does for its own pre-dispatch checks. A plain "409 means not started" rule would be unsafe: by my reading of the route's catch block, an error raised after the tool started is answered with 409 as well.

F2. Invalid tool arguments reach the caller without the reason. For write_file with a relative path the worker answers 409 managed_runtime_identity_conflict with File path must be absolute: relative/path.txt. The Broker forwards the code and replaces the message with Managed Runtime control request failed (HTTP 409).; the TypeScript client keeps only status and code. An unknown tool name gives the same code. The worker used this one code for 10 different reasons in my runs. A Harness that wants to hand a validation error back to the model cannot tell it from an identity conflict. An oversized argument is renamed on the way: the worker says 400 managed_runtime_provider_invalid, the caller sees 400 managed_runtime_attestation_invalid.

F3. The worker keeps every released provider Session. One worker, 1,500 Runtime Sessions acquired, used and released in sequence: RSS 161 → 316 MiB, 278 MiB after 15 s idle, about 80 KiB per released Session. The same count through the raw path on the same build stays flat (152 → 146 MiB). ManagedRuntimeProviderWorker.sessions never drops an entry, and the entry holds the Session's Config, tool set and file history. After release the Broker answers 404 runtime_session_not_found for history, status and cancel, so nothing can read that state any more. With one Runtime Session per turn this grows with conversation length.

Note 1. Approval is the caller's duty, not the worker's. On boot v1 prepare reports defaultPermission: "ask", yet prepare → preflight → start with no confirm writes the file (B.1b). That is core's existing contract ("the private parent calls this only after permission"). I mention it because "boot v1 uses DEFAULT approval" can be read as the worker enforcing it.

Note 2. The argument bound is 256 KiB, not the 1 MiB envelope. A 255 KiB write_file goes through all steps (the largest response is 511 KiB); 256 KiB is refused at prepare by core's MAX_JSON_BYTES. This equals the raw path's payload bound, so the two paths agree.

Merge order. A trial merge with current main (3f5ae3ffeb), #12855 and #12870 is clean. With #12848 it conflicts in managed-context-worker.ts, RuntimeBrokerHttpServer.java and RuntimeBrokerService.java; with the W0e drafts #12839, #12865 and #12869 it conflicts in RuntimeBrokerService.java and WorkspaceRuntimeTest.java. Whichever lands second needs a manual resolution and a rerun of the fault gates.

Not covered

Linux and Windows; MariaDB (CI covers it); Broker restart with live provider Sessions and prepared-state recovery after a worker restart, which the PR lists as out of scope; a product flow that drives the provider path with a model, because no consumer of that path exists yet. The provider probes are therefore at the level of the provider API.

Reproduce

Rig, probes, mutant list and raw logs: pr12868/harness and pr12868/results. Order: spring.sh with the proxy, then s4-ab, s1-contract, s2-faults, s5-raw, s9-author-driver-mysql, s6-hosted-real-model, s3-retention, s8-limits, mutate.

中文版

在 f67bde6b41 上的真实环境验证

结论:从真实环境验证的角度可以合并,没有阻塞问题。 Reviewer Test Plan 的每一条都在 PR 自带证据没有覆盖的环境里成立:MySQL 8.4.7(PR 用的是 H2)、由生产 local-process provisioner 拉起的 worker(PR 用的是静态租约)、以及真实模型驱动的 Hosted Workspace 循环。下面另有三条后续观察(F1–F3)和两条说明,都不改变 PR 声明的任何结果。

跑了什么 结果
provider 链路 A/B:merge-base 233d49b5cc 对比本 head 基线:每个控制都是 501 runtime_session_verb_unsupported · 本 head:全部 ok,19 个 provider 请求到达 worker
在真实栈上执行 Reviewer Test Plan(boot v1 与 boot v2) 76/76 项通过
故障注入:execute 响应丢失、worker 被 SIGKILL、并发竞争、拒绝 15/15 项通过
真实模型跑 Hosted Workspace 循环(#12831 raw 路径),3 轮 5 次工具执行 SETTLED/success,从未进入 recovery-blocked
PR 自带的 Hosted driver 原样接到 MySQL 8.4.7 HOSTED_WORKSPACE_TOOLS_OK,7 次执行 SETTLED/success
本 head 上的仓库自带套件(JDK 21、MySQL 8.4.7) 全绿,见下表
变异矩阵:新增生产代码的 40 个单行变异体 仓库单测杀掉 33 个,7 个存活

环境

请求链路上全部是发布产物:构建出来的 TypeScript BrokerManagedRuntimeProvider(packages/cli/dist)、Spring managed-agent-server jar 及其内嵌 Runtime Broker 与生产 HttpRuntimeTransport / WorkspaceRuntimeTransport、经 Flyway 迁移的 MySQL 8.4.7、由 local-process provisioner 启动的打包 worker(dist/cli.js managed-runtime-worker)。Session 通过公开的 /v1/agents/sessions API 创建:普通 Session 对应 boot v1,Workspace Session 对应 boot v2(含上下文安装、激活与存储持有)。

只有两样是装置自带的:一个很小的鉴权适配器(给公开 API 指定 actor),以及 Broker 与 worker 之间的 HTTP 正向代理(JVM -Dhttp.proxyHost,nonProxyHosts 置空)。代理记录每一个 Broker → worker 请求,并负责注入故障。副作用用「追加一行」的 shell 命令计数,记录直接从 MySQL 读取。

1. 修改前后对比

见上方图 1。同一个脚本、同一个 MySQL、两个构建。基线上生产 HTTP transport 对所有 Session 动词返回 501,普通 Session 连 acquire 都失败并停在 ACQUIRING;本 head 上九个控制、七字段预留、start 和 release 都对真实 worker 完成。

2. Reviewer Test Plan

测试计划条目 实测
准备和持久预留不改文件;审批、preflight、start 后只执行一次 A.3–A.9:预留后只有一条 PREPARED 记录(dispatch generation 0),文件不存在;start 后文件追加一行。再做两次 provider start、两次裸 :start、一次重启后的 provider,得到同一个执行:1 条记录、generation 1、线上 1 次 execute、1 行。24 个并发 start 仍然是 1 次(H.1)
保存的 reference 恰有七个身份字段,不含工具名和参数 A.5/A.6 及图 3 的记录转储。参数文本只出现在一个请求里(prepare),MySQL 中没有
raw 预留仍需匹配的 payloadJson;混用两种契约在派发前失败 C.1 provider 预留 + payload → 409 runtime_execution_conflict,0 个 worker 请求 · C.5 raw 预留不带 payload → 400 runtime_payload_invalid,0 个 worker 请求 · C.4 即时路由 + 七字段 reference → 400
取消后释放:等待清理、拒绝活跃工作、保留证据、拒绝新工作与冲突 owner D.1–D.14。存在未取消的预留或命令在运行时,release 返回 409 runtime_session_busy;取消运行中的命令 60 ms 内结算为 cancelled,之后没有迟到的副作用;已释放的 Session 拒绝 acquire 与 control;其他 Harness Session 或其他 turn kind 得到 409 runtime_session_conflict
执行响应丢失保持 UNKNOWN,不重放 F.1/F.2:命令只跑了一次,记录为 UNKNOWN;重试、重启后的 provider、inspect、reconcile 都没有产生第二次 execute
Hosted Workspace 循环保持原行为:历史重载、start 回执丢失、Shell 拒绝、未知结果阻塞 第 4 节:PR 自带 driver 接 MySQL,外加真实模型

3. 故障注入

worker 被杀后,对已预留调用的 start 失败关闭,连 worker 请求都不会发出,也不会从保存的 reference 重建任何东西,与设计文档写明的开放边界一致。boot v1 下,下一个 Runtime Session 会得到替代 worker 并正常工作。

4. 真实模型跑 Hosted Workspace 循环

本 PR 也改了 raw-only Session 的释放路径:现在先由 worker 确认 release,再停用 Workspace。我用两次运行覆盖它。

HostedWorkspaceToolTurnIT 写死使用内存 H2,所以我把 PR 自带的 driver 原样接到 MySQL 8.4.7 后端的 Spring 上。它输出 HOSTED_WORKSPACE_TOOLS_OK(两个 Workspace、两轮工具、start 回执丢失、重放、Shell 拒绝、未知结果阻塞),7 次执行 SETTLED/success,两次 worker release 确认之后都跟着 Workspace 停用。

然后用 qwen3.8-max 跑同一个循环:三轮读 / 写 / 编辑,5 次工具执行全部 SETTLED/success,保存的 reference 仍是 raw 形状,每个工具轮都以 control[release] → /v3/activation 收尾,每轮结束后存储持有都被释放。

5. 仓库自带套件

套件 结果
runtime-broker 单测 384 个,0 失败,跳过 1 个(可选的真实 worker 测试)
runtime-broker 故障门禁(-Pfault-gates) 28/28:context 12、lost response 7、process crash 6、concurrency/storage 2、control 1
runtime-broker MySQL 集成,MySQL 8.4.7 2/2
managed-agent-server 单测 + MySQL 集成 + Checkstyle 138/138 与 10/10,0 个违规
HostedHarnessMySqlIT(MySQL 8.4.7)、HostedWorkspaceToolTurnIT(自带 H2) 1/1 与 1/1
runtime-broker Checkstyle 通过
cli src/serve(245 个文件) 11,054 通过;5 个文件里 37 个失败,这些文件本 PR 都没有改
core managed-tool* 189/189

这 5 个失败的 cli 文件是我本机环境的问题:ssh-workspace、session-pr-backfill、session-pr-refresh、aone-mrs 在 merge-base 构建上失败的是同样的 36 个用例,web-shell-pairing 重跑通过。CI 在本 head 上是绿的(23 个检查通过),机器人评审也批准了这个提交。

6. 变异矩阵

40 个单行变异体中,仓库单测杀掉 33 个。存活的 7 个都是「第二道」守卫,针对同一输入的第一道守卫在线路的另一端。

  • J4、J5、J13、J14:reserve 与 start 请求体的封闭字段集、控制操作里的未知字段、Broker 侧的文件历史 owner 检查。我把这四个变异编进同一个 server jar 重跑探针,B.7、B.9、C.2、C.3 在该 jar 上失败,说明真实栈能发现单测发现不了的问题。补四个 HTTP 层的小测试就能钉住。
  • T2、T8、T18:provider release 时的 closeSessionAdmission(该 Session 已被 providerSessions 围住)、worker 对「仍有进行中控制」的 release 拒绝(Broker 会先拒绝)、worker 对超限响应返回 413(transport 自己也有上限)。

J2 变异的是本 PR 之前就有的守卫,列出来是因为新的「禁止混用」规则依赖它。全量运行时 hosted-harness-session.test.ts 出现过 4 个失败,那是高负载下的抖动,不算击杀:应用同样的变异后单独跑该文件是 16/16 通过。

机器人评审提出的点,逐条实测

机器人评审是读代码得出的,并点名「缺一次端到端运行」是它唯一的缺口。本报告就是这次运行。它的其余各点与实测对照如下:

机器人的点 实测
1. worker 错误码把不同失败压成一个 成立,而且比它说的更宽:409 managed_runtime_identity_conflict 下出现了 10 种不同原因,并且 Broker 会丢掉原因文本(F2)
2. observePreparedCancellation 固定 25 ms 轮询、无退避 正常情况下到达 worker 的只有一次 cancel 和一次 status(D.4)。始终不结算的 worker 没有测
3. 已释放的 provider Session 存活整个 worker 代际 实测约每个已释放 Session 80 KiB(F3)。评审认为保留这些状态是为了释放后的观察,但实际做不到:释放后 history、status、cancel 在 Broker 处都返回 404 runtime_session_not_found,provider 在本地就拒绝,到达 worker 的请求为 0
4. 被取消的 shell 结果现在报告 cancelled 在真实 worker 上观察到:取消运行中的命令后结算为 cancelled(D.8)
悬而未决的问题:JsonCodec 写 null 是否改变已持久化的字节? 对 local-process provisioner 来说不会。merge-base 构建写入的 4 行和本构建写入的 44 行里,resource_handle_json 都是 {"provider":"local-process"}。provision seed 是六个非空字段。我还让本构建直接接管 merge-base 构建写过的库:Flyway 保持在第 15 级,两种 boot 的 provider 链路都通过,服务端日志没有 ERROR 行。static 与 Kubernetes provisioner 没有测

观察(不阻塞)

F1. worker 的明确拒绝被记成 UNKNOWN,此后 Session 永远无法释放。 worker 用 409 加原因拒绝 execute 时,Broker 记为 UNKNOWN,与响应丢失的处理完全一样。UNKNOWN 算作活跃工作,:resolve 返回 501,HTTP 面也没有 reconcile 路由,所以之后的 release 一直返回 409 runtime_session_busy。在 boot v2 下存储持有不会释放:同一 Workspace 的下一个 Runtime Session 和兄弟 Harness Session 都得到 409 workspace_busy。

这是继承来的,不是本 PR 引入的。merge-base 构建在 raw 路径上行为相同(Managed Runtime does not admit this tool. → UNKNOWN → release 被拒、存储被占)。本 PR 增加的是会落入这条规则的拒绝种类,我观察到三种新的:preflight has not permitted execution、invocation cancelled、Managed Runtime protocol conflicts(raw 调用进入 provider 持有的 Session,C.6)。也就是说「拒绝混合准入」确实成立且没有副作用,但代价是这个 Session 报废。可以考虑的方向:让 worker 给「派发前的拒绝」单独一个 code,transport 把它结算成 not_started,WorkspaceRuntimeTransport.execute 对自己的派发前检查已经是这样做的。直接规定「409 就是未执行」不安全:按我对该路由 catch 块的阅读,工具启动之后抛出的错误同样以 409 返回。

F2. 工具参数非法时,调用方拿不到原因。 write_file 传相对路径时,worker 返回 409 managed_runtime_identity_conflict 和 File path must be absolute: relative/path.txt;Broker 保留 code,但把消息换成 Managed Runtime control request failed (HTTP 409).;TypeScript 客户端只保留 status 和 code。未知工具名得到的也是同一个 code。在我的运行里,worker 用这一个 code 表达了 10 种不同原因。Harness 想把校验错误反馈给模型时,无法把它和身份冲突区分开。参数超限时 code 还会在途中改名:worker 说的是 400 managed_runtime_provider_invalid,调用方看到的是 400 managed_runtime_attestation_invalid。

F3. worker 保留了每一个已释放的 provider Session。 同一个 worker 上顺序 acquire、使用、release 1,500 个 Runtime Session:RSS 161 → 316 MiB,空闲 15 秒后 278 MiB,约每个已释放 Session 80 KiB。同一构建上用 raw 路径跑同样次数则保持平稳(152 → 146 MiB)。ManagedRuntimeProviderWorker.sessions 从不删除条目,条目里挂着该 Session 的 Config、工具集和文件历史。释放后 Broker 对 history、status、cancel 都返回 404 runtime_session_not_found,这些状态已经没有任何途径可以读到。按每轮一个 Runtime Session 计算,占用随对话长度增长。

说明 1. 审批由调用方负责,worker 不强制。 boot v1 下 prepare 报告 defaultPermission: "ask",但不调用 confirm、直接 prepare → preflight → start 也会写入文件(B.1b)。这符合 core 既有约定("the private parent calls this only after permission")。之所以提出来,是因为「boot v1 使用 DEFAULT approval」容易被理解成 worker 会强制审批。

说明 2. 参数上限是 256 KiB,不是 1 MiB 的信封上限。 255 KiB 的 write_file 能走完全部步骤(最大响应 511 KiB);256 KiB 在 prepare 处被 core 的 MAX_JSON_BYTES 拒绝。这与 raw 路径的 payload 上限相同,两条路径一致。

合并顺序。 与当前 main(3f5ae3ffeb)、#12855、#12870 试合并无冲突。与 #12848 在 managed-context-worker.ts、RuntimeBrokerHttpServer.java、RuntimeBrokerService.java 冲突;与 W0e 草稿 #12839、#12865、#12869 在 RuntimeBrokerService.java 和 WorkspaceRuntimeTest.java 冲突。后合入的一方需要手工解决并重跑故障门禁。

未覆盖

Linux 与 Windows;MariaDB(CI 已覆盖);带着存活 provider Session 重启 Broker、worker 重启后的已准备状态恢复(PR 已列为范围外);由模型驱动 provider 路径的产品流程(该路径目前还没有消费方),所以 provider 探针停留在 provider API 这一层。

复现

装置、探针、变异体清单和原始日志:pr12868/harness 与 pr12868/results。顺序:带代理启动 spring.sh,然后依次是 s4-ab、s1-contract、s2-faults、s5-raw、s9-author-driver-mysql、s6-hosted-real-model、s3-retention、s8-limits、mutate。

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head f67bde6b

Not approving, on coverage: this is a cross-language protocol change with roughly 1,500 lines of production code on two sides of a wire — a TypeScript protocol module (428) and worker (475), a Java ProviderRuntimeProtocol (188), HttpRuntimeTransport (+134/-15), RuntimeBrokerService (+136/-47), the tool executor (+58/-8) and the provider (+30/-11). I verified one layer of it myself and could not reach the rest inside my budget, and the layer I did not reach is the one holding the lease and release ordering.

No historical blocker, and the failed check is not one

There are no inline comments and therefore no review threads, and the only review is a bot approval at this head, so nothing has ever been filed against this PR. Triage stage 2 at this exact head reports "No critical blockers", and the maintainer's real-stack verification at this head — MySQL 8.4.7 rather than H2, workers from the production local-process provisioner rather than a static lease, and a live model on the Hosted path — reports "ready to merge from the real-stack side. No blocking finding."

review-pr shows fail on this head. That lane is not a gate in any state, and nothing else has failed: 24 checks pass, including the full Java matrix, Runtime Broker and Managed Agent MariaDB, Hosted no-tool processes / MySQL 8.4, Real daemon E2E / Java 11, Test (ubuntu-latest, Node 22.x), Lint & Static and the integration lanes. I am not treating the failure as a finding, and I am not treating the green board as coverage of the code below either.

What I verified in the code myself

The Java side of the wire contract is closed-world, and closed in the way that matters:

private static final Set<String> REFERENCE_FIELDS = Set.of("sessionId",
        "promptId", "callId", "capabilityDigest", "policyRevision",
        "invocationId", "argsDigest");

static boolean isReference(Map<String, Object> reference) {
    return reference != null && reference.keySet().equals(REFERENCE_FIELDS);
}

keySet().equals(...) is an exact match, not a subset test, so the claim that a provider reference carries seven identity fields and no tool name or arguments is enforced by the shape check rather than by convention — an extra key is a rejection, and so is a missing one. control() dispatches on kind with an explicit required set and an explicit optional set per operation, rejects any key in neither, and rejects an unknown kind; confirm restricts outcome to a fixed nine-value enumeration and phase to permission or preflight. bind-history does more than check shape — it requires binding.ownerRuntimeSessionId to equal the caller's sessionId and binding.ownerSessionId to equal the harnessSessionId, so a worker cannot bind another session's history through this route. Control and history payloads are bounded at 1 MiB and 8 MiB.

In RuntimeBrokerService, references now pass through immutableMap(...) before use, the provider path validates through ProviderRuntimeProtocol.reference(...) at reserve time and again against the stored record at start time, and the pre-existing error taxonomy is intact — runtime_execution_conflict for a start against something not reserved for deferred dispatch, the 256 KiB runtime_payload_invalid bound, runtime_idempotency_conflict when a payload differs from its reserved digest, and a new explicit "Deferred execution requires payloadJson". Re-validating the stored reference at start is what makes the description's claim structural: a raw four-field reservation cannot satisfy the seven-field check, so a retry cannot switch protocols.

What I did not reach

RuntimeBrokerService's release path — the claim that release requires the worker to close admission before Workspace activation and storage ownership are released is the one with data-corruption potential, and I did not read it. HttpRuntimeTransport's 134 added lines. The whole TypeScript side: managed-runtime-provider-protocol.ts, managed-runtime-provider-worker.ts, the executor's +58/-8, the context and attestation workers, and managed-tool-runtime.ts. And the ambiguity discipline the test plan names — that a lost execution response stays UNKNOWN without replay — which I saw reflected in the cancel/status error mapping but did not trace. The 561-line fixture corpus and its TS test are the mechanism that would catch a two-sided divergence, and both sides' lanes are green, but I did not read them.

What closes this

A pass over the release and admission ordering in RuntimeBrokerService, the transport's new control route, and the TypeScript protocol/worker pair against the Java validator — specifically that the two sides agree on the seven-field reference, the operation kinds and the confirm outcome enumeration, and that a lost start or control response stays UNKNOWN rather than being replayed.

Verdict: COMMENT — No Critical found, and none has ever been filed here. The Java wire contract verifies as closed-world with an exact seven-field reference, an enumerated confirm outcome, an ownership cross-check on history binding and bounded payloads, and the stored reference is re-validated at start so protocols cannot be switched on retry. What blocks an approval is that the release/admission ordering, the transport and the entire TypeScript half of the protocol are outside what I read in budget, and for a two-sided contract I am not willing to approve one side on the strength of the other.

@wenshao
wenshao dismissed a stale review via f0f217d September 28, 2026 01:49
@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the code review and real-stack findings in f0f217dfa52fe00dea669b9c9ffdad05c2efb832.

F2 is fixed. Unknown tools and invalid tool construction now return 400 managed_runtime_tool_invalid; unsupported provider profiles return 501 managed_runtime_provider_unsupported. Known provider errors retain their status, code and bounded reason through Java and TypeScript. Java only accepts the matching status/code, closed JSON body and expected response headers; malformed or unrecognized responses keep the generic error. The TypeScript client only displays a bounded reason from the Broker's closed three-field JSON error envelope. These changes preserve execution uncertainty and replay rules.

E2E report: the final rebuilt node dist/cli.js worker, production Java Broker, built TypeScript client and H2 journal passed 39 existing check groups plus 8 targeted groups (including two startup checks), without mocked control/error responses. Relative WriteFile paths and unknown tools reach the caller with the original reason and tool_invalid; oversized arguments retain provider_invalid instead of becoming an attestation error; a real unsupported boot-v2 profile retains 501 provider_unsupported. Correcting the arguments in the same Session then prepares and executes successfully. A deliberate execute without preflight still becomes UNKNOWN: a repeated start is rejected, exactly one execute reaches the worker, and no file is written.

Mutation survivors are covered. Four HTTP regression tests now reject extra reserve/start fields, unknown control fields and foreign history owners before any transport effect or reservation mutation. In isolated source copies, removing each J4/J5/J13/J14 guard makes its corresponding test fail with HTTP 200 instead of 400; each run has one assertion failure and zero errors. The intact HTTP suite passes all 11 tests.

Other comments are recorded explicitly:

  • F1 — deferred: HTTP execute rejection conservatively becoming UNKNOWN is inherited behavior. Error text alone cannot prove dispatch never occurred. Safe pre-dispatch evidence and settlement need a separate change; release can still return runtime_session_busy and Workspace ownership can remain held. The new E2E deliberately confirms this limitation rather than treating it as fixed.
  • F3 — deferred: retained heavyweight released-Session state is real. Independent reproduction retained about 100.70 MiB of post-GC heap after 1500 Sessions; idle/GC did not remove it. However, the private worker still serves current-turn status/cancel and bound history after release, while Broker HTTP rejects them. Immediate deletion/disposal would remove existing evidence. A follow-up needs lighter retained evidence and an explicit retention policy; this revision does not claim bounded memory.
  • Approval and size limits — clarified in both design languages: boot-v1 DEFAULT approval relies on the caller to enforce confirmation; arguments retain the existing 256 KiB canonical-JSON bound independently of the 1 MiB control envelope / 8 MiB history envelope.
  • 25 ms polling and a sibling JSON schema — deferred: bounded backoff and schema consistency are useful follow-ups. Current cancellation observation is serial and bounded by the operation lease, and both languages already enforce the closed contract through parsers/shared fixtures. No correctness fix was established for these suggestions.
  • JsonCodec persistence — no confirmed regression: current provision seeds use six non-null fields; local-process handles use non-null provider data. The reported MySQL old-database adoption and comparison support that existing path. Static provisioning does not persist a handle here, and this repository has no Kubernetes provisioner implementation. This is not a claim about arbitrary external handle extensions.
  • Cancelled Shell / consumers: cancellation remains an explicit cancelled outcome with failure-hook behavior, covered through the private provider runtime. This does not enable the separately gated Hosted product path.

Final validation: full build, typecheck and bundle; 136 core + 892 CLI tests; all 390 Broker tests with zero failures and one existing optional bundle-provisioner test skipped; changed-file ESLint/Prettier, Java Checkstyle and diff checks. The initial legacy interop run collided with the TypeScript rebuild before worker readiness; the post-build full Java suite passed that exact test. Existing feature-stage Hosted integration evidence remains in the earlier E2E report; it was not rerun for this diagnostic-only revision.

Before this commit, three fresh full open-ended and reverse audit rounds completed. Round 1 tightened the TypeScript error envelope and reset the clean count; rounds 2 and 3 independently found no new actionable issues on the unchanged patch. The audits traced the lifecycle/ownership, control/execute separation, error consumers and test oracles; no unresolved Critical was identified. Deferred suggestions above remain visible.

The reviewed and E2E-verified patch against 1b179d35a3 has SHA-256 f065e042080ba0e13884fbac30c02304b9839781ce547268540340e89e91fcef. No deferred item above is silently counted as fixed.

wenshao pushed a commit to wenshao/qwen-code that referenced this pull request Sep 28, 2026
@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification, round 2: f0f217dfa5

Verdict: ready to merge. No blocking finding. The fix commit does what its description says: F2 is fixed on the real stack, the four Java mutation survivors are pinned by tests, and nothing else moved. One residual of the F2 fix is worth a small follow-up (R1 below); a candidate patch with a test is attached. It does not block.

Round 1 was at f67bde6b41. The head now adds a merge of main and f0f217dfa5. Same rig as round 1: built TypeScript provider → Spring-embedded Broker on MySQL 8.4.7 → bundled workers, with the JVM HTTP proxy between Broker and worker.

What was run Result
F2: 10 refused controls, three hops, previous head vs this head previous head: reason kept 0/10, one code renamed · this head: status, code and reason kept 10/10
Reason bound (R1) over 4,096 characters or with NUL the code is lost again; the candidate keeps it in all 11 cases
12 worker answers on the production transport, 11 of them rewritten on the wire the untouched one and the 2 conforming rewrites are forwarded; the other 9 fall back to the generic answer
A well-formed refusal substituted on execute record UNKNOWN, 1 dispatch, 1 effect, no replay
Round-1 probes rerun 76/76 and 15/15
New: lost control replies, enumeration agreement 13/13
Hosted Workspace loop PR's own driver on MySQL: HOSTED_WORKSPACE_TOOLS_OK · live model, 3 turns: 5 executions SETTLED/success
Repository suites, this head and trial merge with current main all green on both; CI on this head has 23 passed checks
Mutation matrix, 61 mutants 56 killed, 5 survive (round 1: 33 of 40)

What the head contains

The merge commit 1b179d35a3 has the same tree as a mechanical merge of f67bde6b41 with 0a136f89d7 (b8f0272e), so nothing was changed by hand in it. It brings in #12863 and #12864, both test-only. f0f217dfa5 touches 12 files, 447 lines added.

1. F2 is fixed

Error envelope before and after

After a refusal the Session stays usable. Two refused preparations, then corrected arguments under the same call identity: prepare, reserve, approve, preflight and start succeed, one row, one effect, release works (Q.1–Q.3 on both boots).

2. R1: the code is lost again when the reason is long

Reason bound

The Broker forwards a provider error only when the reason has at most 4,096 characters and no NUL. That bound is right. The gap is on the sending side: validation messages echo the caller's input, and the worker sends them at any length. A model-written shell script of 5,969 characters that trips the sleep guard gives a reason of 6,242 characters. The Broker then drops the code together with the text, and the caller sees 400 managed_runtime_attestation_invalid with no reason, which is the round-1 behaviour. A path with a NUL does the same.

Candidate: bound the reason in the worker and keep the code. 33 changed lines in managed-runtime-provider-worker.ts plus one test; Java is unchanged. On the real stack it keeps managed_runtime_tool_invalid in all 11 cases. The worker test file passes 11/11 with it, and ESLint, Prettier and the cli typecheck are clean.

Candidate patch (production part)
+const PROVIDER_REASON_LIMIT = 4096;
+const PROVIDER_REASON_OMISSION = ' … [reason truncated]';
+
+/**
+ * The Broker forwards a reason only when it fits its bound and holds no NUL;
+ * otherwise it drops the code together with the text. A reason that echoes
+ * caller input can exceed the bound, so bound it here and keep the code.
+ */
+function providerReason(error: unknown): string {
+  const text =
+    error instanceof Error && error.message
+      ? error.message
+      : 'Managed Runtime provider operation failed.';
+  const clean = text.replaceAll('\0', '�');
+  if (clean.length <= PROVIDER_REASON_LIMIT) return clean;
+  let end = PROVIDER_REASON_LIMIT - PROVIDER_REASON_OMISSION.length;
+  const last = clean.charCodeAt(end - 1);
+  // Never split a surrogate pair.
+  if (last >= 0xd800 && last <= 0xdbff) end--;
+  return clean.slice(0, end) + PROVIDER_REASON_OMISSION;
+}

The four error: error.message sites in the route's catch block become error: providerReason(error). Full patch with the test: candidate-bounded-reason.patch.

3. The Broker trusts only a closed, matching error

Tampered answers

This is the description's claim "Java only accepts the matching status/code, closed JSON body and expected response headers", checked on the production transport with a real worker behind it. The last block matters most: a well-formed refusal in place of the answer of an execute that really ran settles nothing.

4. Regression and new probes

Regression

Two things here are new and answer the review by qqqys, which asked for the parts it could not reach:

Asked for Measured
Release and admission ordering Round 1 A.12: control[release] precedes /v3/activation on the wire. New K.5: when the release acknowledgement is lost, the Session is RELEASING, storage is still held, a control in between gets 409 runtime_session_not_ready, and a second release completes and drops the storage
The transport's control route Figures 1 and 3
Both sides agree on the seven-field reference, the operation kinds and the confirm outcomes L.1–L.3: core has 10 outcomes, both sides accept the same 9 and refuse restore_previous; values outside the enumeration are refused before any worker request. Only manifest and history are admitted bare; transport-only and unknown kinds never reach the worker. Reference shape: round 1 A.5 and the stored row
A lost start or control response is never replayed Lost execute reply: UNKNOWN, one dispatch (F.1/F.2). Lost control reply (K.1–K.4): the caller gets 503 managed_runtime_unavailable; the same request again succeeds for bind-history, manifest, begin-turn, prepare, confirm and preflight, and the invocation then executes once

The author did not rerun the Hosted integration for this revision, so both Hosted runs in the figure are on this head.

5. Mutation matrix

Mutation matrix

56 of 61 mutants are killed by the repository's unit suites. Round 1 was 33 of 40.

  • The four Java survivors of round 1 (J4, J5, J13, J14) are killed by the four new HTTP tests.
  • 21 new mutants on the lines the fix adds: 19 killed. The two survivors, N17 and N18, are both about which code a 409 carries. Removing the ManagedToolConflictError branch, or turning the catch-all back into managed_runtime_identity_conflict, fails no test. No test asserts the code of a 409 refusal, which is the same area as the label note below.
  • T2, T8 and T18 survive as in round 1. They are second guards behind a first guard on the other side of the wire.
  • The other 33 round-1 mutants were rerun on this head and are all killed.

Failures in hosted-harness-session.test.ts appeared in 15 of the full TypeScript runs while the host was under load. They are not counted as kills: with T2, T14 and N19 applied, that file passes 16/16 on its own.

Status of the round-1 items

Item Status on this head
F1 definite refusal → UNKNOWN → Session and storage pinned Deferred by the author. Remeasured, unchanged (E.1/E.2): release 409 runtime_session_busy, next Runtime Session and sibling Harness Session 409 workspace_busy. The design document now states this
F2 reason and code lost Fixed, with the residual R1
F3 worker keeps released Sessions Deferred by the author, who reproduced about 100 MiB of retained heap after 1,500 Sessions. After release the Broker still answers 404 for history, status and cancel, and no request reaches the worker. The design document now states the retention
Approval is the caller's duty; 256 KiB argument bound Both written into the design document in both languages. Remeasured: 255 KiB passes, 256 KiB is refused, now as managed_runtime_provider_invalid
Four Java mutation survivors Killed by the four new HTTP tests

Two small notes on the new codes. Neither blocks.

  • Labels are partly swapped. A tampered argsDigest, a changed decision and changed arguments are real identity mismatches, and they now carry the catch-all managed_runtime_provider_operation_failed. A background shell and a media context on the wrong tool are input problems, and they still carry managed_runtime_identity_conflict. The reason text makes each one readable.
  • The reason is forwarded verbatim, including newlines and ESC sequences (R5). A consumer that prints or logs it should treat it as untrusted text.

Merge order

A trial merge with current main (c767078ff7) is clean, and on that merged tree the Broker unit suite (390), the fault gates (28), the server suites and the Hosted ITs pass, including the Hosted reply-loss gate that #12873 added. #12855 and #12870 merge cleanly on top. #12848 conflicts in 4 files; #12839, #12865 and #12869 conflict in 3 files each (broker-managed-runtime-provider.ts, RuntimeBrokerService.java, WorkspaceRuntimeTest.java).

Not covered

Linux and Windows; MariaDB; Broker restart with live provider Sessions; a worker that never settles a prepared cancellation; the unsupported boot-v2 profile on the real stack, which the Spring server cannot produce (its 501 was checked on the transport by substitution, and the worker branch by the unit test).

Reproduce

Round-2 rig, probes, candidate and logs: pr12868/r2. New probes are s11-error-envelope and s12-lost-control; proxy.mjs gained a rewrite fault.

中文版

真实环境验证,第二轮:f0f217dfa5

结论:可以合并,没有阻塞问题。 修复提交做到了描述里说的事:F2 在真实栈上已修复,四个 Java 变异存活项已被测试钉住,其余行为没有变化。F2 的修复留有一个残留(下文 R1),值得一个小的后续改动,附了带测试的候选补丁,不阻塞合并。

第一轮验证的是 f67bde6b41。现在的 head 多了一次 main 合并和 f0f217dfa5。装置与第一轮相同:构建出来的 TypeScript provider → Spring 内嵌 Broker(MySQL 8.4.7)→ 打包 worker,Broker 与 worker 之间挂着 JVM HTTP 代理。

跑了什么 结果
F2:10 个被拒绝的控制,三跳,上一个 head 对比本 head 上一个 head:原因保留 0/10,一个错误码被改名 · 本 head:status、code、原因全部保留 10/10
原因长度上限(R1) 超过 4,096 字符或含 NUL 时错误码再次丢失;候选补丁在 11 个用例里全部保住
生产 transport 上的 12 个 worker 响应,其中 11 个在线路上被改写 未改动的那个和 2 个合规的改写被转发;其余 9 个回退成通用响应
在 execute 上替换成一个格式正确的拒绝 记录为 UNKNOWN,派发 1 次,副作用 1 次,不重放
重跑第一轮探针 76/76 与 15/15
新增:控制响应丢失、枚举一致性 13/13
Hosted Workspace 循环 PR 自带 driver 接 MySQL:HOSTED_WORKSPACE_TOOLS_OK · 真实模型 3 轮:5 次执行 SETTLED/success
仓库套件:本 head,以及与当前 main 的试合并树 两者全绿;本 head 的 CI 有 23 项检查通过
变异矩阵,61 个变异体 杀掉 56 个,存活 5 个(第一轮为 40 个中的 33 个)

head 里有什么

合并提交 1b179d35a3 的树与 f67bde6b41、0a136f89d7 的机械合并结果(b8f0272e)完全相同,说明合并里没有手工改动。它带入了 #12863 和 #12864,都只改测试。f0f217dfa5 改了 12 个文件,新增 447 行。

1. F2 已修复

见图 1。拒绝之后 Session 仍然可用:两次被拒绝的 prepare 之后,用同一个调用身份换上正确参数,prepare、预留、审批、preflight、start 全部成功,一条记录、一次副作用,release 正常(Q.1–Q.3,两种 boot 都是)。

2. R1:原因过长时错误码再次丢失

见图 2。Broker 只在原因不超过 4,096 字符且不含 NUL 时才转发 provider 错误,这个上限本身是对的。缺口在发送端:校验消息会回显调用方的输入,而 worker 不限长度地发出去。模型写的一段 5,969 字符的 shell 脚本触发 sleep 守卫后,原因有 6,242 字符。Broker 于是把错误码连同文本一起丢掉,调用方看到的是不带原因的 400 managed_runtime_attestation_invalid,也就是第一轮的行为。路径里带 NUL 也一样。

候选做法:在 worker 端给原因设上限并保留错误码。managed-runtime-provider-worker.ts 改动 33 行,外加一个测试,Java 不用动。在真实栈上 11 个用例全部保住 managed_runtime_tool_invalid。加上补丁后 worker 测试文件 11/11 通过,ESLint、Prettier 和 cli 的 typecheck 都干净。补丁正文见上面英文部分的折叠块,完整补丁见 candidate-bounded-reason.patch。

3. Broker 只信任封闭且匹配的错误

见图 3。这对应描述里的声明「Java 只接受匹配的 status/code、封闭 JSON 体和预期响应头」,这次是在生产 transport 上、后面接着真实 worker 验证的。最后一块最重要:对一次确实执行过的 execute,把响应替换成格式正确的拒绝,不会结算任何东西。

4. 回归与新增探针

见图 4。其中两项是新增的,用来回应 qqqys 的评审里点名没读到的部分:

评审要求 实测
释放与准入的顺序 第一轮 A.12:线上 control[release] 先于 /v3/activation。新增 K.5:release 确认丢失时,Session 处于 RELEASING,存储仍被持有,其间的控制请求得到 409 runtime_session_not_ready,第二次 release 完成并释放存储
transport 的控制路由 图 1 和图 3
两端在七字段 reference、操作种类、confirm 结果枚举上一致 L.1–L.3:core 有 10 个结果值,两端接受相同的 9 个并都拒绝 restore_previous;枚举外的值在到达 worker 之前就被拒绝。不带字段时只有 manifest 和 history 被接受;仅限 transport 的种类和未知种类都到不了 worker。reference 形状见第一轮 A.5 和落库记录
start 或控制响应丢失后不重放 execute 响应丢失:UNKNOWN,派发一次(F.1/F.2)。控制响应丢失(K.1–K.4):调用方得到 503 managed_runtime_unavailable;同一请求重发后,bind-history、manifest、begin-turn、prepare、confirm、preflight 都成功,随后该调用只执行一次

作者这一版没有重跑 Hosted 集成,所以图里的两次 Hosted 运行都是在本 head 上做的。

5. 变异矩阵

61 个变异体中,仓库单测杀掉 56 个。第一轮是 40 个中的 33 个。

  • 第一轮存活的四个 Java 变异体(J4、J5、J13、J14)被新增的四个 HTTP 测试杀掉。
  • 针对修复新增代码的 21 个新变异体:杀掉 19 个。存活的 N17 和 N18 都与「409 带哪个错误码」有关:去掉 ManagedToolConflictError 分支,或者把兜底错误码改回 managed_runtime_identity_conflict,都没有任何测试失败。没有测试断言 409 拒绝的错误码,这与下文「标签」那条说明是同一块区域。
  • T2、T8、T18 与第一轮一样存活。它们是第二道守卫,第一道在线路另一端。
  • 第一轮其余 33 个变异体在本 head 上重跑,全部被杀。

宿主高负载期间,全量 TypeScript 运行里有 15 次出现了 hosted-harness-session.test.ts 的失败,这些不计入击杀:应用 T2、T14、N19 后单独跑该文件都是 16/16 通过。

第一轮各项的现状

条目 在本 head 上的状态
F1 明确拒绝 → UNKNOWN → Session 和存储被钉住 作者延后处理。复测无变化(E.1/E.2):release 返回 409 runtime_session_busy,下一个 Runtime Session 和兄弟 Harness Session 返回 409 workspace_busy。设计文档现在写明了这一点
F2 原因和错误码丢失 已修复,留有残留 R1
F3 worker 保留已释放的 Session 作者延后处理,并自行复现了 1,500 个 Session 后约 100 MiB 的常驻堆。释放后 Broker 对 history、status、cancel 仍返回 404,没有请求到达 worker。设计文档现在写明了这种保留
审批由调用方负责;参数上限 256 KiB 两点都已写入两种语言的设计文档。复测:255 KiB 通过,256 KiB 被拒绝,现在的错误码是 managed_runtime_provider_invalid
四个 Java 变异存活项 被新增的四个 HTTP 测试杀掉

关于新错误码有两点小说明,都不阻塞。

  • 标签有一部分是反的。 被篡改的 argsDigest、改变的审批决定、改变的参数是真正的身份不匹配,现在带的是兜底的 managed_runtime_provider_operation_failed;后台 shell 和用错工具的 media context 属于输入问题,带的仍是 managed_runtime_identity_conflict。好在原因文本能让每一种都读得明白。
  • 原因是原样转发的,包括换行和 ESC 序列(R5)。打印或记录它的消费方应当把它当作不可信文本。

合并顺序

与当前 main(c767078ff7)试合并无冲突;在合并后的树上,Broker 单测(390)、故障门禁(28)、server 套件和 Hosted IT 全部通过,其中包括 #12873 新增的 Hosted 响应丢失门禁。#12855 和 #12870 叠上去也无冲突。#12848 有 4 个文件冲突;#12839、#12865、#12869 各有 3 个文件冲突(broker-managed-runtime-provider.ts、RuntimeBrokerService.java、WorkspaceRuntimeTest.java)。

未覆盖

Linux 与 Windows;MariaDB;带着存活 provider Session 重启 Broker;始终不结算已准备取消的 worker;真实栈上的「不支持的 boot-v2 profile」,Spring 服务端造不出这种情况(它的 501 是在 transport 上用替换响应验证的,worker 分支由单测覆盖)。

复现

第二轮的装置、探针、候选补丁和日志:pr12868/r2。新增探针是 s11-error-envelope 和 s12-lost-control;proxy.mjs 增加了 rewrite 故障。

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@wenshao
wenshao dismissed a stale review via ac0b6cf September 28, 2026 05:32
@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Resolved the three conflicts with main d66fdadd2794 in merge commit ac0b6cf966.

The merge preserves provider/raw protocol separation, bounded error reasons, main’s terminal/runtime-loss receipts, and Workspace loss fences. Repeated cancellation of a prepared provider invocation still waits for worker cleanup rather than acknowledging only the Broker’s settled record. Both English and Chinese designs and the PR description now distinguish terminal receipts from live Session/history controls.

Validation on the exact committed tree:

  • Build, typecheck, bundle, ESLint/Prettier and Java Checkstyle passed.
  • CLI: 115 tests; Broker: 402 tests with no skips; Workspace/schema: 17 tests; packaged Hosted Workspace tool loop: 1 integration test, all passed.
  • Real built TS provider → production Java Broker → bundled worker: 40 provider groups and 8 error/recovery groups passed (four readiness groups included). The added released-receipt probe verifies unchanged Session state/version, persisted execution and file effects, with zero additional worker requests. It uses H2 execution storage with static placement and in-memory Session metadata; the added real-chain probe covers SETTLED, while service/recovery tests cover ABANDONED.
  • Two consecutive full conflict-resolution audits, each open-ended and reverse-checked, found no new Critical issues. The committed tree exactly matches the audited/tested tree; source and build fingerprints were unchanged during E2E.

Per the requested Critical-only policy after round 5, the non-blocking R1 oversized/NUL-reason suggestion in the latest real-stack review, along with its error-label, terminal-escaping and remaining mutation-coverage suggestions, is deferred to follow-up work. Earlier F1/F3 limitations remain documented and are not claimed fixed.

已解决 3 处冲突并合入 main;连续两轮无方向审计与反向审计未发现新增 Critical。上述本地构建、测试和真实链路验证全部通过;非阻塞建议按第 5 轮后仅修复 Critical 的要求保留为后续工作。

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

wenshao pushed a commit to wenshao/qwen-code that referenced this pull request Sep 28, 2026
@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification, round 3: ac0b6cf966

Verdict: ready to merge. No blocking finding in this PR. The head adds one commit, the merge of main after #12839, and that merge was resolved by hand. Both seams of the resolution behave as the merge note says, and nothing from rounds 1 and 2 moved.

Two things to know before merging:

Same rig as before: built TypeScript provider → Spring-embedded Broker on MySQL 8.4.7 → bundled workers, with the JVM HTTP proxy between Broker and worker.

What was run Result
Anatomy of the merge commit not mechanical: 3 conflicted files, plus 2 design documents and 1 test file edited by hand; no production line outside the conflicts
Seam 1: terminal receipts after release, repeated cancellation of a prepared call (M1, M2) 10/10. Previous head: everything after release 404. This head: GET, cancel and the same reservation answer 200 from the record, 0 worker requests
The hand-written guard (M4, M5), this head vs a jar with mutant G4 this head 12/12; with the mutant a finished call answers 503 to a later cancellation
Seam 2: worker killed with one invocation reserved this head answers exactly as main does, line for line, on the raw and on the provider path
Rounds 1 and 2 probes rerun 76/76, 12/12, 7/7 (20 of 20 refusals keep code and reason), 13/13
Hosted Workspace loop PR's own driver on MySQL: HOSTED_WORKSPACE_TOOLS_OK, 7 executions SETTLED/success · live model, 3 turns: 5 executions SETTLED/success
Repository suites on this head Broker 402, fault gates 29, Broker MySQL IT 2, server 140 + 10, Hosted ITs 4, core 190: all pass
Trial merge with current main textually clean; Broker 402, fault gates 29 and core 190 pass; server and Hosted suites cannot start (two V16 on main)
Trial merge with main + #12900 clean; Broker 407, server 154 + 11, Hosted 154 + 5 and core 190 pass; the jar starts; on the real stack A–D 76/76 and M1, M2, M4, M5 22/22
Mutation matrix, 70 mutants 64 killed, 6 survive (round 2: 56 of 61)

What the head contains (figure 1)

Anatomy of the merge commit

The round-2 merge had the same tree as a mechanical merge. This one does not, so I read it with git show --remerge-diff, which prints only what the recorded merge changes against a mechanical one. The production decisions are all in RuntimeBrokerService.java:

  • Reservations (createExecution, prepareExecution) now go through main's receipt lookup, with this PR's validated reference.
  • cancelExecution answers a terminal record from storage, except a prepared provider call whose Session is not yet released, which still goes to the worker.
  • That exception is decided by a new helper, isPreparedProviderCancellation(): settled, cancel requested, dispatch generation 0, provider reference.

1. Seam 1: terminal receipts reach the provider path (figure 2)

Receipts and repeated cancellation

Claim in the merge note Measured here
main's terminal receipts are preserved After release, GET, cancel and a repeated reservation with the same key and reference return 200 with the same execution id. 0 worker requests, Session row unchanged (RELEASED, same version)
Receipts stay bound Same key with another invocation: 409 runtime_idempotency_conflict. Another Harness Session: 409 runtime_execution_conflict
Live Session and history controls stay distinct from receipts After release a new reservation, start and control[history] are still 404
Repeated cancellation of a prepared provider call still waits for the worker Before release every cancellation reaches the worker: 1st cancel,status, 2nd cancel, three concurrent ones all 200 cancelled, no effect. After release the cancellation is answered from the record
The released-receipt probe ran on H2 with static placement and in-memory Session metadata Same result here on MySQL 8.4.7 with the Spring server's placement and JDBC Session metadata, boot v1 and boot v2
The real chain covers SETTLED; ABANDONED is covered by service tests Same limit here: this rig did not produce an ABANDONED record. The client side was checked against a stand-in peer (N.1): a terminal runtime_lost answer becomes unknown / terminal, a plain one stays unknown

2. The guard written during the merge (figure 6)

Dispatch generation guard

isPreparedProviderCancellation() has four conditions. Two mutants each remove one. Removing the provider-reference test (G5) is killed by the repository's tests. Removing getDispatchGeneration() == 0 (G4) is not: all 402 Broker tests pass with it.

So I compiled G4 into a server jar and ran two probes on both jars:

  • M4, a call that was dispatched and then cancelled. This head answers a repeated cancellation from the record. The mutant asks the worker again each time. The answer is the same, 200 cancelled.
  • M5, a call that finished before the worker heard the cancellation (the cancel request is delayed 5 s on the wire; the command takes 2 s). The record is SETTLED/success with cancel requested. This head answers a later cancellation 200 success. The mutant answers 503 runtime_execution_cancel_failed, every time, because the worker reports success and the prepared-cancellation path accepts only cancelled or not started.

The guard is right. A service test for the M5 case (settled, success, cancel requested, generation 1, cancel again → stored record, transport not called) would pin it.

3. Seam 2: worker death under main's Runtime-loss fences (figure 3)

Worker death

The worker is killed with SIGKILL while one invocation is reserved. On every build nothing executes and nothing is replayed.

Previous head f0f217dfa5 main d66fdadd27 (raw path) This head, raw and provider path
Placement FAILED LOST LOST
Replacement worker started, generation 2 none none
GET of the execution 409 runtime_broker_execution_unknown 409 runtime_admission_closed 409 runtime_admission_closed
Release 409 runtime_session_busy 503 runtime_reconciliation_required 503 runtime_reconciliation_required
Next Runtime Session 409 workspace_busy 503 runtime_broker_runtime_lost 503 runtime_broker_runtime_lost
Sibling Harness Session 409 workspace_busy 409 runtime_placement_recovery_required 409 runtime_placement_recovery_required

The change is inherited from #12839. The merge neither weakened nor extended it. For the rig this meant one fresh database and Broker per worker-death scenario, because a lost placement refuses later placements.

4. Regression (figure 4)

Regression

The cli src/serve failures in the figure are local. CI on this head is green: 21 checks pass, 4 are skipped, none fail. That includes Test (ubuntu-latest, Node 22.x).

5. Mutation matrix (figure 5)

Mutation matrix

Nine new mutants (G1–G9) edit the lines written or relied on while resolving the merge: 8 are killed, G4 survives (section 2). The 61 mutants of rounds 1 and 2 were rerun on the merged code: 56 killed, and the same five survive (T2, T8, T18, N17, N18). No verdict changed, so the merge lost none of the behaviour the tests pin.

Kills are attributed per test from the JSON reports. Failures of hosted-harness-session.test.ts in full runs under load are not counted: N17 had only such failures, and with N17 applied that file passes 16/16 when run alone.

6. Current main carries two V16 migrations (figure 7)

Two V16 migrations on main

Found while trial-merging this head with current main. #12839 merged at 05:14Z with V16__runtime_loss_evidence.sql. This head's merge commit was made at 05:32Z. #12881 merged at 05:37Z with V16__managed_session_operation.sql. git sees two different file names and reports nothing; Flyway refuses to start.

#12900 moves the later migration to V17. A trial merge of main (fc4e01b9fc) + #12900 + this head has no conflict. On that tree the Broker unit suite (407), the server suite (154 and 11 ITs), the Hosted suite (154 and 5 ITs) and core (190) pass. Its server jar starts on a fresh MySQL database and applies versions 15, 16 and 17 in order. On that jar the reviewer test plan passes 76/76 and the merge probes 22/22.

Status of earlier items

Item Status on this head
F1 definite refusal → UNKNOWN → Session and storage pinned Deferred by the author. Remeasured, unchanged (E.1/E.2): release 409 runtime_session_busy; :resolve is still 501, and reconcileExecution still has no HTTP route
F3 worker keeps released Sessions Deferred by the author. Not remeasured; the worker code did not change in this head
R1 reason over 4,096 characters or with NUL loses the code Deferred by the author. Remeasured, unchanged: 4 of 11 long or NUL reasons lose the code. The round-2 candidate patch still applies
Swapped error labels, reason forwarded verbatim Deferred by the author. Unchanged (mutants N17 and N18 still survive)
256 KiB argument bound Unchanged: 255 KiB passes, 256 KiB is refused with managed_runtime_provider_invalid

One small interaction between the two sides of the merge. It does not block. The client attaches a reason only when the error body has exactly three keys. main added an optional fourth key, details. An error with details therefore keeps its code and never its reason (N.2, stand-in peer). Today the only such answer is the ABANDONED one, which the client translates to unknown / terminal, so no caller sees a difference yet.

Merge order

With Result of a trial merge with this head
current main (fc4e01b9fc) clean; needs #12900 to start
#12900, #12855, #12896 clean
#12848 4 files conflict: managed-context-worker.ts, RuntimeBrokerHttpServer.java, RuntimeBrokerService.java, RuntimeBrokerHttpServerTest.java
#12894 6 files conflict: the four above, WorkspaceRuntimeTest.java, JsonCodec.java
drafts #12865, #12869 5 and 14 files conflict

Not covered

Linux and Windows; MariaDB; an ABANDONED record on the real stack, which needs the recovery flow of #12839; Broker restart with live provider Sessions; a worker that never settles a prepared cancellation; F3 memory growth was not remeasured.

Reproduce

Round-3 rig, probes, mutants and logs: pr12868/r3. New probes are s13-merge.mjs (M1–M5), kill-scenario.sh and s14-details-envelope.mjs. remerge-r3.diff is the output of git show --remerge-diff.

中文版

真实环境验证,第三轮:ac0b6cf966

结论:可以合并,本 PR 没有阻塞问题。 head 只多了一个提交,即合入 #12839 之后的 main,而且这次合并是手工解决的。解决方案的两处接缝,实测行为都与合并说明一致;第一轮和第二轮验证过的行为没有变化。

合并前需要知道两件事:

装置与前两轮相同:构建出来的 TypeScript provider → Spring 内嵌 Broker(MySQL 8.4.7)→ 打包 worker,Broker 与 worker 之间挂着 JVM HTTP 代理。

跑了什么 结果
合并提交的构成 不是机械合并:3 个冲突文件,另有 2 份设计文档和 1 个测试文件被手工修改;冲突之外没有改动任何生产代码
接缝 1:释放后的终态回执、对已准备调用的重复取消(M1、M2) 10/10。上一个 head:释放后全部 404。本 head:GET、取消、同一预留从记录返回 200,到达 worker 的请求为 0
手写守卫(M4、M5),本 head 对比编入 G4 变异的 jar 本 head 12/12;变异 jar 对一个已完成的调用再次取消时返回 503
接缝 2:预留了一个调用时杀掉 worker 本 head 的响应与 main 逐行相同,raw 路径和 provider 路径都是
重跑第一、二轮探针 76/76、12/12、7/7(20 个拒绝全部保留错误码和原因)、13/13
Hosted Workspace 循环 PR 自带 driver 接 MySQL:HOSTED_WORKSPACE_TOOLS_OK,7 次执行 SETTLED/success · 真实模型 3 轮:5 次执行 SETTLED/success
本 head 上的仓库套件 Broker 402、故障门禁 29、Broker MySQL IT 2、server 140 + 10、Hosted IT 4、core 190,全部通过
与当前 main 试合并 文本无冲突;Broker 402、故障门禁 29、core 190 通过;server 和 Hosted 套件无法启动(main 上有两个 V16)
与 main + #12900 试合并 无冲突;Broker 407、server 154 + 11、Hosted 154 + 5、core 190 全部通过;jar 能启动;真实栈上 A–D 76/76,M1、M2、M4、M5 22/22
变异矩阵,70 个变异体 杀掉 64 个,存活 6 个(第二轮为 61 个中的 56 个)

head 里有什么

见图 1。第二轮的合并提交与机械合并的树完全相同,这一轮不是,所以我用 git show --remerge-diff 来读它,这个命令只输出「实际记录的合并」相对「机械合并」的差异。生产代码上的决策都在 RuntimeBrokerService.java 里:

  • 预留(createExecution、prepareExecution)现在走 main 的回执查找,使用的是本 PR 校验过的 reference。
  • cancelExecution 对终态记录直接从存储返回,例外是 Session 尚未释放的「已准备的 provider 调用」,它仍然要去问 worker。
  • 这个例外由新的辅助函数 isPreparedProviderCancellation() 判定:已结算、已请求取消、派发代数为 0、reference 是 provider 形态。

1. 接缝 1:终态回执到达 provider 路径

见图 2。

合并说明里的声明 这里的实测
保留了 main 的终态回执 释放之后,GET、取消、以及用同一 key 和 reference 的重复预留都返回 200,执行 id 相同。到达 worker 的请求为 0,Session 行不变(RELEASED,版本相同)
回执仍然有绑定 同一 key 换一个 invocation:409 runtime_idempotency_conflict。换一个 Harness Session:409 runtime_execution_conflict
存活 Session 的控制、history 与回执是两回事 释放之后,新预留、start、control[history] 仍是 404
对已准备 provider 调用的重复取消仍然等待 worker 释放之前每次取消都到达 worker:第 1 次 cancel,status,第 2 次 cancel,三个并发取消全部 200 cancelled,没有副作用。释放之后取消从记录返回
释放后回执的探针跑在 H2、静态 placement、内存 Session 元数据上 这里在 MySQL 8.4.7、Spring server 的 placement、JDBC Session 元数据上得到相同结果,boot v1 和 boot v2 都是
真实链路覆盖 SETTLED;ABANDONED 由 service 测试覆盖 这里有同样的限制:本装置没有造出 ABANDONED 记录。客户端一侧用替身对端检查过(N.1):终态 runtime_lost 响应被报告为 unknown / terminal,普通的仍是 unknown

2. 合并时手写的守卫

见图 6。isPreparedProviderCancellation() 有四个条件,有两个变异体各去掉其中一个。去掉 provider reference 判断的(G5)被仓库测试杀掉。去掉 getDispatchGeneration() == 0 的(G4)没有被杀:带着它 402 个 Broker 测试全部通过。

于是我把 G4 编进一个 server jar,在两个 jar 上各跑两个探针:

  • M4,一个被派发后又被取消的调用。本 head 对重复取消直接从记录回答;变异 jar 每次都再去问 worker。两者的响应相同,都是 200 cancelled。
  • M5,一个在 worker 收到取消之前就已完成的调用(取消请求在线路上被延迟 5 秒,命令本身耗时 2 秒)。记录是 SETTLED/success 且带取消请求标记。本 head 对之后的取消回答 200 success。变异 jar 每次都回答 503 runtime_execution_cancel_failed,因为 worker 报告的是 success,而「已准备取消」路径只接受 cancelled 或 not started。

守卫本身是对的。补一个针对 M5 情形的 service 测试(已结算、success、已请求取消、代数为 1,再次取消 → 返回存储记录且不调用 transport)就能钉住它。

3. 接缝 2:main 的 Runtime 丢失围栏下的 worker 死亡

见图 3。在预留了一个调用时用 SIGKILL 杀掉 worker。所有构建上都没有任何执行,也没有重放。

上一个 head f0f217dfa5 main d66fdadd27(raw 路径) 本 head,raw 与 provider 路径
placement FAILED LOST LOST
替补 worker 启动了,代数 2 没有 没有
对该执行的 GET 409 runtime_broker_execution_unknown 409 runtime_admission_closed 409 runtime_admission_closed
释放 409 runtime_session_busy 503 runtime_reconciliation_required 503 runtime_reconciliation_required
下一个 Runtime Session 409 workspace_busy 503 runtime_broker_runtime_lost 503 runtime_broker_runtime_lost
兄弟 Harness Session 409 workspace_busy 409 runtime_placement_recovery_required 409 runtime_placement_recovery_required

这个变化继承自 #12839,合并既没有削弱它也没有扩大它。对装置的影响是:每个 worker 死亡场景都要用全新的数据库和 Broker,因为 placement 丢失后会拒绝之后的 placement。

4. 回归

见图 4。图里 cli src/serve 的失败是本机环境问题。本 head 的 CI 是绿的:21 项检查通过,4 项跳过,没有失败项,其中包括 Test (ubuntu-latest, Node 22.x)。

5. 变异矩阵

见图 5。新增的 9 个变异体(G1–G9)改的是解决合并时写下或依赖的那些行:杀掉 8 个,G4 存活(见第 2 节)。第一、二轮的 61 个变异体在合并后的代码上重跑:杀掉 56 个,存活的仍是那五个(T2、T8、T18、N17、N18)。没有任何判定发生变化,也就是说合并没有丢掉任何已被测试钉住的行为。

击杀是按 JSON 报告逐个测试归因的。高负载下全量运行时 hosted-harness-session.test.ts 的失败不计入:N17 只有这类失败,而带着 N17 单独跑这个文件是 16/16 通过。

6. 当前 main 上有两个 V16 迁移

见图 7。这是在把本 head 与当前 main 试合并时发现的。#12839 在 05:14Z 合入,带来 V16__runtime_loss_evidence.sql。本 head 的合并提交生成于 05:32Z。#12881 在 05:37Z 合入,带来 V16__managed_session_operation.sql。git 看到的是两个不同的文件名,不会报任何冲突;Flyway 则拒绝启动。

#12900 把后合入的迁移挪到了 V17。main(fc4e01b9fc)+ #12900 + 本 head 的试合并没有冲突。在这棵树上,Broker 单测(407)、server 套件(154 加 11 个 IT)、Hosted 套件(154 加 5 个 IT)和 core(190)全部通过。它的 server jar 能在全新的 MySQL 数据库上启动,并按顺序应用 15、16、17 三个版本。在这个 jar 上,评审测试计划 76/76 通过,合并探针 22/22 通过。

前两轮各项的现状

条目 在本 head 上的状态
F1 明确拒绝 → UNKNOWN → Session 和存储被钉住 作者延后处理。复测无变化(E.1/E.2):release 返回 409 runtime_session_busy;:resolve 仍是 501,reconcileExecution 仍然没有 HTTP 路由
F3 worker 保留已释放的 Session 作者延后处理。本轮没有复测;这个 head 没有改动 worker 代码
R1 原因超过 4,096 字符或含 NUL 时丢失错误码 作者延后处理。复测无变化:11 个过长或含 NUL 的原因里有 4 个丢失错误码。第二轮的候选补丁仍然适用
错误标签部分颠倒、原因原样转发 作者延后处理。无变化(变异体 N17、N18 仍然存活)
参数上限 256 KiB 无变化:255 KiB 通过,256 KiB 被拒绝,错误码 managed_runtime_provider_invalid

合并两侧之间有一处小的相互作用,不阻塞。客户端只在错误体恰好有三个键时才附上原因;main 增加了可选的第四个键 details。因此带 details 的错误会保留错误码,但永远不带原因(N.2,替身对端)。目前唯一带 details 的响应是 ABANDONED,客户端会把它转换成 unknown / terminal,所以现在还没有调用方能看到差别。

合并顺序

对象 与本 head 试合并的结果
当前 main(fc4e01b9fc) 无冲突;需要 #12900 才能启动
#12900、#12855、#12896 无冲突
#12848 4 个文件冲突:managed-context-worker.ts、RuntimeBrokerHttpServer.java、RuntimeBrokerService.java、RuntimeBrokerHttpServerTest.java
#12894 6 个文件冲突:上面四个,加上 WorkspaceRuntimeTest.java、JsonCodec.java
草稿 #12865、#12869 分别有 5 个和 14 个文件冲突

未覆盖

Linux 与 Windows;MariaDB;真实栈上的 ABANDONED 记录(需要 #12839 的恢复流程);带着存活 provider Session 重启 Broker;始终不结算已准备取消的 worker;F3 的内存增长本轮没有复测。

复现

第三轮的装置、探针、变异体和日志:pr12868/r3。新增探针是 s13-merge.mjs(M1–M5)、kill-scenario.sh 和 s14-details-envelope.mjs。remerge-r3.diff 是 git show --remerge-diff 的输出。

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review comments addressed on cc06ee0e4e.

R1-30 and R1-32 are fixed with focused regression tests; the mixed R1-2 comment has a separate evidence-based reply. The remaining 27 inline Suggestions were checked against the current code and are recorded below under the requested policy to land only Critical fixes after round 5.

Comment Disposition
R1-6, R1-12 Deferred: synchronize older bilingual design documents that still describe unsupported Session verbs/local acknowledgements.
R1-7, R1-9 Deferred: the immediate raw API remains a deliberate compatibility path. Provider and deferred raw reservations keep payloads separate; removing the old API is outside these fixes.
R1-8 Deferred: commit the full provider real-chain fixture as a maintained CI lane and add immediate raw fault coverage. Local real-chain evidence is not represented as an existing committed CI test.
R1-10 Deferred: refine policy-refusal error labels; refusals remain enforced today.
R1-11 Deferred: align uppercase UUID handling across Java and TS. This currently fails closed; do not narrow unrelated opaque Session IDs.
R1-14 Deferred: precompile Java regexes/remove repeated scans; no current correctness failure established.
R1-16 Deferred: cover the two remaining allowed error-code/status pairs explicitly.
R1-18 Deferred: add a dedicated before-network malformed-reference guard test; the production guard is present.
R1-20, R1-53 Deferred: enumerate all confirmation outcomes in the shared corpus. The current nine-value sets agree; these comments duplicate the same coverage request.
R1-21 Not accepted as a bypass of the current contract: the private provider explicitly requires its caller to enforce permission confirmation. Adding worker-side enforcement would change that contract; further policy/coverage work is deferred.
R1-22 Deferred: strengthen read/shell output assertions using the provider result shape, rather than the raw protocol responseParts shape.
R1-24 Deferred: remove the redundant turnStarted branch; the active-work guard remains live.
R1-25 Deferred: broaden v3 and post-await race mutation coverage. Both raw protocols retain the actual admission guards, including during this rollback fix.
R1-26 Partly stale: current release can retry without requireReadySession and clears its failed release promise; loss recovery requires durable stopped-writer evidence. Unknown cleanup must retain storage ownership. Additional false/timeout/recovery coverage is deferred.
R1-42 Deferred: update the transport method documentation to describe both existing raw and provider result shapes.
R1-43 Deferred: normalize direct Java API null-input failures. The HTTP boundary already rejects these invalid fields before reservation or dispatch.
R1-44 Deferred: report oversized control requests consistently instead of retryable transport failure. This change does not enlarge wire limits.
R1-46 Deferred: extend provider request/response and transport-only operation guard coverage. Existing closed-envelope and history-owner guards remain present.
R1-49 Deferred: add service-layer no-write assertions for bad references. Existing HTTP extra-field tests deliberately pin the separate envelope guard.
R1-51 Deferred: align prepare/start malformed-envelope error naming; both currently reject before effects.
R1-57 Deferred: pin stable idempotency keys with a golden fixture; current field-ordered hashing is deterministic.
R1-58 Deferred with the ID-invariant portion of R1-2. No current live write escape was reproduced; hashing core IDs alone would also require preserving the wire/core identity relation.
R1-60 Deferred: reduce redundant canonicalization/hashing while preserving independently callable parser bounds.
R1-61 Deferred: pin the declaration/per-kind maximum relationship. The declared 8 MiB already equals the maximum admitted kind; it is not a claim that every kind accepts 8 MiB.

Additional verification comments: the lone-surrogate TS/Java divergence, synthetic cancelled/error combination coverage, G4 mutation coverage and previously disclosed F1/F3/long-reason limitations remain follow-ups. The proposed correction from “402 / zero skips” to one skip does not apply to our recorded run: we passed -Dqwen.runtime.worker.bundle=/Users/wenshao/git/qwen-code-k1/dist/cli.js, and the saved Maven result is Tests run: 402, Failures: 0, Errors: 0, Skipped: 0. The reviewer’s run without that property legitimately skipped its optional worker test; these are different runs. We have not rerun that whole Java suite for this TS-only revision.

Validation of this committed tree: build/typecheck/bundle and ESLint/Prettier passed; 136 core tests and 859 CLI tests passed. Three isolated source mutations were each caught by the target regression assertions. Real bundled-worker verification passed (three corrected-behavior groups plus two unchanged diagnostics), followed by 40/40 provider and 8/8 error/recovery groups through the production Java Broker, H2 execution storage and real worker. The 40/8 totals include four readiness groups. All recorded source, scripts and built artifacts were unchanged during the final runs. The first real-worker attempt exposed a fixture collision from reusing raw callId across Sessions; the fixture was corrected, evidence retained, and the final runs passed without another product change.

Three complete open-ended and reverse audit rounds preceded this commit. Round 1 found a stale test expectation; after correcting it and rerunning validation, rounds 2 and 3 were clean. The committed tree exactly matches the audited/tested tree. These checks do not claim live-model, new-platform, scale, or runtime-loss recovery coverage.

本轮修复 R1-30、R1-32;R1-2 拆分核验后按非阻塞后续项记录,其他 27 条建议逐项保留上表结论。没有扩大已声明的权限、容量或兼容性契约。

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@doudouOUC doudouOUC left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed the current head and the updated discussions. The previous provider response-loss read and cancellation findings are fixed: the broker now uses the same provider/v3 observability predicate for HTTP observation and physical cancellation, and the new HTTP test checks that an UNKNOWN provider execution sends a worker-side cancel without redispatch. I found no remaining blocking issue in this update. Locally, RuntimeBrokerHttpServerTest passes (19 tests) and RuntimeBrokerServiceTest passes (116 tests); the remaining CI jobs are still running.

R1-11 remains a non-blocking suggestion: an uppercase UUID accepted in the provider envelope conflicts with core's lowercased identity on identity-bearing controls. Given the number of review rounds, please track that edge case in a follow-up issue or PR; this approval does not mark it fixed.

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head ebea694e4e59172c7a17577d528a1fea5553f590

Verdict: COMMENT, on budget rather than on a newly proven defect. One previously confirmed Critical is now fixed, and I verified the fix by reading the code at this head. Four historical blocker threads remain unconfirmed by me, and the full-diff scan of a 50-file / +8310−245 change did not fit inside this run's time budget, so under this channel's rules I cannot issue an Approve.

R4-1 — released Runtime Session never freed: FIXED at this head

The unbounded-retention Critical I confirmed at 5c0c9bf3 is resolved in packages/cli/src/serve/managed-runtime-provider-worker.ts, and it is resolved in the way the finding required — bounded retention with a terminal tombstone, not a dispose-at-release one-liner:

  • RETAINED_RELEASED_SESSIONS = 8 caps how many released Sessions keep heavy state.
  • The release arm of control now pushes the entry onto this.released and awaits retireReleased(), which evicts entries beyond the cap, oldest idle first (findIndex(entry => entry.pending === 0)), so a Session with a call in flight is never retired underneath it.
  • retire(entry) runs runtime.dispose(), history.drain() and config.shutdown({ shutdownTelemetry: false, skipSessionWriter: true }) as independent steps that each execute even if an earlier one throws, collecting errors into an AggregateError, then sets entry.value = undefined and entry.ready = undefined. It is memoised through entry.retirement ??=, so concurrent callers await one retirement instead of racing several.
  • The pinned post-release observability survives: the 8 retained Sessions stay fully answerable, and an evicted one becomes a tombstone that the session?.closed && !session.ready branch answers with { state: 'unknown' } for status/cancel while still raising a conflict for result-bearing operations. A failed retirement is logged and never fails the release, so the eviction path cannot turn a successful release into an error.
  • forgetSession now also drops the per-session published model (unregisterSessionModel) alongside the project dir, which closes the second process-global accumulation the original finding named.

What is left in this.sessions per evicted turn is a small tombstone object, not a Config plus runtime plus file history. That is a deliberate documented trade-off and I am not filing it as a blocker.

Blocker threads I could not confirm at this head

The round-4 automated review disclosed that it never re-checked the PR's existing blocker threads, and I had no remaining budget to read them here. They are therefore unconfirmed, not cleared:

  • R1-32 — the provider route turning the protocol's { state: 'unknown' } answer into a 409 that reports the caller's reference as forged. Worth a specific look now, because the tombstone branch above is exactly the code path that produces { state: 'unknown' }, and eviction makes that answer reachable in normal operation rather than only after a failure.
  • R2-1 — whether the terminal-cancel fast path refreshes from the persisted session record.
  • R2-2 — the per-kind response cap applied to result-bearing operations whose producers legitimately emit more.
  • R3-9 — the result fitter measuring the overage in bytes.
  • The Shell-result-over-limit finding raised by @doudouOUC at cc06ee0e.

For the record on other reviewers' verdicts: @doudouOUC approved this head and reported the provider response-loss read and the cancellation case fixed, with RuntimeBrokerHttpServerTest (19) and RuntimeBrokerServiceTest (116) passing locally. I did not independently re-read those two call chains, so I am not carrying that approval as my own confirmation. R1-11 (uppercase UUID accepted in the provider envelope versus core's lowercased identity on identity-bearing controls) is explicitly non-blocking and already deferred to a follow-up; it does not affect this verdict.

CI

No check at this head reports failure or cancellation. Test (ubuntu-latest, Node 22.x) is still running, and Hosted process fault gates / MySQL 8.4 / Java 21, OpenTUI no-flicker gate and TUI parity snapshots (ink vs opentui) are queued; pending state is not treated as a gate here. The Desktop Shell matrix failure I noted at 5c0c9bf3 no longer appears — both Desktop Shell (ubuntu-22.04) and Desktop Shell (windows-2022) are green.

Next step: the remaining work is confirmation, not new fixes. Re-run the review round at the head that will actually merge so the five threads above get ruled on, and add a regression test that acquires and releases more than RETAINED_RELEASED_SESSIONS distinct runtimeSessionIds and asserts the evicted Sessions' Config was shut down and their runtime payload no longer retained — that is the test which keeps R4-1 fixed.

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification, round 7: ebea694e4e, merged as 1f1209bc70

The PR was merged while this round ran. Two findings of the round are on main now and need a follow-up: R7-1, and the smaller R7-2. 9cb9dc86e8 introduced both, the head had them, and main a63157304a has them. R7-1: a start the worker refuses never returns to the caller. Nothing in production calls the provider yet, so nothing that runs today is affected. One condition in the Broker removes it (candidate B, below); with it every probe of all seven rounds passes, and the patch applies to main. Everything else the round measured is in order.

The squash commit carries the files of the head byte for byte, except config.ts, which main had changed as well. On main a63157304a, one commit after the merge, the start of R7-1 has no answer after 20 s while the worker is asked 335 times; with candidate B it is refused in 0.1 s, Y1 to Y8 pass 21 of 21, and the Broker suite passes with 514 tests.

The head is the round-6 head plus seven commits. The branch moved five times while this round ran. Each state was built and measured, and every number below is from the head unless it names another build. 2 reviewers have approved this head; all 96 review threads are resolved; the review decision on the PR reads APPROVED. None of the reviews names R7-1. Earlier rounds: 1, 2, 3, 4, 5, 6.

R7-1. A start the worker refuses never returns

A call is prepared and reserved, and the caller starts it before preflight has permitted it. The worker refuses the execute with 409, as it should.

50fb28301e, before The head, and main after the merge
The caller's startExecution refused after 0.1 s, 409 runtime_broker_execution_unknown no answer; still waiting when the probe stopped after 20 s. On 9cb9dc86e8 a probe without a limit waited 4 min 26 s, until it was stopped by hand
Requests to the worker meanwhile 1 execute 1 execute, then status about 17 times a second for as long as the provider lives: 314 in 20 s, and 4,549 in those 4 min 26 s
The Broker, asked directly start 409, read 409 start 200 prepared, read 200 prepared

Since 9cb9dc86e8 the Broker asks the worker about a provider execution whose record is UNKNOWN, and answers 200 with the state the worker names. For a lost answer that is right: the worker says executing or settled. For a refused start the worker says prepared. Nothing will start that call, because the Broker never dispatches a second time, so the provider client, which polls every 50 ms until the state is settled, polls without an end. Only disposing the provider stops it.

The same request pattern is what a Hosted turn would send if it started a call it had not yet been allowed to run. The provider has no production caller today, so nothing that runs is affected. But before this commit the mistake failed closed in a tenth of a second, and now it hangs the turn and loads the worker. It also shows in earlier probes: E.2 of round 1 gets no answer, and checks of groups B and C that count requests to the worker find the polling of an unrelated call (section 3).

ebea694e4e softens one consequence and leaves the cause: a caller that cancels now reaches the worker, the record settles as not_started, and the Session is released. The start itself still does not return.

Candidate B is one condition in RuntimeBrokerHttpServer.observedExecutionEnvelope: a provider execution the worker still holds as prepared was never started, so the Broker keeps the answer it gave before, 409 runtime_broker_execution_unknown. With it on this head: the start is refused after 0.1 s with 409, and the worker is asked about the call once, not 17 times a second; Y1 to Y8 pass 21 of 21, every probe of rounds 1 to 6 passes, the Broker suite passes with 514 tests, and Checkstyle is clean. Its test fails without the condition. Lost answers still settle (Y6), and the cancellation of ebea694e4e still works.

Candidate B
diff --git a/packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServer.java b/packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServer.java
index 1ddaf1d258..772e1655c9 100644
--- a/packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServer.java
+++ b/packages/sdk-java/runtime-broker/src/main/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServer.java
@@ -332,7 +332,12 @@ public final class RuntimeBrokerHttpServer implements AutoCloseable {
             String runtimeSessionId, ExecutionReconciliation observation) {
         ToolExecutionRecord record = observation.getRecord();
         String state = observation.getRuntimeState();
-        if (record.getState() == ToolExecutionRecord.State.UNKNOWN
+        // A provider call the worker still holds as prepared was never
+        // started, and nothing dispatches it a second time: it stays UNKNOWN
+        // for the caller, who would otherwise wait for it for ever.
+        boolean neverStarted = "prepared".equals(state)
+                && ProviderRuntimeProtocol.isReference(record.getReference());
+        if (record.getState() == ToolExecutionRecord.State.UNKNOWN && !neverStarted
                 && ("prepared".equals(state) || "executing".equals(state)
                     || "cancel_requested".equals(state))) {
             Map<String, Object> response = envelope(harnessSessionId, runtimeSessionId,
diff --git a/packages/sdk-java/runtime-broker/src/test/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServerTest.java b/packages/sdk-java/runtime-broker/src/test/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServerTest.java
index 2be20af7c3..53464b377b 100644
--- a/packages/sdk-java/runtime-broker/src/test/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServerTest.java
+++ b/packages/sdk-java/runtime-broker/src/test/java/com/alibaba/qwen/code/runtimebroker/RuntimeBrokerHttpServerTest.java
@@ -226,6 +226,34 @@ class RuntimeBrokerHttpServerTest {
         }
     }
 
+    @Test
+    void providerStartTheWorkerNeverBeganStaysUnknownForTheCaller() throws Exception {
+        String runtime = "550e8400-e29b-41d4-a716-446655440303";
+        try (Fixture fixture = new Fixture()) {
+            fixture.service.acquire("harness", runtime, "bootstrap").toCompletableFuture().join();
+            Map<String, Object> reference = Map.of("sessionId", runtime, "promptId", "turn",
+                    "callId", "worker-call", "capabilityDigest", "a".repeat(64),
+                    "policyRevision", "policy", "invocationId", "invocation", "argsDigest", "b".repeat(64));
+            String id = fixture.service.prepareExecution("harness", runtime, "provider", reference)
+                    .toCompletableFuture().join().getExecutionCallId();
+            Map<String, Object> start = Map.of("protocolVersion", 1, "requestId", "start",
+                    "harnessSessionId", "harness", "runtimeSessionId", runtime);
+            // The worker refused the execute: it still holds the call as prepared.
+            fixture.transport.runtimeStatus = Map.of("state", "prepared");
+            for (int attempt = 0; attempt < 2; attempt++) {
+                HttpResponse<String> started = fixture.post("/executions/" + id + ":start", start);
+                assertEquals(409, started.statusCode(), started.body());
+                assertEquals("runtime_broker_execution_unknown",
+                        JSON.parseObject(started.body()).getString("code"));
+            }
+            HttpRequest read = HttpRequest.newBuilder(fixture.uri("/executions/" + id
+                    + "?requestId=read&harnessSessionId=harness&runtimeSessionId=" + runtime))
+                    .header("Authorization", "Bearer secret").GET().build();
+            assertEquals(409, fixture.client.send(read, HttpResponse.BodyHandlers.ofString()).statusCode());
+            assertEquals(1, fixture.transport.executions.get());
+        }
+    }
+
     @Test
     void providerObservesTheOriginalExecutionAfterResponseLoss() throws Exception {
         // Provider references carry a UUID Runtime Session id.

R7-2. After a worker is lost, an inspection throws where it answered "unknown"

Smaller, from the same commit, and for the same follow-up. A worker is killed while a command runs. The record is UNKNOWN, as before, nothing is started twice, and the storage stays held. What the caller is told changed with 9cb9dc86e8:

50fb28301e, before The head, and main after the merge
The start the caller waited on 409 runtime_broker_execution_unknown 409 runtime_execution_evidence_unavailable
A read of the execution 409 runtime_broker_execution_unknown 409 runtime_admission_closed, "Runtime generation is no longer live"
inspectExecution of the provider {"outcome":"unknown"} throws 409 runtime_admission_closed, not retryable

The Broker now tries to ask the worker, finds it gone, and answers with the error of that attempt. The provider turns only runtime_broker_execution_unknown into {outcome: 'unknown'}, so a caller that recovers by inspecting gets an exception where it got an answer. The state on the stack is as safe as it was; what moved is what that caller must handle. There is no candidate for this one: whether the Broker's answer or the client's reading should move is a decision of the design.

The rest of the round

Open after round 6 This head
R4-1, F3 of round 1: a worker keeps what a released Session held Closed on the path this PR adds. After 300 Sessions the worker was 286 MiB larger; now it is not larger after 300, nor after 1,000 (section 1)
F1 of round 1: a definite refusal leaves the record UNKNOWN, and the Session and its storage pinned Closed on the provider path by ebea694e4e. The caller cancels, the record settles as not_started, release answers 200 and the storage is free (section 5)
K7 and K41: no test holds an id with a letter outside ASCII Closed. Both allow-list tests now hold letters and digits outside ASCII; the mutants K7 and K41 are killed (section 6)
G4, R4-7: the guard of the round-3 merge has no test Closed. G4 is killed by repeatedCancellationOfADispatchedProviderExecutionReturnsItsReceipt (section 6)
The 503 after a Broker restart never ends with the default provisioner Unchanged; the update now says so. With the durable provisioner of #12865 it ends (section 3)
Opened by the merge with main, closed by 50fb28301e Merge commit This head
A tool input with an unpaired surrogate raw path refused, provider path ran it with ? in its place refused on both paths
A Shell call with a directory outside the Session's workspace ran in the directory of another storage refused, also through .. and through a symbolic link
A status cursor in a number form that is not an exact integer accepted refused
Added by the commits after that Before This head
9cb9dc86e8: the answer to execute is lost record UNKNOWN for good, Session and storage pinned the Broker asks the worker, the record settles success; one dispatch, one effect
9cb9dc86e8: a link moved out of the workspace after prepare the command ran outside the call settles as an error and does not run
760174073b: artifacts that cannot fit the result did not fit, the model's text was replaced by the stub artifacts dropped, the model's text whole, the result fits
ebea694e4e: cancelling a call whose answer was lost, or that was refused 409, nothing sent to the worker 200, the worker is asked, the record settles

What remains besides R7-1 and R7-2. None of it is urgent, and none of it came with this PR:

  1. The raw path keeps about 325 KiB per released Session, on main too. It is the path the Hosted harness uses today, and this PR does not touch it: the fix bounds what the provider route keeps, the executor's journal is outside the diff. It deserves an issue of its own (section 1).
  2. 16 of 200 mutants survive, 7 of them new. The new ones are edges of the retirement, of the response buffer, of the cursor and of the directory check; none is a rule the real stack was seen to break (section 6).
  3. One fault gate that came with main fails on a loaded host, on main as well. DurableLocalRuntimeFaultGateTest, run by itself on a host under load, passes 2 of 3 times on the head, 1 of 3 on main; two builds at once fail every time. The whole profile passed 44 of 44 on its first run on the head (section 3).

Same rig as before: built TypeScript provider → Spring-embedded Broker on MySQL 8.4.7 → bundled workers, with the JVM HTTP proxy between Broker and worker; a trial merge with main be1ebc74d7, made before the merge, as a second arm; main a63157304a after the merge as a third; candidate B on the head and on main; a Linux container for the durable provisioner.

What was run Result
Worker memory after 300 released Sessions, 200 KiB per turn, provider path previous head: 177 → 463 MiB · this head: 179 → 148 MiB
The same after 1,000 Sessions this head: 179 → 124 MiB
The same on the raw path this head: 165 → 262 MiB · main: 162 → 257 MiB
Time of a release at the worker previous head: median 1 ms, slowest 24 ms · this head, 1,300 releases: median 1 ms, slowest 36 ms
What a released Session answers (Z1), both boot versions previous head: all 12 in full · this head: 8 in full, 4 tombstones, 8/8
A live Session while others are retired (Z2) 4/4: its prepared call runs, a running command finishes
32 Sessions, six at a time, on one worker (Z3) 3/3: 32 of 32
warm with a Harness Session id outside the allow-list (Z4) previous head: 503, retryable · this head: 400, 4/4
The rules of main, lost answers, refused starts (Y1 to Y8) this head, trial merge and main after the merge: 19/21, the two of R7-1 fail · with candidate B, on the head and on main: 21/21
Probes of rounds 1 to 6 this head: 75/76, 10/12, 7/7, 13/13, 22/22, 30/30, 24/24, 10/10, 9/9, 6/6; what fails is R7-1 · main after the merge: 75/76, 10/12, the rest in full · with candidate B, on the head and on main: 76/76, 12/12 and the rest in full
Broker replaced, durable provisioner, Linux 4 of 4 scenarios, 5/5 each: acquire in 0.1 to 0.2 s, cancellation confirmed
Worker killed under a running command, macOS and Linux 4/4 each: not started again, storage held until the old writer is known to have stopped
Hosted loop repository ITs on MySQL 8.4.7 pass · live model, 3 turns: 7 executions settled, 6 success and 1 error (a file the model looked for before it wrote it)
Repository suites this head: Broker 513, fault gates 44, Broker MySQL ITs 3, server 188 + 15 ITs, Hosted ITs 10, core 194 · trial merge: the same, and 11 Hosted ITs · all pass, Checkstyle clean
Mutation matrix 200 mutants run on this head: 184 killed, 16 survive. 57 are new: 50 killed, 7 survive

What the head contains (figure 1)

What the head contains

1. What a worker keeps after a release (figure 2)

Retention

One worker, Sessions one after another, each used for one write_file of 200 KiB and released, as a Hosted turn does. The memory is the resident set of the worker process, at the first release and 15 s after the last.

Path Build Sessions Worker RSS Per released Session
provider previous head 300 177 → 463 MiB 980 KiB
provider e94f781523, the fix alone 300 178 → 149 MiB none
provider this head 300 179 → 148 MiB none
provider this head 1,000 179 → 124 MiB none
raw this head 300 165 → 262 MiB 332 KiB
raw main 12793013c4 300 162 → 257 MiB 325 KiB
Claim in the update Measured here
The worker keeps the full state of up to eight released Sessions of 12 released Sessions the last 8 answer status, cancel and history in full (section 2)
The oldest one not being observed is retired the 4 oldest are tombstones
No release leaves more than eight released Sessions retained memory does not grow over 1,000 Sessions: the samples taken while it runs lie between 156 and 231 MiB, and 15 s after the last release the worker holds 124 MiB
A release answers only after the retirement it triggered it costs nothing that shows: 1,300 releases, median 1 ms, 99th percentile 8 ms, slowest 36 ms
Tombstones still grow by one per released Session not visible at 1,000 Sessions

The raw rows are the same on this head and on main. That retention is older than this PR and lives in the executor, which the fix does not touch. It matters more than the provider rows did, because the Hosted harness runs on the raw path today and the provider has no production caller yet.

2. What a released Session still answers (figure 3)

Released Sessions

Asked of a released Session, at the worker One of the last eight An older one
status of its call settled, in full unknown
cancel of its call settled unknown
history 200 409
prepare of new work 409 409

The Broker sends none of these to a released Session; the rig did, with the Broker's headers. Through the Broker nothing changes: the execution of a retired Session reads 200 settled/success from the Broker's journal, and a repeated release answers 200.

Retirement shuts the Config of the retired Session down. It does not reach the others on the same worker:

  • A Session acquired first, with a call prepared and approved, runs that call after twelve others were released and four retired, and then runs another.
  • A Shell command that runs while those twelve come and go finishes with all its output.
  • 32 Sessions, six at a time, all run and release; no worker answer is a refusal.

Boot v2 admits one tool turn per Workspace at a time (409 workspace_busy), so the last two ran on boot v1.

warm now refuses a Harness Session id outside the allow-list with 400 runtime_broker_invalid_request, before a worker is started. The previous head answered a retryable 503 runtime_scope_resolution_failed.

3. Regression, the restart and the suites (figure 4)

Regression

Regression. On this head, on the trial merge and on main after the merge the probes of rounds 1 to 6 pass except where R7-1 reaches them. E.2 starts a call the worker refused and gets no answer in 30 s, on boot v1 and on boot v2. Checks of round 1 that count the requests which reach the worker in a window find a status request of the call that B.1 started before and that is still polling: C.5 in these runs, on the head, on the trial merge and on main; which check it hits differs from run to run. With candidate B all of them pass, on the head and on main.

Three checks state a rule that 9cb9dc86e8 changed by design, and were rewritten to the new rule: F.1 and F.2 of round 1 and S.exec of round 2 said that the record of a call whose answer was lost stays UNKNOWN. It now settles from the result the worker kept. The logs of the earlier builds in this round carry the old wording.

The restart. As in round 6. With the default provisioner the repeated cancellation answers a retryable 503, acquire answers 503 after 120 s, and that repeats. On Linux, with the durable provisioner, which the head now has from main, acquire adopts the worker in 0.1 to 0.2 s and the cancellation is confirmed: 4 of 4 scenarios, 5/5 each.

A worker killed under a running command. On both rigs the record becomes UNKNOWN, the call the caller waited on answers 409, start sent again is refused, and release answers 503 until the Broker knows the old writer stopped; the storage stays held. The worker is gone, so there is nobody to ask: the new reconciliation changes the code of the answers (R7-2), not the state. The client's message carries the Broker's reason. The answer has no details on either rig, so the terminal ABANDONED answer was not produced on the real stack, and R4-3 rests on its unit test and its mutant (section 6). On Linux the command's own process outlives the worker and finishes its work once; the Broker's refusal to give the storage away fits that.

Suites and CI. The repository suites pass on the head and on the trial merge (figure 4). CI on the head: 22 checks pass, none fails; assign was still queued when this was written. No suite of the repository starts a call the worker refuses and waits for it, which is why R7-1 shows in none of them.

One fault gate depends on the load of the host. DurableLocalRuntimeFaultGateTest came with #12865; this PR changes none of it. While other runs of this round loaded the host, the fault-gate profile failed in that class on two earlier states of the branch (1 and 3 of 44); on the head and on the trial merge it passed 44 of 44 on its first run. Run by itself and repeated, with the load average between 13 and 54 on ten cores, the class passes 2 of 3 times on the head, 1 of 3 on main 12793013c4, and 1 of 4 on the main before #12975; with two builds running it at once it fails every time, on the branch as on main. The failures are the same two assertions throughout: absentWorkerAfterRestartLeavesItsDomainPinned gets runtime_provision_fenced where it expects runtime_broker_runtime_lost, and adoptedWorkerCanCancelItsOriginalActiveCall sees no cancellation reach the worker. One run names the cause, Runtime recovery claim expired.

That answers a question the review bot left open. In CI the job Hosted process fault gates failed once on this head, in this class and with this error (closingAnObserverCannotKillTheSharedWorker, 503 runtime_provision_fenced, Runtime recovery claim expired), and passed when it was run again. The bot would not call it a flake, because 760174073b and ebea694e4e had landed since the last green run. In the runs above the same error, runtime_provision_fenced, fails other tests of the same class on main without either commit, and on the main before #12975. It is a lease that expires on a busy host; neither commit is needed for it. The test CI named did not fail here.

4. The merge with main, and the rules main had meanwhile (figure 5)

The merge

main had moved by 38 commits since this branch last merged it. Three files conflicted: two design documents and RuntimeBrokerService.java. There, #12975 had changed the deferred start, which this PR had moved into a branch of its own; the merge carried the two checks of #12975 into that branch by hand. No migration version exists twice.

The part written by hand is right. A deferred start whose payload holds an unpaired surrogate, in a string of the input, in a key of it or in the tool name, answers 400 runtime_payload_invalid on the head as on main, and nothing is sent to the worker. Before the merge all three were sent, with ? in place of the surrogate. The reservation of a refused start stands until the caller cancels it; then the Session is released and its storage is free. The three mutants of those lines (M1 to M3) are killed by the tests #12975 brought.

What merging could not bring. The provider path is this branch's own; rules main gave the raw path do not reach it by merging. 50fb28301e names three and adds them. Each was measured on the merge commit 174f974e3e and on the head:

Through the provider 174f974e3e, the merge This head
prepare of ls data-\ud800.txt the worker received ls data-?.txt, ran it, and listed data-1.txt and data-2.txt 400 runtime_control_operation_invalid, nothing sent
prepare of write_file with content a\ud800b, or with a key k\ud800 the file holds 61 3f 62 0a; the key arrived as k? 400, nothing sent
a Shell call with directory in another storage ran there 400 managed_runtime_tool_invalid
the same directory, reached through .. or through a symbolic link inside the workspace ran there 400 managed_runtime_tool_invalid
the worker answers a cancellation with lastSeq written as 65540S, 260B or 1.0000000000000001D 200, read as a number 400 managed_runtime_attestation_invalid; the cancellation can be repeated and the Session released

A surrogate pair passes both paths unchanged and is written as the four bytes of its code point. A Shell call with a directory inside the workspace runs there. A cursor written as 4 is accepted. On the raw path the call with a directory outside the workspace does not run, before and after.

The cursor forms are not something a worker of this repository writes; the proxy of the rig rewrote the worker's answer to produce them.

When the trial merge was made, main was six commits further than the head held (#12987, #12990, #13025, #13021, #13022, #12945). None touches the Broker, the provider or its worker. The trial merge was clean, and on it the probes of all rounds give the results of the head.

5. Lost answers, moved links, a refused start, artifacts (figure 6)

The later commits

A lost answer settles (Y6). The proxy drops the worker's answer to execute. The command ran once. The start answers 200 executing; a read asks the worker and answers 200 settled/success; start sent again and cancel answer the settled result with nothing sent to the worker; release succeeds and the storage is free. One execute on the wire. Before 9cb9dc86e8 the read, the repeated start and the cancel answered 409, release answered 409 too, and the Session and its storage stayed held.

A link moved after prepare (Y7). The directory of a Shell call is a link to a directory inside the workspace. After prepare, reserve and preflight the link is pointed at the directory of another storage. On 50fb28301e the command ran there. On the head the call settles as an error and does not run.

A refused start (Y8) is R7-1, above. The figure shows it on the commit before, on 9cb9dc86e8, on the head, on main after the merge, and on main with candidate B.

Cancellation reaches the worker. On 9cb9dc86e8 the caller of a refused start had no way out: cancel answered 200 prepared and did nothing, the record stayed UNKNOWN, release answered 409, the storage stayed held. On the head cancel answers 200 settled/not_started, the record is settled, release answers 200. That closes F1 of round 1 on the provider path. On the raw path a tool v2 record keeps its UNKNOWN, as before.

Artifacts. The fitter is a function, so this probe calls it as built, without a stack. With artifacts of 1.5 MiB beside 1,000 characters for the model, 9cb9dc86e8 returned a result of 1,536 KiB that did not fit the limit of 1 MiB, with the model's text replaced by the stub and the hook result dropped. The head drops the artifacts and keeps the text and the hook result whole. Artifacts that fit are kept; a structured display goes before them; a result that already fits is returned unchanged. The fuzz of round 6 passes, 8/8.

6. Mutation matrix (figure 7)

Mutation matrix

200 mutants ran on this head: 184 killed, 16 survive.

57 are new: L1 to L35 edit lines e94f781523 adds, M1 to M3 the lines the merge wrote by hand, P1 to P9 the lines of 50fb28301e, Q1 to Q6 those of 9cb9dc86e8, R1 and R2 those of 760174073b, S1 and S2 those of ebea694e4e. 50 are killed. What the later commits add is pinned where it matters: a provider operation that is not checked, or checked for prepare only, or for the input only (P1 to P3); a cursor read the old way (P6); a directory outside the workspace, at prepare and before the run, and links not followed (P7, Q2, Q3, Q5); a lost answer that is not reconciled and a cancellation that does not reach the worker (Q1, S1, S2); artifacts that are not dropped (R1). 7 survive:

  • L11, a retirement that failed reports nothing. No test makes a retirement fail and looks at what is reported.
  • L20, a release retires at most one Session. No test has two released Sessions past the bound at one release.
  • L34, the response buffer grows to the needed size only. The bytes are the same; only the number of copies differs.
  • L35, the ownership check of a terminal record no longer compares the generation. No test reads a terminal record of another generation through this check.
  • P5, a status cursor beyond 2^53 - 1 is accepted. No test sends a cursor beyond the largest safe integer.
  • Q4, a relative directory is admitted. Equivalent: core refuses a relative directory before this line is reached.
  • Q6, a directory that cannot be resolved is admitted. No test names a directory that does not exist.

No mutant stands for R7-1: there the line is as the commit meant it, and the tests of the commit pin exactly that (state: executing). The case the tests do not have is the worker that answers prepared; the test of candidate B is that case.

Of the 150 earlier mutants, seven edit lines that no longer exist (I9, I23, I24, I25, I26, I28, I2). The other 143 were rerun: 134 killed, 9 survive (H3, H8, I6, K36, K39, N17, T18, T2, T8), the same nine as on e94f781523. Three verdicts changed against round 6, all to killed: K7 and K41, the rule admits letters outside ASCII, and G4, the guard of the round-3 merge.

TypeScript mutants of the provider protocol and its worker ran against the eleven test files that exercise them; every survivor was then confirmed against all twenty. Failures of hosted-harness-session.test.ts are not counted as kills. The TypeScript mutants ran on 760174073b, whose TypeScript sources and tests are the ones of the head; the Java mutants ran on the head. The host was under load while the matrix ran, so every kill was compared with the tests that killed the same mutant on e94f781523.

Status of earlier items

Item Status on this head
R4-1, F3: a worker keeps released Sessions Closed on the provider path (section 1). The raw path is unchanged, on main too
F1 of round 1, definite refusal → UNKNOWN → Session and storage pinned Closed on the provider path (section 5): cancel settles the record. The start that precedes it is R7-1
Round 6: the 503 after a restart with the default provisioner Unchanged; named in the update
Round 6: K7, K41 Closed: both are killed
G4, the untested guard of the merge (round 3), R4-7 Closed: the mutant is killed
Reason over 4,096 characters or with NUL loses the code Deferred. Unchanged: 4 of 11
Confirmation details of a write over a large file are refused Recorded by the author as follow-up. Unchanged and harmless (X6, 2/2)

The merge

Merged as 1f1209bc70 on 29 September, 16:37 UTC, by squash onto main 8c914ebe03. The trial merge of this round was made with main be1ebc74d7, two commits earlier; it was clean and ran as a second arm. main is at a63157304a, one commit after the merge (#12955); Y1 to Y8 and the probes of rounds 1 to 6 were run on it as well.

Not covered

Windows; MariaDB; a live-model Hosted Shell turn (the ITs cover the Shell turn with a fake model); the 501 branch of a refused acquire; an ABANDONED record on the real stack; more than eight runtimes alive while retirements are in flight (the author records it); a retirement that hangs; an unpaired surrogate in modification, in the payload of confirm or in bind-history on the real stack (the unit test of 50fb28301e covers them); artifacts from a real tool on the real stack (the probe calls the fitter as built).

Reproduce

Round-7 rig, probes, mutants and logs: pr12868/r7. New probes are s21-r7.mjs (Z1 to Z4), s22-loss-in-flight.mjs, s23-merge-r7.mjs (Y1 to Y8) and fit-artifacts.mjs; s3-retention.mjs and s3b-retention-raw.mjs measure the memory; candidate-B.patch is the candidate.

中文版

真实环境验证,第七轮:ebea694e4e,已合入为 1f1209bc70

本轮进行期间 PR 已经合入。本轮有两个发现现在存在于 main 上,需要后续修复:R7-1,以及较小的 R7-2。 两者都由 9cb9dc86e8 引入,head 上有,main a63157304a 上也有。R7-1:worker 拒绝执行的调用,调用方的 start 永远不返回。目前没有任何生产代码调用 provider,所以今天在跑的东西不受影响。在 Broker 里加一个条件就能消除(见下面的候选 B);加上它之后七轮的全部探针都通过,补丁也能直接应用到 main。本轮测量的其余内容都没有问题。

squash 提交里的文件与 head 逐字节相同,只有 config.ts 例外——main 自己也改过它。在合入之后又多一个提交的 main a63157304a 上,R7-1 里的 start 在 20 秒后仍无应答,期间 worker 被问了 335 次;加上候选 B 之后它在 0.1 秒内被拒绝,Y1 到 Y8 为 21/21,Broker 套件 514 个测试通过。

本 head 是第六轮的 head 加七个提交。本轮进行期间分支变动了五次;每个状态都各自构建并实测过,下面的数字除非写明是别的构建,否则都来自 head。已有 2 位评审人批准了这个 head;96 个评审 thread 全部已解决;PR 上的评审结论是 APPROVED。这些评审都没有提到 R7-1。前几轮:第一轮、第二轮、第三轮、第四轮、第五轮、第六轮。

R7-1:worker 拒绝的 start 永不返回

一个调用已经 prepare 并预留,调用方在 preflight 放行之前就 start。worker 以 409 拒绝执行,这是正确的。

50fb28301e(之前) head,以及合入之后的 main
调用方的 startExecution 0.1 秒后被拒绝,409 runtime_broker_execution_unknown 没有应答;探针 20 秒后停止时仍在等待。在 9cb9dc86e8 上,一个不设上限的探针等了 4 分 26 秒,直到被手动停止
这期间发给 worker 的请求 1 次 execute 1 次 execute,随后只要 provider 还活着就以每秒约 17 次的频率发 status:20 秒内 314 次,那 4 分 26 秒里 4,549 次
直接问 Broker start 409,read 409 start 200 prepared,read 200 prepared

从 9cb9dc86e8 起,对于记录为 UNKNOWN 的 provider 执行,Broker 会去问 worker,并以 200 返回 worker 报告的状态。对于应答丢失的情形这是对的:worker 会说 executing 或 settled。但对于被拒绝的 start,worker 说的是 prepared。这个调用永远不会被启动,因为 Broker 从不二次派发;而 provider 客户端每 50 毫秒轮询一次、直到状态变成 settled,于是它会无止境地轮询下去。只有销毁 provider 才能让它停下。

如果 Hosted 的某个回合在调用尚未获准时就 start,发出的正是同样的请求序列。provider 目前没有生产调用方,所以今天在跑的东西不受影响。但在这个提交之前,这种错误会在十分之一秒内以失败告终;现在它会挂住整个回合,并给 worker 持续加压。它也影响到了更早的探针:第一轮的 E.2 得不到应答;B 组和 C 组里统计发往 worker 请求数的检查,会数到另一个不相干调用的轮询请求(第 3 节)。

ebea694e4e 缓解了其中一个后果,但没有消除原因:调用方现在发起的 cancel 能到达 worker,记录会结算为 not_started,Session 可以释放。但 start 本身仍然不返回。

候选 B 是 RuntimeBrokerHttpServer.observedExecutionEnvelope 里的一个条件:worker 仍然标记为 prepared 的 provider 执行其实从未启动,因此 Broker 保持它原来的应答 409 runtime_broker_execution_unknown。在本 head 上加上它之后:start 在 0.1 秒 后以 409 被拒绝,worker 只被问一次而不是每秒 17 次;Y1 到 Y8 为 21/21,第一到第六轮的全部探针通过,Broker 套件 514 个测试通过,Checkstyle 干净。去掉这个条件,它的测试就会失败。应答丢失的情形仍然能结算(Y6),ebea694e4e 的取消也仍然有效。补丁见英文部分的折叠块。

R7-2:worker 丢失之后,inspection 从回答“unknown”变成了抛异常

这个问题较小,来自同一个提交,可以放在同一个后续修复里。命令执行过程中 worker 被杀。记录和之前一样是 UNKNOWN,没有任何调用被启动两次,存储也一直被占。但从 9cb9dc86e8 起,调用方得到的应答变了:

50fb28301e(之前) head,以及合入之后的 main
调用方正在等待的那个 start 409 runtime_broker_execution_unknown 409 runtime_execution_evidence_unavailable
读取这次执行 409 runtime_broker_execution_unknown 409 runtime_admission_closed,“Runtime generation is no longer live”
provider 的 inspectExecution {"outcome":"unknown"} 抛出 409 runtime_admission_closed,不可重试

Broker 现在会先尝试去问 worker,发现它已经不在了,于是把这次尝试的错误作为应答返回。而 provider 只会把 runtime_broker_execution_unknown 转成 {outcome: 'unknown'},所以靠 inspection 做恢复的调用方,原来能拿到一个答案,现在拿到的是异常。栈上的状态和之前一样安全;变的是这个调用方需要处理的东西。这一项没有给候选补丁:该改的是 Broker 的应答还是客户端的解读,属于设计上的取舍。

本轮的其余内容

第六轮之后仍未解决的 本 head
R4-1,即第一轮的 F3:worker 保留已释放 Session 持有的内容 在本 PR 新增的路径上已关闭。 之前 300 个 Session 之后 worker 大了 286 MiB;现在 300 个之后不变大,1,000 个之后也不变大(第 1 节)
第一轮的 F1:明确的拒绝会让记录停在 UNKNOWN,Session 和存储被钉住 ebea694e4e 在 provider 路径上关闭了它。 调用方取消后记录结算为 not_started,release 返回 200,存储被释放(第 5 节)
K7 和 K41:没有测试用到带非 ASCII 字母的 ID 已关闭。 两边的白名单测试现在都包含非 ASCII 的字母和数字;变异体 K7 和 K41 被杀掉了(第 6 节)
G4,即 R4-7:第三轮合并时手写的守卫没有测试 已关闭。 G4 被 repeatedCancellationOfADispatchedProviderExecutionReturnsItsReceipt 杀掉(第 6 节)
Broker 重启之后,默认 provisioner 下 503 永不结束 没有变化;更新说明现在已经写明。配合 #12865 的持久 provisioner 可以结束(第 3 节)
合入 main 之后出现、由 50fb28301e 关闭的 合并提交 本 head
含未配对代理项的工具输入 raw 路径被拒绝;provider 路径把该字符换成 ? 后执行了 两条路径都被拒绝
Shell 调用的目录在 Session 工作区之外 在另一个存储的目录里执行了 被拒绝,经由 .. 或符号链接也一样
状态游标写成了不是精确整数的数字形式 被接受 被拒绝
之后几个提交新增的 之前 本 head
9cb9dc86e8:execute 的应答丢失 记录永久停在 UNKNOWN,Session 和存储被钉住 Broker 去问 worker,记录结算为 success;只派发一次、只生效一次
9cb9dc86e8:prepare 之后符号链接被改指到工作区之外 命令在工作区之外执行了 调用以错误结算,没有执行
760174073b:放不下的 artifacts 结果超出上限,给模型的文本被换成占位文字 artifacts 被丢弃,给模型的文本完整保留,结果不超限
ebea694e4e:取消一个应答丢失或被拒绝的调用 409,没有任何请求发往 worker 200,会去问 worker,记录得到结算

除 R7-1 和 R7-2 之外还剩下的事项。都不紧急,也都不是本 PR 带来的:

  1. raw 路径上每个已释放的 Session 大约留下 325 KiB,在 main 上也一样。 这是 Hosted harness 目前使用的路径,本 PR 没有改动它:这次修复限制的是 provider 路由保留的内容,executor 的日志不在改动范围内。它值得单独开一个 issue(第 1 节)。
  2. 200 个变异体里有 16 个存活,其中 7 个是新增的。新增的存活者都落在退役、响应缓冲区、游标和目录检查的边角上,没有一个对应真实栈上观察到被打破的规则(第 6 节)。
  3. 随 main 带进来的一个故障门禁在高负载的宿主机上会失败,在 main 上也一样。 在有负载的宿主机上单独运行 DurableLocalRuntimeFaultGateTest,head 上 3 次通过 2 次,main 上 3 次通过 1 次;两个构建同时跑则每次都失败。整个 profile 在 head 上第一次运行就 44 个全部通过(第 3 节)。

装置与之前相同:构建出来的 TypeScript provider → Spring 内嵌 Broker(MySQL 8.4.7)→ 打包 worker,Broker 与 worker 之间挂着 JVM HTTP 代理;合入之前与 main be1ebc74d7 做的试合并作为第二个对照臂;合入之后的 main a63157304a 作为第三个;候选 B 分别加在 head 和 main 上;用 Linux 容器测试持久 provisioner。

跑了什么 结果
provider 路径,释放 300 个 Session 之后的 worker 内存,每回合写 200 KiB 上一个 head:177 → 463 MiB · 本 head:179 → 148 MiB
同上,1,000 个 Session 本 head:179 → 124 MiB
同上,raw 路径 本 head:165 → 262 MiB · main:162 → 257 MiB
worker 上一次 release 的耗时 上一个 head:中位 1 ms,最慢 24 ms · 本 head,1,300 次释放:中位 1 ms,最慢 36 ms
已释放的 Session 还能回答什么(Z1),两种 boot 版本 上一个 head:12 个全部完整回答 · 本 head:8 个完整回答,4 个墓碑,8/8
其他 Session 退役时,存活的 Session(Z2) 4/4:它准备好的调用正常执行,正在运行的命令正常结束
同一个 worker 上 32 个 Session,每次并行 6 个(Z3) 3/3:32 个全部成功
用白名单之外的 Harness Session ID 调用 warm(Z4) 上一个 head:可重试的 503 · 本 head:400,4/4
main 的规则、应答丢失、被拒绝的 start(Y1 到 Y8) 本 head、试合并、合入之后的 main:19/21,失败的两条就是 R7-1 · 加上候选 B(head 和 main 上):21/21
第一到第六轮的探针 本 head:75/76、10/12、7/7、13/13、22/22、30/30、24/24、10/10、9/9、6/6;失败的都源于 R7-1 · 合入之后的 main:75/76、10/12,其余全过 · 加上候选 B(head 和 main 上):76/76、12/12,其余全过
Broker 被替换,持久 provisioner,Linux 4 个场景全部 5/5:acquire 0.1 到 0.2 秒完成,取消被确认
命令执行中 worker 被杀,macOS 与 Linux 各 4/4:调用没有被再次启动,在确认旧写入方已停止之前存储一直被占
Hosted 循环 仓库自带 IT 在 MySQL 8.4.7 上通过 · 真实模型 3 轮:7 次执行全部结算,6 次 success、1 次 error(模型在写文件之前先去读了那个文件)
仓库套件 本 head:Broker 513、故障门禁 44、Broker MySQL IT 3、server 188 + 15 个 IT、Hosted IT 10、core 194 · 试合并:相同,Hosted IT 为 11 个 · 全部通过,Checkstyle 干净
变异矩阵 在本 head 上实跑 200 个变异体:杀掉 184 个,存活 16 个。其中 57 个是新增的:杀掉 50 个,存活 7 个

head 里有什么

见图 1。

1. 释放之后 worker 保留了什么

见图 2。一个 worker 上依次跑 Session,每个 Session 执行一次 200 KiB 的 write_file 然后释放,与 Hosted 的一个回合相同。内存取的是 worker 进程的常驻内存,分别在第一次释放时和最后一次释放 15 秒之后。

路径 构建 Session 数 worker RSS 每个已释放的 Session
provider 上一个 head 300 177 → 463 MiB 980 KiB
provider e94f781523,只有修复 300 178 → 149 MiB 无
provider 本 head 300 179 → 148 MiB 无
provider 本 head 1,000 179 → 124 MiB 无
raw 本 head 300 165 → 262 MiB 332 KiB
raw main 12793013c4 300 162 → 257 MiB 325 KiB
更新说明里的声明 这里的实测
worker 最多完整保留八个已释放的 Session 释放 12 个 Session 之后,最后 8 个对 status、cancel、history 都能完整回答(第 2 节)
最早的、且没有正在被观察的那个会被退役 最早的 4 个变成了墓碑
任何一次释放之后,被保留的已释放 Session 不超过八个 跑 1,000 个 Session 内存没有增长:运行期间的采样在 156 到 231 MiB 之间,最后一次释放 15 秒之后 worker 占 124 MiB
release 要等它触发的退役完成之后才应答 看不出额外代价:1,300 次释放,中位 1 ms,99 分位 8 ms,最慢 36 ms
墓碑仍然随每个已释放的 Session 增加一个 在 1,000 个 Session 的规模下看不出来

raw 路径那两行在本 head 和 main 上是一样的。这部分留存早于本 PR,存在于 executor 里,这次修复没有碰它。它比 provider 那几行更值得关注,因为 Hosted harness 目前跑的就是 raw 路径,而 provider 还没有生产调用方。

2. 已释放的 Session 还能回答什么

见图 3。

在 worker 上对已释放的 Session 发起 最后八个之一 更早的
对它的调用做 status settled,内容完整 unknown
对它的调用做 cancel settled unknown
history 200 409
prepare 新的工作 409 409

Broker 不会向已释放的 Session 发送这些请求;这些是装置带着 Broker 的请求头直接发的。经过 Broker 时没有任何变化:已退役 Session 的执行记录仍然从 Broker 自己的日志里读出 200 settled/success,重复的 release 返回 200。

退役会关闭被退役 Session 的 Config,但不会波及同一个 worker 上的其他 Session:

  • 最先 acquire 的那个 Session,带着一个已经 prepare 并批准的调用;在另外十二个 Session 被释放、其中四个被退役之后,它照常执行了这个调用,随后又执行了一个。
  • 在这十二个 Session 来来去去的过程中一直运行的一条 Shell 命令正常结束,输出完整。
  • 32 个 Session 每次并行 6 个,全部执行并释放;worker 没有返回任何拒绝。

boot v2 同一个 Workspace 同一时刻只允许一个工具回合(409 workspace_busy),所以后两项是在 boot v1 上跑的。

warm 现在会用 400 runtime_broker_invalid_request 拒绝白名单之外的 Harness Session ID,而且发生在启动任何 worker 之前。上一个 head 返回的是可重试的 503 runtime_scope_resolution_failed。

3. 回归、重启与套件

见图 4。

回归。 在本 head、试合并以及合入之后的 main 上,第一到第六轮的探针除了被 R7-1 波及的之外全部通过。E.2 对一个被 worker 拒绝的调用发起 start,30 秒内得不到应答,boot v1 和 boot v2 都一样。第一轮里统计“这段时间内到达 worker 的请求数”的检查,会数到 B.1 先前 start 的那个调用仍在发出的 status 轮询:在 head、试合并和 main 上的这几次运行里撞上的是 C.5;具体撞上哪一条,每次运行不一样。加上候选 B 之后它们全部通过,head 和 main 上都是。

有三条检查陈述的是被 9cb9dc86e8 有意改变的规则,已改写为新规则:第一轮的 F.1、F.2 和第二轮的 S.exec 原先断言“应答丢失的调用其记录保持 UNKNOWN”;现在它会按 worker 保留的结果结算。本轮更早几个构建的日志里保留的是旧措辞。

重启。 与第六轮相同。默认 provisioner 下,重复的取消返回可重试的 503,acquire 在 120 秒后返回 503,如此反复。在 Linux 上用持久 provisioner(head 现在已经从 main 得到它),acquire 在 0.1 到 0.2 秒内接管 worker,取消得到确认:4 个场景全部 5/5。

命令执行中 worker 被杀。 在两套装置上都是:记录变成 UNKNOWN,调用方等待的那个调用返回 409,再次发送 start 被拒绝,release 返回 503,直到 Broker 确认旧的写入方已经停止;存储一直被占。worker 已经不在了,没有对象可问:新的对账逻辑改变的是应答里的错误码(见 R7-2),不是状态。客户端的错误信息里带着 Broker 给出的原因。两套装置上的响应都不带 details,所以真实栈上没有产生终态的 ABANDONED 响应,R4-3 目前依靠它的单元测试和对应的变异体(第 6 节)。在 Linux 上,命令自己的进程比 worker 活得久,并且把它的工作完成了一次;Broker 拒绝放开存储与此相符。

套件与 CI。 仓库套件在 head 和试合并上都通过(图 4)。head 的 CI:22 项检查通过,没有失败的;写这份报告时 assign 还在排队。仓库里没有任何套件会对一个被 worker 拒绝的调用发起 start 并等待它,所以 R7-1 在这些套件里都不会暴露。

有一个故障门禁依赖宿主机的负载。 DurableLocalRuntimeFaultGateTest 是 #12865 带来的,本 PR 没有改动它。在本轮其他运行给宿主机加负载的时候,fault-gate profile 在分支的两个更早状态上都在这个类里失败过(44 个里分别失败 1 个和 3 个);在 head 和试合并上它都是第一次运行就 44 个全部通过。在十核机器、负载均值 13 到 54 的条件下单独反复运行这个类:head 上 3 次通过 2 次,main 12793013c4 上 3 次通过 1 次,#12975 之前的 main 上 4 次通过 1 次;两个构建同时运行时每次都失败,分支和 main 都一样。失败的始终是同样两个断言:absentWorkerAfterRestartLeavesItsDomainPinned 期望 runtime_broker_runtime_lost 却得到 runtime_provision_fenced,adoptedWorkerCanCancelItsOriginalActiveCall 没有看到取消请求到达 worker。有一次运行给出了原因:Runtime recovery claim expired。

这也回答了评审机器人留下的一个问题。CI 里 Hosted process fault gates 这个 job 在本 head 上失败过一次,失败在同一个类、同一个错误上(closingAnObserverCannotKillTheSharedWorker、503 runtime_provision_fenced、Runtime recovery claim expired),重跑之后通过。机器人不愿把它当成偶发失败,理由是上一次绿色运行之后合入了 760174073b 和 ebea694e4e。而在上面的运行里,同一个错误 runtime_provision_fenced 在不含这两个提交的 main 上、以及 #12975 之前的 main 上,也会让同一个类里的其他测试失败。它是繁忙宿主机上过期的一个租约,与这两个提交无关。CI 点名的那个测试在这里没有失败过。

4. 与 main 的合并,以及 main 期间新增的规则

见图 5。

自本分支上一次合入以来,main 前进了 38 个提交。有三个文件冲突:两份设计文档和 RuntimeBrokerService.java。在后者里,#12975 修改了 deferred start,而本 PR 之前已经把这段代码挪进了自己的一个分支;合并时是手工把 #12975 的两处检查搬进这个分支的。没有重复的迁移版本号。

手写的部分是对的。 deferred start 的 payload 里只要含有未配对代理项——无论在输入的字符串里、输入的键里还是工具名里——head 上都返回 400 runtime_payload_invalid,与 main 一致,并且不会向 worker 发送任何请求。合并之前这三种都会被发出去,代理项被换成 ?。被拒绝的 start 对应的预留会一直保留到调用方取消它;取消之后 Session 可以释放,存储也被释放。这几行上的三个变异体(M1 到 M3)都被 #12975 带来的测试杀掉。

合并带不过来的部分。 provider 路径是本分支自己的;main 给 raw 路径加的规则不会因为合并而自动落到它上面。50fb28301e 列出了三条并补上。每一条都分别在合并提交 174f974e3e 和 head 上实测过:

经由 provider 174f974e3e,合并提交 本 head
prepare 一条 ls data-\ud800.txt worker 收到的是 ls data-?.txt,执行了它,并列出了 data-1.txt 和 data-2.txt 400 runtime_control_operation_invalid,什么都没有发出
prepare 一个 write_file,内容是 a\ud800b,或带一个键 k\ud800 文件内容是 61 3f 62 0a;键到达时变成 k? 400,什么都没有发出
Shell 调用的 directory 在另一个存储里 在那里执行了 400 managed_runtime_tool_invalid
同一个目录,经由 .. 或经由工作区内的符号链接到达 在那里执行了 400 managed_runtime_tool_invalid
worker 回答取消请求时,把 lastSeq 写成 65540S、260B 或 1.0000000000000001D 200,被当成一个数字读了 400 managed_runtime_attestation_invalid;取消可以重发,Session 可以释放

成对的代理项在两条路径上都原样通过,并按其码点写成四个字节。目录在工作区之内的 Shell 调用在那里正常执行。写成 4 的游标被接受。在 raw 路径上,目录在工作区之外的调用前后都不会执行。

这些游标形式并不是本仓库的 worker 会写出来的;它们是装置的代理改写 worker 的响应制造出来的。

做试合并的时候,main 比 head 已含的那个多六个提交(#12987、#12990、#13025、#13021、#13022、#12945),都不涉及 Broker、provider 及其 worker。试合并没有冲突,在它上面所有轮次的探针给出的结果都与 head 一致。

5. 应答丢失、链接改指、被拒绝的 start、artifacts

见图 6。

丢失的应答能结算(Y6)。 代理丢掉 worker 对 execute 的应答。命令执行了一次。start 返回 200 executing;read 会去问 worker 并返回 200 settled/success;再次发送 start 和 cancel 都返回已结算的结果,并且不向 worker 发送任何请求;release 成功,存储被释放。链路上只有一次 execute。在 9cb9dc86e8 之前,read、再次发送的 start 和 cancel 都返回 409,release 也返回 409,Session 和存储一直被占。

prepare 之后链接被改指(Y7)。 Shell 调用的目录是一个指向工作区内某目录的符号链接。在 prepare、预留和 preflight 之后,把链接改指到另一个存储的目录。在 50fb28301e 上,命令在那里执行了。在 head 上,调用以错误结算,没有执行。

被拒绝的 start(Y8) 就是上面的 R7-1。图里依次是:之前的提交、9cb9dc86e8、head、合入之后的 main,以及 main 加候选 B。

取消能到达 worker。 在 9cb9dc86e8 上,被拒绝的 start 的调用方没有出路:cancel 返回 200 prepared 却什么都没做,记录停在 UNKNOWN,release 返回 409,存储被占。在 head 上,cancel 返回 200 settled/not_started,记录得到结算,release 返回 200。这在 provider 路径上关闭了第一轮的 F1。在 raw 路径上,tool v2 的记录仍然保持 UNKNOWN,与之前一样。

artifacts。 裁剪逻辑是一个函数,所以这个探针直接调用构建产物里的函数,不需要起栈。artifacts 为 1.5 MiB、给模型的文本为 1,000 个字符时,9cb9dc86e8 返回的结果是 1,536 KiB,超出了 1 MiB 的上限,而且给模型的文本被换成了占位文字,hook 结果被丢弃。head 丢弃的是 artifacts,文本和 hook 结果完整保留。放得下的 artifacts 会被保留;结构化的显示内容先于它们被处理;本来就不超限的结果原样返回。第六轮的 fuzz 通过,8/8。

6. 变异矩阵

见图 7。在本 head 上实跑了 200 个变异体:杀掉 184 个,存活 16 个。

其中 57 个是新增的:L1 到 L35 改的是 e94f781523 新增的行,M1 到 M3 改的是合并时手写的行,P1 到 P9 对应 50fb28301e,Q1 到 Q6 对应 9cb9dc86e8,R1、R2 对应 760174073b,S1、S2 对应 ebea694e4e。杀掉 50 个。后面几个提交新增的内容在关键处都被测试钉住了:provider operation 不做检查、只检查 prepare、只检查输入(P1 到 P3);游标按旧方式读取(P6);目录在工作区之外——prepare 时、执行前,以及不跟随符号链接(P7、Q2、Q3、Q5);丢失的应答不对账、取消到不了 worker(Q1、S1、S2);artifacts 不被丢弃(R1)。存活 7 个:

  • L11,退役失败时不上报。没有测试让退役失败并检查上报的内容。
  • L20,一次释放最多只退役一个 Session。没有测试在一次释放时让两个已释放的 Session 同时超出上限。
  • L34,响应缓冲区只按需增长、不再翻倍。结果的字节相同,只是复制次数不同。
  • L35,终态记录的归属检查不再比较 generation。没有测试经由这个检查去读另一个 generation 的终态记录。
  • P5,接受超过 2^53 - 1 的状态游标。没有测试发送超过最大安全整数的游标。
  • Q4,放行相对路径的目录。等价变异体:core 会在走到这一行之前就拒绝相对路径。
  • Q6,放行无法解析的目录。没有测试使用不存在的目录。

R7-1 没有对应的变异体:那一行代码正是提交想要的样子,提交自带的测试钉住的也正是这一点(state: executing)。测试里缺的是 worker 回答 prepared 的情形;候选 B 的测试补的就是它。

此前的 150 个变异体里,有 7 个改的行已经不存在了(I9、I23、I24、I25、I26、I28、I2)。其余 143 个重跑了一遍:杀掉 134 个,存活 9 个(H3、H8、I6、K36、K39、N17、T18、T2、T8),与 e94f781523 上是同样的九个。相对第六轮有三个判定发生变化,都变成了被杀:K7 和 K41(规则放行非 ASCII 字母),以及 G4(第三轮合并时的守卫)。

provider 协议及其 worker 上的 TypeScript 变异体,先用真正覆盖它们的 11 个测试文件跑;每个存活者随后都用全部 20 个文件复核过。hosted-harness-session.test.ts 的失败不计为击杀。TypeScript 变异体是在 760174073b 上跑的,它的 TypeScript 源码和测试与 head 完全相同;Java 变异体是在 head 上跑的。矩阵运行期间宿主机有负载,所以每一个击杀都与 e94f781523 上杀掉同一个变异体的测试做了比对。

前几轮各项的现状

条目 在本 head 上的状态
R4-1,即 F3:worker 保留已释放的 Session 在 provider 路径上已关闭(第 1 节)。raw 路径没有变化,main 上也一样
第一轮的 F1:明确拒绝 → UNKNOWN → Session 和存储被钉住 在 provider 路径上已关闭(第 5 节):cancel 能让记录结算。它之前的那个 start 就是 R7-1
第六轮:默认 provisioner 下重启之后的 503 没有变化;更新说明里已写明
第六轮:K7、K41 已关闭:两个都被杀掉
G4,合并时手写、没有测试的守卫(第三轮),即 R4-7 已关闭:这个变异体被杀掉
原因超过 4,096 字符或含 NUL 时丢失错误码 延后处理。无变化:11 个里有 4 个
覆盖大文件时的确认详情被拒绝 作者记录为后续事项。无变化,且无害(X6,2/2)

合入情况

9 月 29 日 16:37(UTC)以 squash 方式合入,提交为 1f1209bc70,基于 main 8c914ebe03。本轮的试合并用的是更早两个提交的 main be1ebc74d7;它没有冲突,并且作为第二个对照臂实跑过。main 现在在 a63157304a,即合入之后又多了一个提交(#12955);Y1 到 Y8 以及第一到第六轮的探针也在它上面跑过。

未覆盖

Windows;MariaDB;真实模型驱动的 Hosted Shell 回合(IT 用假模型覆盖了 Shell 回合);acquire 被拒的 501 分支;真实栈上的 ABANDONED 记录;退役进行中时同时存活的 runtime 超过八个的情形(作者已记录);退役挂住不返回的情形;真实栈上 modification、confirm 的 payload 或 bind-history 里的未配对代理项(50fb28301e 的单元测试覆盖了它们);真实栈上由真实工具产生的 artifacts(探针直接调用构建产物里的裁剪函数)。

复现

第七轮的装置、探针、变异体和日志:pr12868/r7。新增探针是 s21-r7.mjs(Z1 到 Z4)、s22-loss-in-flight.mjs、s23-merge-r7.mjs(Y1 到 Y8)和 fit-artifacts.mjs;内存由 s3-retention.mjs 和 s3b-retention-raw.mjs 测量;candidate-B.patch 是候选补丁。

qwen-code-dev-bot added a commit that referenced this pull request Sep 29, 2026
… delivery

Merge origin/main into the durable remote Shell result delivery branch.

main's generic Broker provider controls (#12868) and this branch's durable
Tool v3 publication both rewrote the same prepared-execution and start
validation, so the resolutions are semantic rather than adjacent:

- The HTTP field allow-lists stay exact, and now admit toolProtocol and
  publicationId on a prepared execution, plus publicationId and
  publicationToken on a start that carries a payload.
- prepareExecution keeps main's immutable-copy and provider-reference
  structure, including its stage-channel rejection of a null-valued
  reference, with the deferred_v3 reservation branch inside the raw-tool
  path.
- startExecution keeps main's session lock and readiness re-check around
  the synchronous validation; only the publication-install continuation
  runs after the lock is released, and beginDispatch stays idempotent.
- Cancellation keeps main's prepared-provider route and this branch's
  deferred_v3 flag, and short-circuits an UNKNOWN record on both
  observableAfterLoss() and the durable reservation.
- The managed context worker registers main's provider route with this
  branch's per-call publisher selection, and both new tests on each side
  are kept.
he-yufeng pushed a commit to he-yufeng/qwen-code that referenced this pull request Sep 30, 2026
…ad of prepared (QwenLM#13064)

* fix(runtime-broker): answer a refused provider start as unknown

Since 9cb9dc8 (QwenLM#12868), the Runtime Broker reconciled an UNKNOWN
provider record against the worker and answered 200 with the observed
state prepared for a start the worker had refused. Nothing dispatches
such a call a second time, so the provider client's startExecution
polled every 50 ms without an end. Keep a provider execution the worker
still holds as prepared answered as 409 runtime_broker_execution_unknown,
as before the change; lost answers keep settling from the worker's
retained result.

Fixes QwenLM#13059

* test(runtime-broker): pin the cancel answer for a refused provider start

The neverStarted guard serves the :cancel route through the same
envelope as :start and the read; cover that combination: an UNKNOWN
provider record whose worker still answers prepared is answered 409
runtime_broker_execution_unknown, while the physical cancellation still
reaches the worker. Both mutations redden the assertion: removing the
guard, and narrowing it to the start route.

Refs QwenLM#13059

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Partially reviewed — gaps disclosed.

Not reviewed: verification (Step 4) — the verifier agent stalled on every launch attempt across ~2h (the provider subagent lane, including single-agent waves and after the worktree repair); all 21 findings remain unverified and terminal-only.

Not reviewed: reverse audit of chunk 14 (round 2) — the auditor stalled on every attempt; chunk 14's Step 3 territory review and its round-1 audit did run and returned substantive receipts.

Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) and the Test (macos/windows) lanes were skipped in CI, and the Java suites (packages/sdk-java) never ran locally (no mvn/JDK on this machine), so no execution evidence covers the Java half of the diff.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 2)": the worker's ManagedToolRuntime construction site was not read — bindMediaTool being unset rests on an all-packages identifier grep (3 hits, none an assignm…; "agent reverse-audit (round 2)": the zh-CN mirror of the mediaContext sentence (2026-09-27-broker-provider-control.zh-CN.md, ~L70) was not located or verified.; "agent reverse-audit (round 2)": doc claims left unverified against code: "acquire/release results are exactly true ", "every successful response repeats the version, protocol and Session", "B…; chunk 7: none — all checks completed; the vitest run was environment-blocked, not budget-cut, and was replaced by source-level verification as described.; chunk 26: none — all checks completed within budget (10 of ~36 tool calls)., and 5 more.

Not reviewed: "agent verify" — the agent made no tool call: it read nothing.

Mechanism health: this round did not close cleanly, so it withholds the incremental anchor — and the round it recovered had no anchor this round could use either — none at all, one with no certifier, one certified by an identity other than the one this round runs under, or one this round's fetch refused or resolved to the head — so the next review re-reads the whole diff unless recovery grafts an earlier own anchor that the round running it can use onto the complete work list this round leaves behind, and keeps doing so until a round's marker carries an anchor again or a graft lands that the round running it can use. (Stated, not acted on — this changes nothing about what the round posts.)

中文说明

仅完成部分审查,审查缺口已披露。

未审查(原文为英文):verification (Step 4) — the verifier agent stalled on every launch attempt across ~2h (the provider subagent lane, including single-agent waves and after the worktree repair); all 21 findings remain unverified and terminal-only.

未审查(原文为英文):reverse audit of chunk 14 (round 2) — the auditor stalled on every attempt; chunk 14's Step 3 territory review and its round-1 audit did run and returned substantive receipts.

未审查(原文为英文):build-and-test — Integration Tests (CLI, No Sandbox) and the Test (macos/windows) lanes were skipped in CI, and the Java suites (packages/sdk-java) never ran locally (no mvn/JDK on this machine), so no execution evidence covers the Java half of the diff.

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 2)":the worker's ManagedToolRuntime construction site was not read — bindMediaTool being unset rests on an all-packages identifier grep (3 hits, none an assignm…;"agent reverse-audit (round 2)":the zh-CN mirror of the mediaContext sentence (2026-09-27-broker-provider-control.zh-CN.md, ~L70) was not located or verified.;"agent reverse-audit (round 2)":doc claims left unverified against code: "acquire/release results are exactly true ", "every successful response repeats the version, protocol and Session", "B…;chunk 7:none — all checks completed; the vitest run was environment-blocked, not budget-cut, and was replaced by source-level verification as described.;chunk 26:none — all checks completed within budget (10 of ~36 tool calls).,另有 5 条。

未审查:"agent verify"——该 agent 未发起任何工具调用:它什么都没读。

机制健康:本轮未能干净收尾,因而扣留了增量锚点,而它恢复到的那一轮也没有留下本轮可用的锚点——要么完全没有、要么没有认证者、要么由本轮运行身份之外的身份认证、要么被本轮的获取拒绝或解析为头提交——因此下一次评审将重读整个 diff,除非恢复流程把本轮能使用的更早自有锚点嫁接到本轮留下的完整工作清单上;并会一直如此,直到某一轮的标记重新带上锚点,或落地的嫁接能被运行该轮的评审使用。(仅陈述,不据此行动——这不改变本轮发布的任何内容。)

— glm-5.3-flash via Qwen Code /review (v0.24.6)

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Partially reviewed — gaps disclosed.

Not reviewed: verification (Step 4) — the verifier agent stalled on every launch attempt across ~2h (the provider subagent lane, including single-agent waves and after the worktree repair); all 21 findings remain unverified and terminal-only.

Not reviewed: reverse audit of chunk 14 (round 2) — the auditor stalled on every attempt; chunk 14's Step 3 territory review and its round-1 audit did run and returned substantive receipts.

Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) and the Test (macos/windows) lanes were skipped in CI, and the Java suites (packages/sdk-java) never ran locally (no mvn/JDK on this machine), so no execution evidence covers the Java half of the diff.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 2)": the worker's ManagedToolRuntime construction site was not read — bindMediaTool being unset rests on an all-packages identifier grep (3 hits, none an assignm…; "agent reverse-audit (round 2)": the zh-CN mirror of the mediaContext sentence (2026-09-27-broker-provider-control.zh-CN.md, ~L70) was not located or verified.; "agent reverse-audit (round 2)": doc claims left unverified against code: "acquire/release results are exactly true ", "every successful response repeats the version, protocol and Session", "B…; chunk 7: none — all checks completed; the vitest run was environment-blocked, not budget-cut, and was replaced by source-level verification as described.; chunk 26: none — all checks completed within budget (10 of ~36 tool calls)., and 5 more.

Not reviewed: "agent verify" — the agent made no tool call: it read nothing.

Mechanism health: this round did not close cleanly, so it withholds the incremental anchor — and the round it recovered had no anchor this round could use either — none at all, one with no certifier, one certified by an identity other than the one this round runs under, or one this round's fetch refused or resolved to the head — so the next review re-reads the whole diff unless recovery grafts an earlier own anchor that the round running it can use onto the complete work list this round leaves behind, and keeps doing so until a round's marker carries an anchor again or a graft lands that the round running it can use. (Stated, not acted on — this changes nothing about what the round posts.)

中文说明

仅完成部分审查,审查缺口已披露。

未审查(原文为英文):verification (Step 4) — the verifier agent stalled on every launch attempt across ~2h (the provider subagent lane, including single-agent waves and after the worktree repair); all 21 findings remain unverified and terminal-only.

未审查(原文为英文):reverse audit of chunk 14 (round 2) — the auditor stalled on every attempt; chunk 14's Step 3 territory review and its round-1 audit did run and returned substantive receipts.

未审查(原文为英文):build-and-test — Integration Tests (CLI, No Sandbox) and the Test (macos/windows) lanes were skipped in CI, and the Java suites (packages/sdk-java) never ran locally (no mvn/JDK on this machine), so no execution evidence covers the Java half of the diff.

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 2)":the worker's ManagedToolRuntime construction site was not read — bindMediaTool being unset rests on an all-packages identifier grep (3 hits, none an assignm…;"agent reverse-audit (round 2)":the zh-CN mirror of the mediaContext sentence (2026-09-27-broker-provider-control.zh-CN.md, ~L70) was not located or verified.;"agent reverse-audit (round 2)":doc claims left unverified against code: "acquire/release results are exactly true ", "every successful response repeats the version, protocol and Session", "B…;chunk 7:none — all checks completed; the vitest run was environment-blocked, not budget-cut, and was replaced by source-level verification as described.;chunk 26:none — all checks completed within budget (10 of ~36 tool calls).,另有 5 条。

未审查:"agent verify"——该 agent 未发起任何工具调用:它什么都没读。

机制健康:本轮未能干净收尾,因而扣留了增量锚点,而它恢复到的那一轮也没有留下本轮可用的锚点——要么完全没有、要么没有认证者、要么由本轮运行身份之外的身份认证、要么被本轮的获取拒绝或解析为头提交——因此下一次评审将重读整个 diff,除非恢复流程把本轮能使用的更早自有锚点嫁接到本轮留下的完整工作清单上;并会一直如此,直到某一轮的标记重新带上锚点,或落地的嫁接能被运行该轮的评审使用。(仅陈述,不据此行动——这不改变本轮发布的任何内容。)

— glm-5.3-flash via Qwen Code /review (v0.24.6)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

review/self-reported The linked issue was opened by the PR author (self-reported)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Runtime Broker: HttpRuntimeTransport has no Session verbs, so no tool call can be dispatched through it

3 participants