Repository navigation
feat(managed-agent): Recover Workspace holders after trusted local reboot (W0e-3) - #12869
Conversation
W0e-3 E2E and regression reportEnvironment: macOS arm64, JDK 21, Node.js, bundled Runtime worker, isolated MySQL Community 8.4.11. No external model calls. The dedicated MySQL instance was shut down after validation; test processes were checked for leaks. Passing evidence
All final targeted runs had zero failures, errors or skips. Counts refer to distinct cases; targeted reruns are not added twice. This is evidence across the complete runs and affected-case reruns, not a claim that an initially failing invocation was green. What the new tests establish
Failures found and corrected during validationAn initial MySQL run reused a fixed-prefix database from earlier verification; its generation assertion correctly failed. A fresh isolated database passed. The new MySQL fixture initially bypassed Spring's transaction proxy; adding its explicit creation transaction made the real SQL path pass. The initial escaped-writer test checked the test root instead of its actual Workspace; the directory was corrected and the growth baseline moved after confirmed loss. The two affected gates then passed. The first Stage F run passed all 38 unchanged gates; only these two fixture cases required rerunning. Pending physical acceptanceA real Linux host reboot has not been run. macOS uses test-only host identity, the Spring worker chain uses a synthetic boot change, and MySQL tests use synthetic stop evidence. Native Linux identity and a real preserved-disk reboot with an escaped writer remain a separate acceptance gate. This PR remains draft until that gate can be performed. Hosted Turn continuation, public Workspace message enablement and hostile-tool containment are outside this slice. 中文本地完成 Broker 407 项、Spring 145 项常规测试,14 项 MySQL 集成用例,以及 38 项既有真实进程故障门禁。新逃逸写入者的 2 项门禁修复夹具目录后通过;另有 1 项 Spring+H2+真实 worker 恢复链路通过。更新后的维护断言和事务夹具做了定向补验,未重复计入数量。build、typecheck、bundle、Checkstyle 和改动文档格式检查均通过,最终补验无失败或跳过。 验证期间发现并修复了旧测试库复用、原始 MySQL store 夹具缺少事务、逃逸写入者测试目录错误,并将增长测量起点移至 Broker 确认失联之后。真实 SQL 验证了分批终态化、清理重试、迟到 acquire 屏障和保留新 holder;真实子进程验证了 worker 退出后继续写入时不能复用原位置。 未执行真实 Linux 整机重启。 模拟 boot、MySQL 测试证据和真实 worker 链路分别报告,不能合称真实重启验收。PR 保持 draft,待专用 Linux 环境验证后再推进。 |
3433bd8 to
7236322
Compare
66bde47 to
f8bf5d7
Compare
|
Please do not rebase or force-push to an active PR as it invalidates existing review comments. Note for future reference, the bots always squash all changes into a single commit automatically as part of the integration. 中文请勿对活跃的 PR 执行 rebase 或 force-push,因为这会使已有的评审评论失效。另外,供日后参考:作为集成流程的一部分,机器人始终会自动将所有改动压缩(squash)为单个提交。 |
|
Rebased this draft PR onto W0e-2 commit Validation of this exact stacked head: Runtime Broker 410 tests with 0 failures/errors and 1 conditional skip; managed agent server 145 tests with 0 failures/errors; 12 real-process reboot and durable-runtime fault gates passed; Checkstyle reported 0 violations. The earlier MySQL and broader fault-gate results remain recorded above; this rebase added only the W0e-1 Java fix and paired documentation, with no TypeScript changes. This PR stays draft until the physical Linux reboot acceptance gate and the documented pre-upgrade handling of historical seeded |
Real-host verification: the pending Linux reboot acceptance, run for realVerified on a dedicated Linux host at Summary
For the merge decision: from the real-host side the physical acceptance gate is met. I would fix F1 before merging (the change is small), add the two gap tests from section 5, and add the soft-reboot sentence. The pre-upgrade handling of historical All times are UTC, 2026-09-27. 1. Physical reboot acceptance
The JVM ran in 2. ControlsOne more control at
Enabling the option without durable mode stops the server at startup with 3. After the recovery4. F1: the scan refuses healthy requests, and a Hosted Session that hits it is blockedWhat happens:
The window is about 20 ms in every 5 s per binding, so a Session with a few hundred tool turns is likely to hit it. Nothing is wrong with the Runtime when it happens. Candidate (diff, test): when the observation ends with a healthy adoption, complete the reservation with that result instead of the 503. Waiters have already resolved current authorization before they reach An alternative is to skip bindings that are live in this Broker process. That would also remove three row writes per healthy binding every 5 s, but the scan would no longer notice the death of an idle worker. Your call. 5. Suites on Linux and mutation results26 single-site mutants of the PR's production changes at
Full list and verdicts: mutants, pass 1, pass 2. This PR has no Java CI: 6. Smaller notes
What this run does not coverBare metal and x86_64 (the host is an aarch64 VM, see below), MariaDB, more than one Broker on the host, hostile same-UID tools, remote writers, restored or cloned disks, Kubernetes. I did not interrupt the server between the holder clear and the final retirement on the real host. That path is covered by the unit tests and the mutation results only. Rig and method
中文版真机验证:把「待完成的 Linux 重启验收」真正跑了一遍验证对象是当前 head 结论
合并建议: 从真机角度看,物理验收门已经满足。建议合并前修掉 F1(改动很小),补上两个缺口测试,并在 README 加一句 soft-reboot 说明。作者列出的「历史 时间均为 UTC,2026-09-27。 1. 物理重启验收(图 1)
JVM 时区是 2. 对照组(图 2)在
只开启恢复选项而不开启 durable 模式时,服务在启动阶段失败,报 3. 恢复之后(图 3)旧回执保持为事实;之后能启动什么由当前授权决定。 4. F1:扫描拒绝健康请求,撞上的 Hosted Session 被阻塞(图 4)过程:
每个 binding 每 5 秒有约 20 ms 的窗口,所以一个有几百个工具回合的 Session 很可能撞上。发生时 Runtime 本身没有任何问题。 候选修复:观察以健康接管结束时,用该结果完成预留,而不是返回 503。等待者在进入 另一种做法是跳过本 Broker 进程内存活的 binding。这样还能省掉每个健康 binding 每 5 秒 3 次行写入,但扫描将不再发现空闲 worker 的死亡。由作者决定。 5. Linux 套件与变异测试(图 5)针对本 PR 生产代码改动做了 26 个单点变异体,依次用单元套件、故障门禁、真实 worker IT、MySQL IT 去跑:
本 PR 没有 Java CI: 6. 其他说明
本轮没有覆盖物理机和 x86_64(主机是 aarch64 虚拟机)、MariaDB、同机多个 Broker、恶意同 UID 工具、远程写入者、恢复或克隆的磁盘、Kubernetes。没有在真机上做「holder 清除之后、最终退休之前」的中断,这条路径只有单元测试和变异结果覆盖。 装置与方法
|
7236322 to
68d763d
Compare
f8bf5d7 to
8d87878
Compare
68d763d to
5248a8f
Compare
8d87878 to
e7f4ae4
Compare
5248a8f to
1e93de6
Compare
e7f4ae4 to
713f79a
Compare
1e93de6 to
36f87cd
Compare
713f79a to
9a1de09
Compare
Real-host verification, round 2: head
|
| Test plan item | Round 1 | Round 2 |
|---|---|---|
| Interrupt cleanup after the holder clear | unit tests and mutation only | real host, two crashes (R14) |
| Retry keeps the original receipt identity | not checked on the real host | 134 of 134 identical |
| Remaining pins are released only after cleanup completes | not checked on the real host | session and slot stay until the retirement commits |
| More than one transaction's worth of executions | 130 + in-flight converge | same, and the crash fell between the two batches (101 done, 31 open) |
2. F1 at the new head
The base brought a retry for prepare in the Harness (hosted-workspace-broker.ts), for transport errors only. acquire is not retried, and a 503 answer is not a transport error, so the chain in round 1 is unchanged: JdbcRuntimeBindingRepository.java:340 → RuntimeBrokerService.java:1549 / :1606 / :825 → hosted-workspace-tool-turn.ts:185-197.
Candidate for this head: diff, test.
3. What lands on main
- A plain merge of this head into
main(6017f11d) conflicts in 9 files. That is the usual effect of a squash-merged base, not a problem in this PR. - Replaying W0e-2 and W0e-3 onto
mainis clean except forManagedAgentMySqlIT.java:main(feat(managed-agent): Make Session close, archive and delete durable operations (Stage D4) #12881) and this PR each add a test at@Order(8). Keeping both and renumbering one resolves it. - The result has one migration per version, V1 to V17. This PR adds none.
- With that jar and
main's bundle, a real reboot recovered 8 of 8 generations 22.0 s after the reboot command (R15).
Mutation re-check. The six mutants that survived every stage in round 1 still do at this head (unit suites of both modules, fault gates, real-worker IT, both MySQL ITs). The control mutant "the same boot counts as a reboot" is still killed by both LocalRebootFaultGateTest cases and WorkspaceRecoveryWorkerIT. RebootRecoveryGapTest passes 2 of 2 at the head and fails under each of its two mutants.
4. Notes
- A
SETTLEDreceipt does not mean the file is on disk after a power cut. In R9 a tool call wrote a file 15 s before the power cut and its receipt isSETTLED / success. After the reboot the file was empty (ext4, the tool does not fsync). This is ordinary page-cache loss and not caused by this PR. It may deserve a sentence in the design, because recovery treatsSETTLEDas a known outcome. - New base behaviour, checked on the real host. After a worker died before READY, a second Session of the same Workspace now starts on the same boot (the feat(managed-agent): Add W0e terminal recovery fences #12839 change to
blocksPlacement). Both bindings were recovered after the reboot. - Flaky under load. My first run on the rebased tree had 7 gate failures and one unit error while the rig server and its workers were running on the same VM and the machine under it was at load 50 to 60. A rerun on the quiet VM was green. In a same-conditions A/B the head had 0 failures in 4 gate runs and the rebased tree 1 in 4 (
adoptedWorkerCanCancelItsOriginalActiveCall); the Broker code of both trees is identical except forBrokerValues.
What this run does not cover
Same as round 1, without the interruption item: bare metal and x86_64, MariaDB, more than one Broker on the host, hostile same-UID tools, remote writers, restored or cloned disks, Kubernetes.
Rig changes in this round
- Worker and Harness bundle rebuilt from
9a1de09e(pnpm install,npm run build,npm run bundle). For the rebased tree the bundle was built frommain. - The rebased tree is
main6017f11dplus the two stacked commits replayed withgit merge-tree --merge-base, with the one conflict resolved as described above. - Crash injection:
CREATE TRIGGER … BEFORE UPDATE ON qwen_runtime_binding_slot … SLEEP(25)whenactive_binding_idbecomesNULL. The watcher readsperformance_schema.processlistand the holder rows, then kills the server's main PID. - Everything else is as in round 1. Scripts, raw results and the candidate are in the round-2 bundle.
中文版
真机验证第二轮:head 9a1de09e
承接第一轮(head f8bf5d74)。同一台专用 Linux 主机,方法相同。时间均为 UTC,2026-09-28。
与第一轮相比改了什么。 这次推送是一次变基。W0e-3 提交和它下面的 W0e-2 提交的 git patch-id 都与之前相同。底座多了 #12839 的两个提交和两次合入 main:4 个 Java 文件、47 个其他文件,其中包括 Hosted Harness 的 Broker 客户端。因为这次 TypeScript 有变化,我用这个 head 重新构建了 worker 和 Harness 的 bundle,并在它上面全部重跑。
结论
- 第一轮的结果在新 head 上复现。 断电和
systemctl reboot之后,所有保存的代数都只靠扫描被回收。两个对照组(选项关闭、machine-id不同)保持阻塞。 - 新增覆盖:清理过程中途崩溃。 在 holder 清除已提交、最终退休仍在进行时,把服务杀掉两次。重启后全部收敛,134 条回执的标识全部不变。
- F1 仍未解决。 选项开启时 470 个健康回合有 5 个被拒,关闭时 411 个里 0 个。Hosted Session 和之前一样被阻塞。第一轮的候选修复不用修改就能应用到这个 head,并消除问题(433 个里 0 个)。
- 要落到
main上,这个栈需要变基,因为 feat(managed-agent): Add W0e terminal recovery fences #12839 是以 squash 方式合入的。把两个提交重放到main上只有一处测试文件冲突。结果通过全部套件,真实重启后也能恢复。
合并建议: 与第一轮相同。物理验收门在这个 head 上同样满足。F1 是我建议合并前先修的一项。第一轮提的两个缺口测试和 soft-reboot 说明仍是待采纳的建议。
1. 重启、对照组、清理中途崩溃(图 1)
崩溃点的放置方式:用一个 MySQL 触发器延迟「释放 placement slot」的那条 UPDATE,这条语句只有最终退休步骤才会发出。产品代码没有改动。只有当 MySQL 显示该语句被挂起、且 holder 行已经清空时,才用 SIGKILL 杀掉服务。
| 测试计划条目 | 第一轮 | 第二轮 |
|---|---|---|
| holder 清除后中断清理 | 只有单元测试和变异结果 | 真机,两次崩溃(R14) |
| 重试保留原回执标识 | 未在真机上检查 | 134 条全部一致 |
| 剩余占用只在清理完成后释放 | 未在真机上检查 | session 和 slot 一直保留到退休提交 |
| 超过单事务批量的执行 | 130 条加在途调用全部收敛 | 同样收敛,且崩溃正好落在两批之间(101 条已完成,31 条未完成) |
2. 新 head 上的 F1(图 2)
底座给 Harness 的 prepare 加了重试(hosted-workspace-broker.ts),只针对传输错误。acquire 没有重试,而 503 响应也不属于传输错误,所以第一轮描述的链路没有变化:JdbcRuntimeBindingRepository.java:340 → RuntimeBrokerService.java:1549 / :1606 / :825 → hosted-workspace-tool-turn.ts:185-197。
3. 落到 main 上会是什么样(图 3)
- 把这个 head 直接合并进
main(6017f11d)有 9 个文件冲突。这是底座被 squash 合入后的常见现象,不是本 PR 的问题。 - 把 W0e-2、W0e-3 重放到
main上,除ManagedAgentMySqlIT.java外都是干净的:main(feat(managed-agent): Make Session close, archive and delete durable operations (Stage D4) #12881)和本 PR 各自在@Order(8)加了一个测试。两个都保留、其中一个改序号即可。 - 结果里每个迁移版本号只有一个文件,V1 到 V17。本 PR 没有新增迁移。
- 用这棵树的 jar 和
main的 bundle 做真实重启(R15):重启命令发出后 22.0 秒,8 个代数全部回收。
变异复查。 第一轮里六个「全部存活」的变异体在这个 head 上依然通过全部六个阶段(两个模块的单元套件、故障门禁、真实 worker IT、两个 MySQL IT)。对照变异体「同一 boot 被当成重启」仍被 LocalRebootFaultGateTest 的两个用例和 WorkspaceRecoveryWorkerIT 杀掉。RebootRecoveryGapTest 在 head 上 2/2 通过,在对应的两个变异体下各自失败。
4. 其他说明
SETTLED回执不代表断电后文件已落盘。 R9 里一次工具调用在断电前 15 秒写了一个文件,回执是SETTLED / success。重启后该文件是空的(ext4,工具没有 fsync)。这是普通的页缓存丢失,不是本 PR 引起的。因为恢复逻辑把SETTLED当作已知结果,设计文档里也许值得加一句。- 底座的新行为,已在真机上确认。 worker 在 READY 之前死亡后,同一 Workspace 的第二个 Session 现在可以在同一 boot 内启动(feat(managed-agent): Add W0e terminal recovery fences #12839 对
blocksPlacement的修改)。重启后两个 binding 都被回收。 - 负载下的抖动。 我第一次跑重放树时,装置的服务和 worker 正在同一台 VM 上运行,底层机器的负载在 50 到 60,结果出现 7 个门禁失败和 1 个单元错误。在安静的 VM 上重跑全绿。同条件 A/B 里,head 的 4 次门禁运行 0 失败,重放树 4 次里 1 次失败(
adoptedWorkerCanCancelItsOriginalActiveCall);两棵树的 Broker 代码除BrokerValues外完全相同。
本轮没有覆盖
与第一轮相同,去掉「中断清理」一项:物理机和 x86_64、MariaDB、同机多个 Broker、恶意同 UID 工具、远程写入者、恢复或克隆的磁盘、Kubernetes。
本轮装置的变化
- worker 和 Harness 的 bundle 由
9a1de09e重新构建(pnpm install、npm run build、npm run bundle)。重放树使用由main构建的 bundle。 - 重放树是
main6017f11d加上用git merge-tree --merge-base重放的两个提交,唯一的冲突按上文方式解决。 - 崩溃注入:
CREATE TRIGGER … BEFORE UPDATE ON qwen_runtime_binding_slot … SLEEP(25),在active_binding_id变为NULL时触发。观察脚本读取performance_schema.processlist和 holder 行,然后杀掉服务的主进程。 - 其余与第一轮相同。脚本、原始结果和候选补丁都在第二轮证据包里(链接见英文版)。
36f87cd to
4136721
Compare
9a1de09 to
e867a99
Compare
bceaa60 to
8c2b626
Compare
Real-host verification, round 3: head
|
| # | Finding (round) | Severity | Status at 8c2b626c |
|---|---|---|---|
| F1 | Scan refuses healthy requests; Hosted Session blocked (r1, r2) | worth fixing before merge | stands — re-measured: 2 of 417 healthy turns refused with the option on (r2: 5/470), 0 of 418 with it off; Hosted chain recovery-blocked on the aligned acquire. The r1/r2 candidate applies cleanly (offset 58) and clears it: 0 of 406, Hosted chain completes, all suites + 3/3 candidate tests pass. |
| — | Two gap tests (RebootRecoveryGapTest) open suggestion (r1) |
suggestion | superseded in part — the W0e-2 base's new fault-gate coverage now kills M01 and M25 in-tree (mutation matrix below). The gap test still passes 2/2 at head; M16 (fail-open recoverResources default) remains without an in-tree pin. |
| — | Soft-reboot README sentence (r1) | suggestion | stands — no README change in this push. |
| — | SETTLED receipt ≠ on-disk after power cut (r2 note) |
docs note | stands — design doc unchanged on this point; harmless page-cache reality, worth the sentence. |
| — | Stack needs a rebase onto main (r2) | process | fixed — the stack was rebased; git merge-tree against current main a77d80d19f is conflict-free. |
Central claim — reboot recovery on a real host, re-measured at this head
Same dedicated-VM method as rounds 1–2, rebuilt here from scratch (KVM/qemu, Debian 12 arm64, kernel 6.1.0-53-cloud, JDK 21, MySQL 8.4, worker bundle built from this head; systemd unit, non-root, KillMode=process). Three runs:
| run | scenario | result |
|---|---|---|
| R20b | hands-off systemctl reboot, 8 generations on 6 storages (incl. 1 worker-only death, 1 interrupted startup, 1 access-revoked), 130 PREPARED + in-flight + settled receipts |
pass — boot id changed (17 s down); all 8 RELEASED 24 s after the new boot's server start; holders 6→0; 132 nonterminal → ABANDONED without results; SETTLED keep results; revoked arm stays 409; authorized warm creates generation 2 (2.2 s); old Runtime Session 404 |
| R21 | power cut (SIGKILL qemu), then boot with option OFF, then foreign machine-id, then recovery | pass — option OFF: nothing moved (6 holders, 132 executions intact), authorized warm 503 runtime_broker_reconcile_timeout. Foreign machine-id: identical block. Original machine-id restored: all 8 RELEASED in 7.7 s, evidence names the pre-cut boot |
| R22b | reboot, then 2× SIGKILL between the committed holder clear and the final retirement (MySQL trigger holds the slot-freeing UPDATE 25 s) | pass — crash 1 hit the first retirement, crash 2 hit the retried one; converged alone; 134/134 receipts identical, 132 ABANDONED without result, 2 SETTLED kept; 0 worker processes left |
F1 at this head (unchanged mechanism, re-measured)
The scan still observes every healthy durable binding every 5 s and still completes the reservation with 503 runtime_reconciliation_required; the Hosted chain still turns that into a recovery-blocked Session. Numbers: 2 of 417 turns refused (option on) vs 0 of 418 (off) vs 0 of 406 (candidate). The aligned-delivery probe shows the exact window: acquire delivered when the scan holds the binding → 503 at head, 200 with the candidate.
Suites and mutation re-check
- Head: broker unit 428, server unit 178, WorkspaceRecoveryWorkerIT 1, broker MySQL IT 3, server MySQL IT 14 — all green incl. checkstyle. Candidate: same, plus the F1 test and both gap tests pass 3/3 (at head the F1 test errors as designed).
- Mutation matrix (7 mutants from the r1/r2 survivor set + control, six stages each): M01 and M25 are now killed by the base's strengthened durable-startup fault gate — two of the six round-1 survivors now have in-tree pins. M11 (claim ignores the Runtime Session row), M16 (fail-open default), M19 (expected generation ignored), M21 (overlapping scan batches) still survive everything; adjudication unchanged from round 2 (M16 is the one guarding a safety property).
- One load flake:
ProcessCrashFaultGateTest.aLostJournalEndsPollingWithoutReleasingTheWriterDomainfailed once while the host ran the VM plus two Maven containers (runtime_provision_fencedsurfaced beforeruntime_broker_runtime_lost); 5/5 green isolated at head. The mutation runs show the same two codes are order-sensitive at that boundary — worth a look from whoever owns the startup fence, not a blocker.
What changed since round 2 (scope of this round)
git range-diff r2-head..head: the W0e-3 commit touches only the IT fork timeout (180→600 s) and one test's @Order (8→11, base added tests). The base gained 56ada7e510 and 430a76eac9 (durable startup reconciliation termination; legacy startup fence limited to its slot) plus docs, and main advanced ~460 files. Everything above was re-run, not carried over.
Not covered
Bare metal and x86_64 (host is an aarch64 VM), MariaDB, more than one Broker per host, hostile same-UID tools, remote writers, restored/cloned disks, Kubernetes, claim-expiry retry as a dedicated real-host scenario (unit + mutation coverage only). Round 1's SETTLED-vs-fsync note stands unaddressed in the docs.
Methodology
Head worktree at 8c2b626c (pnpm install, npm run build, npm run bundle → worker bundle; Maven in a JDK 21 container → qwen-managed-agent-server jar, sha256 831f89e9…, candidate jar 8109c4aa… = head + the r2 F1 candidate re-applied with offset 58). VM: qemu/KVM arm64, Debian 12 cloud image, MySQL 8.4 in docker, server jar under systemd with the rig auth filter. Power cut = SIGKILL on the qemu process; reboot = systemctl reboot inside the guest. Harness scripts from the round-1/2 evidence bundle, unchanged except a guard in the crash watcher (skip kill while systemd reports MainPID 0 — a rig bug that ate round-3's first crash attempt; product untouched). Scripted assertions: assert.mjs re-validates the saved raw logs (38 checks), plus 8 scripted mutation verdicts, 5 isolated gate re-runs, 12 suite-stage exit checks. One scripted check failed once (the loaded gate run) and passed 5/5 isolated — disclosed above, counted as the flake it is.
Raw logs, harness and assertion script: in the round-3 bundle (assert.mjs re-validates the saved logs; results/ holds the raw per-run output).
For the merge decision: unchanged from rounds 1–2 — the physical acceptance gate is met again at this head, on a fresh independent rig. F1 remains the one item to fix before the option is used with Hosted turns; the attached candidate is now verified against three consecutive heads. The stack merges into current main (a77d80d19f) conflict-free; #12839 is already in, so the order left is #12865 → this PR.
…boot (cherry picked from commit 8c2b626)
8c2b626 to
55ded4c
Compare
|
Follow-up to the round 3 real-host review: F1 is fixed in Because #12865 was squash-merged, I replayed the unchanged W0e-3 commit onto current Current-head verification: Broker unit/H2 430 tests, 0 failures/errors and 1 conditional skip; managed-agent server unit/H2 184 passed; real-process LocalReboot and DurableLocalRuntime fault gates 12/12 passed; Spring/H2/real-worker recovery integration 1/1 passed; Java Checkstyle, repository build, typecheck, bundle and changed Markdown formatting passed. The reviewer physically reboot-tested the earlier W0e-3 head; this new SHA has not had another physical reboot run. There are 0 unresolved inline threads. I am moving the PR to ready for review so exact-head CI and review can run. |
|
[codex] Thanks — agreed on the CI blocker. Fixed in
The targeted hosted IT passed (1/1), the guard's hosted and non-hosted family checks passed in an isolated verification, and the repository build and typecheck passed. The new push should trigger the PR CI that did not run on the draft head. Resolved 0/0 inline threads. |
…ge with main
Real-host verification, round 4: head
|
| Run | What it shows |
|---|---|
| R16, power cut | Both controls stay blocked for the whole observation (option off, foreign machine-id). With the option on and the original machine-id, 8 of 8 generations are released 3.6 s after the server starts. |
R17, R20, R21, R22, systemctl reboot |
Hands-off. All 8 generations are released 18.2 s to 26.8 s after the reboot command. R20 uses the jar of 62584d31, R21 and R22 use merges with main. |
| R19, reboot and two crashes | The server is killed twice while MySQL holds the final retirement statement. After the restarts everything converges, and all 134 receipts keep their identity. |
3. The CI guard, before and after the rename
- At
55ded4c3the guard fails in both jobs with the two messages the bot predicted. In the Hosted job the guard step comes before the fault gates step, so CI would not have run the gates either. - At
62584d31both guards pass. Before the push I had tested the bot's suggested fix as an A/B arm, and62584d31is file for file that tree (10019 files, 0 differences). - First CI run.
SDK Java(8 of 8 jobs) andQwen Code CIare green. The test counts in the CI logs equal the counts of my replays. - Why CI was missing before.
55ded4c3was pushed while the PR was a draft on a stacked base. Changing the base and marking the PR ready do not start these workflows.
4. Merge with today's main
main has moved eight commits since the PR's base (99adce25 → 1b696297). Under packages/sdk-java they change Hosted tests, SessionEventHub, the V15 migration and the OpenAPI file.
| Check | Result on the merge with 1b696297 |
|---|---|
git merge of 62584d31 |
No conflict. 18 migrations (V1 to V18), one per version. This PR adds none. |
| Server jar | 610 entries. The 8 that differ from the jar of 62584d31 are main's changes. No recovery or Broker class differs. |
Hosted job, with main's bundle |
10 of 10 integration tests pass, the guard passes |
| Real reboot with that jar and bundle (R22) | 8 of 8 generations released, hands-off, 26.8 s (the guest took longer to boot) |
The same checks passed an hour earlier on main b1eb94da (R21, 18.2 s). main moved again while I was testing, so I repeated them.
5. Mutation results and a second look at RuntimeRecoveryEvidence.matches
Correction to round 3. Round 3 listed M01 (rebooted() ignores the host id) and M25 (evidence matches any lease) as killed by DurableLocalRuntimeFaultGateTest.absentWorkerAfterRestartLeavesItsDomainPinned. That was wrong:
- The gate builds its Brokers with trusted reboot recovery off, so it never calls
rebooted(). - In this round both mutants pass all six stages. The runs were sequential on the quiet host, and a stage counts as a kill only if it fails twice.
- The failures in round 3 came from the load-sensitive gate that round 3 itself listed as a flake.
What is pinned now.
| Mutant | Rounds 1 to 3 | This round |
|---|---|---|
M16, the default recoverResources allows managed cleanup |
survives | killed by the author's new RuntimeMaintenanceRecoveryTest case |
| M01, another machine with another boot id counts as a reboot | survives | survives |
| M11, M19, M21 | survive | survive (rated as smaller gaps since round 1) |
M01 is the one survivor that guards a safety property. On the real host the foreign machine-id control stays blocked (R16), so the code is right and only the test is missing. RebootRecoveryGapTest from round 1 covers it: it passes 2 of 2 at 62584d31, fails under M01, and is Checkstyle clean.
RuntimeRecoveryEvidence.matches. The bot asked for a second look at the relaxed lease check.
- With a lease, nothing is relaxed.
matcheshas one caller,RuntimeBindingRecord.requireEvidence. The constructor rejects a lease that does not match the seed before it reaches that call (RuntimeBindingRecord.java:118). So the new clause can never be false there, which is why M25 survives: it is an equivalent mutant, not a test gap. - Without a lease, the evidence is still tied to the seed (request id, runtime instance id, incarnation, lease id, epoch) and to the resource handle.
- On the real host this case is the worker that died before READY. In R16, binding
1ce9de00had no lease and wasRECOVERY_BLOCKEDbefore the power cut. After the reboot it was released with both evidence records (table). Before this PR,matchesreturned false for every binding without a lease.
6. Notes, none blocking
- The scan leaves no trace in the log. In R19 the server reclaimed four generations and was killed twice, and its log has no line from the scan or from the recovery (log). An operator can follow a recovery only in the database. The author deferred logging to a follow-up. This is how that item looks on a real host.
- The scan still observes every healthy binding every 5 s and rewrites its row. After the fix this no longer refuses requests. The author deferred the optimization, and I agree that it can wait.
- SPI.
RuntimeBindingRepositorygains two methods without a default (finishLostRecovery,findRecoveryCandidates), so an implementation outside this repository would stop compiling. Both implementations in the tree are updated, andRuntimeProvisioner.recoverResourceshas a fail-closed default. The package is0.1.0-alpha. Whether that needs a release note is a maintainer call. - The repository's own script tests do not see a guard mismatch.
check-failsafe-reports.test.jsandhosted-process-ci.test.jspass 13 of 13 at55ded4c3, where both guards fail.
What this run does not cover
Bare metal and x86_64, more than one Broker on the host, hostile same-UID tools, remote writers, restored or cloned disks, Kubernetes. On MariaDB only the two integration test classes of the CI job ran. The real-host runs use MySQL 8.4.
Rig and method in this round
- Host. The same VM as in rounds 1 and 2: Ubuntu 24.04.4, kernel 6.8.0-117, aarch64, ext4. The server jar runs under systemd with MySQL 8.4.11 and production
HostIdentity.linux(). Round 3 used a different VM. - Bundle. This PR changes no TypeScript. The worker and Harness bundle was built from
55ded4c3, and for R21 and R22 from the merge results. - CI replays. The commands are taken from
.github/workflows/sdk-java.yml. The MariaDB job ran in a Linux container on the dedicated host (MariaDB 10.11.18). The Hosted job ran on macOS (JDK 21.0.12, Node 22.23.2, MySQL 8.4.7), because the Hosted tests need the repository'snode_modules. GitHub CI runs both on Ubuntu. - One rig error. My first MariaDB replay of the renamed tree used the database of the previous arm, and
JdbcRuntimeBrokerMySqlITfailed withexpected: <[1]> but was: <[2]>. The contract test writes with a fixed prefix and needs a database that no earlier run has used. CI starts a fresh MariaDB for every job. As a control, the unchanged55ded4c3tree fails the same way on the used database. On a fresh database the replay passes. - Crash watcher. In the first attempt (R18) the second kill hit the restarted server while MySQL still held the statement of the first one. The watcher now requires the held statement to be younger than the running server. R19 is the run with both kills inside a held retirement. R18 converged as well and is in the bundle.
- Mutation. Run at
55ded4c3.62584d31has the same sources except for the name of one test class. - Everything is in the bundle: harness and raw results.
中文版
真机验证第四轮:head 62584d31
承接第三轮(head 8c2b626c)。使用第一、二轮的同一台专用 Linux 主机。时间均为 UTC,2026-09-29。
与第三轮相比改了什么。 分支现在直接基于 main,共三个提交:
- W0e-3 提交,
git patch-id与8c2b626c时相同; 55ded4c3:F1 修复、两个新测试和文档更新;62584d31:给一个测试类改名,并删掉 pom 里的两行。
我先测的是 55ded4c3,测试途中作者推送了 62584d31。两个 head 构建出的服务 jar 里 537 个 class 文件逐字节相同,所以结果可以沿用。我另外在新 head 上补跑了一次重启和两个 Java CI 作业。
结论
- F1 已修复。 选项开启时,1383 个健康回合 0 个被拒(第一轮 871 个里 6 个,第二轮 470 个里 5 个)。第一到三轮里以 Session 被阻塞告终的 Hosted 链路现在能跑完。不该放行的请求仍然被拒绝。
- 这份代码通过了物理重启验证。 一次断电加两个对照组、四次免人工重启、一次在清理过程中崩溃两次的重启,全部收敛。其中一次免人工重启用的是
62584d31构建的 jar,补上了作者提到的「新 SHA 还没做过物理重启」这个缺口。 - triage 机器人指出的 CI 阻塞项确实存在,
62584d31已修复。 在55ded4c3上,所有 Maven 步骤都通过,但两个 Java 作业的报告守卫都失败。在62584d31上两个守卫都通过,我的回放和本 PR 第一次真实 CI 的结果一致。 - 与当前
main(1b696297)合并没有冲突。 Hosted 作业在合并结果上通过,用合并后的 jar 和main的 bundle 做真实重启,8 个代数全部回收。 - 对第三轮的一处更正: M01 和 M25 并没有被仓库里的测试杀掉。详见第 5 节。
合并建议: 我没有发现阻塞合并的问题。机器人的 CHANGES_REQUESTED 评审针对的是 55ded4c3,它提出的唯一阻塞项已经修复。还有三条建议,都不阻塞:补一个「换了机器」场景的测试(已附上)、给扫描加日志、以及作者已说明延后处理的几项。
1. F1 已修复(图 1)
- 修复方式就是第一轮的候选方案。观察结果健康时,
recoverBinding用接管到的上下文完成它的 reservation。其他结果仍然以503 runtime_reconciliation_required结束。 - Hosted 链路。 投递方式与第一到三轮相同:在扫描持有该 binding 的 claim 时,把未经修改的 acquire 请求送到 Broker。三个回合全部完成,Session 没有进入恢复阻塞状态。
- 另一面。 把 worker 杀掉后连续发出 733 个
warm和acquire请求,全部被拒,binding 变为LOST且没有停止证据,也没有启动新的 worker。撤销访问权限后,977 个请求全部返回409 workspace_unavailable,binding 保持不变。 - 针对修复的测试。 把修复回退会被新增的
MaintenanceProbeWaiterTest抓到。修复还有两个变异体存活,但都不代表缺测试:一个是等价变异体,另一个会在adoptObservation里被再次拒绝。
2. 物理重启(图 2)
| 运行 | 说明 |
|---|---|
| R16,断电 | 两个对照组(选项关闭、machine-id 不同)在整个观察期内保持阻塞。选项开启且 machine-id 是原来的值时,服务启动后 3.6 秒,8 个代数全部回收。 |
R17、R20、R21、R22,systemctl reboot |
全程无人工介入。重启命令发出后 18.2 到 26.8 秒,8 个代数全部回收。R20 用 62584d31 的 jar,R21 和 R22 用与 main 合并后的结果。 |
| R19,重启加两次崩溃 | 在 MySQL 挂起最终退休语句时把服务杀掉两次。重启后全部收敛,134 条回执的标识全部不变。 |
3. 改名前后的 CI 守卫(图 3)
- 在
55ded4c3上,两个作业的守卫都失败,报错信息与机器人预测的两条一致。Hosted 作业里守卫步骤排在故障门禁步骤之前,所以 CI 连门禁也不会跑。 - 在
62584d31上,两个守卫都通过。作者推送之前,我已经把机器人建议的修法当作 A/B 的一个分支测过,62584d31与那棵树逐文件相同(10019 个文件,0 处差异)。 - 第一次 CI。
SDK Java(8 个作业)和Qwen Code CI全绿。CI 日志里的测试数与我回放得到的数字相同。 - 之前为什么没有 CI。
55ded4c3推送时 PR 还是草稿,而且 base 是堆叠分支。修改 base 和标记为 ready 都不会触发这些工作流。
4. 与当前 main 合并
main 在 PR 的 base 之后又前进了八个提交(99adce25 → 1b696297)。在 packages/sdk-java 下,这些提交改了 Hosted 测试、SessionEventHub、V15 迁移和 OpenAPI 文件。
| 检查项 | 与 1b696297 合并后的结果 |
|---|---|
合并 62584d31 |
无冲突。18 个迁移(V1 到 V18),每个版本号一个。本 PR 没有新增迁移。 |
| 服务 jar | 610 个条目。与 62584d31 的 jar 相比有 8 个不同,都是 main 的改动。恢复和 Broker 相关的 class 没有差异。 |
Hosted 作业(使用 main 的 bundle) |
10 个集成测试全部通过,守卫通过 |
| 用这个 jar 和 bundle 做真实重启(R22) | 8 个代数全部回收,无人工介入,26.8 秒(这次虚拟机启动较慢) |
一小时前在 main b1eb94da 上做的同样检查也通过(R21,18.2 秒)。测试期间 main 又前进了,所以我重做了一遍。
5. 变异结果,以及对 RuntimeRecoveryEvidence.matches 的复查
对第三轮的更正。 第三轮把 M01(rebooted() 不检查 host id)和 M25(证据匹配任意 lease)记为被 DurableLocalRuntimeFaultGateTest.absentWorkerAfterRestartLeavesItsDomainPinned 杀死。这个结论是错的:
- 这个门禁在创建 Broker 时没有开启可信重启恢复,所以根本不会调用
rebooted()。 - 本轮两个变异体都通过了全部六个阶段。运行方式是在安静的主机上串行执行,并且一个阶段要连续失败两次才算杀死。
- 第三轮看到的失败来自那个对负载敏感的门禁,第三轮自己也把它列为抖动。
现在哪些已被钉住。
| 变异体 | 第一到三轮 | 本轮 |
|---|---|---|
M16:recoverResources 的默认实现允许 managed 清理 |
存活 | 被作者新增的 RuntimeMaintenanceRecoveryTest 用例杀死 |
| M01:另一台机器、另一个 boot id 也被当成重启 | 存活 | 存活 |
| M11、M19、M21 | 存活 | 存活(第一轮起就归为较小的缺口) |
M01 是存活者里唯一守护安全性质的一个。真机上 machine-id 不同的对照组保持阻塞(R16),所以代码是对的,缺的只是测试。第一轮附上的 RebootRecoveryGapTest 覆盖了这个场景:在 62584d31 上 2/2 通过,在 M01 下失败,Checkstyle 无告警。
RuntimeRecoveryEvidence.matches。 机器人希望有人再看一眼被放宽的 lease 检查。
- 有 lease 时没有任何放宽。
matches只有一个调用方RuntimeBindingRecord.requireEvidence。构造函数在走到这个调用之前,就已经拒绝了与 seed 不匹配的 lease(RuntimeBindingRecord.java:118)。所以新加的子句在那里不可能为假,这就是 M25 存活的原因:它是等价变异体,不是测试缺口。 - **没有 lease 时,**证据仍然和 seed(request id、runtime instance id、incarnation、lease id、epoch)以及 resource handle 绑定。
- 真机上这种情况对应「worker 在 READY 之前死亡」。R16 里 binding
1ce9de00没有 lease,断电前是RECOVERY_BLOCKED。重启后它带着两条证据被回收(表格链接见英文版)。在本 PR 之前,只要 binding 没有 lease,matches就返回 false。
6. 其他说明(均不阻塞)
- 扫描在日志里不留痕迹。 R19 里服务回收了四个代数,还被杀掉两次,但日志里没有任何一行来自扫描或恢复(日志链接见英文版)。运维只能从数据库里看到恢复过程。作者已说明把日志放到后续处理。这里给出的是这一项在真机上的表现。
- 扫描仍然每 5 秒观察一次所有健康的 binding,并重写它们的行。修复之后这不再导致请求被拒。作者把这项优化延后,我同意可以延后。
- SPI。
RuntimeBindingRepository新增了两个没有默认实现的方法(finishLostRecovery、findRecoveryCandidates),仓库外的实现会因此无法编译。仓库内的两个实现都已更新,RuntimeProvisioner.recoverResources有 fail-closed 的默认实现。包版本是0.1.0-alpha。是否需要写进 release note 由维护者决定。 - 仓库自带的脚本测试发现不了守卫不匹配。
check-failsafe-reports.test.js和hosted-process-ci.test.js在55ded4c3上 13/13 通过,而那时两个守卫都是失败的。
本轮没有覆盖
物理机和 x86_64、同机多个 Broker、恶意同 UID 工具、远程写入者、恢复或克隆的磁盘、Kubernetes。MariaDB 上只跑了 CI 作业里的两个集成测试类。真机运行使用 MySQL 8.4。
本轮的装置和方法
- 主机。 与第一、二轮相同的 VM:Ubuntu 24.04.4,内核 6.8.0-117,aarch64,ext4。服务 jar 由 systemd 运行,搭配 MySQL 8.4.11 和生产的
HostIdentity.linux()。第三轮用的是另一台 VM。 - Bundle。 本 PR 没有改动 TypeScript。worker 和 Harness 的 bundle 由
55ded4c3构建,R21 和 R22 使用由合并结果构建的 bundle。 - CI 回放。 命令取自
.github/workflows/sdk-java.yml。MariaDB 作业在专用主机上的 Linux 容器里运行(MariaDB 10.11.18)。Hosted 作业在 macOS 上运行(JDK 21.0.12、Node 22.23.2、MySQL 8.4.7),因为 Hosted 测试需要仓库的node_modules。GitHub CI 的两个作业都跑在 Ubuntu 上。 - 一处装置错误。 我第一次在 MariaDB 上回放改名后的树时,复用了上一个分支用过的数据库,
JdbcRuntimeBrokerMySqlIT报expected: <[1]> but was: <[2]>。这个契约测试用固定前缀写数据,需要一个没被之前的运行用过的数据库。CI 为每个作业启动全新的 MariaDB。对照:未改动的55ded4c3在用过的数据库上同样失败。换成全新的数据库后回放通过。 - 崩溃观察脚本。 第一次尝试(R18)里,第二次 kill 打在了重启后的服务上,而当时 MySQL 挂起的还是第一个服务的语句。现在脚本要求被挂起的语句比当前运行的服务更新。R19 是两次 kill 都落在被挂起的退休语句期间的那次运行。R18 同样收敛,结果也在证据包里。
- 变异测试在
55ded4c3上运行。62584d31与它相比,源码只差一个测试类的名字。 - 脚本和原始结果都在第四轮证据包里(链接见英文版)。
|
[codex] Thanks for round 4 and for correcting the M01/M25 mutation result. I updated the PR description in both languages to cite the current-head physical reboot, power-cut, F1 load and CI evidence, and removed the obsolete retest caveat. I checked that |
|
@qwen-code /triage |
qqqys
left a comment
There was a problem hiding this comment.
Independent Critical-only review — head 62584d315d93392232b012edd38ba1147ea580a0
Verdict: COMMENT. The one blocking issue ever filed against this PR is confirmed fixed at this head, and I found no Critical in the production code I read. What keeps this off the Approve path is coverage: RuntimeBrokerService.java (+129/−21), the file that holds the recovery entry point, is outside what I could read inside this review's budget.
Historical blocking issue — verified FIXED
The CHANGES_REQUESTED at 55ded4c3 named one mechanical blocker: scripts/check-failsafe-reports.js partitions *IT.java by class name (^Hosted.*IT$ is the hosted family, everything else is not), and the new WorkspaceRecoveryWorkerIT was named for the non-hosted family while pom.xml scheduled it into hosted-harness-mysql, so both Java jobs failed in opposite directions. The prescribed fix was to rename the class and drop both pom hunks.
At this head the file is service/HostedWorkspaceRecoveryWorkerIT.java (+115), no pom.xml appears anywhere in the 39-file diff, and every Java job is green — ubuntu-latest / Java 11, 17 and 21, windows-latest / Java 21, macos-latest / Java 21, Runtime Broker and Managed Agent MariaDB / Java 21, Hosted process fault gates / MySQL 8.4 / Java 21 and Real daemon E2E / Java 11. The prior head had never compiled in CI because every push landed while the PR was a draft; this head has. Both halves of the fix were taken and the symptom is gone.
There is no other blocking finding in this PR's history: no [Critical] inline comment exists, and the maintainer's long review at this exact head concludes that nothing blocks the merge, with three remaining items explicitly non-blocking.
What I verified myself
releaseLost is fenced far more tightly than a holder release needs to be, and fails closed at every step. It refuses up front unless the binding is managed-context, LOST, has stopped writers, and has an operation owner. Inside one transaction it takes FOR UPDATE on the binding row and re-checks binding id, generation, binding_state = 'LOST', tenant, workspace, storage id, record_version, operation_owner, operation_generation, both evidence columns non-null, and an operation lease that has not expired — measuring expiry against the database's own clock read in the same query (UNIX_TIMESTAMP() plus EXTRACT(MICROSECOND …)), so no application/DB clock skew can widen the window. It then requires zero unsettled executions for that binding and generation, locks the lease row, validates that holder_key equals the digest of bindingId, generation and session id, re-checks the lease against a freshly queried DB clock, and issues a conditional UPDATE naming all five columns with a changed != 1 refusal. Every mismatch throws the non-retryable 409 rather than proceeding, and all SQL is parameterized.
claim gained two precondition reads that can only narrow it. Both new FOR UPDATE queries require an exactly-one matching row — the binding READY and not draining with tenant, workspace, workspace generation and storage id all equal, and the session in ACQUIRING or READY with matching harness session — and throw unavailable() otherwise. A claim that previously succeeded on a stale or draining binding now fails closed.
The recovery split preserves existing callers exactly. recoverLost now delegates with holdersCleared = !expected.getRequest().isManagedContext(), and the guard became if (!holdersCleared || !current.hasStoppedWriters() || …) return current;. For a non-managed context holdersCleared is true, so !holdersCleared is false and the condition reduces to the original expression — unchanged behaviour. For a managed context it is false, so recovery returns early and the holder must be released through WorkspaceExecutionStore.releaseLost before the new finishLostRecovery completes it. That is the right direction: the scan cannot clear a workspace holder implicitly.
The coordinator is bounded and cannot wedge. scan() is single-flight behind running.compareAndSet(false, true) and returns immediately once closed. It reads a batch of 8 with a keyset cursor that advances only on a full page and resets on a short one, isolates each candidate with exceptionally(error -> null) plus a catch (RuntimeException) around the synchronous call, and releases running in both whenComplete and the outer catch — including the empty-batch case, where allOf on an empty array completes at once. A failing candidate therefore cannot abort the batch, and a failing batch cannot leave the flag set.
Gates I could not confirm inside budget
RuntimeBrokerService.java(+129/−21) — the largest production change and the one holdingrecoverBinding, the method the coordinator calls per candidate, plus the maintenance and evidence handling around it. Not read.LocalProcessRuntimeProvisioner.java(+33/−8) — the writer-stop evidence a reboot recovery depends on. Not read.EmbeddedRuntimeBroker.java(+20/−2),WorkspaceRuntimeProvisioner.java(+15/−1),ManagedAgentProperties.java(+9) and theapplication.ymldefault — the wiring and scheduling of the scan. Not read.findRecoveryCandidatesinJdbcRuntimeBindingRepositorybeyond itsrequireIdvalidation, andInMemoryRuntimeBindingRepository(+31/−1), which must mirror the JDBC semantics for the unit tests to mean anything.
CI
Every check passes at this head except review-pr, which failed after 6h0m45s — a timeout of the review automation rather than a build or test failure. Under this channel's rules review-pr is not a gating check in any state, and no other job failed, so I attribute nothing to this PR. Note that reviewDecision still reads CHANGES_REQUESTED because the bot's review at 55ded4c3 was never dismissed; its blocker is fixed, so that review needs dismissing by its author before the decision clears.
Next step: items 1 and 2 would close this. recoverBinding is the single method that ties the coordinator to the fenced store calls I verified, so reading it is what would let me confirm that a reboot recovery cannot adopt a binding whose writers have not provably stopped.
wenshao
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed. Suggestions are inline.
Not explored to full depth (tool budget reached): "agent reverse-audit (round 5)": did not execute the new tests (H2 WorkspaceRecoveryTest and the MySQL channel) — the assertion walk is source-level only; I did not build or run mvn in the …; "agent reverse-audit (round 5)": WorkspaceRecoveryContract.java:87 — did not attribute WorkspaceRuntimeResolver.resolve 's refusal to the absent-mounts configuration rather than to the grant…; "agent reverse-audit (round 5)": did not read HostedWorkspaceRecoveryWorkerIT end to end to decide whether the real WorkspaceRuntimeProvisioner.recoverResources managed branch is covered ou…; chunk 7: mapSeed 's handling of a null provision_seed_ciphertext ( JdbcRuntimeBindingRepository.java:817-832 ) was not read line-by-line; it only matters if the quali….
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.
中文说明
仅完成部分审查,审查缺口已披露。 建议见行内评论。
未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 5)":did not execute the new tests (H2 WorkspaceRecoveryTest and the MySQL channel) — the assertion walk is source-level only; I did not build or run mvn in the …;"agent reverse-audit (round 5)":WorkspaceRecoveryContract.java:87 — did not attribute WorkspaceRuntimeResolver.resolve 's refusal to the absent-mounts configuration rather than to the grant…;"agent reverse-audit (round 5)":did not read HostedWorkspaceRecoveryWorkerIT end to end to decide whether the real WorkspaceRuntimeProvisioner.recoverResources managed branch is covered ou…;chunk 7:mapSeed 's handling of a null provision_seed_ciphertext ( JdbcRuntimeBindingRepository.java:817-832 ) was not read line-by-line; it only matters if the quali…。
未审查:反向审计——在 5 轮的反审轮数上限内未收敛。
— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)
| ContextBinding changedGeneration = new ContextBinding(binding.getTenantId(), "other-workspace", 2, | ||
| binding.getStorageId(), "child", binding.getContextConfigRef(), 1); | ||
| assertBusy(() -> otherClient.claim(changedGeneration, loser)); | ||
| assertUnavailable(() -> otherClient.claim(changedGeneration, loser)); |
There was a problem hiding this comment.
[Suggestion] R1-1: The rewritten fencing assertions no longer witness the generation or storage clause they are named for.
changedGeneration and independentStorage are both built with workspaceId = "other-workspace", which alone mismatches the stored qwen_runtime_binding row, so both claims are refused by the workspace-id comparison in WorkspaceExecutionStore.claim's new live predicate — the generation and storage comparisons this same diff added are never the deciding clause.
Measured with the clauses muted out of claim in a scratch tree, WorkspaceRuntimeTest alone:
baseline Tests run: 13, Failures: 0
generation clause deleted Tests run: 13, Failures: 0
generation + storage deleted Tests run: 13, Failures: 0
So a change that admits a stale-generation or foreign-storage claim ships with the suite green.
Suggested fix: make each binding differ from the stored row in exactly one field — keep tenant/workspace/storage and change only the generation for changedGeneration; keep everything and change only the storage id for independentStorage. Each variant is then refused by exactly one clause of the predicate.
WorkspaceExecutionStore.claim's live predicate also requires the binding row to be READY and not draining (WorkspaceExecutionStore.java:96-108), so the variants must keep the fixture's other columns.
Acceptance criterion: WorkspaceRuntimeTest.serializesIndependentSqlClientsByStorageAndFencesStaleRelease must go red when && binding.getStorageId().equals(row.getString("storage_id")) (or the generation clause) is deleted from claim's live predicate; today it stays green.
中文说明
[Suggestion] 改写后的 fencing 断言不再能见证它们名字所指的 generation/storage 子句。
changedGeneration 和 independentStorage 都用 workspaceId = "other-workspace" 构造,仅这一项就与存储的 qwen_runtime_binding 行不匹配,因此两次 claim 都是被 WorkspaceExecutionStore.claim 新增 live 谓词里的 workspace id 比较拒绝的 —— 本次 diff 同时加入的 generation 与 storage 比较从来不是决定性子句。
在 scratch 树里把 claim 的这两个子句逐个删除后实测(只跑 WorkspaceRuntimeTest):
基线 Tests run: 13, Failures: 0
删除 generation 子句 Tests run: 13, Failures: 0
删除 generation + storage Tests run: 13, Failures: 0
也就是说,一个允许过期 generation 或外来 storage 的 claim 会在测试全绿的情况下合入。
建议:让每个 binding 只在一个字段上与存储行不同 —— changedGeneration 保留 tenant/workspace/storage,只把 generation 改成 2;independentStorage 只把 storage id 改成 "other-storage"。这样每个变体都恰好被谓词中的一个子句拒绝。
WorkspaceExecutionStore.claim 的 live 谓词还要求该行是 READY 且未在 drain(WorkspaceExecutionStore.java:96-108),因此变体必须保留 fixture 的其他列。
验收标准:从 claim 的 live 谓词中删除 && binding.getStorageId().equals(row.getString("storage_id"))(或 generation 子句)时,WorkspaceRuntimeTest.serializesIndependentSqlClientsByStorageAndFencesStaleRelease 必须变红;目前它仍然全绿。
— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)
| assertTrue(bindings.findRecoveryCandidates("local-process", null, 100).stream() | ||
| .anyMatch(record -> record.getBindingId().equals(fixture.binding.getBindingId()))); | ||
| assertTrue(bindings.findRecoveryCandidates("unsupported-provider", null, 100).isEmpty()); | ||
| assertTrue(bindings.findRecoveryCandidates("local-process", fixture.binding.getBindingId(), 1).stream() |
There was a problem hiding this comment.
[Suggestion] R1-2: The cursor-exclusivity assertion runs on an empty page, so allMatch is vacuously true.
The repository holds exactly one record at this point — the fixture's own binding — and the exclusive cursor excludes it, so the stream is empty and the assertion cannot fail.
pageSize=0 allMatch=true totalRows=1
Measured harm: with findRecoveryCandidates degraded to return an empty page whenever afterBindingId != null, this class stays 4/4 green — the documented "ordered by binding ID after the exclusive cursor" contract is unchecked in the direction that matters.
Suggested fix: insert a second candidate that sorts above the cursor and assert the page exactly:
var page = bindings.findRecoveryCandidates("local-process", fixture.binding.getBindingId(), 1);
assertEquals(1, page.size());
assertTrue(page.get(0).getBindingId().compareTo(fixture.binding.getBindingId()) > 0);or assert isEmpty() explicitly so the expectation is stated rather than vacuous.
findRecoveryCandidates rejects a batch size outside 1..100 (InMemoryRuntimeBindingRepository.java:186), so the strengthened assertion must not use limit 0 to force an empty page.
Acceptance criterion: with a second candidate present, deleting the cursor's > 0 direction (or the whole clause) in either repository's findRecoveryCandidates must turn the rewritten assertion red.
中文说明
[Suggestion] cursor 排他性断言跑在空页上,因此 allMatch 恒为真。
此时仓库里只有一条记录 —— fixture 自己的 binding —— 而排他 cursor 把它排除在外,所以流为空,断言不可能失败。
pageSize=0 allMatch=true totalRows=1
实测危害:把 findRecoveryCandidates 改成「afterBindingId != null 时永远返回空页」后,该测试类仍然 4/4 全绿 —— 文档承诺的「按 binding ID 排在排他 cursor 之后」在最关键的方向上无人校验。
建议:插入一条排序在 cursor 之后的候选,并精确断言整页内容(见上方代码),或者显式断言 isEmpty(),把期望写出来而不是留成空断言。
findRecoveryCandidates 会拒绝 1..100 之外的批量大小(InMemoryRuntimeBindingRepository.java:186),因此加强后的断言不能用 limit 0 来制造空页。
验收标准:在有第二条候选的情况下,删除任一仓库 findRecoveryCandidates 中 cursor 的 > 0 方向(或整个子句)时,改写后的断言必须变红。
— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)
| var pending = new CompletableFuture<Void>(); | ||
| try (var service = service(bindings, sessions, executions, provisioner(saved -> pending), Duration.ofMillis(100))) { | ||
| var recovery = service.recoverBinding(lost.getBindingId(), lost.getGeneration()).toCompletableFuture(); | ||
| assertThrows(java.util.concurrent.ExecutionException.class, () -> recovery.get(3, TimeUnit.SECONDS)); |
There was a problem hiding this comment.
[Suggestion] R1-3: The late-cleanup-completion test never reaches the expired-claim guard it is named for.
provisioner(saved -> pending) returns the test's own future, and safeStage returns that stage unchanged, so cleanupLost's orTimeout(operationLeaseDuration…) mutates the same CompletableFuture. With this test's 100 ms lease, recovery.get(3 s) returns only after that timeout has already completed it:
lease=PT0.1S outcome=recovery-failed:TimeoutException pendingDoneBefore=true completeDelivered=false
lease=PT30S outcome=recovery-still-pending pendingDoneBefore=false completeDelivered=true
So pending.complete(null) is a no-op and the assertions after it pass identically — delete those lines and the test stays green. A regression letting an expired claim retire the binding (e.g. removing the hasLiveOperationAt clause) would not be caught by the test whose name claims to pin exactly that.
Suggested fix: have the cleanup callback itself outlive the lease and then return a completed future, so the chain proceeds synchronously to the claim check:
RuntimeProvisioner provisioner = provisioner(saved -> {
Thread.sleep(leaseMillis * 2); // the lease lapses while cleanup runs
return CompletableFuture.completedFuture(null);
});then assert the failure cause is RuntimeBrokerException with code runtime_provision_fenced, and that the binding is still LOST with its session still active.
The guard is only reachable if the callback completes synchronously: recoverBinding wraps the whole operation with a second orTimeout(operationLeaseDuration…) (RuntimeBrokerService.java:1658).
Acceptance criterion: deleting the liveness clause || !current.hasLiveOperationAt(clock.instant()) from the in-memory repository must turn the rewritten test red.
中文说明
[Suggestion] 「清理迟到完成」这个测试根本没有走到它名字所指的过期 claim 守卫。
provisioner(saved -> pending) 返回的是测试自己的 future,而 safeStage 原样返回该 stage,所以 cleanupLost 的 orTimeout(operationLeaseDuration…) 修改的就是同一个 CompletableFuture。在本测试 100ms 的 lease 下,recovery.get(3s) 返回时该超时早已把它完成:两行实测见上。
因此 pending.complete(null) 是空操作,其后的断言只是照常通过 —— 删掉这几行测试依然全绿。一个让过期 claim 退休 binding 的回归(例如删除 hasLiveOperationAt 子句)不会被这个自称钉住它的测试抓到。
建议:让清理回调自身活过 lease 之后再返回已完成的 future,使调用链同步走到 claim 检查(见上方代码),然后断言失败原因是 code 为 runtime_provision_fenced 的 RuntimeBrokerException,且 binding 仍为 LOST、其 session 仍然活跃。
该守卫只有在回调同步完成时才可达:recoverBinding 还用第二个 orTimeout(operationLeaseDuration…) 包住整个操作(RuntimeBrokerService.java:1658)。
验收标准:删除内存仓库中的活跃性子句 || !current.hasLiveOperationAt(clock.instant()) 时,改写后的测试必须变红。
— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)
| if (broker.isDurableLocalProcess()) { | ||
| requireRecoveryDirectoryOutsideWorkspaces(broker, stateDirectory); | ||
| return LocalProcessRuntimeProvisioner.durable(command, stateDirectory, transport); | ||
| return LocalProcessRuntimeProvisioner.durable(command, stateDirectory, transport, |
There was a problem hiding this comment.
[Suggestion] R1-4: The trusted-local-reboot-recovery option is never enabled through EmbeddedRuntimeBroker in any test, so reverting this wiring keeps every suite green.
setTrustedLocalRebootRecovery(true) appears in exactly one test (EmbeddedRuntimeBrokerTest, and only to assert the negative guard), and every test of the trusted behaviour builds the provisioner by hand — FaultGateRig, FaultGateBroker, DurableLocalProcessRuntimeProvisionerTest, LocalRebootTestSupport — bypassing this composition root. Sweep of test-side setter sites: 1 of 1, negative guard only.
So reverting these two lines to the retained 3-argument durable(...) overload (which compiles) would mean observeDurable never returns WRITERS_STOPPED: a deployed service would keep every rebooted generation pinned forever, while unit, H2, MySQL, fault-gate and hosted jobs all stay green. The same applies to the coordinator construction a few lines above — drop it and recovery is null and the untested recoverSavedRuntimes() silently does nothing every 5 s.
Suggested fix: add a managed-agent-server test that constructs EmbeddedRuntimeBroker with trusted reboot recovery on, durable local-process on, a real state directory and a saved same-host/different-boot record, drives recoverSavedRuntimes(), and asserts the binding reaches RELEASED; or expose the constructed provisioner so the flag can be asserted directly.
Note the trusted path cannot be composed off Linux (LocalRuntimeStore.HostIdentity.linux()), and the guard at EmbeddedRuntimeBroker.java:178-181 already refuses the flag without durable local-process — so the new test needs a Linux/CI lane.
Acceptance criterion: that test must go red when this call is reverted to the 3-arg overload, or when the coordinator construction is dropped.
中文说明
[Suggestion] 没有任何测试通过 EmbeddedRuntimeBroker 打开 trusted-local-reboot-recovery,所以把这处接线改回去所有测试套件依然全绿。
setTrustedLocalRebootRecovery(true) 在整个仓库只出现在一个测试里(EmbeddedRuntimeBrokerTest,而且只断言反向守卫),而所有验证可信行为的测试都是手工构造 provisioner —— FaultGateRig、FaultGateBroker、DurableLocalProcessRuntimeProvisionerTest、LocalRebootTestSupport —— 绕开了这个组合根。测试侧 setter 站点实测:1 处,且只有反向守卫。
因此把这两行改回保留的 3 参 durable(...) 重载(它能编译)后,observeDurable 永远不会返回 WRITERS_STOPPED:部署中的服务会把每个重启后的 generation 永久钉住,而单元、H2、MySQL、fault-gate 与 hosted 各作业全部保持绿色。上面几行的 coordinator 构造同理 —— 去掉它 recovery 为 null,未被测试覆盖的 recoverSavedRuntimes() 每 5 秒静默空转。
建议:新增一个 managed-agent-server 测试,在可信重启恢复与持久本地进程均开启、有真实状态目录和「同宿主不同 boot」的已保存记录下构造 EmbeddedRuntimeBroker,驱动 recoverSavedRuntimes() 并断言 binding 到达 RELEASED;或者把构造出的 provisioner 暴露出来以便直接断言该开关。
注意可信路径无法在非 Linux 上组装(LocalRuntimeStore.HostIdentity.linux()),且 EmbeddedRuntimeBroker.java:178-181 的守卫已拒绝「无持久本地进程」时开启该开关,因此新测试需要 Linux/CI 通道。
验收标准:当这处调用被改回 3 参重载、或 coordinator 构造被删除时,新测试必须变红。
— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)
| assertThat(jdbc.queryForObject("SELECT COUNT(*) FROM managed_workspace_execution_lease WHERE binding_id = ?", | ||
| Long.class, original.getBindingId())).isZero(); | ||
| assertThat(fixture.bindings.findActive(original.getRequest())).isNull(); | ||
| assertThat(worker.isAlive()).isFalse(); |
There was a problem hiding this comment.
[Suggestion] R1-5: This assertion cannot fail, so the "no replacement worker was started" claim is unwitnessed.
worker is the ProcessHandle of the phase-1 worker that this test already killed and awaited 30 lines earlier:
worker.destroyForcibly(); worker.onExit().get(5, TimeUnit.SECONDS);
...
assertThat(worker.isAlive()).isFalse();Measured: after that pair worker.isAlive() is false, and a restarted process is alive under a different ProcessHandle. So make the restored Broker's recovery start a worker — the regression this line appears to guard — and line 103 stays green while the reader is told the no-restart property is asserted here.
Suggested fix: assert the property on something the code under test controls — capture the live worker pids before and after the release (ProcessHandle.current().descendants(), or a scan for a live node … managed-runtime-worker) and assert the set is unchanged — or drop the line and let the phase-2 must-not-run command carry the claim.
The phase-2 broker is built with the command List.of("must-not-run"), so a started worker would already make the provisioner throw — that is the weaker witness that exists today.
Acceptance criterion: with the replacement assertion, having recoverBinding start a runtime for the released binding must turn it red; today line 103 cannot distinguish that mutation.
中文说明
[Suggestion] 这个断言不可能失败,因此「没有启动替代 worker」这条声明无人见证。
worker 是本测试在 30 行之前就已经杀掉并等待结束的 phase-1 worker 的 ProcessHandle(见上方代码)。
实测:那一对调用之后 worker.isAlive() 为 false,而重启出来的进程是活的、且持有另一个 ProcessHandle。所以让恢复后的 Broker 启动一个 worker(正是这行看似要防的回归),第 103 行依然全绿,而读者却以为「不重启 worker」这条性质在这里被断言了。
建议:把断言建立在被测代码能控制的对象上 —— 在释放前后采集存活的 worker pid(ProcessHandle.current().descendants(),或扫描存活的 node … managed-runtime-worker)并断言集合不变 —— 或者删掉这一行,让 phase-2 的 must-not-run 命令承担该声明。
phase-2 的 broker 用命令 List.of("must-not-run") 构造,因此真启动了 worker 本来就会让 provisioner 抛错 —— 这是目前已有的、更弱的见证。
验收标准:换成新断言后,让 recoverBinding 为已释放的 binding 启动 runtime 必须使它变红;目前第 103 行区分不出这个变异。
— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)
| assertEquals(RuntimeBindingRecord.State.LOST, bindings.recoverLost(sessions, executions, lost).getState()); | ||
| assertTrue(executions.hasActiveByBinding(lost.getBindingId(), lost.getGeneration())); | ||
| assertEquals(1, sessions.countActiveByBinding(lost.getBindingId(), lost.getGeneration())); | ||
| assertThrows(RuntimeBrokerException.class, () -> bindings.completeSessionRelease(sessions, fixture.session)); |
There was a problem hiding this comment.
[Suggestion] R1-6: This refusal comes from the generation-state guard, not from a live-holder check, so the assertion cannot fail for the reason its position implies.
completeSessionRelease performs one guard per repository — RuntimeAdmission.requireRelease(binding, session) — which reads only binding/session identity, scope and state; no disjunct consults hasActiveByBinding or countActiveByBinding. Measured, the identical call with this fixture:
[READY, 205 executions, 1 session] -> RETURNED RELEASED
[LOST, 205 executions, 1 session] -> THREW runtime_reconciliation_required
[READY, 0 executions, 1 session] -> RETURNED RELEASED
So a change that let a release complete while executions for the generation are still live is not caught here; the only place holders fence a transition is recoverLost/finishLostRecovery.
Suggested fix: keep the assertion — it does pin that a managed LOST generation cannot be released through ordinary completeSessionRelease — but re-anchor it: assert the error code (runtime_reconciliation_required) so the pinned rule is explicit, and add the control it is missing (a READY-state binding with live executions for the same generation).
The LEGACY case must keep failing (ProcessCrashFaultGateTest asserts runtime_reconciliation_required for a restarted-Broker release), so add the managed-shape control rather than loosening that assertion.
Acceptance criterion: the re-anchored assertion must go red when the managed LOST refusal is removed from RuntimeAdmission.requireRelease.
中文说明
[Suggestion] 这个拒绝来自 generation 状态守卫,而不是「持有者仍然活跃」的检查,因此该断言不可能以它所在位置暗示的原因失败。
completeSessionRelease 在每个仓库里只有一个守卫 —— RuntimeAdmission.requireRelease(binding, session) —— 它只读取 binding/session 身份、scope 与状态;没有任何分支去查 hasActiveByBinding 或 countActiveByBinding。用本 fixture 做同一个调用,实测三行结果见上。
也就是说,一个「该 generation 仍有活跃执行却允许 release 完成」的改动在这里抓不到;真正用持有者栅栏住状态跃迁的地方是 recoverLost/finishLostRecovery。
建议:保留该断言 —— 它确实钉住了「managed LOST generation 不能通过普通 completeSessionRelease 释放」—— 但重新锚定:断言错误 code(runtime_reconciliation_required)把被钉住的规则写明确,并补上它缺失的对照(同 generation、READY 状态且存在活跃执行的 binding)。
LEGACY 场景必须继续失败(ProcessCrashFaultGateTest 对重启后 Broker 的 release 断言 runtime_reconciliation_required),所以应新增 managed 形态的对照,而不是放宽那条断言。
验收标准:当 RuntimeAdmission.requireRelease 中 managed LOST 的拒绝被移除时,重新锚定的断言必须变红。
— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)
| throw new AssertionError("Maintenance cannot call a dead worker: " + method); | ||
| }); | ||
| return new RuntimeBrokerService(id -> { throw new AssertionError("Maintenance cannot resolve current grants"); }, | ||
| provisioner, transport, bindings, sessions, executions, "maintenance", lease, lease); |
There was a problem hiding this comment.
[Suggestion] R1-7: The refusal half of the new recoverBinding is pinned by no test.
Every call site in the tree passes a generation read from the same record and uses a single owner id, and no test races a second owner: 14 recoverBinding( sites, zero with a non-matching generation, and RuntimeRecoveryCoordinatorTest mocks the service. Measured — each of the three refusals neutralized in turn, whole module re-run:
baseline Tests run: 430, Failures: 0, Errors: 0, Skipped: 2
precondition narrowed to `record == null` Tests run: 430, Failures: 0, Errors: 0, Skipped: 2
`putIfAbsent` refusal removed Tests run: 430, Failures: 0, Errors: 0, Skipped: 2
rival-broker refusal degraded to a plain findById Tests run: 430, Failures: 0, Errors: 0, Skipped: 2
So the guard that keeps a loser Broker from starting a second cleanup of the same generation can be dropped without a red test — the declared retryable 503 runtime_reconcile_in_progress would degrade to an unclassified failure inside the recovery stage.
Suggested fix: add two cases reusing this file's Fixture/provisioner/service helpers — one calling recoverBinding(lost.getBindingId(), lost.getGeneration() + 1) (and a mismatched provisioner kind) asserting runtime_broker_recovery_blocked with the binding still LOST; and one where a second RuntimeBrokerService with a different owner id holds the claim, asserting runtime_reconcile_in_progress while that claim is live.
The rival case needs a different owner id, not a second service with the same one — the claim is taken as claimOperation(bindingId, brokerOwnerId, operationLeaseDuration) (RuntimeBrokerService.java:88,1613).
Acceptance criterion: the new cases must go red when the generation/provisioner precondition, the putIfAbsent refusal, or the rival-owner refusal is removed.
中文说明
[Suggestion] 新增的 recoverBinding 的「拒绝分支」没有任何测试覆盖。
仓库里每个调用点都传入从同一条记录读出的 generation,并使用同一个 owner id,也没有测试制造第二个 owner 竞争:全树 14 处 recoverBinding(,没有一处传入不匹配的 generation,而 RuntimeRecoveryCoordinatorTest 是对 service 打桩。实测:把这三个拒绝分支逐个去掉后整个模块重跑,结果与基线一模一样(Tests run: 430, Failures: 0, Errors: 0, Skipped: 2)。
因此「阻止落败的 Broker 对同一 generation 发起第二次清理」的守卫可以在没有测试变红的情况下被删除 —— 声明为可重试的 503 runtime_reconcile_in_progress 会退化成恢复阶段里一个未分类的失败。
建议:复用本文件的 Fixture/provisioner/service 辅助方法补两个用例 —— 一个调用 recoverBinding(lost.getBindingId(), lost.getGeneration() + 1)(以及 provisioner kind 不匹配),断言 runtime_broker_recovery_blocked 且 binding 仍为 LOST;另一个用不同 owner id 的第二个 RuntimeBrokerService 持有 claim,断言在其存活期间得到 runtime_reconcile_in_progress。
竞争用例必须用不同的 owner id,而不是同一个 owner id 的第二个 service —— claim 是通过 claimOperation(bindingId, brokerOwnerId, operationLeaseDuration) 取得的(RuntimeBrokerService.java:88,1613)。
验收标准:当 generation/provisioner 前置条件、putIfAbsent 拒绝、或对手 owner 拒绝被移除时,新用例必须变红。
— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)
| fixture.session.sessionId()); | ||
| jdbc.update("UPDATE managed_workspace_registry SET storage_id = 'changed', state = 'REMOVED' WHERE tenant_id = ?", | ||
| fixture.tenant); | ||
| var noMounts = new WorkspaceRuntimeResolver(store, authority, new ManagedAgentProperties()); |
There was a problem hiding this comment.
[Suggestion] R1-8: The phase-2 authorization canary sits on a dependency recoverBinding cannot reach, so a tolerated authorization consultation leaves this IT green.
RuntimeBrokerService.recoverBinding never reads its sessionResolver — that field is read only in resolveScope, reached from the acquire/warm paths — and the phase-2 broker here is never asked to acquire or warm, so the AssertionError canary cannot fire for this test's only call. The objects the release path does hold are noMounts and authority, and both are observed only through the outcome.
Measured, injecting the shape this canary is meant to exclude (resolveScope(...).exceptionally(error -> null) into cleanupLost):
baseline consulted=0 state=RELEASED loss=true
mutant consulted=1 state=RELEASED loss=true
and RuntimeMaintenanceRecoveryTest stays Tests run: 4, Failures: 0 under the mutant, because safeStage catches Error as well as RuntimeException (RuntimeBrokerService.java:2574). The canaries make a cheating recovery fail loudly only when the consultation does not tolerate the refusal.
Suggested fix: put the canary on the seam the recovery path actually crosses. Build the phase-2 WorkspaceRuntimeResolver around a spy of the store or of the execution authority and, after recoverBinding, assert the authorizing interaction never happened:
verify(authority, never()).authorize(any());WorkspaceRuntimeResolver is final, so instrument a collaborator it delegates to (AgentStateStore sessions, WorkspaceExecutionStore authority) or spy the concrete instance.
Acceptance criterion: with the spy assertion in place, inserting noMounts.resolve(...) or authority.authorize(...) into the LOST/release branch must turn the IT red; today that mutation leaves the whole IT green.
中文说明
[Suggestion] phase-2 的「授权探测哨」挂在了 recoverBinding 结构上无法触及的依赖上,因此一次「容错吞掉拒绝」的授权查询会让这个 IT 保持全绿。
RuntimeBrokerService.recoverBinding 从不读取它的 sessionResolver —— 该字段只在 resolveScope 里被读,而那只从 acquire/warm 路径到达 —— 而这里 phase-2 的 broker 从未被要求 acquire 或 warm,所以这个测试唯一的一次调用不可能触发 AssertionError 哨兵。释放路径真正持有的对象只有 noMounts 与 authority,而两者都只通过最终结果被观察。
实测:注入这个哨兵本该排除的形态(把 resolveScope(...).exceptionally(error -> null) 放进 cleanupLost)后,基线与变异体的四个可观测量完全一致(见上),且变异体下 RuntimeMaintenanceRecoveryTest 仍是 Tests run: 4, Failures: 0 —— 因为 safeStage 同时捕获 Error 与 RuntimeException(RuntimeBrokerService.java:2574)。只有当授权查询不容忍拒绝时,哨兵才会「大声失败」。
建议:把哨兵挪到恢复路径真正跨越的接缝上。用 store 或执行授权对象的 spy 构造 phase-2 的 WorkspaceRuntimeResolver,并在 recoverBinding 之后断言授权交互从未发生(见上方 verify(authority, never()).authorize(any());)。
WorkspaceRuntimeResolver 是 final,因此需要插桩它委托的协作者(AgentStateStore sessions、WorkspaceExecutionStore authority),或对具体实例做 spy。
验收标准:在加入 spy 断言后,把 noMounts.resolve(...) 或 authority.authorize(...) 插入 LOST/release 分支必须让该 IT 变红;目前这个变异下整个 IT 仍然全绿。
— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)
…M#12977) * feat(managed-agent): Recover Workspace holders after trusted local reboot (cherry picked from commit 8c2b626) * fix(runtime-broker): Serve healthy waiters during maintenance * fix(sdk-java): Route recovery IT to hosted CI (QwenLM#12869) * fix(sdk-java): Add audited Hosted Workspace operator recovery * codex: address PR review feedback (QwenLM#12977) * fix(sdk-java): Validate operator evidence after recovery prepare * fix(sdk-java): Fence operator recovery against concurrent placement * fix(sdk-java): Upgrade fault-gate H2 test dependency
















What this PR does
Adds opt-in recovery for trusted local Runtime workloads after the original Linux host reboots. Recovery observes the saved generation, preserves terminal uncertainty for lost execution results, clears only its original Workspace storage holder, and retires the generation after cleanup succeeds. A bounded background scan can make progress even after the actor loses access or the product Session is deleted. Maintenance never starts a replacement worker or replays an execution.
This is W0e-3. Its prerequisites #12839 and #12865 are merged. This PR now targets main and contains only this recovery slice plus the healthy-request maintenance fix. Independent real-host round 4 acceptance covers the current head
62584d31: healthy-request load, physical reboot, power-cut recovery, crash retries, and both Java CI families passed.Why it's needed
Durable adoption preserves a live worker across Broker restart, but worker death alone cannot prove that detached descendants stopped writing. Reclamation therefore needs separate physical evidence and cleanup of the original holder. Current Session authorization and placement selection cannot safely substitute for the saved ownership during that cleanup.
Reviewer Test Plan
How to verify
62584d31and on a merge withmainat1b696297.Evidence (Before & After)
No UI change. Before: live workers could be adopted, but trusted reboot had no production stop-proof producer, original-holder cleanup or authorization-independent scan. After: portable real-worker and SQL tests exercise the recovery chain and prove worker-only death stays blocked. Synthetic boot transitions alone do not certify physical reboot; see the independent real-host round 4 report.
Validation details are recorded in the separate E2E report comment.
Tested on
62584d31Environment (optional)
macOS arm64, JDK 21, Node.js, local MySQL Community 8.4.11, bundled Runtime worker; independent Debian 12 and Ubuntu 24.04 aarch64 VMs for physical reboot acceptance. No external model service is used by the new gates.
Risk & Scope
Design: English, 简体中文. Both versions describe the same decisions, constraints and physical acceptance evidence.
Linked Issues
Related to #12380, #12670 and #12766. These broader issues are not automatically closed by this slice.
中文说明
本 PR 的改动
增加显式启用的可信本地 Runtime 宿主重启恢复。恢复观察原保存代数,保留丢失执行结果的终态不确定性,只清理其原 Workspace 存储持有者,并在清理成功后退休原代数。有界后台扫描可在 actor 被撤权或产品 Session 删除后继续推进。维护过程不启动替代 worker,也不重放执行。
这是 W0e-3。前置 #12839 和 #12865 已合并;本 PR 现面向 main,只包含本切片及健康请求与维护扫描并发时的修复。独立真机第四轮验收已覆盖当前 head
62584d31:健康请求负载、物理重启、断电恢复、崩溃重试及两组 Java CI 均通过。为什么需要
持久接管可以在 Broker 重启后保留存活 worker,但仅 worker 死亡不能证明逃逸后代停止写入。因此回收需要独立物理证据和原 holder 清理。清理期间不能用当前 Session 授权和位置选择替代保存的原归属。
评审测试计划
如何验证
62584d31和与main的1b696297合并结果上完成此物理门禁。前后证据
无 UI 变更。此前可接管存活 worker,但可信重启没有生产 stop 证据来源、原 holder 清理和独立于授权的扫描。现在,可移植真实 worker 与 SQL 测试覆盖恢复链路,并证明仅 worker 死亡继续阻断。仅模拟 boot 变化不能证明实际物理重启;独立真机第四轮报告提供了物理证据。
验证详情见独立 E2E 报告评论。
测试平台
62584d31上独立完成真实重启、断电、F1 负载及 CI 验收环境
macOS arm64、JDK 21、Node.js、本地 MySQL Community 8.4.11、打包 Runtime worker;独立物理重启验收使用 Debian 12 和 Ubuntu 24.04 aarch64 虚拟机。新增门禁不调用外部模型服务。
风险与范围
设计:English、简体中文。两种语言的决策、约束和物理验收证据一致。
关联 Issue
关联 #12380、#12670、#12766。本切片不自动关闭这些范围更大的 issue。