Skip to content

feat(managed-agent): Recover Workspace holders after trusted local reboot (W0e-3) - #12869

Merged
doudouOUC merged 3 commits into
mainfrom
codex/managed-workspace-w0e-reboot
Sep 29, 2026
Merged

doudouOUC merged 3 commits into
mainfrom
codex/managed-workspace-w0e-reboot

Conversation

@doudouOUC

@doudouOUC doudouOUC commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Adds opt-in recovery for trusted local Runtime workloads after the original Linux host reboots. Recovery observes the saved generation, preserves terminal uncertainty for lost execution results, clears only its original Workspace storage holder, and retires the generation after cleanup succeeds. A bounded background scan can make progress even after the actor loses access or the product Session is deleted. Maintenance never starts a replacement worker or replays an execution.

This is W0e-3. Its prerequisites #12839 and #12865 are merged. This PR now targets main and contains only this recovery slice plus the healthy-request maintenance fix. Independent real-host round 4 acceptance covers the current head 62584d31: healthy-request load, physical reboot, power-cut recovery, crash retries, and both Java CI families passed.

Why it's needed

Durable adoption preserves a live worker across Broker restart, but worker death alone cannot prove that detached descendants stopped writing. Reclamation therefore needs separate physical evidence and cleanup of the original holder. Current Session authorization and placement selection cannot safely substitute for the saved ownership during that cleanup.

Reviewer Test Plan

How to verify

  • With recovery disabled, retain the existing behavior. Reject reboot recovery unless durable local provisioning is also enabled.
  • While maintenance observes a healthy binding, concurrent authorized warm and acquire must succeed; uncertain or lost observations must remain blocked.
  • With an intact original registration on the same host but a changed boot, persist the original tombstone and both loss and stop evidence before clearing any holder. Missing or mismatched identity must remain blocked.
  • Kill only the worker while its detached child keeps writing: the original placement stays pinned, receipts remain uncertain, and no replacement is launched.
  • Revoke the creator's access, delete the product Session and change the Registry or mount configuration. Maintenance should still clean only the original saved holder. An old cleanup retry must preserve a newer holder; a late acquire must not restore the lost holder.
  • Interrupt cleanup after the holder clear, or expire its claim. Retrying must preserve original receipt identity and release the remaining pins only after cleanup completes. More than one transaction's worth of executions must converge without replay.
  • On a dedicated supported Linux host, preserve disk and SQL across a real reboot, check the old writer nonce stops permanently, then admit a new authorized generation. An independent reviewer completed this physical gate on the current head 62584d31 and on a merge with main at 1b696297.

Evidence (Before & After)

No UI change. Before: live workers could be adopted, but trusted reboot had no production stop-proof producer, original-holder cleanup or authorization-independent scan. After: portable real-worker and SQL tests exercise the recovery chain and prove worker-only death stays blocked. Synthetic boot transitions alone do not certify physical reboot; see the independent real-host round 4 report.

Validation details are recorded in the separate E2E report comment.

Tested on

OS Status
🍏 macOS ✅ Java/H2/MySQL and real process tests; test-only boot identities
🪟 Windows N/A — durable local recovery is Linux-only
🐧 Linux ✅ Independent real reboot, power-cut, F1 load and CI acceptance at 62584d31

Environment (optional)

macOS arm64, JDK 21, Node.js, local MySQL Community 8.4.11, bundled Runtime worker; independent Debian 12 and Ubuntu 24.04 aarch64 VMs for physical reboot acceptance. No external model service is used by the new gates.

Risk & Scope

  • Main risk or tradeoff: recovery authority spans the Broker and Workspace persistence. Maintainer review is requested for this cross-package feature and the cleanup contract. The supported deployment is one Linux host, administrator-managed persistent local disks and trusted tools.
  • Not validated / out of scope: bare metal and x86_64; hostile same-UID tools; remote writers, disk/VM snapshot restoration and external jobs that recreate writers; Kubernetes; Hosted Turn continuation/history settlement; public bound-message enablement. The full Workspace execution capability remains false.
  • Breaking changes / migration notes: the option defaults off and requires durable local mode. Existing SQL tables suffice; this slice adds no migration. Embedders with custom binding repositories must implement bounded candidate discovery and finalization after holder cleanup. Managed provisioners without the cleanup callback fail closed. The prerequisite evidence migration in feat(managed-agent): Add W0e terminal recovery fences #12839 is merged.

Design: English, 简体中文. Both versions describe the same decisions, constraints and physical acceptance evidence.

Linked Issues

Related to #12380, #12670 and #12766. These broader issues are not automatically closed by this slice.

中文说明

本 PR 的改动

增加显式启用的可信本地 Runtime 宿主重启恢复。恢复观察原保存代数,保留丢失执行结果的终态不确定性,只清理其原 Workspace 存储持有者,并在清理成功后退休原代数。有界后台扫描可在 actor 被撤权或产品 Session 删除后继续推进。维护过程不启动替代 worker,也不重放执行。

这是 W0e-3。前置 #12839 和 #12865 已合并;本 PR 现面向 main,只包含本切片及健康请求与维护扫描并发时的修复。独立真机第四轮验收已覆盖当前 head 62584d31:健康请求负载、物理重启、断电恢复、崩溃重试及两组 Java CI 均通过。

为什么需要

持久接管可以在 Broker 重启后保留存活 worker,但仅 worker 死亡不能证明逃逸后代停止写入。因此回收需要独立物理证据和原 holder 清理。清理期间不能用当前 Session 授权和位置选择替代保存的原归属。

评审测试计划

如何验证

  • 默认关闭时保留现有行为;未启用本地持久 provisioning 时拒绝开启重启恢复。
  • 维护观察健康 Binding 时,并发的已授权 warm 和 acquire 应成功;不确定或失联观察仍须阻断。
  • 原登记完整、同宿主 boot 改变时,先持久化原墓碑和 loss/stop 证据,再清除 holder。身份缺失或不匹配必须继续阻断。
  • 仅杀死 worker、让其分离子进程持续写入:原位置保持占用,回执保持不确定,不启动替代者。
  • 撤销创建者权限、删除产品 Session、改变 Registry 或 mount 配置后,维护仍只清理原保存 holder。旧清理重试必须保留新 holder;迟到 acquire 不能恢复已失联的原 holder。
  • holder 清理提交后中断流程,或让 claim 过期;重试必须保留原回执身份,只在清理完成后释放剩余占用。超过单事务批量的执行应逐批收敛,不能重放。
  • 在专用受支持 Linux 宿主上保留磁盘和 SQL 进行真实重启,验证旧写入者 nonce 永久停止,再准入新授权代数。独立评审者已在当前 head 62584d31 和与 main 的 1b696297 合并结果上完成此物理门禁。

前后证据

无 UI 变更。此前可接管存活 worker,但可信重启没有生产 stop 证据来源、原 holder 清理和独立于授权的扫描。现在,可移植真实 worker 与 SQL 测试覆盖恢复链路,并证明仅 worker 死亡继续阻断。仅模拟 boot 变化不能证明实际物理重启;独立真机第四轮报告提供了物理证据。

验证详情见独立 E2E 报告评论。

测试平台

OS 状态
🍏 macOS ✅ Java/H2/MySQL 和真实进程测试;使用仅测试可用的 boot 身份
🪟 Windows N/A — 本地持久恢复仅支持 Linux
🐧 Linux ✅ 62584d31 上独立完成真实重启、断电、F1 负载及 CI 验收

环境

macOS arm64、JDK 21、Node.js、本地 MySQL Community 8.4.11、打包 Runtime worker;独立物理重启验收使用 Debian 12 和 Ubuntu 24.04 aarch64 虚拟机。新增门禁不调用外部模型服务。

风险与范围

  • 主要风险或取舍:恢复权限跨 Broker 和 Workspace 持久层,请 maintainer 评审跨包功能及清理契约。支持的部署为 Linux 单宿主、管理员管理的本地持久磁盘及可信工具。
  • 未验证或不在范围:物理机和 x86_64;恶意同 UID 工具;远程写入者、磁盘/VM 快照恢复、重新创建写入者的外部任务;Kubernetes;Hosted Turn 接续和历史结算;公开绑定消息入口。完整 Workspace 执行能力仍为 false。
  • 兼容性和迁移:默认关闭且要求持久本地模式。复用已有 SQL 表,本切片无迁移。自定义 Binding 仓库需实现有界候选发现和 holder 清理后的最终释放;没有清理回调的 managed provisioner 保持阻断。feat(managed-agent): Add W0e terminal recovery fences #12839 的前置证据迁移已合并。

设计:English、简体中文。两种语言的决策、约束和物理验收证据一致。

关联 Issue

关联 #12380、#12670、#12766。本切片不自动关闭这些范围更大的 issue。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

W0e-3 E2E and regression report

Environment: macOS arm64, JDK 21, Node.js, bundled Runtime worker, isolated MySQL Community 8.4.11. No external model calls. The dedicated MySQL instance was shut down after validation; test processes were checked for leaks.

Passing evidence

Group Result
Broker ordinary Java/H2 tests 407 passed; latest maintenance assertions additionally rerun in 3 targeted cases
Managed Agent ordinary Java/H2 tests 145 passed; updated transaction fixture additionally rerun in its H2 case
Broker MySQL integration 3 passed in a fresh database; updated candidate-query assertions rerun in the maintenance case
Managed Agent MySQL integration 10 existing cases passed; new holder-recovery case passed after fixing its transaction fixture
Existing real-process Stage F gates 38 passed, covering adoption, context installation, process loss, response loss, concurrency and storage
New escaped-writer gates 2 passed after fixing their fixture path; both boot protocols keep the original placement pinned while a real detached child continues writing after confirmed loss
Spring + H2 + real worker 1 passed through actual context acquisition, worker loss, original-holder cleanup and retirement after a synthetic boot transition
Build / typecheck / bundle Passed
Java Checkstyle / changed Markdown and YAML formatting Passed

All final targeted runs had zero failures, errors or skips. Counts refer to distinct cases; targeted reruns are not added twice. This is evidence across the complete runs and affected-case reruns, not a claim that an initially failing invocation was green.

What the new tests establish

  • A same-host changed boot emits matching journal-loss and writer-stop evidence only after durable retirement of the original registration. Interrupted startup does not need an invented lease.
  • More than 100 executions become terminal in bounded transactions. Unknown outcomes remain ABANDONED without replay or fabricated results.
  • Cleanup failures and expired claims retain remaining pins. A committed holder clear can be retried, and an old cleanup cannot erase a different current holder.
  • Actual MySQL and H2 row locking rejects a late acquire after the loss fence. Cleanup still uses original ownership after grants are revoked, the product Session is deleted, and Registry/mount data changes.
  • The real escaped child is alive and its file grows after the restarted Broker has confirmed the original Runtime is lost. No replacement worker is launched.
  • The independent scanner has a rotating cursor and does not overlap batches. Maintenance never consults the current Session resolver or provisions a replacement.

Failures found and corrected during validation

An initial MySQL run reused a fixed-prefix database from earlier verification; its generation assertion correctly failed. A fresh isolated database passed. The new MySQL fixture initially bypassed Spring's transaction proxy; adding its explicit creation transaction made the real SQL path pass. The initial escaped-writer test checked the test root instead of its actual Workspace; the directory was corrected and the growth baseline moved after confirmed loss. The two affected gates then passed. The first Stage F run passed all 38 unchanged gates; only these two fixture cases required rerunning.

Pending physical acceptance

A real Linux host reboot has not been run. macOS uses test-only host identity, the Spring worker chain uses a synthetic boot change, and MySQL tests use synthetic stop evidence. Native Linux identity and a real preserved-disk reboot with an escaped writer remain a separate acceptance gate. This PR remains draft until that gate can be performed. Hosted Turn continuation, public Workspace message enablement and hostile-tool containment are outside this slice.

中文

本地完成 Broker 407 项、Spring 145 项常规测试,14 项 MySQL 集成用例,以及 38 项既有真实进程故障门禁。新逃逸写入者的 2 项门禁修复夹具目录后通过;另有 1 项 Spring+H2+真实 worker 恢复链路通过。更新后的维护断言和事务夹具做了定向补验,未重复计入数量。build、typecheck、bundle、Checkstyle 和改动文档格式检查均通过,最终补验无失败或跳过。

验证期间发现并修复了旧测试库复用、原始 MySQL store 夹具缺少事务、逃逸写入者测试目录错误,并将增长测量起点移至 Broker 确认失联之后。真实 SQL 验证了分批终态化、清理重试、迟到 acquire 屏障和保留新 holder;真实子进程验证了 worker 退出后继续写入时不能复用原位置。

未执行真实 Linux 整机重启。 模拟 boot、MySQL 测试证据和真实 worker 链路分别报告,不能合称真实重启验收。PR 保持 draft,待专用 Linux 环境验证后再推进。

@github-actions

Copy link
Copy Markdown
Contributor

Please do not rebase or force-push to an active PR as it invalidates existing review comments. Note for future reference, the bots always squash all changes into a single commit automatically as part of the integration.

中文

请勿对活跃的 PR 执行 rebase 或 force-push,因为这会使已有的评审评论失效。另外,供日后参考:作为集成流程的一部分,机器人始终会自动将所有改动压缩(squash)为单个提交。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Rebased this draft PR onto W0e-2 commit 7236322a4, which now includes the W0e-1 transient attestation fix. New head: f8bf5d746. git range-diff shows the W0e-3 commit unchanged; the only rebase conflict in the stack was resolved in the paired W0e-2 design documents.

Validation of this exact stacked head: Runtime Broker 410 tests with 0 failures/errors and 1 conditional skip; managed agent server 145 tests with 0 failures/errors; 12 real-process reboot and durable-runtime fault gates passed; Checkstyle reported 0 violations. The earlier MySQL and broader fault-gate results remain recorded above; this rebase added only the W0e-1 Java fix and paired documentation, with no TypeScript changes. This PR stays draft until the physical Linux reboot acceptance gate and the documented pre-upgrade handling of historical seeded FAILED bindings are complete.

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

Real-host verification: the pending Linux reboot acceptance, run for real

Verified on a dedicated Linux host at f8bf5d74 (current head). I started at 66bde476; the stack was rebased while I was testing. The W0e-3 commit has the same git patch-id on both heads and the TypeScript tree is unchanged, so I kept the first results and re-ran everything that matters on the new head.

Summary

  • The reboot recovery works on a real host. Four real reboots of the same machine (two systemctl reboot, two power cuts): every saved generation was reclaimed by the scheduled scan alone, 21 s after the host went down in the hands-off runs. The escaped writers stopped with the old boot, receipts kept their meaning, only the original holders were cleared, and no worker was started by the recovery.
  • The controls stay blocked: option off, a foreign machine-id, worker-only death, systemd soft-reboot, and missing / damaged / foreign registration records.
  • One thing to fix before anyone enables the option with Hosted turns (F1): the 5-second scan also observes healthy bindings, and a request that arrives during that observation gets 503 runtime_reconciliation_required. The Hosted Harness turns a failed acquire into a recovery-blocked Session. 6 of 871 healthy turns hit it. A 13-line candidate and a unit test are attached.
  • With the option off (the default) nothing changes: 0 of 847 turns refused, and the saved generations stay blocked after a reboot exactly as before.

For the merge decision: from the real-host side the physical acceptance gate is met. I would fix F1 before merging (the change is small), add the two gap tests from section 5, and add the soft-reboot sentence. The pre-upgrade handling of historical FAILED bindings that the author lists as still open belongs to #12839 and was not part of this run. The merge order stays #12839 → #12865 → this PR.

All times are UTC, 2026-09-27.

1. Physical reboot acceptance

Real reboot runs

Reviewer test plan item Result on the real host
Same host, changed boot: tombstone + loss and stop evidence, then holder cleanup Pass in R1, R2, R7. Every recovered registration is RETIRED and still carries the original boot id. Both evidence records name the original machine, boot and resource.
Worker-only death stays pinned, no replacement Pass. The scan marks the binding LOST with loss evidence only. The holder stays, the escaped writer keeps writing (+4200 B in 3 s), warm gets 503 runtime_broker_runtime_lost, and no worker is launched. After the reboot the same binding is released and keeps its first loss evidence.
Access revoked, Session deleted, Registry and mount changed Pass (arm C). The original holder is cleared. warm and acquire for that Session are refused with 409 workspace_unavailable, and no new generation appears.
More than one transaction's worth of executions Pass. 130 PREPARED plus the in-flight calls all become ABANDONED without a result. SETTLED receipts keep their results. Nothing is replayed.
Interrupted startup without a lease Pass (arm E). A worker that died before READY is RECOVERY_BLOCKED and refuses warm on the same boot. After the reboot it is released.
Records written just before a power cut Pass (arm G). The binding was warmed 0.5 s before the power cut. Its registration survived and it was released.
Dedicated host, real reboot, new authorized generation Pass. An authorized warm creates generation 2 in 543 ms, acquire claims the holder again, a new call runs, and the old Runtime Session answers 404.

The JVM ran in Asia/Shanghai and MySQL in UTC on purpose. The lease checks in releaseLost were not affected.

2. Controls

Controls

One more control at f8bf5d74 (run R8, records changed while the server was stopped, then a real reboot):

Registration record at reboot After the reboot
deleted RECOVERY_BLOCKED, no evidence, holder kept, warm → 409 runtime_broker_resource_conflict
overwritten with garbage same
hostId edited to another machine same
untouched RELEASED with both evidence records

Enabling the option without durable mode stops the server at startup with Trusted reboot recovery requires durable local-process provisioning.

3. After the recovery

After the recovery

4. F1: the scan refuses healthy requests, and a Hosted Session that hits it is blocked

Healthy traffic

What happens:

  1. findRecoveryCandidates includes READY bindings (JdbcRuntimeBindingRepository.java:340), so each scan calls recoverBinding for every healthy durable binding.
  2. recoverBinding registers a reservation in bindingOperations while it claims, observes, re-attests and rewrites the binding (RuntimeBrokerService.java:1549). It always completes that reservation with 503 runtime_reconciliation_required, also when the observation found the binding healthy (:1606).
  3. ensureBinding hands the reservation to any request that arrives meanwhile (:825). That covers warm and the first acquire of a Runtime Session.
  4. The Hosted Harness treats a failed acquire as possibly effective (hosted-workspace-tool-turn.ts:183-197) and raises HostedToolRecoveryRequiredError. The Session becomes recovery-blocked and every later prompt gets 409 hosted_turn_recovery_required.

The window is about 20 ms in every 5 s per binding, so a Session with a few hundred tool turns is likely to hit it. Nothing is wrong with the Runtime when it happens.

Candidate (diff, test): when the observation ends with a healthy adoption, complete the reservation with that result instead of the 503. Waiters have already resolved current authorization before they reach ensureBinding, so this is the same hand-over reconcileBinding already does. With it: 0 of 837 turns refused, the Hosted chain completes under the same delivery, all suites pass, and a real reboot still recovers (R6).

An alternative is to skip bindings that are live in this Broker process. That would also remove three row writes per healthy binding every 5 s, but the scan would no longer notice the death of an idle worker. Your call.

5. Suites on Linux and mutation results

Suites and mutation

26 single-site mutants of the PR's production changes at f8bf5d74, each run against the unit suites, then the fault gates, the real-worker IT and the MySQL ITs:

  • 15 killed. 13 by the unit suites. Two only by the real-process tests: "the same boot counts as a reboot" (both LocalRebootFaultGateTest cases and WorkspaceRecoveryWorkerIT) and "cleanup reported without clearing the SQL holder" (WorkspaceRecoveryWorkerIT).

  • 5 survive because a second guard covers the same path. The stale-version and other-claim checks in releaseLost mask each other; removing both is caught by WorkspaceRecoveryTest and ManagedAgentMySqlIT.

  • 6 survive everything. Two of them guard safety properties:

    • rebooted() without the host check: another machine with another boot id then produces stop evidence.
    • the default recoverResources allowing managed cleanup, which the description says must fail closed.

    The PR code is right in both cases (see the controls), but no test would notice a regression. RebootRecoveryGapTest pins both: 2 of 2 pass at the head, and each fails under its mutant.
    The other four are smaller: claim ignoring the Runtime Session row, recoverBinding ignoring the expected generation, overlapping scan batches, and evidence matching any lease.

Full list and verdicts: mutants, pass 1, pass 2.

This PR has no Java CI: sdk-java.yml only runs for PRs whose base is main, and this one is stacked. main has moved three commits since the merge base, none of them under packages/sdk-java; git merge-tree is clean and the merged tree has one migration per version (V1 to V16).

6. Smaller notes

  • Soft-reboot (docs). systemctl soft-reboot (systemd 254 and later, so Ubuntu 24.04) kills every writer but keeps the kernel boot id. The bindings become LOST without stop evidence and stay pinned until a full reboot. That is the safe side. One sentence in the README would save an operator some confusion.
  • Flaky tests seen under load, none touched by this PR. ManagedAgentServerIntegrationTest failed in 2 of 11 runs with a different test each time. ConcurrencyStorageFaultGateTest.aDatabaseOutageDuringTheResultCommitNeverSettles failed in 1 of 8 runs.
  • Context for the receipts. A foreground call sent through POST /executions that runs longer than 30 s is reported UNKNOWN even while a client polls it. This happens with the option off too, so it is not caused by this PR. It is why the in-flight calls in R1 and R2 were already UNKNOWN before the host went down.

What this run does not cover

Bare metal and x86_64 (the host is an aarch64 VM, see below), MariaDB, more than one Broker on the host, hostile same-UID tools, remote writers, restored or cloned disks, Kubernetes. I did not interrupt the server between the holder clear and the final retirement on the real host. That path is covered by the unit tests and the mutation results only.

Rig and method
  • Host. A VM created only for this run (Lima/colima, Apple Virtualization, aarch64, 4 vCPU, 8 GiB), Ubuntu 24.04.4 LTS, kernel 6.8.0-117, ext4 root disk. Reboots were systemctl reboot inside the guest, or a power cut by killing the hypervisor process. It is a VM, not bare metal. The kernel, systemd, /etc/machine-id and boot_id behave as on any Linux host.
  • Server. The PR's qwen-managed-agent-server jar as a systemd unit (non-root user, KillMode=process, Restart=on-failure), Temurin 21.0.9, MySQL 8.4.11 with its data directory on the root disk, Node 22.23.2, and the worker bundle built from the stack's TypeScript. State directory and Workspace roots on the root disk. Production HostIdentity.linux(), no test identity.
  • Driving it. Workspace Sessions are created through the public API. The reference server ships no authentication adapter, so a rig-only servlet filter supplies the actor. warm, acquire and executions go to the embedded Broker's HTTP routes with its bearer token. The chain in section 4 uses the packaged Hosted Harness and the repository's scripted model server.
  • Escaped writer. Started by a real run_shell_command call: a detached Node child that appends and fsyncs one nonce line every 200 ms into the Workspace.
  • F1 delivery. For the Hosted chain a transparent proxy between Harness and Broker delays the unchanged acquire request until SQL shows the scan's claim on that binding. The 3/592 and 6/871 numbers come from undelayed back-to-back turns.
  • One rig problem. After the first in-guest reboot the VM's Docker data root moved to a fresh data disk, so I re-created the MySQL container by hand in R1. Its data directory was untouched. R6, R7 and R8 needed no manual step.
  • Everything is in the bundle: harness, raw results, candidate.
中文版

真机验证:把「待完成的 Linux 重启验收」真正跑了一遍

验证对象是当前 head f8bf5d74,环境是一台专用 Linux 主机。我从 66bde476 开始测,途中整个栈被变基。W0e-3 这个提交在两个 head 上的 git patch-id 相同,TypeScript 树没有变化,所以第一轮结果保留,关键项目在新 head 上全部重跑。

结论

  • 重启恢复在真实主机上成立。 同一台机器真实重启 4 次(2 次 systemctl reboot,2 次断电):所有保存的代数都只靠定时扫描被回收,无人工干预的两次在主机宕掉后 21 秒完成。逃逸写入者随旧 boot 停止,回执语义不变,只清除原 holder,恢复过程没有启动任何 worker。
  • 对照组全部保持阻塞:选项关闭、machine-id 不同、仅 worker 死亡、systemd soft-reboot,以及登记记录缺失、损坏或属于别的机器。
  • 开启该选项并配合 Hosted 回合使用前需要修一个问题(F1): 5 秒一次的扫描也会观察健康的 binding,观察期间到达的请求会收到 503 runtime_reconciliation_required。Hosted Harness 把 acquire 失败当作需要恢复,Session 因此被阻塞。871 个健康回合里有 6 个撞上。附 13 行候选修复和一个单元测试。
  • 选项关闭(默认)时行为不变:847 个回合 0 个被拒;重启后保存的代数仍像以前一样保持阻塞。

合并建议: 从真机角度看,物理验收门已经满足。建议合并前修掉 F1(改动很小),补上两个缺口测试,并在 README 加一句 soft-reboot 说明。作者列出的「历史 FAILED binding 的升级前处理」属于 #12839,本轮没有覆盖。合并顺序仍是 #12839 → #12865 → 本 PR。

时间均为 UTC,2026-09-27。

1. 物理重启验收(图 1)

评审测试计划条目 真机结果
同宿主、boot 改变:墓碑 + loss/stop 证据,再清 holder R1、R2、R7 通过。恢复后的登记记录都是 RETIRED,且仍保留原 boot id。两条证据都指向原机器、原 boot、原资源。
仅 worker 死亡保持占用,不启动替代者 通过。扫描把 binding 标为 LOST,只有 loss 证据。holder 保留,逃逸写入者继续写(3 秒 +4200 B),warm 得到 503 runtime_broker_runtime_lost,没有新 worker。重启后该 binding 被释放,第一次的 loss 证据不变。
撤权、删除 Session、改 Registry 和 mount 通过(C 臂)。原 holder 被清除。该 Session 的 warm、acquire 被 409 workspace_unavailable 拒绝,没有产生新代数。
超过单事务批量的执行 通过。130 条 PREPARED 加上在途调用全部变为 ABANDONED 且没有结果。SETTLED 回执保留结果。没有任何重放。
启动中断、没有 lease 通过(E 臂)。READY 之前死亡的 worker 是 RECOVERY_BLOCKED,同一 boot 内 warm 被拒。重启后被释放。
断电前一刻写入的记录 通过(G 臂)。binding 在断电前 0.5 秒完成 warm,登记记录存活,随后被释放。
专用主机、真实重启、新的授权代数 通过。授权的 warm 在 543 ms 内创建第 2 代,acquire 重新占有 holder,新调用执行成功,旧 Runtime Session 返回 404。

JVM 时区是 Asia/Shanghai,MySQL 是 UTC,这是有意不对齐的。releaseLost 里的 lease 判断没有受影响。

2. 对照组(图 2)

在 f8bf5d74 上另做了一组对照(R8:停掉服务后改动登记记录,再真实重启):

重启时的登记记录 重启后
被删除 RECOVERY_BLOCKED,无证据,holder 保留,warm → 409 runtime_broker_resource_conflict
被垃圾内容覆盖 同上
hostId 被改成另一台机器 同上
未改动 RELEASED,两条证据齐全

只开启恢复选项而不开启 durable 模式时,服务在启动阶段失败,报 Trusted reboot recovery requires durable local-process provisioning。

3. 恢复之后(图 3)

旧回执保持为事实;之后能启动什么由当前授权决定。

4. F1:扫描拒绝健康请求,撞上的 Hosted Session 被阻塞(图 4)

过程:

  1. findRecoveryCandidates 包含 READY 状态(JdbcRuntimeBindingRepository.java:340),所以每次扫描都会对每个健康的 durable binding 调用 recoverBinding。
  2. recoverBinding 在认领、观察、重新认证、回写期间往 bindingOperations 放一个预留(RuntimeBrokerService.java:1549),并且总是以 503 runtime_reconciliation_required 结束它,即使观察结果是健康的(:1606)。
  3. ensureBinding 会把这个预留交给期间到达的请求(:825),影响 warm 和 Runtime Session 的首次 acquire。
  4. Hosted Harness 认为失败的 acquire 可能已经生效(hosted-workspace-tool-turn.ts:183-197),抛出 HostedToolRecoveryRequiredError。Session 进入 recovery-blocked,之后每个 prompt 都是 409 hosted_turn_recovery_required。

每个 binding 每 5 秒有约 20 ms 的窗口,所以一个有几百个工具回合的 Session 很可能撞上。发生时 Runtime 本身没有任何问题。

候选修复:观察以健康接管结束时,用该结果完成预留,而不是返回 503。等待者在进入 ensureBinding 之前已经完成了当前授权解析,这与 reconcileBinding 现有的交接方式相同。修复后:837 个回合 0 个被拒,Hosted 链路在同样的投递时机下正常完成,全部套件通过,真实重启仍能恢复(R6)。

另一种做法是跳过本 Broker 进程内存活的 binding。这样还能省掉每个健康 binding 每 5 秒 3 次行写入,但扫描将不再发现空闲 worker 的死亡。由作者决定。

5. Linux 套件与变异测试(图 5)

针对本 PR 生产代码改动做了 26 个单点变异体,依次用单元套件、故障门禁、真实 worker IT、MySQL IT 去跑:

  • 15 个被杀。 13 个被单元套件杀掉。2 个只被真实进程测试杀掉:「同一 boot 被当成重启」和「报告清理完成但没清 SQL holder」。

  • 5 个存活,但同一路径有第二道守卫。 releaseLost 里的版本检查和认领检查互相遮蔽;两个一起去掉会被 WorkspaceRecoveryTest 和 ManagedAgentMySqlIT 抓到。

  • 6 个全部存活。 其中两个守护的是安全性质:

    • rebooted() 去掉宿主检查:另一台机器加另一个 boot id 会产生 stop 证据。
    • 默认的 recoverResources 放行 managed 清理,而描述里说必须保持阻断。

    PR 代码在这两点上都是对的(见对照组),但回归时没有测试会发现。RebootRecoveryGapTest 把两者钉住:head 上 2/2 通过,在对应变异体下各自失败。
    其余四个较小:claim 忽略 Runtime Session 行、recoverBinding 忽略期望代数、扫描批次可重叠、证据匹配任意 lease。

本 PR 没有 Java CI:sdk-java.yml 只对 base 为 main 的 PR 运行,而这是堆叠 PR。main 自合并基以来前进了 3 个提交,都不涉及 packages/sdk-java;git merge-tree 无冲突,合并树里每个迁移版本号只有一个(V1 到 V16)。

6. 其他说明

  • soft-reboot(文档)。 systemctl soft-reboot(systemd 254 起,Ubuntu 24.04 可用)会杀掉所有写入者,但内核 boot id 不变。binding 变成没有 stop 证据的 LOST,一直占用到下一次完整重启。这是安全的一侧,README 里加一句可以省去运维的困惑。
  • 负载下的偶发失败,都与本 PR 无关。 ManagedAgentServerIntegrationTest 11 次里失败 2 次,每次是不同的用例。ConcurrencyStorageFaultGateTest.aDatabaseOutageDuringTheResultCommitNeverSettles 8 次里失败 1 次。
  • 回执状态的背景。 经 POST /executions 发出的前台调用超过 30 秒后会被报告为 UNKNOWN,即使有客户端在轮询。选项关闭时同样如此,所以不是本 PR 引起的。R1、R2 里的在途调用在宕机前已是 UNKNOWN,原因在此。

本轮没有覆盖

物理机和 x86_64(主机是 aarch64 虚拟机)、MariaDB、同机多个 Broker、恶意同 UID 工具、远程写入者、恢复或克隆的磁盘、Kubernetes。没有在真机上做「holder 清除之后、最终退休之前」的中断,这条路径只有单元测试和变异结果覆盖。

装置与方法

  • 主机。 专为本轮创建的虚拟机(Lima/colima,Apple Virtualization,aarch64,4 vCPU,8 GiB),Ubuntu 24.04.4 LTS,内核 6.8.0-117,根盘 ext4。重启是在客户机内执行 systemctl reboot,断电是杀掉虚拟机进程。它是虚拟机而不是物理机;内核、systemd、/etc/machine-id 和 boot_id 的行为与普通 Linux 主机一致。
  • 服务。 用本 PR 的 qwen-managed-agent-server jar 以 systemd 单元运行(非 root 用户,KillMode=process,Restart=on-failure),Temurin 21.0.9,MySQL 8.4.11(数据目录在根盘),Node 22.23.2,worker bundle 由该栈的 TypeScript 构建。状态目录和 Workspace 根目录都在根盘。使用生产的 HostIdentity.linux(),没有测试身份。
  • 驱动方式。 Workspace Session 通过公开 API 创建。参考服务不带认证适配器,所以用一个仅供装置使用的 servlet filter 提供 actor。warm、acquire、执行请求带 bearer token 直接发往内嵌 Broker 的 HTTP 路由。第 4 节的链路使用打包后的 Hosted Harness 和仓库自带的脚本化模型服务。
  • 逃逸写入者。 由一次真实的 run_shell_command 调用启动:一个分离的 Node 子进程,每 200 ms 向 Workspace 追加一行 nonce 并 fsync。
  • F1 的投递时机。 Hosted 链路里,Harness 与 Broker 之间放了一个透明代理,它把原样的 acquire 请求延迟到 SQL 显示扫描已认领该 binding 时再转发。3/592 和 6/871 这两组数字来自没有任何延迟的连续回合。
  • 一个装置问题。 第一次客户机内重启后,虚拟机的 Docker 数据根被换到了一块新数据盘上,所以 R1 里我手工重建了 MySQL 容器,数据目录没有动过。R6、R7、R8 不需要人工步骤。
  • 脚本、原始结果和候选补丁都在证据包里(链接见英文版)。

@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-adoption branch from 7236322 to 68d763d Compare September 28, 2026 02:25
@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-reboot branch from f8bf5d7 to 8d87878 Compare September 28, 2026 02:25
@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-adoption branch from 68d763d to 5248a8f Compare September 28, 2026 02:35
@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-reboot branch from 8d87878 to e7f4ae4 Compare September 28, 2026 02:35
@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-adoption branch from 5248a8f to 1e93de6 Compare September 28, 2026 02:41
@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-reboot branch from e7f4ae4 to 713f79a Compare September 28, 2026 02:42
@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-adoption branch from 1e93de6 to 36f87cd Compare September 28, 2026 03:53
@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-reboot branch from 713f79a to 9a1de09 Compare September 28, 2026 03:53
@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

Real-host verification, round 2: head 9a1de09e

Follows round 1 (head f8bf5d74). Same dedicated Linux host, same method. All times are UTC, 2026-09-28.

What changed since round 1. This push is a rebase. The W0e-3 commit and the W0e-2 commit below it have the same git patch-id as before. The base gained two #12839 commits and two merges of main: 4 Java files and 47 other files, among them the Hosted Harness Broker client. Because the TypeScript changed this time, I rebuilt the worker and Harness bundle from this head and re-ran everything on it.

Summary

  • Round 1 reproduces at the new head. Power cut and systemctl reboot: every saved generation is reclaimed by the scan alone. Both controls (option off, foreign machine-id) stay blocked.
  • New coverage: a crash in the middle of the cleanup. The server was killed twice while its holder clear was already committed and its final retirement was still in flight. After the restarts everything converged, and all 134 receipts kept their identity.
  • F1 is still open. 5 of 470 healthy turns were refused with the option on, 0 of 411 with it off. The Hosted Session is blocked as before. The round-1 candidate applies to this head without change and removes the problem (0 of 433).
  • To land on main the stack needs a rebase, because feat(managed-agent): Add W0e terminal recovery fences #12839 went in as a squash. Replaying the two commits onto main gives one conflict in a test file. The result passes all suites and recovers after a real reboot.

For the merge decision: unchanged from round 1. The physical acceptance gate is met at this head too. F1 is the one item I would fix before merging. The two gap tests and the soft-reboot sentence from round 1 are still open suggestions.

1. Reboots, controls and a crash inside the cleanup

Reboots and crash injection

How the crash was placed: a MySQL trigger delays the UPDATE that frees the placement slot, which only the final retirement issues. The product is unchanged. The server was killed with SIGKILL only while MySQL showed that statement as held and the holder row was already cleared.

Test plan item Round 1 Round 2
Interrupt cleanup after the holder clear unit tests and mutation only real host, two crashes (R14)
Retry keeps the original receipt identity not checked on the real host 134 of 134 identical
Remaining pins are released only after cleanup completes not checked on the real host session and slot stay until the retirement commits
More than one transaction's worth of executions 130 + in-flight converge same, and the crash fell between the two batches (101 done, 31 open)

2. F1 at the new head

Healthy traffic

The base brought a retry for prepare in the Harness (hosted-workspace-broker.ts), for transport errors only. acquire is not retried, and a 503 answer is not a transport error, so the chain in round 1 is unchanged: JdbcRuntimeBindingRepository.java:340 → RuntimeBrokerService.java:1549 / :1606 / :825 → hosted-workspace-tool-turn.ts:185-197.

Candidate for this head: diff, test.

3. What lands on main

Suites, merge and mutation

  • A plain merge of this head into main (6017f11d) conflicts in 9 files. That is the usual effect of a squash-merged base, not a problem in this PR.
  • Replaying W0e-2 and W0e-3 onto main is clean except for ManagedAgentMySqlIT.java: main (feat(managed-agent): Make Session close, archive and delete durable operations (Stage D4) #12881) and this PR each add a test at @Order(8). Keeping both and renumbering one resolves it.
  • The result has one migration per version, V1 to V17. This PR adds none.
  • With that jar and main's bundle, a real reboot recovered 8 of 8 generations 22.0 s after the reboot command (R15).

Mutation re-check. The six mutants that survived every stage in round 1 still do at this head (unit suites of both modules, fault gates, real-worker IT, both MySQL ITs). The control mutant "the same boot counts as a reboot" is still killed by both LocalRebootFaultGateTest cases and WorkspaceRecoveryWorkerIT. RebootRecoveryGapTest passes 2 of 2 at the head and fails under each of its two mutants.

4. Notes

  • A SETTLED receipt does not mean the file is on disk after a power cut. In R9 a tool call wrote a file 15 s before the power cut and its receipt is SETTLED / success. After the reboot the file was empty (ext4, the tool does not fsync). This is ordinary page-cache loss and not caused by this PR. It may deserve a sentence in the design, because recovery treats SETTLED as a known outcome.
  • New base behaviour, checked on the real host. After a worker died before READY, a second Session of the same Workspace now starts on the same boot (the feat(managed-agent): Add W0e terminal recovery fences #12839 change to blocksPlacement). Both bindings were recovered after the reboot.
  • Flaky under load. My first run on the rebased tree had 7 gate failures and one unit error while the rig server and its workers were running on the same VM and the machine under it was at load 50 to 60. A rerun on the quiet VM was green. In a same-conditions A/B the head had 0 failures in 4 gate runs and the rebased tree 1 in 4 (adoptedWorkerCanCancelItsOriginalActiveCall); the Broker code of both trees is identical except for BrokerValues.

What this run does not cover

Same as round 1, without the interruption item: bare metal and x86_64, MariaDB, more than one Broker on the host, hostile same-UID tools, remote writers, restored or cloned disks, Kubernetes.

Rig changes in this round
  • Worker and Harness bundle rebuilt from 9a1de09e (pnpm install, npm run build, npm run bundle). For the rebased tree the bundle was built from main.
  • The rebased tree is main 6017f11d plus the two stacked commits replayed with git merge-tree --merge-base, with the one conflict resolved as described above.
  • Crash injection: CREATE TRIGGER … BEFORE UPDATE ON qwen_runtime_binding_slot … SLEEP(25) when active_binding_id becomes NULL. The watcher reads performance_schema.processlist and the holder rows, then kills the server's main PID.
  • Everything else is as in round 1. Scripts, raw results and the candidate are in the round-2 bundle.
中文版

真机验证第二轮:head 9a1de09e

承接第一轮(head f8bf5d74)。同一台专用 Linux 主机,方法相同。时间均为 UTC,2026-09-28。

与第一轮相比改了什么。 这次推送是一次变基。W0e-3 提交和它下面的 W0e-2 提交的 git patch-id 都与之前相同。底座多了 #12839 的两个提交和两次合入 main:4 个 Java 文件、47 个其他文件,其中包括 Hosted Harness 的 Broker 客户端。因为这次 TypeScript 有变化,我用这个 head 重新构建了 worker 和 Harness 的 bundle,并在它上面全部重跑。

结论

  • 第一轮的结果在新 head 上复现。 断电和 systemctl reboot 之后,所有保存的代数都只靠扫描被回收。两个对照组(选项关闭、machine-id 不同)保持阻塞。
  • 新增覆盖:清理过程中途崩溃。 在 holder 清除已提交、最终退休仍在进行时,把服务杀掉两次。重启后全部收敛,134 条回执的标识全部不变。
  • F1 仍未解决。 选项开启时 470 个健康回合有 5 个被拒,关闭时 411 个里 0 个。Hosted Session 和之前一样被阻塞。第一轮的候选修复不用修改就能应用到这个 head,并消除问题(433 个里 0 个)。
  • 要落到 main 上,这个栈需要变基,因为 feat(managed-agent): Add W0e terminal recovery fences #12839 是以 squash 方式合入的。把两个提交重放到 main 上只有一处测试文件冲突。结果通过全部套件,真实重启后也能恢复。

合并建议: 与第一轮相同。物理验收门在这个 head 上同样满足。F1 是我建议合并前先修的一项。第一轮提的两个缺口测试和 soft-reboot 说明仍是待采纳的建议。

1. 重启、对照组、清理中途崩溃(图 1)

崩溃点的放置方式:用一个 MySQL 触发器延迟「释放 placement slot」的那条 UPDATE,这条语句只有最终退休步骤才会发出。产品代码没有改动。只有当 MySQL 显示该语句被挂起、且 holder 行已经清空时,才用 SIGKILL 杀掉服务。

测试计划条目 第一轮 第二轮
holder 清除后中断清理 只有单元测试和变异结果 真机,两次崩溃(R14)
重试保留原回执标识 未在真机上检查 134 条全部一致
剩余占用只在清理完成后释放 未在真机上检查 session 和 slot 一直保留到退休提交
超过单事务批量的执行 130 条加在途调用全部收敛 同样收敛,且崩溃正好落在两批之间(101 条已完成,31 条未完成)

2. 新 head 上的 F1(图 2)

底座给 Harness 的 prepare 加了重试(hosted-workspace-broker.ts),只针对传输错误。acquire 没有重试,而 503 响应也不属于传输错误,所以第一轮描述的链路没有变化:JdbcRuntimeBindingRepository.java:340 → RuntimeBrokerService.java:1549 / :1606 / :825 → hosted-workspace-tool-turn.ts:185-197。

3. 落到 main 上会是什么样(图 3)

  • 把这个 head 直接合并进 main(6017f11d)有 9 个文件冲突。这是底座被 squash 合入后的常见现象,不是本 PR 的问题。
  • 把 W0e-2、W0e-3 重放到 main 上,除 ManagedAgentMySqlIT.java 外都是干净的:main(feat(managed-agent): Make Session close, archive and delete durable operations (Stage D4) #12881)和本 PR 各自在 @Order(8) 加了一个测试。两个都保留、其中一个改序号即可。
  • 结果里每个迁移版本号只有一个文件,V1 到 V17。本 PR 没有新增迁移。
  • 用这棵树的 jar 和 main 的 bundle 做真实重启(R15):重启命令发出后 22.0 秒,8 个代数全部回收。

变异复查。 第一轮里六个「全部存活」的变异体在这个 head 上依然通过全部六个阶段(两个模块的单元套件、故障门禁、真实 worker IT、两个 MySQL IT)。对照变异体「同一 boot 被当成重启」仍被 LocalRebootFaultGateTest 的两个用例和 WorkspaceRecoveryWorkerIT 杀掉。RebootRecoveryGapTest 在 head 上 2/2 通过,在对应的两个变异体下各自失败。

4. 其他说明

  • SETTLED 回执不代表断电后文件已落盘。 R9 里一次工具调用在断电前 15 秒写了一个文件,回执是 SETTLED / success。重启后该文件是空的(ext4,工具没有 fsync)。这是普通的页缓存丢失,不是本 PR 引起的。因为恢复逻辑把 SETTLED 当作已知结果,设计文档里也许值得加一句。
  • 底座的新行为,已在真机上确认。 worker 在 READY 之前死亡后,同一 Workspace 的第二个 Session 现在可以在同一 boot 内启动(feat(managed-agent): Add W0e terminal recovery fences #12839 对 blocksPlacement 的修改)。重启后两个 binding 都被回收。
  • 负载下的抖动。 我第一次跑重放树时,装置的服务和 worker 正在同一台 VM 上运行,底层机器的负载在 50 到 60,结果出现 7 个门禁失败和 1 个单元错误。在安静的 VM 上重跑全绿。同条件 A/B 里,head 的 4 次门禁运行 0 失败,重放树 4 次里 1 次失败(adoptedWorkerCanCancelItsOriginalActiveCall);两棵树的 Broker 代码除 BrokerValues 外完全相同。

本轮没有覆盖

与第一轮相同,去掉「中断清理」一项:物理机和 x86_64、MariaDB、同机多个 Broker、恶意同 UID 工具、远程写入者、恢复或克隆的磁盘、Kubernetes。

本轮装置的变化

  • worker 和 Harness 的 bundle 由 9a1de09e 重新构建(pnpm install、npm run build、npm run bundle)。重放树使用由 main 构建的 bundle。
  • 重放树是 main 6017f11d 加上用 git merge-tree --merge-base 重放的两个提交,唯一的冲突按上文方式解决。
  • 崩溃注入:CREATE TRIGGER … BEFORE UPDATE ON qwen_runtime_binding_slot … SLEEP(25),在 active_binding_id 变为 NULL 时触发。观察脚本读取 performance_schema.processlist 和 holder 行,然后杀掉服务的主进程。
  • 其余与第一轮相同。脚本、原始结果和候选补丁都在第二轮证据包里(链接见英文版)。

@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-adoption branch from 36f87cd to 4136721 Compare September 28, 2026 15:41
@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-reboot branch from 9a1de09 to e867a99 Compare September 28, 2026 15:46
@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-reboot branch 2 times, most recently from bceaa60 to 8c2b626 Compare September 28, 2026 17:18
@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

Real-host verification, round 3: head 8c2b626c

Follows round 2 (head 9a1de09e) and round 1. Fresh dedicated aarch64 KVM VM (Debian 12), same method; everything re-measured at this head — nothing carried over from earlier rounds.

Verdict: findings — 62 pass / 1 fail (the fail is an adjudicated load flake, re-run 5/5 green). Verified head: 8c2b626c74f88fd39ffffd3e168099cab488041d (W0e-3), base 430a76eac9 (W0e-2 stack on main a765229c0a). All times UTC, 2026-09-28.

中文摘要

结论:findings(62 通过 / 1 失败——失败项已裁定为主机负载抖动,安静环境 5/5 全绿)。

与第 2 轮(9a1de09e)相比,本 head 的 W0e-3 提交只改了两处测试基建(IT 超时 180s→600s、一个 MySQL 测试的 @Order 编号);W0e-3 生产代码没有变化。底座 W0e-2 新增两个生产修复(56ada7e510 终止失败的 durable 启动协调、430a76eac9 把未准入的旧启动栅栏限定到其槽位),main 也前进了很多。

  • 物理验收在新 head 重新实测通过:R20b 真实重启(systemctl reboot)8/8 代数被回收,holders 6→0,132 条在途执行全部 ABANDONED 且无结果,回执身份不变;R21 断电后 option OFF 和 foreign machine-id 两个对照全部保持阻塞,恢复原 machine-id 后 7.7 秒收敛;R22b 在 holder 清除已提交、退休未完成之间 SIGKILL 两次(含重试中的那次),134/134 回执身份不变并全部收敛。
  • F1 仍然存在:开启选项时 417 个健康回合有 2 个被 503 runtime_reconciliation_required 拒绝;Hosted 链路在对齐注入下 turn 2 收到 503、Session 进入 recovery-blocked,之后所有 prompt 409。第 2 轮的候选补丁在新 head 原样适用(偏移 58 行),同样流量下 406/406 零拒绝、Hosted 链路全程完成。
  • 变异复测:第 1/2 轮的 6 个幸存者中,M01(rebooted 不查 hostId)和 M25(证据匹配任意 lease)现在被底座新增的故障门禁 DurableLocalRuntimeFaultGateTest.absentWorkerAfterRestartLeavesItsDomainPinned 杀死;M11、M16、M19、M21 仍然存活(与第 2 轮判定一致:M16 是 fail-open 默认的安全性质,仍靠随附的 RebootRecoveryGapTest 钉住,head 上 2/2 通过)。控制变异体 M02 仍被杀。
  • 套件:broker 单元 428、server 单元 178、fault gates 41、WorkspaceRecoveryWorkerIT、两个 MySQL IT 在新 head 全绿(含 checkstyle);候选构建同样全绿,且 F1 测试 + 2 个 gap 测试在候选上 3/3 通过。
  • 落地:当前 main(a77d80d19f)git merge-tree 试合并无冲突;迁移版本一一对应(V1–V14、V16–V18,本 PR 不新增)。第 2 轮「需要 rebase」的提示已被作者解决。
  • 建议:合并前修 F1(候选补丁已两轮验证);M16 的 in-tree 钉测试仍建议补。负载下 ProcessCrashFaultGateTest 与 DurableLocalRuntimeFaultGateTest 各抖过一次(runtime_provision_fenced vs runtime_broker_runtime_lost 的判定先后),安静环境 5/5 通过,与底座新增的启动栅栏判定有关,值得作者留意但不阻塞。

未覆盖:裸金属与 x86_64(本机为 aarch64 VM)、MariaDB、同机多 Broker、恶意同 UID 工具、远程写入者、磁盘/VM 快照恢复、Kubernetes、claim 过期重试的专项真机项(仅单元/变异覆盖)。

Previous-finding status (rounds 1–2 → this head)

# Finding (round) Severity Status at 8c2b626c
F1 Scan refuses healthy requests; Hosted Session blocked (r1, r2) worth fixing before merge stands — re-measured: 2 of 417 healthy turns refused with the option on (r2: 5/470), 0 of 418 with it off; Hosted chain recovery-blocked on the aligned acquire. The r1/r2 candidate applies cleanly (offset 58) and clears it: 0 of 406, Hosted chain completes, all suites + 3/3 candidate tests pass.
— Two gap tests (RebootRecoveryGapTest) open suggestion (r1) suggestion superseded in part — the W0e-2 base's new fault-gate coverage now kills M01 and M25 in-tree (mutation matrix below). The gap test still passes 2/2 at head; M16 (fail-open recoverResources default) remains without an in-tree pin.
— Soft-reboot README sentence (r1) suggestion stands — no README change in this push.
— SETTLED receipt ≠ on-disk after power cut (r2 note) docs note stands — design doc unchanged on this point; harmless page-cache reality, worth the sentence.
— Stack needs a rebase onto main (r2) process fixed — the stack was rebased; git merge-tree against current main a77d80d19f is conflict-free.

Central claim — reboot recovery on a real host, re-measured at this head

Same dedicated-VM method as rounds 1–2, rebuilt here from scratch (KVM/qemu, Debian 12 arm64, kernel 6.1.0-53-cloud, JDK 21, MySQL 8.4, worker bundle built from this head; systemd unit, non-root, KillMode=process). Three runs:

run scenario result
R20b hands-off systemctl reboot, 8 generations on 6 storages (incl. 1 worker-only death, 1 interrupted startup, 1 access-revoked), 130 PREPARED + in-flight + settled receipts pass — boot id changed (17 s down); all 8 RELEASED 24 s after the new boot's server start; holders 6→0; 132 nonterminal → ABANDONED without results; SETTLED keep results; revoked arm stays 409; authorized warm creates generation 2 (2.2 s); old Runtime Session 404
R21 power cut (SIGKILL qemu), then boot with option OFF, then foreign machine-id, then recovery pass — option OFF: nothing moved (6 holders, 132 executions intact), authorized warm 503 runtime_broker_reconcile_timeout. Foreign machine-id: identical block. Original machine-id restored: all 8 RELEASED in 7.7 s, evidence names the pre-cut boot
R22b reboot, then 2× SIGKILL between the committed holder clear and the final retirement (MySQL trigger holds the slot-freeing UPDATE 25 s) pass — crash 1 hit the first retirement, crash 2 hit the retried one; converged alone; 134/134 receipts identical, 132 ABANDONED without result, 2 SETTLED kept; 0 worker processes left

reboot acceptance

power cut controls

crash in cleanup

F1 at this head (unchanged mechanism, re-measured)

F1 healthy traffic

The scan still observes every healthy durable binding every 5 s and still completes the reservation with 503 runtime_reconciliation_required; the Hosted chain still turns that into a recovery-blocked Session. Numbers: 2 of 417 turns refused (option on) vs 0 of 418 (off) vs 0 of 406 (candidate). The aligned-delivery probe shows the exact window: acquire delivered when the scan holds the binding → 503 at head, 200 with the candidate.

Suites and mutation re-check

suites and mutation

  • Head: broker unit 428, server unit 178, WorkspaceRecoveryWorkerIT 1, broker MySQL IT 3, server MySQL IT 14 — all green incl. checkstyle. Candidate: same, plus the F1 test and both gap tests pass 3/3 (at head the F1 test errors as designed).
  • Mutation matrix (7 mutants from the r1/r2 survivor set + control, six stages each): M01 and M25 are now killed by the base's strengthened durable-startup fault gate — two of the six round-1 survivors now have in-tree pins. M11 (claim ignores the Runtime Session row), M16 (fail-open default), M19 (expected generation ignored), M21 (overlapping scan batches) still survive everything; adjudication unchanged from round 2 (M16 is the one guarding a safety property).
  • One load flake: ProcessCrashFaultGateTest.aLostJournalEndsPollingWithoutReleasingTheWriterDomain failed once while the host ran the VM plus two Maven containers (runtime_provision_fenced surfaced before runtime_broker_runtime_lost); 5/5 green isolated at head. The mutation runs show the same two codes are order-sensitive at that boundary — worth a look from whoever owns the startup fence, not a blocker.

What changed since round 2 (scope of this round)

git range-diff r2-head..head: the W0e-3 commit touches only the IT fork timeout (180→600 s) and one test's @Order (8→11, base added tests). The base gained 56ada7e510 and 430a76eac9 (durable startup reconciliation termination; legacy startup fence limited to its slot) plus docs, and main advanced ~460 files. Everything above was re-run, not carried over.

Not covered

Bare metal and x86_64 (host is an aarch64 VM), MariaDB, more than one Broker per host, hostile same-UID tools, remote writers, restored/cloned disks, Kubernetes, claim-expiry retry as a dedicated real-host scenario (unit + mutation coverage only). Round 1's SETTLED-vs-fsync note stands unaddressed in the docs.

Methodology

Head worktree at 8c2b626c (pnpm install, npm run build, npm run bundle → worker bundle; Maven in a JDK 21 container → qwen-managed-agent-server jar, sha256 831f89e9…, candidate jar 8109c4aa… = head + the r2 F1 candidate re-applied with offset 58). VM: qemu/KVM arm64, Debian 12 cloud image, MySQL 8.4 in docker, server jar under systemd with the rig auth filter. Power cut = SIGKILL on the qemu process; reboot = systemctl reboot inside the guest. Harness scripts from the round-1/2 evidence bundle, unchanged except a guard in the crash watcher (skip kill while systemd reports MainPID 0 — a rig bug that ate round-3's first crash attempt; product untouched). Scripted assertions: assert.mjs re-validates the saved raw logs (38 checks), plus 8 scripted mutation verdicts, 5 isolated gate re-runs, 12 suite-stage exit checks. One scripted check failed once (the loaded gate run) and passed 5/5 isolated — disclosed above, counted as the flake it is.

Raw logs, harness and assertion script: in the round-3 bundle (assert.mjs re-validates the saved logs; results/ holds the raw per-run output).

For the merge decision: unchanged from rounds 1–2 — the physical acceptance gate is met again at this head, on a fresh independent rig. F1 remains the one item to fix before the option is used with Hosted turns; the attached candidate is now verified against three consecutive heads. The stack merges into current main (a77d80d19f) conflict-free; #12839 is already in, so the order left is #12865 → this PR.

@doudouOUC
doudouOUC force-pushed the codex/managed-workspace-w0e-reboot branch from 8c2b626 to 55ded4c Compare September 29, 2026 01:14
@doudouOUC
doudouOUC changed the base branch from codex/managed-workspace-w0e-adoption to main September 29, 2026 01:15
@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Follow-up to the round 3 real-host review: F1 is fixed in 55ded4c31. A deterministic test first reproduced the healthy warm receiving 503 runtime_reconciliation_required while maintenance observed its binding. A successful, attested READY observation now completes that binding's waiting authorized requests with the adopted context; uncertain, lost, conflicting and failed observations remain blocked. The test now passes for concurrent warm and acquire, and a separate regression pins the fail-closed behavior when a managed provisioner has no holder-cleanup callback.

Because #12865 was squash-merged, I replayed the unchanged W0e-3 commit onto current main (the stable patch-id matches the original), then pushed the two-commit branch with an exact-SHA force lease and retargeted this PR to main. Its diff is now W0e-3 plus this fix, without the merged W0e-2 changes. The bilingual design and server README now record the independent Linux reboot acceptance at 8c2b626c, distinguish that evidence from current-head tests, and clarify soft reboot and file durability limits.

Current-head verification: Broker unit/H2 430 tests, 0 failures/errors and 1 conditional skip; managed-agent server unit/H2 184 passed; real-process LocalReboot and DurableLocalRuntime fault gates 12/12 passed; Spring/H2/real-worker recovery integration 1/1 passed; Java Checkstyle, repository build, typecheck, bundle and changed Markdown formatting passed. The reviewer physically reboot-tested the earlier W0e-3 head; this new SHA has not had another physical reboot run. There are 0 unresolved inline threads. I am moving the PR to ready for review so exact-head CI and review can run.

@doudouOUC
doudouOUC marked this pull request as ready for review September 29, 2026 01:18
@doudouOUC doudouOUC self-assigned this Sep 29, 2026
@doudouOUC

Copy link
Copy Markdown
Collaborator Author

[codex] Thanks — agreed on the CI blocker. Fixed in 62584d315.

Review item Action
Failsafe family mismatch Renamed the integration test to HostedWorkspaceRecoveryWorkerIT and removed both one-off POM selectors. The existing Hosted*IT rules now place it in the hosted job.
Scan of healthy bindings Deferring the optimization: filtering by current host/boot identity needs a separate design for how that trusted identity reaches the candidate query.
Per-binding recovery error logging Deferring observability changes to a follow-up; recovery remains fail-closed and retries on the next scan.
SDK SPI compatibility Leaving the alpha SDK source-compatibility and release-note call to the maintainer.

The targeted hosted IT passed (1/1), the guard's hosted and non-hosted family checks passed in an isolated verification, and the repository build and typecheck passed. The new push should trigger the PR CI that did not run on the draft head. Resolved 0/0 inline threads.

@doudouOUC
doudouOUC enabled auto-merge September 29, 2026 03:37
wenshao added a commit to wenshao/qwen-code that referenced this pull request Sep 29, 2026
wenshao added a commit to wenshao/qwen-code that referenced this pull request Sep 29, 2026
@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

Real-host verification, round 4: head 62584d31

Follows round 3 (head 8c2b626c). Same dedicated Linux host as rounds 1 and 2. All times are UTC, 2026-09-29.

What changed since round 3. The branch now sits on main and has three commits:

  • the W0e-3 commit, with the same git patch-id as at 8c2b626c;
  • 55ded4c3, the F1 fix with two new tests and the doc updates;
  • 62584d31, which renames one test class and removes two pom lines.

I tested 55ded4c3 first, and 62584d31 was pushed while I was working. The server jars of both heads have the same 537 class files, byte for byte, so the results carry over. I also ran a reboot and both Java CI jobs on the new head.

Summary

  • F1 is fixed. With the option on, 0 of 1383 healthy turns were refused (round 1: 6 of 871, round 2: 5 of 470). The Hosted chain that ended in a blocked Session in rounds 1 to 3 now completes. Requests that must not be served are still refused.
  • Physical reboots pass with this code. A power cut with both controls, four hands-off reboots and a reboot with two crashes inside the cleanup all converge. One of the hands-off reboots used the jar built from 62584d31, which closes the gap the author noted for the new SHA.
  • The CI blocker found by the triage bot was real, and 62584d31 fixes it. At 55ded4c3 every Maven step passes and the report guard fails in both Java jobs. At 62584d31 both guards pass, in my replays and in the first real CI run of this PR.
  • The merge with today's main (1b696297) is clean. The Hosted job passes on the merge result, and a real reboot with that jar and main's bundle recovers 8 of 8 generations.
  • One correction to round 3: M01 and M25 are not killed by the tests in the tree. Details are in section 5.

For the merge decision: I found nothing that blocks the merge. The bot's CHANGES_REQUESTED review is on 55ded4c3, and its one blocker is fixed. Three suggestions remain, none of them blocking: a test for the wrong-host case (attached), log lines for the scan, and the items the author deferred.

1. F1 is fixed

Healthy traffic

  • The fix is the candidate from round 1. recoverBinding completes its reservation with the adopted context when the observation ends healthy. Every other outcome still ends with 503 runtime_reconciliation_required.
  • Hosted chain. I used the same delivery as in rounds 1 to 3: the unchanged acquire request reaches the Broker while the scan holds its claim on that binding. All three turns complete and the Session is not recovery-blocked.
  • The other side. After its worker was killed, 733 back-to-back warm and acquire requests were all refused, the binding became LOST without stop evidence, and no worker was started. With the access revoked, 977 requests were all refused with 409 workspace_unavailable and the binding was left untouched.
  • Tests for the fix. Reverting the fix is caught by the new MaintenanceProbeWaiterTest. Two more mutants of the fix survive, and neither points to a missing test: one is equivalent, the other is refused a second time inside adoptObservation.

2. Physical reboots

Reboots

Run What it shows
R16, power cut Both controls stay blocked for the whole observation (option off, foreign machine-id). With the option on and the original machine-id, 8 of 8 generations are released 3.6 s after the server starts.
R17, R20, R21, R22, systemctl reboot Hands-off. All 8 generations are released 18.2 s to 26.8 s after the reboot command. R20 uses the jar of 62584d31, R21 and R22 use merges with main.
R19, reboot and two crashes The server is killed twice while MySQL holds the final retirement statement. After the restarts everything converges, and all 134 receipts keep their identity.

3. The CI guard, before and after the rename

CI, merge and mutation

  • At 55ded4c3 the guard fails in both jobs with the two messages the bot predicted. In the Hosted job the guard step comes before the fault gates step, so CI would not have run the gates either.
  • At 62584d31 both guards pass. Before the push I had tested the bot's suggested fix as an A/B arm, and 62584d31 is file for file that tree (10019 files, 0 differences).
  • First CI run. SDK Java (8 of 8 jobs) and Qwen Code CI are green. The test counts in the CI logs equal the counts of my replays.
  • Why CI was missing before. 55ded4c3 was pushed while the PR was a draft on a stacked base. Changing the base and marking the PR ready do not start these workflows.

4. Merge with today's main

main has moved eight commits since the PR's base (99adce25 → 1b696297). Under packages/sdk-java they change Hosted tests, SessionEventHub, the V15 migration and the OpenAPI file.

Check Result on the merge with 1b696297
git merge of 62584d31 No conflict. 18 migrations (V1 to V18), one per version. This PR adds none.
Server jar 610 entries. The 8 that differ from the jar of 62584d31 are main's changes. No recovery or Broker class differs.
Hosted job, with main's bundle 10 of 10 integration tests pass, the guard passes
Real reboot with that jar and bundle (R22) 8 of 8 generations released, hands-off, 26.8 s (the guest took longer to boot)

The same checks passed an hour earlier on main b1eb94da (R21, 18.2 s). main moved again while I was testing, so I repeated them.

5. Mutation results and a second look at RuntimeRecoveryEvidence.matches

Correction to round 3. Round 3 listed M01 (rebooted() ignores the host id) and M25 (evidence matches any lease) as killed by DurableLocalRuntimeFaultGateTest.absentWorkerAfterRestartLeavesItsDomainPinned. That was wrong:

  • The gate builds its Brokers with trusted reboot recovery off, so it never calls rebooted().
  • In this round both mutants pass all six stages. The runs were sequential on the quiet host, and a stage counts as a kill only if it fails twice.
  • The failures in round 3 came from the load-sensitive gate that round 3 itself listed as a flake.

What is pinned now.

Mutant Rounds 1 to 3 This round
M16, the default recoverResources allows managed cleanup survives killed by the author's new RuntimeMaintenanceRecoveryTest case
M01, another machine with another boot id counts as a reboot survives survives
M11, M19, M21 survive survive (rated as smaller gaps since round 1)

M01 is the one survivor that guards a safety property. On the real host the foreign machine-id control stays blocked (R16), so the code is right and only the test is missing. RebootRecoveryGapTest from round 1 covers it: it passes 2 of 2 at 62584d31, fails under M01, and is Checkstyle clean.

RuntimeRecoveryEvidence.matches. The bot asked for a second look at the relaxed lease check.

  • With a lease, nothing is relaxed. matches has one caller, RuntimeBindingRecord.requireEvidence. The constructor rejects a lease that does not match the seed before it reaches that call (RuntimeBindingRecord.java:118). So the new clause can never be false there, which is why M25 survives: it is an equivalent mutant, not a test gap.
  • Without a lease, the evidence is still tied to the seed (request id, runtime instance id, incarnation, lease id, epoch) and to the resource handle.
  • On the real host this case is the worker that died before READY. In R16, binding 1ce9de00 had no lease and was RECOVERY_BLOCKED before the power cut. After the reboot it was released with both evidence records (table). Before this PR, matches returned false for every binding without a lease.

6. Notes, none blocking

  • The scan leaves no trace in the log. In R19 the server reclaimed four generations and was killed twice, and its log has no line from the scan or from the recovery (log). An operator can follow a recovery only in the database. The author deferred logging to a follow-up. This is how that item looks on a real host.
  • The scan still observes every healthy binding every 5 s and rewrites its row. After the fix this no longer refuses requests. The author deferred the optimization, and I agree that it can wait.
  • SPI. RuntimeBindingRepository gains two methods without a default (finishLostRecovery, findRecoveryCandidates), so an implementation outside this repository would stop compiling. Both implementations in the tree are updated, and RuntimeProvisioner.recoverResources has a fail-closed default. The package is 0.1.0-alpha. Whether that needs a release note is a maintainer call.
  • The repository's own script tests do not see a guard mismatch. check-failsafe-reports.test.js and hosted-process-ci.test.js pass 13 of 13 at 55ded4c3, where both guards fail.

What this run does not cover

Bare metal and x86_64, more than one Broker on the host, hostile same-UID tools, remote writers, restored or cloned disks, Kubernetes. On MariaDB only the two integration test classes of the CI job ran. The real-host runs use MySQL 8.4.

Rig and method in this round
  • Host. The same VM as in rounds 1 and 2: Ubuntu 24.04.4, kernel 6.8.0-117, aarch64, ext4. The server jar runs under systemd with MySQL 8.4.11 and production HostIdentity.linux(). Round 3 used a different VM.
  • Bundle. This PR changes no TypeScript. The worker and Harness bundle was built from 55ded4c3, and for R21 and R22 from the merge results.
  • CI replays. The commands are taken from .github/workflows/sdk-java.yml. The MariaDB job ran in a Linux container on the dedicated host (MariaDB 10.11.18). The Hosted job ran on macOS (JDK 21.0.12, Node 22.23.2, MySQL 8.4.7), because the Hosted tests need the repository's node_modules. GitHub CI runs both on Ubuntu.
  • One rig error. My first MariaDB replay of the renamed tree used the database of the previous arm, and JdbcRuntimeBrokerMySqlIT failed with expected: <[1]> but was: <[2]>. The contract test writes with a fixed prefix and needs a database that no earlier run has used. CI starts a fresh MariaDB for every job. As a control, the unchanged 55ded4c3 tree fails the same way on the used database. On a fresh database the replay passes.
  • Crash watcher. In the first attempt (R18) the second kill hit the restarted server while MySQL still held the statement of the first one. The watcher now requires the held statement to be younger than the running server. R19 is the run with both kills inside a held retirement. R18 converged as well and is in the bundle.
  • Mutation. Run at 55ded4c3. 62584d31 has the same sources except for the name of one test class.
  • Everything is in the bundle: harness and raw results.
中文版

真机验证第四轮:head 62584d31

承接第三轮(head 8c2b626c)。使用第一、二轮的同一台专用 Linux 主机。时间均为 UTC,2026-09-29。

与第三轮相比改了什么。 分支现在直接基于 main,共三个提交:

  • W0e-3 提交,git patch-id 与 8c2b626c 时相同;
  • 55ded4c3:F1 修复、两个新测试和文档更新;
  • 62584d31:给一个测试类改名,并删掉 pom 里的两行。

我先测的是 55ded4c3,测试途中作者推送了 62584d31。两个 head 构建出的服务 jar 里 537 个 class 文件逐字节相同,所以结果可以沿用。我另外在新 head 上补跑了一次重启和两个 Java CI 作业。

结论

  • F1 已修复。 选项开启时,1383 个健康回合 0 个被拒(第一轮 871 个里 6 个,第二轮 470 个里 5 个)。第一到三轮里以 Session 被阻塞告终的 Hosted 链路现在能跑完。不该放行的请求仍然被拒绝。
  • 这份代码通过了物理重启验证。 一次断电加两个对照组、四次免人工重启、一次在清理过程中崩溃两次的重启,全部收敛。其中一次免人工重启用的是 62584d31 构建的 jar,补上了作者提到的「新 SHA 还没做过物理重启」这个缺口。
  • triage 机器人指出的 CI 阻塞项确实存在,62584d31 已修复。 在 55ded4c3 上,所有 Maven 步骤都通过,但两个 Java 作业的报告守卫都失败。在 62584d31 上两个守卫都通过,我的回放和本 PR 第一次真实 CI 的结果一致。
  • 与当前 main(1b696297)合并没有冲突。 Hosted 作业在合并结果上通过,用合并后的 jar 和 main 的 bundle 做真实重启,8 个代数全部回收。
  • 对第三轮的一处更正: M01 和 M25 并没有被仓库里的测试杀掉。详见第 5 节。

合并建议: 我没有发现阻塞合并的问题。机器人的 CHANGES_REQUESTED 评审针对的是 55ded4c3,它提出的唯一阻塞项已经修复。还有三条建议,都不阻塞:补一个「换了机器」场景的测试(已附上)、给扫描加日志、以及作者已说明延后处理的几项。

1. F1 已修复(图 1)

  • 修复方式就是第一轮的候选方案。观察结果健康时,recoverBinding 用接管到的上下文完成它的 reservation。其他结果仍然以 503 runtime_reconciliation_required 结束。
  • Hosted 链路。 投递方式与第一到三轮相同:在扫描持有该 binding 的 claim 时,把未经修改的 acquire 请求送到 Broker。三个回合全部完成,Session 没有进入恢复阻塞状态。
  • 另一面。 把 worker 杀掉后连续发出 733 个 warm 和 acquire 请求,全部被拒,binding 变为 LOST 且没有停止证据,也没有启动新的 worker。撤销访问权限后,977 个请求全部返回 409 workspace_unavailable,binding 保持不变。
  • 针对修复的测试。 把修复回退会被新增的 MaintenanceProbeWaiterTest 抓到。修复还有两个变异体存活,但都不代表缺测试:一个是等价变异体,另一个会在 adoptObservation 里被再次拒绝。

2. 物理重启(图 2)

运行 说明
R16,断电 两个对照组(选项关闭、machine-id 不同)在整个观察期内保持阻塞。选项开启且 machine-id 是原来的值时,服务启动后 3.6 秒,8 个代数全部回收。
R17、R20、R21、R22,systemctl reboot 全程无人工介入。重启命令发出后 18.2 到 26.8 秒,8 个代数全部回收。R20 用 62584d31 的 jar,R21 和 R22 用与 main 合并后的结果。
R19,重启加两次崩溃 在 MySQL 挂起最终退休语句时把服务杀掉两次。重启后全部收敛,134 条回执的标识全部不变。

3. 改名前后的 CI 守卫(图 3)

  • 在 55ded4c3 上,两个作业的守卫都失败,报错信息与机器人预测的两条一致。Hosted 作业里守卫步骤排在故障门禁步骤之前,所以 CI 连门禁也不会跑。
  • 在 62584d31 上,两个守卫都通过。作者推送之前,我已经把机器人建议的修法当作 A/B 的一个分支测过,62584d31 与那棵树逐文件相同(10019 个文件,0 处差异)。
  • 第一次 CI。 SDK Java(8 个作业)和 Qwen Code CI 全绿。CI 日志里的测试数与我回放得到的数字相同。
  • 之前为什么没有 CI。 55ded4c3 推送时 PR 还是草稿,而且 base 是堆叠分支。修改 base 和标记为 ready 都不会触发这些工作流。

4. 与当前 main 合并

main 在 PR 的 base 之后又前进了八个提交(99adce25 → 1b696297)。在 packages/sdk-java 下,这些提交改了 Hosted 测试、SessionEventHub、V15 迁移和 OpenAPI 文件。

检查项 与 1b696297 合并后的结果
合并 62584d31 无冲突。18 个迁移(V1 到 V18),每个版本号一个。本 PR 没有新增迁移。
服务 jar 610 个条目。与 62584d31 的 jar 相比有 8 个不同,都是 main 的改动。恢复和 Broker 相关的 class 没有差异。
Hosted 作业(使用 main 的 bundle) 10 个集成测试全部通过,守卫通过
用这个 jar 和 bundle 做真实重启(R22) 8 个代数全部回收,无人工介入,26.8 秒(这次虚拟机启动较慢)

一小时前在 main b1eb94da 上做的同样检查也通过(R21,18.2 秒)。测试期间 main 又前进了,所以我重做了一遍。

5. 变异结果,以及对 RuntimeRecoveryEvidence.matches 的复查

对第三轮的更正。 第三轮把 M01(rebooted() 不检查 host id)和 M25(证据匹配任意 lease)记为被 DurableLocalRuntimeFaultGateTest.absentWorkerAfterRestartLeavesItsDomainPinned 杀死。这个结论是错的:

  • 这个门禁在创建 Broker 时没有开启可信重启恢复,所以根本不会调用 rebooted()。
  • 本轮两个变异体都通过了全部六个阶段。运行方式是在安静的主机上串行执行,并且一个阶段要连续失败两次才算杀死。
  • 第三轮看到的失败来自那个对负载敏感的门禁,第三轮自己也把它列为抖动。

现在哪些已被钉住。

变异体 第一到三轮 本轮
M16:recoverResources 的默认实现允许 managed 清理 存活 被作者新增的 RuntimeMaintenanceRecoveryTest 用例杀死
M01:另一台机器、另一个 boot id 也被当成重启 存活 存活
M11、M19、M21 存活 存活(第一轮起就归为较小的缺口)

M01 是存活者里唯一守护安全性质的一个。真机上 machine-id 不同的对照组保持阻塞(R16),所以代码是对的,缺的只是测试。第一轮附上的 RebootRecoveryGapTest 覆盖了这个场景:在 62584d31 上 2/2 通过,在 M01 下失败,Checkstyle 无告警。

RuntimeRecoveryEvidence.matches。 机器人希望有人再看一眼被放宽的 lease 检查。

  • 有 lease 时没有任何放宽。 matches 只有一个调用方 RuntimeBindingRecord.requireEvidence。构造函数在走到这个调用之前,就已经拒绝了与 seed 不匹配的 lease(RuntimeBindingRecord.java:118)。所以新加的子句在那里不可能为假,这就是 M25 存活的原因:它是等价变异体,不是测试缺口。
  • **没有 lease 时,**证据仍然和 seed(request id、runtime instance id、incarnation、lease id、epoch)以及 resource handle 绑定。
  • 真机上这种情况对应「worker 在 READY 之前死亡」。R16 里 binding 1ce9de00 没有 lease,断电前是 RECOVERY_BLOCKED。重启后它带着两条证据被回收(表格链接见英文版)。在本 PR 之前,只要 binding 没有 lease,matches 就返回 false。

6. 其他说明(均不阻塞)

  • 扫描在日志里不留痕迹。 R19 里服务回收了四个代数,还被杀掉两次,但日志里没有任何一行来自扫描或恢复(日志链接见英文版)。运维只能从数据库里看到恢复过程。作者已说明把日志放到后续处理。这里给出的是这一项在真机上的表现。
  • 扫描仍然每 5 秒观察一次所有健康的 binding,并重写它们的行。修复之后这不再导致请求被拒。作者把这项优化延后,我同意可以延后。
  • SPI。 RuntimeBindingRepository 新增了两个没有默认实现的方法(finishLostRecovery、findRecoveryCandidates),仓库外的实现会因此无法编译。仓库内的两个实现都已更新,RuntimeProvisioner.recoverResources 有 fail-closed 的默认实现。包版本是 0.1.0-alpha。是否需要写进 release note 由维护者决定。
  • 仓库自带的脚本测试发现不了守卫不匹配。 check-failsafe-reports.test.js 和 hosted-process-ci.test.js 在 55ded4c3 上 13/13 通过,而那时两个守卫都是失败的。

本轮没有覆盖

物理机和 x86_64、同机多个 Broker、恶意同 UID 工具、远程写入者、恢复或克隆的磁盘、Kubernetes。MariaDB 上只跑了 CI 作业里的两个集成测试类。真机运行使用 MySQL 8.4。

本轮的装置和方法

  • 主机。 与第一、二轮相同的 VM:Ubuntu 24.04.4,内核 6.8.0-117,aarch64,ext4。服务 jar 由 systemd 运行,搭配 MySQL 8.4.11 和生产的 HostIdentity.linux()。第三轮用的是另一台 VM。
  • Bundle。 本 PR 没有改动 TypeScript。worker 和 Harness 的 bundle 由 55ded4c3 构建,R21 和 R22 使用由合并结果构建的 bundle。
  • CI 回放。 命令取自 .github/workflows/sdk-java.yml。MariaDB 作业在专用主机上的 Linux 容器里运行(MariaDB 10.11.18)。Hosted 作业在 macOS 上运行(JDK 21.0.12、Node 22.23.2、MySQL 8.4.7),因为 Hosted 测试需要仓库的 node_modules。GitHub CI 的两个作业都跑在 Ubuntu 上。
  • 一处装置错误。 我第一次在 MariaDB 上回放改名后的树时,复用了上一个分支用过的数据库,JdbcRuntimeBrokerMySqlIT 报 expected: <[1]> but was: <[2]>。这个契约测试用固定前缀写数据,需要一个没被之前的运行用过的数据库。CI 为每个作业启动全新的 MariaDB。对照:未改动的 55ded4c3 在用过的数据库上同样失败。换成全新的数据库后回放通过。
  • 崩溃观察脚本。 第一次尝试(R18)里,第二次 kill 打在了重启后的服务上,而当时 MySQL 挂起的还是第一个服务的语句。现在脚本要求被挂起的语句比当前运行的服务更新。R19 是两次 kill 都落在被挂起的退休语句期间的那次运行。R18 同样收敛,结果也在证据包里。
  • 变异测试在 55ded4c3 上运行。62584d31 与它相比,源码只差一个测试类的名字。
  • 脚本和原始结果都在第四轮证据包里(链接见英文版)。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

[codex] Thanks for round 4 and for correcting the M01/M25 mutation result. I updated the PR description in both languages to cite the current-head physical reboot, power-cut, F1 load and CI evidence, and removed the obsolete retest caveat.

I checked that LocalRuntimeStore.rebooted requires the saved host ID to match before a changed boot ID counts as a reboot. The attached M01 test would pin that safety guard; the R16 foreign-host control confirms the current behavior. Since this is a late review round and the gap is test-only, I will carry that regression test into a follow-up issue or PR, alongside the already deferred scan logging and scale work, without widening this PR. The bot's sole CI blocker was fixed in 62584d31; its re-review is still running.

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /triage

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head 62584d315d93392232b012edd38ba1147ea580a0

Verdict: COMMENT. The one blocking issue ever filed against this PR is confirmed fixed at this head, and I found no Critical in the production code I read. What keeps this off the Approve path is coverage: RuntimeBrokerService.java (+129/−21), the file that holds the recovery entry point, is outside what I could read inside this review's budget.

Historical blocking issue — verified FIXED

The CHANGES_REQUESTED at 55ded4c3 named one mechanical blocker: scripts/check-failsafe-reports.js partitions *IT.java by class name (^Hosted.*IT$ is the hosted family, everything else is not), and the new WorkspaceRecoveryWorkerIT was named for the non-hosted family while pom.xml scheduled it into hosted-harness-mysql, so both Java jobs failed in opposite directions. The prescribed fix was to rename the class and drop both pom hunks.

At this head the file is service/HostedWorkspaceRecoveryWorkerIT.java (+115), no pom.xml appears anywhere in the 39-file diff, and every Java job is green — ubuntu-latest / Java 11, 17 and 21, windows-latest / Java 21, macos-latest / Java 21, Runtime Broker and Managed Agent MariaDB / Java 21, Hosted process fault gates / MySQL 8.4 / Java 21 and Real daemon E2E / Java 11. The prior head had never compiled in CI because every push landed while the PR was a draft; this head has. Both halves of the fix were taken and the symptom is gone.

There is no other blocking finding in this PR's history: no [Critical] inline comment exists, and the maintainer's long review at this exact head concludes that nothing blocks the merge, with three remaining items explicitly non-blocking.

What I verified myself

releaseLost is fenced far more tightly than a holder release needs to be, and fails closed at every step. It refuses up front unless the binding is managed-context, LOST, has stopped writers, and has an operation owner. Inside one transaction it takes FOR UPDATE on the binding row and re-checks binding id, generation, binding_state = 'LOST', tenant, workspace, storage id, record_version, operation_owner, operation_generation, both evidence columns non-null, and an operation lease that has not expired — measuring expiry against the database's own clock read in the same query (UNIX_TIMESTAMP() plus EXTRACT(MICROSECOND …)), so no application/DB clock skew can widen the window. It then requires zero unsettled executions for that binding and generation, locks the lease row, validates that holder_key equals the digest of bindingId, generation and session id, re-checks the lease against a freshly queried DB clock, and issues a conditional UPDATE naming all five columns with a changed != 1 refusal. Every mismatch throws the non-retryable 409 rather than proceeding, and all SQL is parameterized.

claim gained two precondition reads that can only narrow it. Both new FOR UPDATE queries require an exactly-one matching row — the binding READY and not draining with tenant, workspace, workspace generation and storage id all equal, and the session in ACQUIRING or READY with matching harness session — and throw unavailable() otherwise. A claim that previously succeeded on a stale or draining binding now fails closed.

The recovery split preserves existing callers exactly. recoverLost now delegates with holdersCleared = !expected.getRequest().isManagedContext(), and the guard became if (!holdersCleared || !current.hasStoppedWriters() || …) return current;. For a non-managed context holdersCleared is true, so !holdersCleared is false and the condition reduces to the original expression — unchanged behaviour. For a managed context it is false, so recovery returns early and the holder must be released through WorkspaceExecutionStore.releaseLost before the new finishLostRecovery completes it. That is the right direction: the scan cannot clear a workspace holder implicitly.

The coordinator is bounded and cannot wedge. scan() is single-flight behind running.compareAndSet(false, true) and returns immediately once closed. It reads a batch of 8 with a keyset cursor that advances only on a full page and resets on a short one, isolates each candidate with exceptionally(error -> null) plus a catch (RuntimeException) around the synchronous call, and releases running in both whenComplete and the outer catch — including the empty-batch case, where allOf on an empty array completes at once. A failing candidate therefore cannot abort the batch, and a failing batch cannot leave the flag set.

Gates I could not confirm inside budget

  1. RuntimeBrokerService.java (+129/−21) — the largest production change and the one holding recoverBinding, the method the coordinator calls per candidate, plus the maintenance and evidence handling around it. Not read.
  2. LocalProcessRuntimeProvisioner.java (+33/−8) — the writer-stop evidence a reboot recovery depends on. Not read.
  3. EmbeddedRuntimeBroker.java (+20/−2), WorkspaceRuntimeProvisioner.java (+15/−1), ManagedAgentProperties.java (+9) and the application.yml default — the wiring and scheduling of the scan. Not read.
  4. findRecoveryCandidates in JdbcRuntimeBindingRepository beyond its requireId validation, and InMemoryRuntimeBindingRepository (+31/−1), which must mirror the JDBC semantics for the unit tests to mean anything.

CI

Every check passes at this head except review-pr, which failed after 6h0m45s — a timeout of the review automation rather than a build or test failure. Under this channel's rules review-pr is not a gating check in any state, and no other job failed, so I attribute nothing to this PR. Note that reviewDecision still reads CHANGES_REQUESTED because the bot's review at 55ded4c3 was never dismissed; its blocker is fixed, so that review needs dismissing by its author before the decision clears.

Next step: items 1 and 2 would close this. recoverBinding is the single method that ties the coordinator to the fenced store calls I verified, so reading it is what would let me confirm that a reboot recovery cannot adopt a binding whose writers have not provably stopped.

@doudouOUC
doudouOUC added this pull request to the merge queue Sep 29, 2026
Merged via the queue into main with commit a8ba9b5 Sep 29, 2026
87 of 88 checks passed

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Partially reviewed — gaps disclosed. Suggestions are inline.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 5)": did not execute the new tests (H2 WorkspaceRecoveryTest and the MySQL channel) — the assertion walk is source-level only; I did not build or run mvn in the …; "agent reverse-audit (round 5)": WorkspaceRecoveryContract.java:87 — did not attribute WorkspaceRuntimeResolver.resolve 's refusal to the absent-mounts configuration rather than to the grant…; "agent reverse-audit (round 5)": did not read HostedWorkspaceRecoveryWorkerIT end to end to decide whether the real WorkspaceRuntimeProvisioner.recoverResources managed branch is covered ou…; chunk 7: mapSeed 's handling of a null provision_seed_ciphertext ( JdbcRuntimeBindingRepository.java:817-832 ) was not read line-by-line; it only matters if the quali….

Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.

中文说明

仅完成部分审查,审查缺口已披露。 建议见行内评论。

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 5)":did not execute the new tests (H2 WorkspaceRecoveryTest and the MySQL channel) — the assertion walk is source-level only; I did not build or run mvn in the …;"agent reverse-audit (round 5)":WorkspaceRecoveryContract.java:87 — did not attribute WorkspaceRuntimeResolver.resolve 's refusal to the absent-mounts configuration rather than to the grant…;"agent reverse-audit (round 5)":did not read HostedWorkspaceRecoveryWorkerIT end to end to decide whether the real WorkspaceRuntimeProvisioner.recoverResources managed branch is covered ou…;chunk 7:mapSeed 's handling of a null provision_seed_ciphertext ( JdbcRuntimeBindingRepository.java:817-832 ) was not read line-by-line; it only matters if the quali…。

未审查:反向审计——在 5 轮的反审轮数上限内未收敛。

— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)

ContextBinding changedGeneration = new ContextBinding(binding.getTenantId(), "other-workspace", 2,
binding.getStorageId(), "child", binding.getContextConfigRef(), 1);
assertBusy(() -> otherClient.claim(changedGeneration, loser));
assertUnavailable(() -> otherClient.claim(changedGeneration, loser));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-1: The rewritten fencing assertions no longer witness the generation or storage clause they are named for.

changedGeneration and independentStorage are both built with workspaceId = "other-workspace", which alone mismatches the stored qwen_runtime_binding row, so both claims are refused by the workspace-id comparison in WorkspaceExecutionStore.claim's new live predicate — the generation and storage comparisons this same diff added are never the deciding clause.

Measured with the clauses muted out of claim in a scratch tree, WorkspaceRuntimeTest alone:

baseline                        Tests run: 13, Failures: 0
generation clause deleted       Tests run: 13, Failures: 0
generation + storage deleted    Tests run: 13, Failures: 0

So a change that admits a stale-generation or foreign-storage claim ships with the suite green.

Suggested fix: make each binding differ from the stored row in exactly one field — keep tenant/workspace/storage and change only the generation for changedGeneration; keep everything and change only the storage id for independentStorage. Each variant is then refused by exactly one clause of the predicate.

WorkspaceExecutionStore.claim's live predicate also requires the binding row to be READY and not draining (WorkspaceExecutionStore.java:96-108), so the variants must keep the fixture's other columns.

Acceptance criterion: WorkspaceRuntimeTest.serializesIndependentSqlClientsByStorageAndFencesStaleRelease must go red when && binding.getStorageId().equals(row.getString("storage_id")) (or the generation clause) is deleted from claim's live predicate; today it stays green.

中文说明

[Suggestion] 改写后的 fencing 断言不再能见证它们名字所指的 generation/storage 子句。

changedGeneration 和 independentStorage 都用 workspaceId = "other-workspace" 构造,仅这一项就与存储的 qwen_runtime_binding 行不匹配,因此两次 claim 都是被 WorkspaceExecutionStore.claim 新增 live 谓词里的 workspace id 比较拒绝的 —— 本次 diff 同时加入的 generation 与 storage 比较从来不是决定性子句。

在 scratch 树里把 claim 的这两个子句逐个删除后实测(只跑 WorkspaceRuntimeTest):

基线                          Tests run: 13, Failures: 0
删除 generation 子句           Tests run: 13, Failures: 0
删除 generation + storage     Tests run: 13, Failures: 0

也就是说,一个允许过期 generation 或外来 storage 的 claim 会在测试全绿的情况下合入。

建议:让每个 binding 只在一个字段上与存储行不同 —— changedGeneration 保留 tenant/workspace/storage,只把 generation 改成 2;independentStorage 只把 storage id 改成 "other-storage"。这样每个变体都恰好被谓词中的一个子句拒绝。

WorkspaceExecutionStore.claim 的 live 谓词还要求该行是 READY 且未在 drain(WorkspaceExecutionStore.java:96-108),因此变体必须保留 fixture 的其他列。

验收标准:从 claim 的 live 谓词中删除 && binding.getStorageId().equals(row.getString("storage_id"))(或 generation 子句)时,WorkspaceRuntimeTest.serializesIndependentSqlClientsByStorageAndFencesStaleRelease 必须变红;目前它仍然全绿。

— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)

assertTrue(bindings.findRecoveryCandidates("local-process", null, 100).stream()
.anyMatch(record -> record.getBindingId().equals(fixture.binding.getBindingId())));
assertTrue(bindings.findRecoveryCandidates("unsupported-provider", null, 100).isEmpty());
assertTrue(bindings.findRecoveryCandidates("local-process", fixture.binding.getBindingId(), 1).stream()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-2: The cursor-exclusivity assertion runs on an empty page, so allMatch is vacuously true.

The repository holds exactly one record at this point — the fixture's own binding — and the exclusive cursor excludes it, so the stream is empty and the assertion cannot fail.

pageSize=0  allMatch=true  totalRows=1

Measured harm: with findRecoveryCandidates degraded to return an empty page whenever afterBindingId != null, this class stays 4/4 green — the documented "ordered by binding ID after the exclusive cursor" contract is unchecked in the direction that matters.

Suggested fix: insert a second candidate that sorts above the cursor and assert the page exactly:

var page = bindings.findRecoveryCandidates("local-process", fixture.binding.getBindingId(), 1);
assertEquals(1, page.size());
assertTrue(page.get(0).getBindingId().compareTo(fixture.binding.getBindingId()) > 0);

or assert isEmpty() explicitly so the expectation is stated rather than vacuous.

findRecoveryCandidates rejects a batch size outside 1..100 (InMemoryRuntimeBindingRepository.java:186), so the strengthened assertion must not use limit 0 to force an empty page.

Acceptance criterion: with a second candidate present, deleting the cursor's > 0 direction (or the whole clause) in either repository's findRecoveryCandidates must turn the rewritten assertion red.

中文说明

[Suggestion] cursor 排他性断言跑在空页上,因此 allMatch 恒为真。

此时仓库里只有一条记录 —— fixture 自己的 binding —— 而排他 cursor 把它排除在外,所以流为空,断言不可能失败。

pageSize=0  allMatch=true  totalRows=1

实测危害:把 findRecoveryCandidates 改成「afterBindingId != null 时永远返回空页」后,该测试类仍然 4/4 全绿 —— 文档承诺的「按 binding ID 排在排他 cursor 之后」在最关键的方向上无人校验。

建议:插入一条排序在 cursor 之后的候选,并精确断言整页内容(见上方代码),或者显式断言 isEmpty(),把期望写出来而不是留成空断言。

findRecoveryCandidates 会拒绝 1..100 之外的批量大小(InMemoryRuntimeBindingRepository.java:186),因此加强后的断言不能用 limit 0 来制造空页。

验收标准:在有第二条候选的情况下,删除任一仓库 findRecoveryCandidates 中 cursor 的 > 0 方向(或整个子句)时,改写后的断言必须变红。

— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)

var pending = new CompletableFuture<Void>();
try (var service = service(bindings, sessions, executions, provisioner(saved -> pending), Duration.ofMillis(100))) {
var recovery = service.recoverBinding(lost.getBindingId(), lost.getGeneration()).toCompletableFuture();
assertThrows(java.util.concurrent.ExecutionException.class, () -> recovery.get(3, TimeUnit.SECONDS));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-3: The late-cleanup-completion test never reaches the expired-claim guard it is named for.

provisioner(saved -> pending) returns the test's own future, and safeStage returns that stage unchanged, so cleanupLost's orTimeout(operationLeaseDuration…) mutates the same CompletableFuture. With this test's 100 ms lease, recovery.get(3 s) returns only after that timeout has already completed it:

lease=PT0.1S  outcome=recovery-failed:TimeoutException  pendingDoneBefore=true   completeDelivered=false
lease=PT30S   outcome=recovery-still-pending            pendingDoneBefore=false  completeDelivered=true

So pending.complete(null) is a no-op and the assertions after it pass identically — delete those lines and the test stays green. A regression letting an expired claim retire the binding (e.g. removing the hasLiveOperationAt clause) would not be caught by the test whose name claims to pin exactly that.

Suggested fix: have the cleanup callback itself outlive the lease and then return a completed future, so the chain proceeds synchronously to the claim check:

RuntimeProvisioner provisioner = provisioner(saved -> {
    Thread.sleep(leaseMillis * 2);   // the lease lapses while cleanup runs
    return CompletableFuture.completedFuture(null);
});

then assert the failure cause is RuntimeBrokerException with code runtime_provision_fenced, and that the binding is still LOST with its session still active.

The guard is only reachable if the callback completes synchronously: recoverBinding wraps the whole operation with a second orTimeout(operationLeaseDuration…) (RuntimeBrokerService.java:1658).

Acceptance criterion: deleting the liveness clause || !current.hasLiveOperationAt(clock.instant()) from the in-memory repository must turn the rewritten test red.

中文说明

[Suggestion] 「清理迟到完成」这个测试根本没有走到它名字所指的过期 claim 守卫。

provisioner(saved -> pending) 返回的是测试自己的 future,而 safeStage 原样返回该 stage,所以 cleanupLost 的 orTimeout(operationLeaseDuration…) 修改的就是同一个 CompletableFuture。在本测试 100ms 的 lease 下,recovery.get(3s) 返回时该超时早已把它完成:两行实测见上。

因此 pending.complete(null) 是空操作,其后的断言只是照常通过 —— 删掉这几行测试依然全绿。一个让过期 claim 退休 binding 的回归(例如删除 hasLiveOperationAt 子句)不会被这个自称钉住它的测试抓到。

建议:让清理回调自身活过 lease 之后再返回已完成的 future,使调用链同步走到 claim 检查(见上方代码),然后断言失败原因是 code 为 runtime_provision_fenced 的 RuntimeBrokerException,且 binding 仍为 LOST、其 session 仍然活跃。

该守卫只有在回调同步完成时才可达:recoverBinding 还用第二个 orTimeout(operationLeaseDuration…) 包住整个操作(RuntimeBrokerService.java:1658)。

验收标准:删除内存仓库中的活跃性子句 || !current.hasLiveOperationAt(clock.instant()) 时,改写后的测试必须变红。

— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)

if (broker.isDurableLocalProcess()) {
requireRecoveryDirectoryOutsideWorkspaces(broker, stateDirectory);
return LocalProcessRuntimeProvisioner.durable(command, stateDirectory, transport);
return LocalProcessRuntimeProvisioner.durable(command, stateDirectory, transport,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-4: The trusted-local-reboot-recovery option is never enabled through EmbeddedRuntimeBroker in any test, so reverting this wiring keeps every suite green.

setTrustedLocalRebootRecovery(true) appears in exactly one test (EmbeddedRuntimeBrokerTest, and only to assert the negative guard), and every test of the trusted behaviour builds the provisioner by hand — FaultGateRig, FaultGateBroker, DurableLocalProcessRuntimeProvisionerTest, LocalRebootTestSupport — bypassing this composition root. Sweep of test-side setter sites: 1 of 1, negative guard only.

So reverting these two lines to the retained 3-argument durable(...) overload (which compiles) would mean observeDurable never returns WRITERS_STOPPED: a deployed service would keep every rebooted generation pinned forever, while unit, H2, MySQL, fault-gate and hosted jobs all stay green. The same applies to the coordinator construction a few lines above — drop it and recovery is null and the untested recoverSavedRuntimes() silently does nothing every 5 s.

Suggested fix: add a managed-agent-server test that constructs EmbeddedRuntimeBroker with trusted reboot recovery on, durable local-process on, a real state directory and a saved same-host/different-boot record, drives recoverSavedRuntimes(), and asserts the binding reaches RELEASED; or expose the constructed provisioner so the flag can be asserted directly.

Note the trusted path cannot be composed off Linux (LocalRuntimeStore.HostIdentity.linux()), and the guard at EmbeddedRuntimeBroker.java:178-181 already refuses the flag without durable local-process — so the new test needs a Linux/CI lane.

Acceptance criterion: that test must go red when this call is reverted to the 3-arg overload, or when the coordinator construction is dropped.

中文说明

[Suggestion] 没有任何测试通过 EmbeddedRuntimeBroker 打开 trusted-local-reboot-recovery,所以把这处接线改回去所有测试套件依然全绿。

setTrustedLocalRebootRecovery(true) 在整个仓库只出现在一个测试里(EmbeddedRuntimeBrokerTest,而且只断言反向守卫),而所有验证可信行为的测试都是手工构造 provisioner —— FaultGateRig、FaultGateBroker、DurableLocalProcessRuntimeProvisionerTest、LocalRebootTestSupport —— 绕开了这个组合根。测试侧 setter 站点实测:1 处,且只有反向守卫。

因此把这两行改回保留的 3 参 durable(...) 重载(它能编译)后,observeDurable 永远不会返回 WRITERS_STOPPED:部署中的服务会把每个重启后的 generation 永久钉住,而单元、H2、MySQL、fault-gate 与 hosted 各作业全部保持绿色。上面几行的 coordinator 构造同理 —— 去掉它 recovery 为 null,未被测试覆盖的 recoverSavedRuntimes() 每 5 秒静默空转。

建议:新增一个 managed-agent-server 测试,在可信重启恢复与持久本地进程均开启、有真实状态目录和「同宿主不同 boot」的已保存记录下构造 EmbeddedRuntimeBroker,驱动 recoverSavedRuntimes() 并断言 binding 到达 RELEASED;或者把构造出的 provisioner 暴露出来以便直接断言该开关。

注意可信路径无法在非 Linux 上组装(LocalRuntimeStore.HostIdentity.linux()),且 EmbeddedRuntimeBroker.java:178-181 的守卫已拒绝「无持久本地进程」时开启该开关,因此新测试需要 Linux/CI 通道。

验收标准:当这处调用被改回 3 参重载、或 coordinator 构造被删除时,新测试必须变红。

— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)

assertThat(jdbc.queryForObject("SELECT COUNT(*) FROM managed_workspace_execution_lease WHERE binding_id = ?",
Long.class, original.getBindingId())).isZero();
assertThat(fixture.bindings.findActive(original.getRequest())).isNull();
assertThat(worker.isAlive()).isFalse();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-5: This assertion cannot fail, so the "no replacement worker was started" claim is unwitnessed.

worker is the ProcessHandle of the phase-1 worker that this test already killed and awaited 30 lines earlier:

worker.destroyForcibly();  worker.onExit().get(5, TimeUnit.SECONDS);
...
assertThat(worker.isAlive()).isFalse();

Measured: after that pair worker.isAlive() is false, and a restarted process is alive under a different ProcessHandle. So make the restored Broker's recovery start a worker — the regression this line appears to guard — and line 103 stays green while the reader is told the no-restart property is asserted here.

Suggested fix: assert the property on something the code under test controls — capture the live worker pids before and after the release (ProcessHandle.current().descendants(), or a scan for a live node … managed-runtime-worker) and assert the set is unchanged — or drop the line and let the phase-2 must-not-run command carry the claim.

The phase-2 broker is built with the command List.of("must-not-run"), so a started worker would already make the provisioner throw — that is the weaker witness that exists today.

Acceptance criterion: with the replacement assertion, having recoverBinding start a runtime for the released binding must turn it red; today line 103 cannot distinguish that mutation.

中文说明

[Suggestion] 这个断言不可能失败,因此「没有启动替代 worker」这条声明无人见证。

worker 是本测试在 30 行之前就已经杀掉并等待结束的 phase-1 worker 的 ProcessHandle(见上方代码)。

实测:那一对调用之后 worker.isAlive() 为 false,而重启出来的进程是活的、且持有另一个 ProcessHandle。所以让恢复后的 Broker 启动一个 worker(正是这行看似要防的回归),第 103 行依然全绿,而读者却以为「不重启 worker」这条性质在这里被断言了。

建议:把断言建立在被测代码能控制的对象上 —— 在释放前后采集存活的 worker pid(ProcessHandle.current().descendants(),或扫描存活的 node … managed-runtime-worker)并断言集合不变 —— 或者删掉这一行,让 phase-2 的 must-not-run 命令承担该声明。

phase-2 的 broker 用命令 List.of("must-not-run") 构造,因此真启动了 worker 本来就会让 provisioner 抛错 —— 这是目前已有的、更弱的见证。

验收标准:换成新断言后,让 recoverBinding 为已释放的 binding 启动 runtime 必须使它变红;目前第 103 行区分不出这个变异。

— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)

assertEquals(RuntimeBindingRecord.State.LOST, bindings.recoverLost(sessions, executions, lost).getState());
assertTrue(executions.hasActiveByBinding(lost.getBindingId(), lost.getGeneration()));
assertEquals(1, sessions.countActiveByBinding(lost.getBindingId(), lost.getGeneration()));
assertThrows(RuntimeBrokerException.class, () -> bindings.completeSessionRelease(sessions, fixture.session));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-6: This refusal comes from the generation-state guard, not from a live-holder check, so the assertion cannot fail for the reason its position implies.

completeSessionRelease performs one guard per repository — RuntimeAdmission.requireRelease(binding, session) — which reads only binding/session identity, scope and state; no disjunct consults hasActiveByBinding or countActiveByBinding. Measured, the identical call with this fixture:

[READY, 205 executions, 1 session] -> RETURNED RELEASED
[LOST,  205 executions, 1 session] -> THREW runtime_reconciliation_required
[READY,   0 executions, 1 session] -> RETURNED RELEASED

So a change that let a release complete while executions for the generation are still live is not caught here; the only place holders fence a transition is recoverLost/finishLostRecovery.

Suggested fix: keep the assertion — it does pin that a managed LOST generation cannot be released through ordinary completeSessionRelease — but re-anchor it: assert the error code (runtime_reconciliation_required) so the pinned rule is explicit, and add the control it is missing (a READY-state binding with live executions for the same generation).

The LEGACY case must keep failing (ProcessCrashFaultGateTest asserts runtime_reconciliation_required for a restarted-Broker release), so add the managed-shape control rather than loosening that assertion.

Acceptance criterion: the re-anchored assertion must go red when the managed LOST refusal is removed from RuntimeAdmission.requireRelease.

中文说明

[Suggestion] 这个拒绝来自 generation 状态守卫,而不是「持有者仍然活跃」的检查,因此该断言不可能以它所在位置暗示的原因失败。

completeSessionRelease 在每个仓库里只有一个守卫 —— RuntimeAdmission.requireRelease(binding, session) —— 它只读取 binding/session 身份、scope 与状态;没有任何分支去查 hasActiveByBinding 或 countActiveByBinding。用本 fixture 做同一个调用,实测三行结果见上。

也就是说,一个「该 generation 仍有活跃执行却允许 release 完成」的改动在这里抓不到;真正用持有者栅栏住状态跃迁的地方是 recoverLost/finishLostRecovery。

建议:保留该断言 —— 它确实钉住了「managed LOST generation 不能通过普通 completeSessionRelease 释放」—— 但重新锚定:断言错误 code(runtime_reconciliation_required)把被钉住的规则写明确,并补上它缺失的对照(同 generation、READY 状态且存在活跃执行的 binding)。

LEGACY 场景必须继续失败(ProcessCrashFaultGateTest 对重启后 Broker 的 release 断言 runtime_reconciliation_required),所以应新增 managed 形态的对照,而不是放宽那条断言。

验收标准:当 RuntimeAdmission.requireRelease 中 managed LOST 的拒绝被移除时,重新锚定的断言必须变红。

— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)

throw new AssertionError("Maintenance cannot call a dead worker: " + method);
});
return new RuntimeBrokerService(id -> { throw new AssertionError("Maintenance cannot resolve current grants"); },
provisioner, transport, bindings, sessions, executions, "maintenance", lease, lease);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-7: The refusal half of the new recoverBinding is pinned by no test.

Every call site in the tree passes a generation read from the same record and uses a single owner id, and no test races a second owner: 14 recoverBinding( sites, zero with a non-matching generation, and RuntimeRecoveryCoordinatorTest mocks the service. Measured — each of the three refusals neutralized in turn, whole module re-run:

baseline                                             Tests run: 430, Failures: 0, Errors: 0, Skipped: 2
precondition narrowed to `record == null`            Tests run: 430, Failures: 0, Errors: 0, Skipped: 2
`putIfAbsent` refusal removed                        Tests run: 430, Failures: 0, Errors: 0, Skipped: 2
rival-broker refusal degraded to a plain findById    Tests run: 430, Failures: 0, Errors: 0, Skipped: 2

So the guard that keeps a loser Broker from starting a second cleanup of the same generation can be dropped without a red test — the declared retryable 503 runtime_reconcile_in_progress would degrade to an unclassified failure inside the recovery stage.

Suggested fix: add two cases reusing this file's Fixture/provisioner/service helpers — one calling recoverBinding(lost.getBindingId(), lost.getGeneration() + 1) (and a mismatched provisioner kind) asserting runtime_broker_recovery_blocked with the binding still LOST; and one where a second RuntimeBrokerService with a different owner id holds the claim, asserting runtime_reconcile_in_progress while that claim is live.

The rival case needs a different owner id, not a second service with the same one — the claim is taken as claimOperation(bindingId, brokerOwnerId, operationLeaseDuration) (RuntimeBrokerService.java:88,1613).

Acceptance criterion: the new cases must go red when the generation/provisioner precondition, the putIfAbsent refusal, or the rival-owner refusal is removed.

中文说明

[Suggestion] 新增的 recoverBinding 的「拒绝分支」没有任何测试覆盖。

仓库里每个调用点都传入从同一条记录读出的 generation,并使用同一个 owner id,也没有测试制造第二个 owner 竞争:全树 14 处 recoverBinding(,没有一处传入不匹配的 generation,而 RuntimeRecoveryCoordinatorTest 是对 service 打桩。实测:把这三个拒绝分支逐个去掉后整个模块重跑,结果与基线一模一样(Tests run: 430, Failures: 0, Errors: 0, Skipped: 2)。

因此「阻止落败的 Broker 对同一 generation 发起第二次清理」的守卫可以在没有测试变红的情况下被删除 —— 声明为可重试的 503 runtime_reconcile_in_progress 会退化成恢复阶段里一个未分类的失败。

建议:复用本文件的 Fixture/provisioner/service 辅助方法补两个用例 —— 一个调用 recoverBinding(lost.getBindingId(), lost.getGeneration() + 1)(以及 provisioner kind 不匹配),断言 runtime_broker_recovery_blocked 且 binding 仍为 LOST;另一个用不同 owner id 的第二个 RuntimeBrokerService 持有 claim,断言在其存活期间得到 runtime_reconcile_in_progress。

竞争用例必须用不同的 owner id,而不是同一个 owner id 的第二个 service —— claim 是通过 claimOperation(bindingId, brokerOwnerId, operationLeaseDuration) 取得的(RuntimeBrokerService.java:88,1613)。

验收标准:当 generation/provisioner 前置条件、putIfAbsent 拒绝、或对手 owner 拒绝被移除时,新用例必须变红。

— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)

fixture.session.sessionId());
jdbc.update("UPDATE managed_workspace_registry SET storage_id = 'changed', state = 'REMOVED' WHERE tenant_id = ?",
fixture.tenant);
var noMounts = new WorkspaceRuntimeResolver(store, authority, new ManagedAgentProperties());

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-8: The phase-2 authorization canary sits on a dependency recoverBinding cannot reach, so a tolerated authorization consultation leaves this IT green.

RuntimeBrokerService.recoverBinding never reads its sessionResolver — that field is read only in resolveScope, reached from the acquire/warm paths — and the phase-2 broker here is never asked to acquire or warm, so the AssertionError canary cannot fire for this test's only call. The objects the release path does hold are noMounts and authority, and both are observed only through the outcome.

Measured, injecting the shape this canary is meant to exclude (resolveScope(...).exceptionally(error -> null) into cleanupLost):

baseline  consulted=0  state=RELEASED  loss=true
mutant    consulted=1  state=RELEASED  loss=true

and RuntimeMaintenanceRecoveryTest stays Tests run: 4, Failures: 0 under the mutant, because safeStage catches Error as well as RuntimeException (RuntimeBrokerService.java:2574). The canaries make a cheating recovery fail loudly only when the consultation does not tolerate the refusal.

Suggested fix: put the canary on the seam the recovery path actually crosses. Build the phase-2 WorkspaceRuntimeResolver around a spy of the store or of the execution authority and, after recoverBinding, assert the authorizing interaction never happened:

verify(authority, never()).authorize(any());

WorkspaceRuntimeResolver is final, so instrument a collaborator it delegates to (AgentStateStore sessions, WorkspaceExecutionStore authority) or spy the concrete instance.

Acceptance criterion: with the spy assertion in place, inserting noMounts.resolve(...) or authority.authorize(...) into the LOST/release branch must turn the IT red; today that mutation leaves the whole IT green.

中文说明

[Suggestion] phase-2 的「授权探测哨」挂在了 recoverBinding 结构上无法触及的依赖上,因此一次「容错吞掉拒绝」的授权查询会让这个 IT 保持全绿。

RuntimeBrokerService.recoverBinding 从不读取它的 sessionResolver —— 该字段只在 resolveScope 里被读,而那只从 acquire/warm 路径到达 —— 而这里 phase-2 的 broker 从未被要求 acquire 或 warm,所以这个测试唯一的一次调用不可能触发 AssertionError 哨兵。释放路径真正持有的对象只有 noMounts 与 authority,而两者都只通过最终结果被观察。

实测:注入这个哨兵本该排除的形态(把 resolveScope(...).exceptionally(error -> null) 放进 cleanupLost)后,基线与变异体的四个可观测量完全一致(见上),且变异体下 RuntimeMaintenanceRecoveryTest 仍是 Tests run: 4, Failures: 0 —— 因为 safeStage 同时捕获 Error 与 RuntimeException(RuntimeBrokerService.java:2574)。只有当授权查询不容忍拒绝时,哨兵才会「大声失败」。

建议:把哨兵挪到恢复路径真正跨越的接缝上。用 store 或执行授权对象的 spy 构造 phase-2 的 WorkspaceRuntimeResolver,并在 recoverBinding 之后断言授权交互从未发生(见上方 verify(authority, never()).authorize(any());)。

WorkspaceRuntimeResolver 是 final,因此需要插桩它委托的协作者(AgentStateStore sessions、WorkspaceExecutionStore authority),或对具体实例做 spy。

验收标准:在加入 spy 断言后,把 noMounts.resolve(...) 或 authority.authorize(...) 插入 LOST/release 分支必须让该 IT 变红;目前这个变异下整个 IT 仍然全绿。

— DeepSeek/deepseek-v4.1-flash@2d473174 via Qwen Code /review (v0.24.6)

JadeCong pushed a commit to CloudEngineHub/qwen-code that referenced this pull request Sep 30, 2026
…M#12977)

* feat(managed-agent): Recover Workspace holders after trusted local reboot

(cherry picked from commit 8c2b626)

* fix(runtime-broker): Serve healthy waiters during maintenance

* fix(sdk-java): Route recovery IT to hosted CI (QwenLM#12869)

* fix(sdk-java): Add audited Hosted Workspace operator recovery

* codex: address PR review feedback (QwenLM#12977)

* fix(sdk-java): Validate operator evidence after recovery prepare

* fix(sdk-java): Fence operator recovery against concurrent placement

* fix(sdk-java): Upgrade fault-gate H2 test dependency
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants