Skip to content

Commit f8bf5d7

Browse files
jinye.djyjinye.djy
authored andcommitted
feat(managed-agent): Recover Workspace holders after trusted local reboot
1 parent 7236322 commit f8bf5d7

39 files changed

Lines changed: 1346 additions & 52 deletions

‎docs/design/2026-09-27-managed-workspace-recovery.md‎

Lines changed: 12 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -2,8 +2,8 @@
22

33
[English](2026-09-27-managed-workspace-recovery.md) | [简体中文](2026-09-27-managed-workspace-recovery.zh-CN.md)
44

5-
Status: W0e-1/2 implementation and proposed W0e-3 follow-up design, 2026-09-27.
6-
Physical reclamation remains follow-up work. Baseline: `e0b8bea9e0ba369a0661bc51cbbb9a27555aff48`.
5+
Status: W0e-1/2/3 source implementation, 2026-09-28.
6+
Dedicated Linux physical reboot acceptance remains pending. Baseline: `e0b8bea9e0ba369a0661bc51cbbb9a27555aff48`.
77

88
Related: [roadmap #12380](https://github.com/QwenLM/qwen-code/issues/12380),
99
[lost executions #12670](https://github.com/QwenLM/qwen-code/issues/12670),
@@ -372,6 +372,16 @@ registration, permanent launch locks, and adoption after Broker restart. The
372372
default ephemeral mode retains the W0e-1 behavior. Neither mode proves stopped
373373
writers after worker-only death; W0e-3 physical reclamation remains separate.
374374

375+
## 6c. W0e-3 implementation
376+
377+
The [trusted local reboot implementation](2026-09-28-local-reboot-recovery.md)
378+
adds separately enabled same-host boot evidence, original-holder cleanup,
379+
late-acquisition fencing and an independent bounded maintenance scan. Cleanup
380+
uses saved physical ownership even after grants, product Session or Registry
381+
change. Loss receipts remain terminal uncertainty. Tests with real workers and
382+
SQL plus synthetic boot identity exercise the chain; a dedicated Linux reboot
383+
is still required before physical acceptance can be claimed.
384+
375385
## 7. Delivery order and boundaries
376386

377387
| Slice | Deliverable | Exit condition |

‎docs/design/2026-09-27-managed-workspace-recovery.zh-CN.md‎

Lines changed: 5 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -2,7 +2,7 @@
22

33
[English](2026-09-27-managed-workspace-recovery.md) | [简体中文](2026-09-27-managed-workspace-recovery.zh-CN.md)
44

5-
状态:W0e-1/2 实现及待评审的 W0e-3 后续设计,2026-09-27。物理回收仍属于后续工作。基线:`e0b8bea9e0ba369a0661bc51cbbb9a27555aff48`。
5+
状态:W0e-1/2/3 源码实现,2026-09-28。专用 Linux 物理重启验收仍待完成。基线:`e0b8bea9e0ba369a0661bc51cbbb9a27555aff48`。
66

77
关联:[路线图 #12380](https://github.com/QwenLM/qwen-code/issues/12380)、[丢失执行 #12670](https://github.com/QwenLM/qwen-code/issues/12670)、[本地 worker #12766](https://github.com/QwenLM/qwen-code/issues/12766)、[W0c-3](2026-09-26-managed-workspace-execution.zh-CN.md) 和 [W0d](managed-workspace-w0d-web-shell-binding.zh-CN.md)。
88

@@ -120,6 +120,10 @@ Flyway V16 和独立 initializer 增加可空证据/放弃字段及 placement
120120

121121
[本地持久接管实现](2026-09-27-local-runtime-adoption.zh-CN.md) 增加显式启用的 Linux 身份存储、带持久 PID/启动 tick 登记的 boot 屏障、永久启动锁和 Broker 重启后的接管。默认临时模式保留 W0e-1 行为。两种模式均不在仅 worker 死亡时证明写入者已停止;W0e-3 物理回收仍单独实现。
122122

123+
## 6c. W0e-3 实现
124+
125+
[可信本地重启实现](2026-09-28-local-reboot-recovery.zh-CN.md) 增加单独启用的同宿主 boot 证据、原 holder 清理、迟到 acquire 屏障和独立有界维护扫描。grant、产品 Session 或 Registry 改变后,清理仍使用原保存物理归属。loss 回执始终表示终态不确定性。真实 worker、SQL 和模拟 boot 身份测试覆盖完整链路;宣称物理验收前仍需专用 Linux 重启验证。
126+
123127
## 7. 交付顺序与边界
124128

125129
| 切片 | 交付内容 | 退出条件 |
Lines changed: 47 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,47 @@
1+
# Trusted Local Reboot Recovery (W0e-3)
2+
3+
[English](2026-09-28-local-reboot-recovery.md) | [简体中文](2026-09-28-local-reboot-recovery.zh-CN.md)
4+
5+
Status: implementation design, following [durable adoption](2026-09-27-local-runtime-adoption.md).
6+
7+
## Problem and scope
8+
9+
A lost Runtime journal can terminate unknown executions without proving that its writers stopped. Workspace storage must remain pinned until both physical evidence and exact holder cleanup exist. Recovery must also progress after the actor loses access or the product Session is deleted. Current authorized warm and acquire routes cannot provide that independent maintenance path.
10+
11+
This slice supports explicitly opted-in trusted workloads on one Linux host with administrator-managed persistent local storage. It does not isolate malicious same-UID tools, remote writers, external jobs that recreate writers, or restore/clone snapshots. Keep machine identity, local records, SQL keys and storage ownership stable. Worker-only death remains insufficient, even when no process is currently visible.
12+
13+
## Trusted reboot proof
14+
15+
The default remains disabled. Enabling trusted local reboot recovery requires durable local provisioning. Production reads machine ID and kernel boot ID; a different boot on the same saved host establishes that writers from the original boot cannot survive within the supported local-storage contract. The original registration must still validate against its complete provision seed, placement and saved handle. Missing, corrupt or old-format records cannot be reconstructed. Same-boot PID/time namespace changes remain uncertain.
16+
17+
Before returning evidence, the provisioner atomically tombstones the original record under its permanent per-seed lock. It returns JOURNAL_LOST and WRITERS_STOPPED for the same original host/boot/resource domain. Earlier process-exit loss evidence stays immutable; later reboot evidence closes only the physical uncertainty. A saved seed and handle suffice for interrupted startup, without inventing a missing endpoint or lease. Observation never relaunches the original seed.
18+
19+
## Ordered cleanup
20+
21+
The Broker first persists loss evidence and abandons at most 100 nonterminal executions per transaction. SETTLED receipts retain their results; ABANDONED remains terminal uncertainty and is never replayed. Loss-only cleanup leaves all physical pins intact.
22+
23+
Once all executions are terminal and writer-stop evidence is durable, a trusted provisioner cleanup callback handles the original saved binding. The Workspace wrapper clears its SQL storage holder using only the original tenant, storage ID, binding ID, generation and holder identity. It does not consult current actor grants, product Session status, Registry, mounts, filesystem or worker HTTP.
24+
25+
The cleanup transaction locks and validates the saved Binding and its live operation claim, then the original holder. It conditionally clears only an exact matching original holder. An absent holder or holder from another generation is already complete; another holder is never erased. Replaying cleanup after a crash is safe. Only after the callback completes does a final bounded transaction release Runtime Sessions, retire the Binding and clear the placement slot. A failed callback, expired claim or partial batch retains the remaining pins.
26+
27+
Late acquisitions must lock Binding, Runtime Session and holder in that order, and recheck READY/not-draining and original Session ownership. They cannot recreate the old holder after the loss fence. Ordinary release cannot bypass cleanup for a managed LOST generation.
28+
29+
## Recovery entry and scheduling
30+
31+
A trusted in-process `recoverBinding(bindingId, expectedGeneration)` entry loads the original saved record. It shares per-binding exclusion and SQL operation claims with interactive work, performs bounded observation and cleanup, and returns the original generation state. It never calls the current Session resolver, constructs a new placement, ensures a new resource, provisions a process or creates a replacement generation.
32+
33+
Spring schedules a bounded scan only when trusted local recovery is enabled. A rotating binding-ID cursor prevents long-lived uncertain records from starving later candidates. Batches cannot overlap and each observation has a deadline. Recovery is independent of active Hosted Turns and current user authorization. Live workers may be adopted, but no new execution starts. A later authorized warm may create the next generation after physical cleanup completes. No additional HTTP route or force-unlock API is introduced.
34+
35+
## Components and compatibility
36+
37+
The local store/provisioner add explicit reboot policy and seed-only evidence matching. Binding repositories expose bounded candidate discovery and a separate finalization step. The Workspace execution store enforces acquisition fencing and exact lost-holder cleanup. Embedded configuration wires the policy, wrapper callback and background coordinator. Existing SQL tables and evidence columns suffice; no migration is added. Default ephemeral behavior and durable live-worker adoption remain unchanged. The public `workspace_context` capability stays false.
38+
39+
## Verification and acceptance
40+
41+
Tests must cover both boot protocols, interrupted registration without a lease, same-host changed boot, wrong host, missing/corrupt records, same-boot PID/time namespace change, worker death with escaped descendants, more than 100 executions, crash after holder clear, stale claims and acquire callbacks, and a newer holder surviving old cleanup. A real SQL chain must recover the original saved Workspace after grant revocation, product Session deletion and Registry/mount changes, without accepting a new unauthorized execution or starting a replacement during maintenance.
42+
43+
Run Java HTTP/H2 tests, the existing MySQL integration profile, real-worker Stage F gates, build/typecheck, formatting and two consecutive clean full-diff audits. Synthetic boot identities prove decisions only. Physical acceptance requires a dedicated supported Linux host: record the original worker and escaped writer, reboot while preserving local disk and SQL, restart the same configuration, then verify old receipts, holder retirement and a new authorized generation. Never reboot a shared development machine as a substitute.
44+
45+
## Open validation
46+
47+
A dedicated rebootable Linux host has been requested and is not yet available. Report production Linux identity and physical reboot evidence separately from portable macOS process tests. W0e source implementation and local tests are not a claim that the physical reboot acceptance gate has passed.
Lines changed: 47 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,47 @@
1+
# 可信本地重启恢复(W0e-3)
2+
3+
[English](2026-09-28-local-reboot-recovery.md) | [简体中文](2026-09-28-local-reboot-recovery.zh-CN.md)
4+
5+
状态:实施设计,接续[持久接管](2026-09-27-local-runtime-adoption.zh-CN.md)。
6+
7+
## 问题与范围
8+
9+
Runtime journal 丢失可以结束未知执行,却不能证明写入者已停止。Workspace storage 必须保留占用,直到物理证据和精确 holder 清理均完成。actor 被撤权或产品 Session 删除后,恢复也必须能推进;现有先授权的 warm/acquire 路径不能提供独立维护入口。
10+
11+
本切片仅支持显式启用的可信工作负载、Linux 单宿主及管理员管理的本地持久存储。不隔离恶意同 UID 工具、远程写入者、重新创建写入者的外部任务或快照恢复/克隆。machine 身份、本地记录、SQL 密钥和存储归属必须稳定。仅 worker 死亡始终不够,即使当前看不到进程。
12+
13+
## 可信重启证明
14+
15+
默认关闭。启用可信本地重启恢复必须同时启用本地持久 provisioning。生产读取 machine ID 和内核 boot ID;在上述本地存储契约内,同一保存宿主的 boot 改变证明原 boot 写入者不再存活。原登记必须仍能验证完整 provision seed、placement 和保存 handle;缺失、损坏或旧格式记录不能重建。同 boot 的 PID/time namespace 变化仍是不确定状态。
16+
17+
产生证据前,provisioner 在永久 per-seed 锁内原子写入原记录墓碑,然后对同一原 host/boot/resource 域返回 JOURNAL_LOST 和 WRITERS_STOPPED。此前进程死亡的 loss 证据保持不变,后来的重启证据只解除物理不确定性。启动中断时,已保存 seed 和 handle 足够核验,无需编造缺失 endpoint 或 lease。观察永不重新启动原 seed。
18+
19+
## 有序清理
20+
21+
Broker 先持久化 loss 证据,每次事务最多放弃 100 个非终态执行。SETTLED 保留结果;ABANDONED 始终表示终态不确定性,不能重放。只有 loss 的清理保留全部物理占用。
22+
23+
全部执行进入终态且 stop 证据持久化后,由可信 provisioner 回调清理原保存 Binding。Workspace wrapper 仅用原 tenant、storage ID、Binding ID、generation 和 holder 身份清理 SQL storage holder,不读取当前 actor grant、产品 Session 状态、Registry、mounts、文件系统或 worker HTTP。
24+
25+
清理事务锁定并验证保存 Binding 与未过期 operation claim,再锁原 holder。只条件清除精确匹配的原 holder;holder 已不存在或属于其他代数时已完成,不能删除其他占用者。清理后崩溃可安全重试。回调完成后,最终有界事务才释放 Runtime Sessions、退休 Binding 和清除 placement slot。回调失败、claim 过期或批处理未完时保留剩余占用。
26+
27+
迟到 acquire 必须依次锁 Binding、Runtime Session、holder,复查 READY、未 draining 和原 Session 归属,不能在 loss 屏障后重新建立旧 holder。managed LOST 代数不能通过普通 release 绕过清理。
28+
29+
## 恢复入口与调度
30+
31+
可信进程内入口 `recoverBinding(bindingId, expectedGeneration)` 读取原保存记录,与交互操作共享 per-binding 排他和 SQL operation claim,执行有界观察与清理,返回原代数状态。不调用当前 Session resolver,不构造新 placement,不 ensure 新资源,不启动进程或创建替代代数。
32+
33+
Spring 仅在启用可信本地恢复时调度有界扫描。轮转的 Binding ID 游标避免长期不确定记录使后续候选饥饿。批次不重叠,每次观察有截止时间。恢复不依赖活跃 Hosted Turn 或当前用户授权。允许接管存活 worker,但不能发起新执行。物理清理完成后,后续通过当前授权的 warm 才能创建下一代。不增加 HTTP 路由或强制解锁 API。
34+
35+
## 组件与兼容性
36+
37+
本地 store/provisioner 增加显式 reboot policy 和无 lease 的 seed 证据匹配。Binding 仓库增加有界候选查询与独立最终释放步骤。Workspace execution store 增加 acquire 屏障及精确失联 holder 清理。嵌入配置连接 policy、wrapper 回调和后台 coordinator。复用已有 SQL 表和证据字段,无需迁移。默认临时行为与持久存活 worker 接管保持不变。公开 `workspace_context` 能力仍为 false。
38+
39+
## 验证与验收
40+
41+
覆盖两种 boot 协议、无 lease 的登记中断、同宿主 boot 改变、错误宿主、缺失/损坏记录、同 boot 的 PID/time namespace 改变、worker 死亡但逃逸子进程存活、超过 100 个执行、清理 holder 后崩溃、过期 claim 与迟到 acquire、新 holder 不被旧清理删除。真实 SQL 链路必须证明撤权、删除产品 Session、修改 Registry/mount 后仍按原保存 Workspace 恢复,且维护过程不能接受新的未授权执行或启动替代 Runtime。
42+
43+
执行 Java HTTP/H2、现有 MySQL 集成通道、真实 worker Stage F、build/typecheck、格式检查和连续两轮干净全量 diff 审计。模拟 boot 身份只能证明决策。物理验收需要专用受支持 Linux 主机:记录原 worker 与逃逸写入者,保留本地磁盘和 SQL 进行重启,以相同配置恢复服务,再核对旧回执、holder 退休和新授权代数。不能重启共享开发机充当验收。
44+
45+
## 待验证事项
46+
47+
已请求专用可重启 Linux 主机,目前尚未获得。生产 Linux 身份和物理重启证据必须与可移植 macOS 进程测试分别报告。W0e 源码实现和本地测试不代表物理重启验收已经通过。

‎packages/sdk-java/managed-agent-server/README.md‎

Lines changed: 18 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -287,6 +287,22 @@ Missing or damaged records and worker death do not authorize replacement;
287287
worker death does not prove escaped writers stopped. No host reboot reclamation
288288
is enabled by this option. Old v1 handles cannot be upgraded by guessing identity.
289289
See the [adoption design](../../../docs/design/2026-09-27-local-runtime-adoption.md).
290+
291+
For trusted same-host Linux reboot recovery, additionally set
292+
`QWEN_MANAGED_AGENT_RUNTIME_TRUSTED_LOCAL_REBOOT_RECOVERY=true`. This requires
293+
durable local mode. A changed kernel boot ID on the original machine can prove
294+
that original local writers stopped; worker-only death still cannot. The
295+
service scans eight saved bindings every five seconds, independently of current
296+
Session grants, and clears only the original SQL holder after all execution
297+
receipts become terminal. Recovery never starts a replacement worker or replays
298+
an unknown execution. A later authorized request may create a new generation.
299+
Keep the same Broker user, local disks, machine identity and SQL keys; remote
300+
writers, restored/cloned snapshots and external jobs that recreate writers are
301+
outside this contract. The option remains disabled by default. The
302+
[reboot recovery design](../../../docs/design/2026-09-28-local-reboot-recovery.md)
303+
distinguishes portable test evidence from the dedicated Linux reboot acceptance
304+
gate, which is still pending.
305+
290306
The Kubernetes adapter's real-cluster fault matrix remains a production gate. This
291307
standalone reference keeps the one configured directory for legacy unbound
292308
Sessions. Persisted bound Sessions use the private Workspace execution path
@@ -330,7 +346,8 @@ responses retain the SQL holder; there is no timeout-based takeover. The
330346
provider and file tools do not confine access to the mount root: Read/Write/Edit
331347
and Shell can reach other paths allowed by the worker's host permissions.
332348
Foreground Shell may create detached descendants. Use this only with trusted
333-
local workloads until physical isolation and W0e cleanup are implemented.
349+
local workloads. The opt-in W0e recovery above handles trusted host reboot; it
350+
does not provide physical isolation or recovery after worker-only death.
334351
Public bound Turn/lifecycle gates and the full Hosted tool loop remain closed.
335352
See the bilingual [execution design](../../../docs/design/2026-09-26-managed-workspace-execution.md)
336353
for the exact boundary.

‎packages/sdk-java/managed-agent-server/pom.xml‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -126,6 +126,7 @@
126126
<configuration>
127127
<excludes>
128128
<exclude>**/Hosted*IT.java</exclude>
129+
<exclude>**/WorkspaceRecoveryWorkerIT.java</exclude>
129130
</excludes>
130131
</configuration>
131132
<executions>
@@ -151,6 +152,7 @@
151152
<configuration>
152153
<includes>
153154
<include>**/Hosted*IT.java</include>
155+
<include>**/WorkspaceRecoveryWorkerIT.java</include>
154156
</includes>
155157
<failIfNoTests>true</failIfNoTests>
156158
<forkedProcessTimeoutInSeconds>180</forkedProcessTimeoutInSeconds>

‎packages/sdk-java/managed-agent-server/src/main/java/com/alibaba/qwen/code/managedagent/config/ManagedAgentProperties.java‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -284,6 +284,7 @@ public static class RuntimeBroker {
284284
private String isolationClass = "session";
285285
private String stateDirectory = "";
286286
private boolean durableLocalProcess;
287+
private boolean trustedLocalRebootRecovery;
287288
private String credentialKeyId = "";
288289
private String credentialKey = "";
289290
private String nodeExecutable = "";
@@ -399,6 +400,14 @@ public void setDurableLocalProcess(boolean durableLocalProcess) {
399400
this.durableLocalProcess = durableLocalProcess;
400401
}
401402

403+
public boolean isTrustedLocalRebootRecovery() {
404+
return trustedLocalRebootRecovery;
405+
}
406+
407+
public void setTrustedLocalRebootRecovery(boolean trustedLocalRebootRecovery) {
408+
this.trustedLocalRebootRecovery = trustedLocalRebootRecovery;
409+
}
410+
402411
public String getStateDirectory() {
403412
return stateDirectory;
404413
}

0 commit comments

Comments
 (0)