Repository navigation
feat(managed-agent): Add durable remote Shell result delivery - #12894
Conversation
O2 本地验收报告(2026-09-28,
|
|
Independent review found three integration issues, fixed in
Post-fix verification: full build, typecheck, bundle, scoped ESLint and formatting passed; the three relevant CLI test files passed 75/75. Independent test-engineer reran the v3 routes, publication and Hosted turn tests (62/62) and Java Workspace Runtime plus publication store tests (31/31), including the 100 MiB worker-host reopen case. The new edge tests use mocked HTTP/Broker/worker behavior. Real private OSS and second-host recovery remain unverified, so this PR stays Draft. |
|
Follow-up verification after merging W0e (
The new W0e fallback test uses controlled HTTP/Broker doubles. A real private OSS bucket and a second host are still unavailable, so this PR remains Draft. No production OSS or cross-host recovery claim is made. |
|
The previous head's Hosted MySQL and Managed Agent MariaDB jobs failed for one deterministic merge-base issue: W0e and the later D4 merge both brought a Head Post-fix local evidence: H2 migration/publication tests 24/24; |
Real-stack verification of #12894 (head
|
|
Thanks for the real-stack report in this comment. I reproduced B1 with a real local Session and B2/B4 with isolated MySQL 8.4 transactions, then pushed the first blocker batch in
Verification for this batch: Core 24/24, CLI 25/25, Java publication store 19/19; repository build, TypeScript typecheck and Java verify passed. An isolated MySQL 8.4 two-transaction reproduction confirmed that The other measured issues in your report remain open, especially lost ingress/admission/receipt replies, finish digest agreement, the misleading Hosted preview, and strict range validation. Real OSS and second-host evidence also remain merge gates. This update does not claim that the PR is ready to merge. The independent migration correction is now on |
Real-stack verification, round 14 — R13-1 fix, real Aliyun OSS and a second host (head
|
| Local path | O2 at 42e2e8b1 |
O2 at af962df7 |
|
|---|---|---|---|
sleep 3; echo slept |
Blocked: sleep 3 followed by: echo slept. … (291 chars) |
Runtime Shell did not start. |
the same 291 chars as the local path |
| Ownership refused in the install window (R5-1) | Workspace execution was refused before dispatch. |
Runtime Shell did not start. |
Workspace execution was refused before dispatch. |
| Real qwen3.8-max, "Wait 3 seconds, then print the current date with the date command. Do it in a single shell command." | 2/2, 3 calls | 2/2, 11 calls each, 105–114 s | 3/3, 3 calls, 23–28 s |
- The real model after the fix. Each run went
sleep 3 && date(blocked), thensleep 3 # intentional-sleep: …, thendate. Each answer explained that the harness had required two calls. - Durability. The reason is stored with the Session. I detached the Session and loaded it on a new Harness process, and the replayed history held the same text.
- Size limit. A command of 200 KB is refused earlier on both paths, with
Hosted tool input exceeds the inline Session Store limit.(turn error, Session not blocked). So reasons of about 5 KB are the realistic way to reach the new branch for reasons over 4096 characters.
R14-1 (low): the head/tail cut can split a surrogate pair
af962df7 keeps errorMessage.slice(0, 2048) and errorMessage.slice(-2048) for reasons over 4096 UTF-16 units. Both cuts count code units, so an emoji or other character outside the BMP that sits on a cut point loses half of its surrogate pair. The Shell tool puts the rest of the command into the reason, so a long command with emoji is enough to trigger it:
| Reason (O2, fake model) | What the model received |
|---|---|
| about 4.9K ASCII | 4123 units with the marker, no lone surrogates |
| emoji whose high surrogate sits at index 2047 | lone U+D83D at 2047 |
| emoji at index 2046 | the pair is kept whole |
| all emoji | lone U+D83D at 2047 and lone U+DE00 at 2075, where the tail starts |
- What happened next: nothing failed on this stack.
- Every turn completed and no Session was blocked.
- The Session Store accepted the message, and the replay on a new Harness process was identical.
- I then loaded the two Sessions with lone surrogates, and the ASCII one as a control, on a Harness backed by real qwen3.8-max. The captured request bodies carried 1 and 2 escaped lone surrogates (0 for the control). DashScope returned 200 and the model explained the block correctly.
- Why it still matters:
- Lone surrogates in tool arguments already block O2 Sessions (fix(managed-agent): Handle unpaired surrogate Shell arguments before dispatch #13010). Here the Harness creates them itself, from valid input.
- A stricter provider or JSON consumer may reject them. I did not test one.
- Test coverage: no test reaches this branch. The marker text
[... error truncated ...]appears only in the source, and the new test uses a 42-character reason. - Suggestion: cut on code-point boundaries.
cutProviderFitSlotinmanaged-runtime-provider-protocol.ts:572already does this: it steps back when a cut would split a pair. Also add a test with a reason over 4096 units that has a surrogate pair at a cut point.
For comparison, the local path gives the model the same all-emoji reason cut to 530 units, ending in U+FFFD. That is main's code, not this PR's.
Real Aliyun OSS (42e2e8b1)
Setup:
- Bucket: a temporary bucket in cn-hangzhou, private and unversioned, with a 1-day lifecycle rule.
- Cleanup: all 1,599 objects (1.59 GB), every version and the bucket itself were deleted afterwards.
- Object store: Spring ran the real
AliyunToolPublicationObjectStorewith the JDK default trust store. - Credentials: they came from a local file and are not in any published file.
- Checking: after each run that checks output bytes, I downloaded every stored object from the bucket and compared its SHA-256 with the catalog and the generator's oracle.
| Scenario | Result |
|---|---|
echo |
pass. 2 objects, 21 B. The small resources stay inline in SQL |
| 100 MiB stdout + 5 MiB stderr, exit 7 | 105 segment objects downloaded and exact; every range read exact; 76 s for the turn (upload at about 12–14 Mbps) |
| Crash after the receipt commit, a lost segment request, a lost admission reply, control characters plus a crash, reload plus a second turn | pass, bytes exact |
| 2 O2 Sessions × 64 MiB at once | exact, 0 deadlocks |
| Preview matrix | same results as with the OSS double |
| Real qwen3.8-max, failing 3,000-line build | 2/2, error quoted |
| Anonymous GET of a segment object | 403 |
| One bit flipped in stored segment 50 | reads return 400. The publication becomes FENCED with quarantined=1 and segment 50 QUARANTINED. Reads of other segments are refused too |
| 1 GiB stdout | stopped at 326 MiB by the default 120 s Shell timeout, because backpressure paces the command to the upload rate. The captured 341,966,848 bytes match the generator's prefix (3341ea68…), and the model is told the command timed out. 300 MiB completed exact |
During the 1 GiB run, the ACK returned one 409 managed_runtime_identity_conflict, and the Harness logged that it can be retried. It did not recur with 300 MiB, with a 5 s timeout over 60 MB, or with a 2 s timeout.
R13-2: bucket versioning turned on while running
The constructor refuses a versioned or non-private bucket at startup, so this case only covers versioning being turned on later. I enabled versioning on the bucket and started a new O2 Shell turn:
- The retries:
requireUnversioned()refused with "Tool publication bucket cannot enforce immutable objects". The refusals were retried 153, 555, 549, 573 and 426 times in five consecutive minutes, about 2,256 in all. Each refusal first callsGetBucketVersioningon OSS (AliyunToolPublicationObjectStore.java:28-35). - How the turn ended: after 241 s the execution outcome became unknown.
/finishedreturned 400, thenclose_not_startedreturned 400, and the Session endedrecoveryBlocked. The publication stayedOPEN/FINISHING. - Existing publications: range reads returned 500
internal_error, and the catalog was unchanged.
Failing closed is correct here. The cost is a four-minute stall, a burst of OSS requests and a blocked Session for every Shell turn until someone fixes the bucket.
Suggestion: treat this refusal as a configuration error rather than a retryable one. The capture would then fail at once with a clear error, or end as a correctable not_started.
Efficiency note: every object PUT and GET also asks OSS for the bucket versioning first. That is one extra round trip per 1 MiB segment.
Second host (42e2e8b1)
Setup: the Harness ran on an Orange Pi 6 Plus (linux-arm64, Node 24.14), built from the same commit with pnpm 11.24.0. Spring, MySQL, the workers and the fake model stayed on the Mac, reached over SSH tunnels.
| Scenario (O2 unless noted) | Result |
|---|---|
| Crash after the receipt commit; Mac Harness killed, the Pi Harness recovers | load after 48 s. The command ran once, bytes exact |
| The same, from the Pi to the Mac | load after 55 s. The command ran once, bytes exact |
| Control characters plus a crash after the receipt, Mac to Pi | pass |
| Clean hand-off, Mac → Pi → Mac | 3 turns, 3 publications REFERENCED. Each turn saw its own output, and the history held 1, 2, then 3 tool results |
| Local capture path (#12848), Harness on the Pi | the first Shell call returns 409 execution_unknown and the Session is blocked. A hand-off from the Mac to the Pi blocks the same way, and loading it back on the Mac returns 409 |
The local path's publisher listens on 127.0.0.1 (hosted-shell-publisher.ts:126), so it needs the Harness and the worker on one host. That is main's design, and it fails closed. Cross-host deployments need O2.
Regression on af962df7 (fake OSS)
- O2:
echowriting to stdout and stderr: 2/2.- 100 MiB stdout + 5 MiB stderr with exit 7: exact.
- Crash after the receipt commit: passes.
- One lost segment request: completes.
- One lost admission reply: completes.
- Control characters plus a crash: passes.
- Reload, then a second Shell turn: passes.
- ACK request lost, then reload: passes.
- Receipt commit delayed 12 s: completes.
- Local path:
echo2/2; 100 MiB + 5 MiB with exit 7; reload, then a second Shell turn. - Mixed: 2 O2 Sessions and 2 local Sessions writing 64 MiB each at the same time: 4/4, with 0 deadlocks.
- Earlier fixes:
- R9-1: invalid arguments get correctable refusals on both paths.
- R2-6: a raw ESC completes.
- R2-5: cancelling before start gives a receipt, and reload works on both paths.
- Gate 2: an absolute
read_filepath gets a correctable refusal on both paths.
- Preview matrix and UTF-8 tail fixtures: the same as round 13 with one exception. The local-path case with 58,000 B stdout and short stderr (the F3 band) showed the truncation marker once instead of twice. That case also did so in round 7.
af962df7does not touch the local path, so I read this as run-to-run variance in F3. - Real qwen3.8-max, failing 3,000-line build: O2 2/2 and local 2/2. The build ran once each time, and the error was quoted.
- Unchanged:
- Gate 1:
loadstill returns 409 after 150 s. - Lone surrogates in arguments still block O2 Sessions.
- On the local path they do not settle within the probe's 120 s window.
- Gate 1:
Round 13 already covered the migrations and the bot merge, and af962df7 does not touch them. I did not rerun 1 GiB (it ran on real OSS above), R6-2 load scaling or the author's six-Workspace driver.
Still open (deferred)
- Gate 1: after a crash following
/admissions/prepare, a new Harness'sloadreturns 409 for 150 s. - R7-1 (fix(serve): Stop renewing settled Hosted Shell publication grants #12957) and R12-1 (Follow up on deferred O2 review suggestions from #12894 #12986): deferred.
af962df7does not touch them. - Lone surrogates in tool arguments (fix(managed-agent): Handle unpaired surrogate Shell arguments before dispatch #13010): they still block the Session.
- Also still deferred: fix(core): Handle partial ANSI and UTF-8 prefixes in Shell previews #12969, F3, the B2/B4 MySQL tests, R2-4 and R4-1 (proposal(managed-agent): Recover expired tool publication candidates safely #13019).
Evidence at wenshao/qwen-code@9623cc38:
r14/results/s20-*.logandr14/results/s2-r14-real-sleep-o2.json: R13-1 and R14-1, with the provider request summaries ins20-r14real2-o.log;r14/results/s15-r14lsi.log: the R5-1 message;r14/results/r14-af962df7.log: the regression;r13b/results/ro-*.log,s1-ro-m100.log,s17-ro-m100.log,s19-rov.logandro-bucket-inventory.txt: real OSS.notes.txtexplains the 1 GiB prefix check and a wrong label ins19-rov.log;r13b/results/xh-42e2e8b1.logands18-xh*.log: the second host;r13b/harness/: the scripts, including the OSS helper and the remote Harness patch.
中文版
真实环境验证第 14 轮 — R13-1 修复、真实阿里云 OSS 与第二台宿主机(head af962df7)
接续第 13 轮和作者的跟进提交 af962df7。本轮分两部分:
af962df7: 只改了hosted-workspace-tool-turn.ts的 O2 未启动分支(7 行,外加一个测试)。我重新打了 bundle,在真实栈上验证 R13-1,并跑了一轮有针对性的回归。Java 侧没有变化,jar 沿用第 13 轮的。- 真实阿里云 OSS 和第二台宿主机(按维护者要求):这部分在
af962df7推送之前、在42e2e8b1上跑完。该提交不涉及发布、对象存储和恢复代码,所以这些结论可以沿用。
结论:
- R13-1: 已修复。O2 下模型现在能看到 Shell 调用未启动的原因。第 13 轮那个任务用真实 qwen3.8-max 重跑,3 次调用、23–28 s 完成(3/3),不再是 11 次盲目的 Shell 调用。
- 真实 OSS: O2 路径在真实 bucket 上成立,字节校验从 bucket 下载回来的每个对象都一致。新问题 R13-2(中,需要 bucket 配置错误才会触发): 服务运行中打开 bucket 版本控制后,每个 Shell 轮次会卡 4 分钟,然后阻塞 Session。
- 第二台宿主机: O2 的崩溃恢复和 Session 交接在 macOS 与 linux-arm64 之间双向可用。
- 新问题 R14-1(低): 新加的长原因首尾截断可能把 emoji 切成孤立代理项。这套栈上没有出故障,修起来很小。
af962df7上的回归: 干净,与第 13 轮一致。- 总体: profile 保持关闭时,没有阻塞合并的问题。在 profile 接入真实 bucket 之前,值得先修 R13-2。
R13-1 修复(af962df7)
| 本地路径 | O2 @ 42e2e8b1 |
O2 @ af962df7 |
|
|---|---|---|---|
sleep 3; echo slept |
Blocked: sleep 3 followed by: echo slept. …(291 字符) |
Runtime Shell did not start. |
与本地路径相同的 291 字符 |
| install 窗口内所有权被拒(R5-1) | Workspace execution was refused before dispatch. |
Runtime Shell did not start. |
Workspace execution was refused before dispatch. |
| 真实 qwen3.8-max,"等 3 秒,然后用 date 命令打印当前日期,用一条 shell 命令完成" | 2/2,3 次调用 | 2/2,每次 11 次调用,105–114 s | 3/3,3 次调用,23–28 s |
- 修复后的真实模型: 每次都是先
sleep 3 && date(被拦),再sleep 3 # intentional-sleep: …,再date;回答里都说明了运行时要求拆成两次调用。 - 持久化: 原因随 Session 存储。detach 后在一个新的 Harness 进程上 load,回放的历史文本一致。
- 长度上限: 200 KB 的命令在两条路径上都会更早被拒,提示
Hosted tool input exceeds the inline Session Store limit.(轮次报错,Session 不阻塞)。所以约 5 KB 的原因是走到新的超长分支(超过 4096 字符)的现实途径。
R14-1(低):首尾截断可能切断代理对
af962df7 对超过 4096 个 UTF-16 单元的原因保留 errorMessage.slice(0, 2048) 和 errorMessage.slice(-2048)。两处都按代码单元计数,所以恰好落在截断点上的 emoji(或其他 BMP 以外的字符)会丢掉代理对的一半。Shell 工具会把命令的剩余部分写进原因,因此一条带 emoji 的长命令就能触发:
| 原因(O2,假模型) | 模型收到的内容 |
|---|---|
| 约 4.9K ASCII | 4123 个单元,带截断标记,无孤立代理项 |
| emoji 的高位代理落在第 2047 位 | 第 2047 位孤立 U+D83D |
| emoji 落在第 2046 位 | 代理对完整保留 |
| 全部是 emoji | 第 2047 位孤立 U+D83D,第 2075 位(尾段开头)孤立 U+DE00 |
- 后续表现: 这套栈上没有出故障。
- 每个轮次都完成了,没有 Session 被阻塞。
- Session Store 接受了这条消息,在新 Harness 进程上回放也一致。
- 我又把带孤立代理项的两个 Session(外加 ASCII 那个作对照)加载到接真实 qwen3.8-max 的 Harness 上。抓到的请求体里分别带 1 个和 2 个转义的孤立代理项(对照为 0),DashScope 返回 200,模型也正确解释了被拦原因。
- 为什么仍值得修:
- 工具参数里的孤立代理项已知会阻塞 O2 Session(fix(managed-agent): Handle unpaired surrogate Shell arguments before dispatch #13010);这里是 Harness 自己从合法输入里制造出来的。
- 更严格的 provider 或 JSON 消费方可能会拒绝。我没有测这类对象。
- 测试覆盖: 没有测试走到这个分支。截断标记
[... error truncated ...]只出现在源码里,新测试用的原因只有 42 个字符。 - 建议: 按码点边界截断。
managed-runtime-provider-protocol.ts:572的cutProviderFitSlot已经这样做了:截断点会切断代理对时就后退一位。再补一个超过 4096 单元、截断点上带代理对的测试。
作为对照,本地路径对同样的全 emoji 原因截到 530 个单元,末尾是 U+FFFD。那是 main 的代码,不属于本 PR。
真实阿里云 OSS(42e2e8b1)
环境:
- bucket: cn-hangzhou 的临时 bucket,私有、未开版本控制,带 1 天的生命周期规则。
- 清理: 结束后删除了全部 1,599 个对象(1.59 GB)、所有版本和 bucket 本身。
- 对象存储: Spring 使用真实的
AliyunToolPublicationObjectStore,信任库是 JDK 默认的。 - 凭据: 来自本地文件,没有出现在任何公开文件里。
- 校验方式: 每次做输出字节校验的运行结束后,把每个存储对象从 bucket 下载回来,按 SHA-256 与目录和生成器的基准比对。
| 场景 | 结果 |
|---|---|
echo |
通过。2 个对象,21 B;小资源仍内联在 SQL 里 |
| 100 MiB stdout + 5 MiB stderr,exit 7 | 105 个分段对象下载比对一致;每次范围读都一致;本轮 76 s(上传约 12–14 Mbps) |
| receipt 提交后崩溃、分段请求丢一次、admission 响应丢一次、控制字符加崩溃、reload 后第二轮 | 通过,字节一致 |
| 2 个 O2 Session 同时各写 64 MiB | 一致,0 死锁 |
| 预览矩阵 | 与 OSS 替身的结果相同 |
| 真实 qwen3.8-max,3,000 行失败构建 | 2/2,都引用了报错 |
| 匿名 GET 分段对象 | 403 |
| 在第 50 段存储对象里翻转一个比特 | 读返回 400。发布变为 FENCED、quarantined=1,第 50 段为 QUARANTINED;其他分段的读取也被拒绝 |
| 1 GiB stdout | 在 326 MiB 处被默认 120 s Shell 超时终止,因为背压把命令节奏压到了上传速率。已捕获的 341,966,848 字节与生成器的前缀一致(3341ea68…),模型被告知命令超时。300 MiB 完整且一致 |
1 GiB 这次运行中 ACK 返回过一次 409 managed_runtime_identity_conflict,Harness 记录为可重试。用 300 MiB、60 MB 加 5 s 超时、2 s 超时都没有复现。
R13-2:运行中打开 bucket 版本控制
构造函数在启动时就会拒绝已开版本控制或非私有的 bucket,所以本场景只覆盖启动之后才打开版本控制的情况。我在 bucket 上打开版本控制,然后新开一个 O2 Shell 轮次:
- 重试:
requireUnversioned()以 "Tool publication bucket cannot enforce immutable objects" 拒绝,拒绝后被不断重试,连续五分钟里每分钟分别是 153、555、549、573、426 次,共约 2,256 次。每次拒绝前都先向 OSS 调一次GetBucketVersioning(AliyunToolPublicationObjectStore.java:28-35)。 - 轮次如何结束: 241 s 后执行结果变为 unknown。
/finished返回 400,接着close_not_started返回 400,Session 最终recoveryBlocked;发布停在OPEN/FINISHING。 - 已有发布: 范围读返回 500
internal_error,目录不变。
这里失败即关闭是对的。代价是在有人修好 bucket 之前,每个 Shell 轮次都要卡 4 分钟、打出一大批 OSS 请求,并阻塞 Session。
建议: 把这种拒绝当作配置错误而不是可重试错误。这样捕获会立即失败并给出清楚的错误,或者以可纠正的 not_started 结束。
效率备注: 每次对象 PUT 和 GET 之前都先向 OSS 查询一次 bucket 版本控制状态,即每个 1 MiB 分段多一次往返。
第二台宿主机(42e2e8b1)
环境: Harness 跑在 Orange Pi 6 Plus 上(linux-arm64,Node 24.14),用同一提交、pnpm 11.24.0 构建。Spring、MySQL、worker 和假模型留在 Mac 上,通过 SSH 隧道访问。
| 场景(未注明即 O2) | 结果 |
|---|---|
| receipt 提交后崩溃:杀掉 Mac 上的 Harness,由 Pi 上的 Harness 恢复 | 48 s 后 load 成功;命令只跑了一次,字节一致 |
| 同上,从 Pi 到 Mac | 55 s 后 load 成功;命令只跑了一次,字节一致 |
| 控制字符加 receipt 后崩溃,Mac → Pi | 通过 |
| 正常交接 Mac → Pi → Mac | 3 轮,3 个发布都是 REFERENCED;每轮都看到自己的输出,历史里的工具结果依次为 1、2、3 个 |
| 本地捕获路径(#12848),Harness 在 Pi 上 | 第一次 Shell 调用返回 409 execution_unknown,Session 被阻塞。从 Mac 交接到 Pi 也同样阻塞,再回 Mac load 返回 409 |
本地路径的 publisher 监听在 127.0.0.1(hosted-shell-publisher.ts:126),所以要求 Harness 和 worker 在同一台机器上。这是 main 的设计,并且是失败即关闭。跨宿主机部署需要 O2。
af962df7 上的回归(fake OSS)
- O2:
echo同时写 stdout 和 stderr:2/2。- 100 MiB stdout + 5 MiB stderr、exit 7:一致。
- receipt 提交后崩溃:通过。
- 分段请求丢一次:完成。
- admission 响应丢一次:完成。
- 控制字符加崩溃:通过。
- reload 后第二个 Shell 轮次:通过。
- ACK 请求丢失后 reload:通过。
- receipt 提交延迟 12 s:完成。
- 本地路径:
echo2/2;100 MiB + 5 MiB、exit 7;reload 后第二个 Shell 轮次。 - 混合: 2 个 O2 Session 和 2 个本地 Session 同时各写 64 MiB:4/4,0 死锁。
- 之前的修复:
- R9-1:非法参数在两条路径上都得到可纠正的拒绝。
- R2-6:原始 ESC 能完成。
- R2-5:启动前取消得到 receipt,两条路径 reload 都正常。
- Gate 2:绝对路径的
read_file在两条路径上都得到可纠正的拒绝。
- 预览矩阵和 UTF-8 尾部夹具: 与第 13 轮一致,只有一个例外:本地路径上 58,000 B stdout 加短 stderr 的用例(F3 区间)截断标记出现 1 次而不是 2 次。这个用例第 7 轮也出现过同样情况。
af962df7不涉及本地路径,所以我判断为 F3 的逐次波动。 - 真实 qwen3.8-max,3,000 行失败构建: O2 2/2、本地 2/2。每次构建只跑一次,报错都被引用。
- 未变:
- Gate 1:
load在 150 s 后仍返回 409。 - 参数里的孤立代理项仍会阻塞 O2 Session。
- 在本地路径上它们在探针的 120 s 窗口内不会结束。
- Gate 1:
迁移和机器人合并第 13 轮已经覆盖,af962df7 也没有涉及。本轮没有重跑 1 GiB(上面在真实 OSS 上跑过)、R6-2 加载规模和作者的六 Workspace 驱动。
仍未解决(已延后)
- Gate 1:
/admissions/prepare之后崩溃,新 Harness 的load在 150 s 内返回 409。 - R7-1(fix(serve): Stop renewing settled Hosted Shell publication grants #12957) 与 R12-1(Follow up on deferred O2 review suggestions from #12894 #12986): 已延后,
af962df7不涉及。 - 工具参数里的孤立代理项(fix(managed-agent): Handle unpaired surrogate Shell arguments before dispatch #13010): 仍会阻塞 Session。
- 其他仍延后的: fix(core): Handle partial ANSI and UTF-8 prefixes in Shell previews #12969、F3、B2/B4 MySQL 测试、R2-4 和 R4-1(proposal(managed-agent): Recover expired tool publication candidates safely #13019)。
证据见 wenshao/qwen-code@9623cc38(清单同英文版)。
Review fixes and main synchronization
Final verification on The MySQL races use two store instances and an immutable object-store fixture. They prove original slot/key/ref/digest/deadline/quota preservation and refusal of late old-epoch installation for segment and terminal variants. This is not real OSS or cross-host evidence; those acceptance requirements remain tracked in #13019 and the O2 rollout verification. Recovery does not reset FINISHING or replace invalid/corrupt results, and does not re-execute Shell. 中文:已采纳并修复两条 Critical,同时解决与主线审批接线的冲突。最终构建、类型检查、bundle、CLI/core/Java 定向回归和真实 MySQL 恢复测试通过;完成两轮无新问题的自审。真实 OSS/跨宿主证据仍需补齐,未把测试替身计作该类验收。 New round-14 feedbackI checked round 14 against this head. Both Suggestions are recorded in #12986 under the repository's review-round cap:
The reviewer's older-head real OSS/two-host tests are useful evidence; they do not verify the newly added expiry recovery protocol. That remaining evidence stays in #13019. Both original inline Critical threads are now resolved (2/2), with zero unresolved review threads. The PR is mergeable; new-head CI and renewed maintainer review are pending. One Qwen Autofix routing run was cancelled before any steps/logs, consistent with its per-PR cancel-in-progress routing; no code failure was observed and no rerun was requested. 中文:新增两项非阻塞建议已在 #12986 记录后续验收,原两条 Critical 已修复并解决。PR 无冲突,等待新 head 的 CI 和维护者复审。 Triage re-run follow-up[codex] The current-head static re-run independently confirms R2-4 and R4-1 are fixed. The PR description is corrected in English and Chinese to V20/V21/V22 and links the current design. It distinguishes the earlier-head real OSS/two-host report from the new recovery protocol's remaining evidence in #13019.
Current live GraphQL returns all 43 review threads, 中文:本轮独立核验确认两条 Critical 已修复;双语 PR 描述及设计链接已更新,非阻塞的响应体清理与 CI 数据库路由已记入 #12986。实时 API 确认 43/43 线程已解决。仍等待维护者复审及当前沙箱验证,不将旧 head 的 OSS 证据计作新增恢复协议验收。 |
|
@qwen-code /triage |
Real-stack verification, round 15 — R4-1 and R2-4 fixes, with real Aliyun OSS and a second host (head
|
| Case | before (af962df7) |
after (0027f7a7) |
candidate (below) |
|---|---|---|---|
| F2: segment PUT takes 9 s once | blocked | blocked, 400 |
recovered, exact |
| F3: one finish read-back GET takes 9 s | FINISHING, blocked |
FINISHING, blocked |
recovered, exact |
| F4 / F5: the same 8 times in a row | blocked | blocked | 3 recoveries, then blocked |
| F6: PUT takes 4.5 s (only the 3 s claim lapses) | blocked | blocked | blocked |
| F7: F2, and the producer loses the reply | blocked | recovered, exact | recovered, exact |
| F8: F3, and the producer loses the reply | blocked | recovered, FINISHED, exact |
recovered, exact |
Relation to #13019: its acceptance scenario is the one where the first PUT response is lost, the deadline expires and the command is not repeated. That now passes on the real stack, including real OSS (F7/F8 here and RF7/RF8 below). R15-1 is the other case, where the response does arrive, as a 400.
Why:
- Server: the in-flight checks throw
IllegalArgumentException("… claim expired"), which becomes400 invalid_request. They are ininstall()(ToolPublicationDataStore.java:1127),installFinish()(:660), the heartbeat (:952) andcheckScanClaim()(:837); the JDI tap showsinstall:1125. Only the start-of-request checks were changed torequireUnexpired(409 managed_tool_publication_operation_expired). - Producer:
remote-shell-result-publication.ts:251-259rethrows any 4xx other than that code unless the observed status isSUCCEEDED. The status at that point is alreadyEXPIRED, but the recovery loop at:272is never reached. - Test coverage: the new Java test holds the first write, moves the deadline back and checks a second store instance. It never asserts what the held request itself answers.
Real Aliyun OSS:
- Bucket: a temporary private bucket, deleted afterwards (367 objects, 372 MB).
- Fault injection: the OSS SDK ignores
https.proxyHost. So the JVM resolved the OSS host names to 127.0.0.1, and a TCP forwarder relayed the TLS bytes unchanged to the real OSS address, holding one chunk for 9 s when armed.
| Case | Result (after) |
|---|---|
| RF2: segment PUT body held 9 s | 400, blocked |
| RF7: RF2, and the producer loses the reply | recovered. The delayed first PUT had already created the object before the retry; the bytes read back exact |
| RF3: finish read-back body held 9 s | 400, FINISHING, blocked |
| RF8: RF3, and the producer loses the reply | recovered, FINISHED, bytes exact |
| RT1: 100 MiB + 5 MiB at 120 s / 30 s, no fault | complete. The finish took 15.4 s; segment requests had a median of 255 ms and a slowest of 5.2 s |
| RT2: the same at 10 s / 5 s, no fault | the stdout seal re-read 100 MiB in 10.5 s, got 400, and the Session was blocked with an incomplete capture |
| RT3: RT2 with the candidate | the seal took 8.0 s. The finish then got 400 after about 10.3 s four times (3 recoveries) and stayed FINISHING, blocked |
What the real-OSS runs show:
- The seal and the finish both re-read the whole stream from OSS.
- From this Mac, 110 MB took 15 s. At the rig's 120 s timeout, a capture of roughly 800 MiB would outlive the finish at that bandwidth, and the rig's execution limit is 2 GiB. In-region bandwidth is higher, so in production the likelier trigger is a stall longer than the remaining deadline.
- Recovery restarts that read from the beginning. So when the read-back itself is longer than the timeout, recovery cannot finish it (RT3).
Candidate (tested, not proposed as the whole fix):
-
The change: one condition in the producer, so that an observed
EXPIREDafter a 4xx goes to the recovery loop:failure.code !== 'managed_tool_publication_operation_expired' && - status['state'] !== 'SUCCEEDED' + status['state'] !== 'SUCCEEDED' && + status['state'] !== 'EXPIRED'
-
What it fixes: single stalls (F2, F3), while staying bounded (F4, F5).
-
What it does not fix:
- A lapse of the claim alone (F6). That needs a retryable code from the server when the deadline has not passed. I did not let the producer retry
RETRYABLEafter a 4xx, because a deterministic refusal would then loop until the 30-minute client deadline. - A seal or finish whose read-back is longer than the timeout (RT3). That needs either verification progress that survives a recovery, or a verification window sized to the bytes.
- A lapse of the claim alone (F6). That needs a retryable code from the server when the deadline has not passed. I did not let the producer retry
-
Suggestion: return
409 managed_tool_publication_operation_expiredfrom the in-flight checks once the deadline has passed, and a distinct retryable code when only the claim lapsed. Also add a test that asserts what the held request answers.
R2-4: resume classification
Setup: Harness A is killed after the receipt commit. Harness B loads the Session once A's writer grant has lapsed (65 s).
| Case | before | after |
|---|---|---|
| The Broker's reply to the resume's acquire is lost | load 200, then recoveryBlocked; prompts get 409 hosted_turn_recovery_required, and a third Harness is needed |
load 503 managed_session_open_failed with the Session not attached; the load 5 s later returns 200 and the turn completes. The command ran once |
| The same, with B on the second host | — | 503, then 200, completes |
| The Workspace lease is handed to another holder first | 200, completes |
200, completes |
- Why the busy case does not trigger: the Broker re-acquires the existing, still-live runtime session without calling
WorkspaceExecutionStore.claim(). That is why the tampered lease is never checked, and why a definite409 workspace_busyduring a resume cannot be produced here. - Releasing A's runtime session by hand only leads to
404 runtime_session_not_foundon the recovery ACK (next section). - Load latency:
loadnow waits for the resume's acquisition before answering.
Pre-existing: Broker restart with an unacknowledged execution
I restarted Spring (the Broker) between A's crash and B's load:
- O2 before: load returns
200, the recovery ACK gets404 runtime_session_not_found, and the Session is blocked. - O2 after: every load takes about 30 s and returns
503(9 loads in 4 minutes). - Local capture path: load returns
409 hosted_turn_recovery_requiredevery time. - The Workspace: in all three, the lease still named the dead runtime session 12 minutes later, and a new Session in that Workspace got
409 workspace_busy.
This is not caused by e177e1da. On O2, though, the receipt is already committed, so a definite 404 on the recovery ACK could count as settled rather than blocking the Session.
Approvals merged into O2 (0027f7a7)
With approvalMode: default, each run_shell_command asks first:
- allow: the command runs once, and the publication is
REFERENCED. - deny: the model sees "The Session owner denied this tool call, so it was not run." There is no Broker call and no publication row.
- Two Shell calls in one reply, allow then deny: the first runs, and the second is refused against its own call id. This holds on O2 and on the local path.
- A 3 s timeout with no answer: the call is refused as expired and not run.
- Harness killed while an approval is pending: load returns
409 hosted_turn_recovery_required. The same happens on the local path, so this is main's D6a behavior.
Second host (0027f7a7 on the Orange Pi)
The Harness ran on linux-arm64 with Node 24; everything else stayed on the Mac:
- Crash after the receipt commit, Mac → Pi and Pi → Mac: load after 56–57 s. The command ran once and the bytes were exact.
- Clean hand-off Mac → Pi → Mac: 3 publications
REFERENCED, and the history held 1, 2, then 3 tool results. - Lost acquire reply during a resume, with B on the Pi:
503, then200, and the turn completes. - Approvals (allow; allow + deny): correct.
Regression on 0027f7a7 (fake OSS, fresh schema)
- O2:
echowriting to stdout and stderr: 2/2.- 100 MiB stdout + 5 MiB stderr with exit 7: exact.
- Crash after the receipt commit: passes.
- One lost segment request: completes.
- One lost admission reply: completes.
- Control characters plus a crash: passes.
- Reload, then a second Shell turn: passes.
- ACK request lost, then reload: passes.
- Receipt commit delayed 12 s: completes.
- Local path:
echo2/2; 100 MiB + 5 MiB with exit 7; reload, then a second Shell turn. - Mixed: 2 O2 Sessions and 2 local Sessions writing 64 MiB each at the same time: 4/4, with 0 deadlocks.
- Earlier fixes: R9-1, R2-6, R2-5 and Gate 2 behave as in round 14.
- Preview matrix and UTF-8 tail fixtures: the same as round 13. The F3-band case that varied in round 14 showed two markers again.
- Real qwen3.8-max, failing 3,000-line build: O2 2/2 and local 2/2. The build ran once each time, and the error was quoted.
- Unchanged:
- Gate 1:
loadstill returns 409 after 150 s. - Lone surrogates in arguments still block O2 Sessions.
- Gate 1:
Not rerun: 1 GiB, R6-2 load scaling and the author's six-Workspace driver. The unstarted-reason code (R13-1, R14-1) is unchanged since round 14.
Merge coordination
Other in-flight PRs add Flyway migrations with the same numbers:
- Stacked on this PR: feat(managed-agent): project durable tool results to WebShell #13037 (V21), feat(managed-agent): Protect Session-owned tool output retirement #13084 (V22) and feat(managed-agent): Collect retired tool output safely #13087 (V23) still carry the old O2 numbers (V19/V20). They have not taken
0027f7a7, so they will collide with this PR's V21/V22 when they do. - Based on main: feat(managed-agent): add verified W1a workspace cold restore #13088 and feat(managed-agent): Implement private Hosted MCP runtime (H1) #12946 each add V21. Whichever lands first forces the others to renumber.
Git does not report these clashes; Flyway fails at startup.
Still open (deferred)
- R7-1 (fix(serve): Stop renewing settled Hosted Shell publication grants #12957) and R12-1 (Follow up on deferred O2 review suggestions from #12894 #12986): deferred.
- Lone surrogates in tool arguments (fix(managed-agent): Handle unpaired surrogate Shell arguments before dispatch #13010): they still block the Session.
- R14-1: the unstarted-reason cut can split a surrogate pair. Unchanged.
- R13-2: versioning turned on while the service runs. Not retested; the object store code is unchanged.
- Also still deferred: fix(core): Handle partial ANSI and UTF-8 prefixes in Shell previews #12969, F3 and the B2/B4 MySQL tests.
Evidence at wenshao/qwen-code@5871c7d5:
results/notes.txtmaps every file to its run;results/r15-r41*.log,jdi-claim-expired-excerpt.txtandcandidate-client-change.diff: R15-1 on the fake OSS;results/ro15*.logandro15-bucket-inventory.txt: real OSS;results/r15-r24*.log: R2-4 and the Broker restarts;results/s22-*.log: approvals;results/xh-0027f7a7.log: the second host;results/flyway-checks-r15.txt: migrations;results/r15-0027f7a7.log: the regression;harness/: the scripts, including the fault proxy, the fake OSS and the TCP forwarder.
中文版
真实环境验证第 15 轮 — R4-1 与 R2-4 修复,含真实阿里云 OSS 与第二台宿主机(head 0027f7a7)
接续第 14 轮和作者的评审修复与 main 同步:
e177e1da: R4-1 的修复为过期的发布操作增加了恢复;R2-4 的修复对 Session 续跑时的 Workspace 拒绝做了分类。0027f7a7: 一次 main 合并,带入了 D6a 审批。
作者说明 R4-1 还缺真实 OSS 和跨宿主证据,本轮两者都覆盖了。我为 0027f7a7 重新打了 bundle 和 jar,同时构建了 af962df7(这两个提交之前的 head)用于 A/B 对照。
结论:
- 迁移: V22 在全新库、main 的 V19 库和第 14 轮的 V21 库上都能应用。
- R4-1:只有生产端丢失应答时才修好了。 新的
/recover路径在 fake OSS 和真实 OSS 上都能用:字节一致,配额不变。- 新问题 R15-1(对启用 O2 而言属于高): 操作在执行过程中超过期限时,服务端返回
400 invalid_request。生产端把它当成最终结果,从不调用/recover,Session 与修复前一样被阻塞。 - 被阻塞的 Session 还会一直占着 Workspace,那里的其他 Session 会得到
409 workspace_busy;80 分钟后我仍看到这种情况。 - 在真实 OSS 上,只要流 seal 或 finish 超过操作超时,不注入任何故障也会出现,因为两者都要把整段 capture 重新读一遍。
- 新问题 R15-1(对启用 O2 而言属于高): 操作在执行过程中超过期限时,服务端返回
- R2-4: Workspace 获取失败的续跑不再挂上一个被阻塞的 Session,下一次 load 就能续完这一轮。R2-4 的确切触发条件(续跑时明确的
409 workspace_busy)在这套栈上无法触发,这个映射仍以单元测试为证据。 - 审批合入 O2: 在 O2 上工作正常。被拒绝或过期的调用什么都不执行、不预留、也不调用 Broker。
- 第二台宿主机(linux-arm64): 全部通过。
- 本轮发现的既有问题: Broker 重启后,最后一次 Shell 调用未确认的 Session 在两条路径上都无法续跑;它残留的 Workspace lease 还会阻塞其他 Session。这在
e177e1da之前就存在。 - 回归: 干净,与第 14 轮一致。
- 总体: profile 保持关闭时,没有阻塞合并的问题。启用 O2 之前需要修复 R15-1,因为 R4-1 只在丢失应答的情况下被解决。
R15-1:执行中超过期限的操作仍会卡死 capture
环境:
- fake OSS,
operation-timeout 6s、claim-timeout 3s。 - 故障代理在指定的发布请求经过时,向对象存储布置一个故障,例如"
segments/stdout/2的对象 PUT 耗时 9 s"。 - "阻塞"指 Session 最终处于
recoveryBlocked。
| 场景 | 修复前(af962df7) |
修复后(0027f7a7) |
候选改动(见下) |
|---|---|---|---|
| F2:分段 PUT 一次耗时 9 s | 阻塞 | 阻塞,400 |
恢复,一致 |
| F3:finish 读回的一次 GET 耗时 9 s | FINISHING,阻塞 |
FINISHING,阻塞 |
恢复,一致 |
| F4 / F5:同样的故障连续 8 次 | 阻塞 | 阻塞 | 3 次恢复后阻塞 |
| F6:PUT 耗时 4.5 s(只有 3 s 的 claim 过期) | 阻塞 | 阻塞 | 阻塞 |
| F7:F2 且生产端丢失应答 | 阻塞 | 恢复,一致 | 恢复,一致 |
| F8:F3 且生产端丢失应答 | 阻塞 | 恢复,FINISHED,一致 |
恢复,一致 |
与 #13019 的关系: 它的验收场景是第一个 PUT 的应答丢失、期限过期、命令不被重复执行。这个场景现在在真实栈上通过了,包括真实 OSS(这里的 F7/F8 和下面的 RF7/RF8)。R15-1 是另一种情况:应答确实回来了,只是一个 400。
原因:
- 服务端: 执行中的检查抛出
IllegalArgumentException("… claim expired"),对外变成400 invalid_request。这些检查位于install()(ToolPublicationDataStore.java:1127)、installFinish()(:660)、heartbeat(:952)和checkScanClaim()(:837);JDI 探针显示抛出点在install:1125。只有请求开始时的检查改成了requireUnexpired(409 managed_tool_publication_operation_expired)。 - 生产端:
remote-shell-result-publication.ts:251-259对这个码以外的 4xx 一律重新抛出,除非查到的状态是SUCCEEDED。此时状态其实已经是EXPIRED,但:272的恢复循环永远走不到。 - 测试覆盖: 新的 Java 测试扣住第一次写入、把期限改到过去,然后检查第二个 store 实例;它没有断言被扣住的那个请求本身返回什么。
真实阿里云 OSS:
- bucket: 临时私有 bucket,结束后已删除(367 个对象、372 MB)。
- 故障注入: OSS SDK 不读取
https.proxyHost。所以让 JVM 把 OSS 域名解析到 127.0.0.1,由一个 TCP 转发器把 TLS 字节原样转给真实 OSS 地址,布置后扣住一个数据块 9 s。
| 场景 | 结果(修复后) |
|---|---|
| RF2:分段 PUT 请求体被扣 9 s | 400,阻塞 |
| RF7:RF2 且生产端丢失应答 | 恢复。重试之前,被延迟的第一次 PUT 已经创建了对象;读回的字节一致 |
| RF3:finish 读回的响应体被扣 9 s | 400,FINISHING,阻塞 |
| RF8:RF3 且生产端丢失应答 | 恢复,FINISHED,字节一致 |
| RT1:100 MiB + 5 MiB,120 s / 30 s,无故障 | 完成。finish 用时 15.4 s;分段请求中位数 255 ms,最慢 5.2 s |
| RT2:同上但 10 s / 5 s,无故障 | stdout seal 用 10.5 s 重读 100 MiB,得到 400,Session 因 capture 不完整被阻塞 |
| RT3:RT2 加候选改动 | seal 用时 8.0 s。之后 finish 四次都在约 10.3 s 后得到 400(3 次恢复),停在 FINISHING,阻塞 |
真实 OSS 的结果说明:
- seal 和 finish 都会把整条流从 OSS 重新读一遍。
- 从这台 Mac 读 110 MB 用了 15 s。按装置的 120 s 超时和这个带宽,约 800 MiB 的 capture 就会让 finish 超过期限,而装置配置的执行上限是 2 GiB。同区域带宽更高,所以生产环境里更可能的触发条件是一次停顿超过了剩余期限。
- 恢复会从头开始读,所以当读回本身就比超时长时,恢复也完成不了(RT3)。
候选改动(已实测,但不是完整修复):
- 改动: 生产端的一个条件,让 4xx 之后查到的
EXPIRED进入恢复循环。 - 能修好的: 单次停顿(F2、F3),并且有上限(F4、F5)。
- 修不好的:
- 只有 claim 过期(F6)。这需要服务端在期限未到时返回一个可重试的码。我没有让生产端在 4xx 之后重试
RETRYABLE,因为遇到确定性的拒绝时会一直循环到客户端 30 分钟的总期限。 - 读回时间长于超时的 seal 或 finish(RT3)。这需要让校验进度在恢复后保留下来,或者让校验窗口随字节数调整。
- 只有 claim 过期(F6)。这需要服务端在期限未到时返回一个可重试的码。我没有让生产端在 4xx 之后重试
- 建议: 期限已过时,执行中的检查返回
409 managed_tool_publication_operation_expired;只有 claim 过期时返回另一个可重试的码。再补一个断言"被扣住的请求本身返回什么"的测试。
R2-4:续跑时的分类
环境: Harness A 在 receipt 提交后被杀;A 的写者授权过期后(65 s),Harness B 加载这个 Session。
| 场景 | 修复前 | 修复后 |
|---|---|---|
| Broker 对续跑 acquire 的应答丢失 | load 200,随后 recoveryBlocked;prompt 得到 409 hosted_turn_recovery_required,需要第三个 Harness |
load 返回 503 managed_session_open_failed,Session 未挂上;5 s 后的 load 返回 200,轮次完成。命令只执行了一次 |
| 同上,B 在第二台宿主机上 | — | 503,然后 200,完成 |
| 先把 Workspace lease 交给另一个持有者 | 200,完成 |
200,完成 |
- 为什么触发不了 busy: Broker 重新获取仍然存活的 runtime session 时,不会调用
WorkspaceExecutionStore.claim()。所以被篡改的 lease 从未被检查,续跑时也就产生不了明确的409 workspace_busy。 - 手动释放 A 的 runtime session: 只会让恢复 ACK 得到
404 runtime_session_not_found(见下一节)。 - load 延迟: load 现在要等续跑的 acquire 完成后才返回。
既有问题:Broker 重启时有未确认的执行
我在 A 崩溃和 B load 之间重启了 Spring(即 Broker):
- O2 修复前: load 返回
200,恢复 ACK 得到404 runtime_session_not_found,Session 被阻塞。 - O2 修复后: 每次 load 约 30 s 后返回
503(4 分钟内 9 次)。 - 本地捕获路径: load 每次都返回
409 hosted_turn_recovery_required。 - Workspace: 三种情况下 lease 12 分钟后仍指向已死的 runtime session,同一 Workspace 的新 Session 得到
409 workspace_busy。
这不是 e177e1da 引起的。不过在 O2 上 receipt 已经提交了,所以恢复 ACK 收到明确的 404 时,可以当作已结算,而不是阻塞 Session。
审批合入 O2(0027f7a7)
在 approvalMode: default 下,每次 run_shell_command 都会先询问:
- 允许: 命令执行一次,发布为
REFERENCED。 - 拒绝: 模型看到 "The Session owner denied this tool call, so it was not run."。没有 Broker 调用,也没有发布行。
- 一条回复里两个 Shell 调用,先允许后拒绝: 第一个执行,第二个按自己的 call id 被拒。O2 和本地路径都成立。
- 3 s 超时无人应答: 该调用以过期为由被拒,不执行。
- 审批待定时杀掉 Harness: load 返回
409 hosted_turn_recovery_required。本地路径也一样,所以这是 main 的 D6a 行为。
第二台宿主机(Orange Pi 上的 0027f7a7)
Harness 跑在 linux-arm64、Node 24 上,其余都在 Mac 上:
- receipt 提交后崩溃,Mac → Pi 与 Pi → Mac: 56–57 s 后 load 成功。命令只执行了一次,字节一致。
- 正常交接 Mac → Pi → Mac: 3 个发布为
REFERENCED,历史中的工具结果依次为 1、2、3 个。 - 续跑时 acquire 应答丢失,B 在 Pi 上:
503,然后200,轮次完成。 - 审批(允许;允许 + 拒绝): 正确。
0027f7a7 上的回归(fake OSS,全新库)
- O2:
echo同时写 stdout 和 stderr:2/2。- 100 MiB stdout + 5 MiB stderr、exit 7:一致。
- receipt 提交后崩溃:通过。
- 分段请求丢一次:完成。
- admission 应答丢一次:完成。
- 控制字符加崩溃:通过。
- reload 后第二个 Shell 轮次:通过。
- ACK 请求丢失后 reload:通过。
- receipt 提交延迟 12 s:完成。
- 本地路径:
echo2/2;100 MiB + 5 MiB、exit 7;reload 后第二个 Shell 轮次。 - 混合: 2 个 O2 Session 与 2 个本地 Session 同时各写 64 MiB:4/4,0 死锁。
- 之前的修复: R9-1、R2-6、R2-5 和 Gate 2 与第 14 轮表现相同。
- 预览矩阵和 UTF-8 尾部夹具: 与第 13 轮相同。第 14 轮波动过的 F3 区间用例,这次截断标记又是两个。
- 真实 qwen3.8-max,3,000 行失败构建: O2 2/2、本地 2/2。每次构建只运行一次,报错都被引用。
- 未变:
- Gate 1:
load在 150 s 后仍返回 409。 - 参数里的孤立代理项仍会阻塞 O2 Session。
- Gate 1:
未重跑:1 GiB、R6-2 加载规模和作者的六 Workspace 驱动。未启动原因相关代码(R13-1、R14-1)自第 14 轮以来没有变化。
合并协调
其他在飞 PR 加了编号相同的 Flyway 迁移:
- 叠在本 PR 上的: feat(managed-agent): project durable tool results to WebShell #13037(V21)、feat(managed-agent): Protect Session-owned tool output retirement #13084(V22)和 feat(managed-agent): Collect retired tool output safely #13087(V23)仍沿用旧的 O2 编号(V19/V20)。它们还没有合入
0027f7a7,合入时会与本 PR 的 V21/V22 冲突。 - 基于 main 的: feat(managed-agent): add verified W1a workspace cold restore #13088 和 feat(managed-agent): Implement private Hosted MCP runtime (H1) #12946 都加了 V21。谁先合入,其他的就要改号。
Git 不会报告这种冲突;Flyway 会在启动时失败。
仍未解决(已延后)
- R7-1(fix(serve): Stop renewing settled Hosted Shell publication grants #12957) 和 R12-1(Follow up on deferred O2 review suggestions from #12894 #12986): 已延后。
- 工具参数里的孤立代理项(fix(managed-agent): Handle unpaired surrogate Shell arguments before dispatch #13010): 仍会阻塞 Session。
- R14-1: 未启动原因的截断可能切断代理对。未变。
- R13-2: 服务运行中打开版本控制。未重测;对象存储代码没有变化。
- 其他仍延后的: fix(core): Handle partial ANSI and UTF-8 prefixes in Shell previews #12969、F3 和 B2/B4 MySQL 测试。
证据见 wenshao/qwen-code@5871c7d5(清单同英文版,results/notes.txt 列出了每个文件对应的运行)。
qqqys
left a comment
There was a problem hiding this comment.
Independent Critical-only review — head 0027f7a7f4d85fc6387fa0fa074fd649ea44cf2f
Verdict: COMMENT. Two Criticals from the maintainer's outstanding CHANGES_REQUESTED have never been adjudicated at this head and I could not verify them within one review budget; a third I did verify, and it is fixed. The diff is 71 files and +13,909/−321, so the current-head scan this channel's Approve path requires is not reachable in the time available. Under the rule that any historical blocking issue which still stands or cannot be confirmed forbids an Approve, this is a COMMENT — not a judgement that the PR is unsound, and I file no new Critical of my own.
The head has moved six times since that review landed (e13b415b → af423eb7 → 299a6fa8 → 42e2e8b1 → af962df7 → e177e1da → 0027f7a7), and no automated round has reviewed anything after 42e2e8b1. My own earlier review at e13b415b is dismissed and is not carried forward.
R2-4 — the second acquire() site missing the definite-refusal contract: FIXED
The finding was that the recovery/open path's acquire() did not carry the contract the class's own execute path implements and its tests pin, so a definite 409 workspace_busy / workspace_unavailable from the broker was indistinguishable from an ambiguous failure: the caller's blanket catch turned a live, actionable workspace conflict into 503 managed_session_open_failed.
At this head both halves of the prescribed fix are present in packages/cli/src/serve/hosted-harness-session.ts. Line 466, on the acquisition path, rethrows the refusal unwrapped instead of wrapping it:
if (isRetryableWorkspaceAcquisition(cause)) throw cause;and the route's catch, at line 852, now classifies before falling through to the 503:
if (isRetryableWorkspaceAcquisition(cause)) {
error(res, 409, cause.code);
} else if (cause instanceof ManagedSessionAlreadyExistsError) {
error(res, 409, 'managed_session_already_exists');
} else if (cause instanceof ManagedSessionNotFoundError) {
error(res, 404, 'managed_session_not_found');
} else {
error(res, 503, 'managed_session_open_failed');
}So a definite refusal reaches the caller as a 409 carrying its own code, and only genuinely ambiguous failures answer 503. The predicate is imported from hosted-workspace-tool-turn.ts — the module that already implemented this contract on the execute path — so the two sites now share one classifier rather than each spelling the condition. Honest limit on this confirmation: I verified the wiring at both call sites and the shared import, but did not read the predicate's body, so I am not ruling on the exact set of codes it matches.
Unconfirmed at this head: R5-1 and R4-1
Both are severity C in the round-5 ledger at e13b415b, and nothing has ruled on them since. Their full text is in that review's inline threads; what the ledger titles name is:
- R5-1 —
WorkspaceRuntimeTransport.java:145: the new grant-mode entry point lets a provably pre-dispatch ownership refusal take a path it should not. This is the newest of the three and sits on the Java transport that this PR's whole delivery mechanism runs through. - R4-1 —
ToolPublicationDataStore.java:549:beginFinishmoves the publication toproducer_phase = 'FINISHING'before the step that the finding objects to, i.e. a phase transition committed ahead of the fact it is meant to follow.
Both are in files this PR adds or heavily changes, and both are the shape where a green suite says nothing: a phase written too early, or a refusal classified on the wrong branch, is exactly what tests written against the intended behaviour pass over. I had no budget to trace either call chain, so they stand as unconfirmed, which under this channel's rules is treated the same as still standing.
Gate not completed: the current-head scan
I read the acquisition and close paths of hosted-harness-session.ts and the review record. I did not scan this diff. For scale, and so the gap is concrete rather than rhetorical: 71 files, +13,909/−321, spanning the durable publication store and its SQL, the Java transport and broker service, the TypeScript runtime and session layers, the OpenAPI contract and its regenerated client, the record contracts and their fixtures, and the design docs in both languages. Nothing in that list was read here beyond the one path above.
Two things in it deserve naming because a green suite does not bound them. Any new migration in this diff should be checked for its release ordering against the migrations still pending on main, since an out-of-order migration is invisible to every test in the PR. And the OpenAPI change should be checked the way the last contract PR was: read x-qwen-implementation-status on each touched operation at this head, so a route the server does not serve cannot be certified as one — planned routes cannot mis-certify shipped behaviour, and the regenerated-client diff shows the real blast radius.
Other reviewers
@chiga0 posted review rounds at af423eb7 and 42e2e8b1, both three heads behind this one, so their findings are likewise unadjudicated here. @wenshao's CHANGES_REQUESTED at e13b415b is the latest maintainer review state on record and has not been superseded by an approval.
CI
Green at this head: Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), Serve A/B, web-shell E2E Smoke, TUI parity snapshots, OpenTUI no-flicker gate, Hosted process fault gates / MySQL 8.4 / Java 21, Runtime Broker and Managed Agent MariaDB / Java 21, Real daemon E2E / Java 11 and the full Java matrix all pass. No review-pr check is reported against this head, so no automated round is pending that would rule on R5-1 and R4-1 by itself.
Next step: the shortest path to an Approve is a ruling on R5-1 and R4-1 at the head that will actually merge — either a fix, or a re-review round that traces both call chains and records them fixed. R2-4 can be closed as fixed on the evidence above. Given this PR has been through five maintainer rounds, it is also worth applying the repository's own guidance about not letting review rounds balloon it: beyond the two Criticals, remaining Suggestions are better deferred to a follow-up issue than folded into another round here.
Follow-up to the Critical-only review at
|
| Item | Status | Evidence |
|---|---|---|
| R5-1: ownership refused before dispatch | Fixed | Rerun at this head: ownership is lost in the publications:install window, and executeV3 settles not_started. The model gets "Workspace execution was refused before dispatch.", the receipt is committed, and a reload returns 200 with the Session not blocked. The Broker → worker ledger has no /v3/execute. The Java transport is unchanged since 42e2e8b1, and the fix had an A/B in round 12. |
| R4-1: an expired operation wedges the capture | Fixed only for lost replies | When the producer loses the reply, /recover works on the fake OSS and on real OSS: bytes exact, quota unchanged, command run once. That is also #13019's acceptance scenario. An operation that expires while in flight still answers 400 invalid_request ("claim expired"), and the producer treats that as final. The segment stays CANDIDATE or the publication stays FINISHING, and the Session is blocked. This is R15-1 in round 15. |
| R2-4: resume classification | Fixed as wired | isRetryableWorkspaceAcquisition (hosted-workspace-tool-turn.ts:63) matches only a 409 workspace_busy or workspace_unavailable. On the real stack, a lost acquire reply during a resume now gives 503 and then 200 on the retried load, and the turn completes (Mac and second host). A definite 409 during a resume could not be produced: the Broker re-acquires a live runtime session without WorkspaceExecutionStore.claim(), so the unit tests remain the evidence for that branch. |
| Migration ordering | No clash with main | V20–V22 apply on main's V19 schema. Other in-flight PRs do collide (#13037, #13084, #13087, #13088, #12946); they are listed in round 15. |
Two corrections to that review:
- Maintainer review state: by the time the review was posted, @wenshao had already approved
0027f7a7(11:20:12Z), so theCHANGES_REQUESTEDate13b415bwas no longer the latest. - OpenAPI: this PR changes no OpenAPI document and no generated client. Its contract files are the internal
managed-tool-publication-v1schema and fixtures, so thex-qwen-implementation-statuscheck does not apply.
Also, before e177e1da the R2-4 failure was a 200 load followed by a blocked Session (prompts got 409 hosted_turn_recovery_required). The 503 managed_session_open_failed answer is new in that commit.
Still open: R15-1 does not block this merge while the profile stays disabled, but it should be fixed before O2 is enabled. It could be tracked in #13019 or in a new issue. Round 15 has the fault matrix, the real-OSS runs and a tested one-line client change, which covers single stalls but not claim-only lapses or read-backs longer than the timeout.
R5-1 rerun logs: s15-r15lsi.log (O2) and s15-r15lsl.log (the local path has no install window, so the command runs normally there).
中文版
关于 0027f7a7 上的 Critical 专项评审:它未能确认的各项
此后 PR 已在这个 head 上获得批准:@wenshao(11:20Z)和 @qqqys(11:36Z)。以下记录那条评审未能确认的各项,在 0027f7a7 的真实栈上的状态:
| 项 | 状态 | 证据 |
|---|---|---|
| R5-1: 派发前所有权被拒 | 已修复 | 在这个 head 上重跑:在 publications:install 窗口内丢失所有权,executeV3 结算为 not_started。模型收到 "Workspace execution was refused before dispatch.",receipt 已提交,reload 返回 200,Session 未被阻塞。Broker → worker 的账本里没有 /v3/execute。Java transport 自 42e2e8b1 起未变,修复在第 12 轮做过 A/B。 |
| R4-1: 过期的操作卡死 capture | 只修好了应答丢失的情况 | 生产端丢失应答时,/recover 在 fake OSS 和真实 OSS 上都能用:字节一致,配额不变,命令只执行一次。这也是 #13019 的验收场景。操作在执行过程中过期时,仍返回 400 invalid_request("claim expired"),生产端把它当成最终结果。分段停在 CANDIDATE 或发布停在 FINISHING,Session 被阻塞。这就是第 15 轮的 R15-1。 |
| R2-4: 续跑时的分类 | 接线已修复 | isRetryableWorkspaceAcquisition(hosted-workspace-tool-turn.ts:63)只匹配 409 的 workspace_busy 或 workspace_unavailable。在真实栈上,续跑时 acquire 应答丢失,现在会先得到 503,重试 load 后得到 200,轮次完成(Mac 和第二台宿主机都成立)。续跑时明确的 409 无法产生:Broker 重新获取仍然存活的 runtime session 时不会调用 WorkspaceExecutionStore.claim(),所以那一支仍以单元测试为证据。 |
| 迁移顺序 | 与 main 不冲突 | V20–V22 能在 main 的 V19 库上应用。其他在飞 PR 之间确有冲突(#13037、#13084、#13087、#13088、#12946),已列在第 15 轮报告中。 |
对那条评审的两处更正:
- 维护者评审状态: 那条评审发出时,@wenshao 已经批准了
0027f7a7(11:20:12Z),所以e13b415b上的CHANGES_REQUESTED已不是最新状态。 - OpenAPI: 这个 PR 没有改动任何 OpenAPI 文档,也没有生成客户端。它的契约文件是内部的
managed-tool-publication-v1schema 和 fixtures,所以x-qwen-implementation-status的检查不适用。
另外,e177e1da 之前 R2-4 的表现是 load 返回 200,随后 Session 被阻塞(prompt 得到 409 hosted_turn_recovery_required)。503 managed_session_open_failed 是那个提交新加的。
仍未解决: profile 保持关闭时,R15-1 不阻塞本次合并,但在启用 O2 之前应当修复。可以放在 #13019 或新开 issue 跟踪。第 15 轮有故障矩阵、真实 OSS 的运行结果,以及一个实测过的一行客户端改动;这个改动能处理单次停顿,但处理不了只有 claim 过期、以及读回时间长于超时的情况。
R5-1 重跑日志:s15-r15lsi.log(O2)和 s15-r15lsl.log(本地路径没有 install 窗口,所以命令正常执行)。










What this PR does
Adds the O2 remote result path for foreground Hosted Shell calls: bounded raw stdout and stderr publication, immutable catalog and object storage, fixed-version range reads, Session receipt admission, Broker and independent worker Tool v3 routing, and Hosted recovery. The private Shell profile stays disabled until deployment supplies explicit capacity and storage settings. The existing no-tool, file-only, v2, and local capture paths remain available.
Why it's needed
O1c can preserve complete Shell output and its receipt when the Session owner and worker share a process. Hosted workers need the same original bytes and decision to survive separate processes and host replacement, without rerunning a command whose side effects may already have happened.
Reviewer Test Plan
How to verify
Evidence (Before & After)
N/A for visual evidence. Local verification passed: full build, typecheck, bundle, scoped source ESLint, formatting, focused core and CLI tests, 21 focused Java publication tests, the runtime Broker suite (374 tests, one skipped), and 11 MySQL 8.4 integration tests on a fresh schema. The 1 GiB incremental H2/file-object-store stress case passed with tail verification. The full Java server suite remains affected by one pre-existing timing-sensitive environment-event test; its isolated class passed 26/26. See the separate acceptance report comment for exact scope and limitations.
Tested on
0172d5eca(round 10 report); current head not covered by that runEnvironment (optional)
Node.js 22, local Java 21, MySQL 8.4, H2, and a faultable local object-store double. The maintainer independently exercised real OSS and a second host on earlier head
42e2e8b1(round 14 report); that run does not cover the later expiry-recovery protocol.Risk & Scope
Linked Issues
Builds on the merged Hosted file-tool loop in #12831 and the O1a/O1b/O1c groundwork.
中文说明
这个 PR 做了什么
为前台 Hosted Shell 调用增加 O2 远程结果路径:有界发布原始 stdout 和 stderr、不可变 catalog 与对象存储、固定版本范围读取、Session 回执接纳、Broker 与独立 worker 的 Tool v3 路由,以及 Hosted 恢复。私有 Shell profile 在部署方显式提供容量和存储配置前保持关闭。现有无工具、仅文件工具、v2 和本地捕获路径仍可使用。
为什么需要
O1c 能在 Session owner 与 worker 同进程时保留完整 Shell 输出和回执。Hosted worker 分处不同进程,需要让原始字节及接纳决定在进程和宿主替换后仍可恢复,同时避免对已经产生副作用的命令再次执行。
评审验证计划
如何验证
证据(前后对比)
非可视界面改动,无截图。已通过本地验证:全仓构建、类型检查、bundle、限定源码范围的 ESLint、格式检查、core 与 CLI 定向测试、21 项 Java publication 定向测试、Runtime Broker 测试套件(374 项,跳过 1 项),以及全新 schema 上的 11 项 MySQL 8.4 集成测试。使用 H2 与本地文件对象存储的 1 GiB 增量压力案例通过尾部校验。完整 Java server 套件仍受一个既有时序敏感的环境事件测试影响;该测试所在类独立运行 26/26 通过。精确范围和限制见单独的验收报告评论。
已测试平台
0172d5eca提交(第 10 轮报告);该次运行未覆盖当前 head环境
Node.js 22、本地 Java 21、MySQL 8.4、H2 和可注入故障的本地对象存储替身。维护者已在较早的
42e2e8b1上独立验证真实 OSS 与第二宿主(第 14 轮报告);该次运行未覆盖后续新增的过期恢复协议。风险与范围
关联事项
基于已合并的 Hosted 文件工具循环 #12831,以及 O1a/O1b/O1c 基础能力。