Repository navigation
feat(serve): add gated Hosted foreground Shell turns - #12848
Conversation
E2E verification reportPASS on macOS, Node.js 22.14.0 and Java 21. The final run uses the packaged CLI, production Java Broker, separate Workspace workers, Spring HTTP Session Store and H2 in MySQL mode with Flyway. The model is a deterministic loopback OpenAI fixture; no real provider is involved.
The real integration runs file-only and Shell Sessions in six distinct Workspaces. It checks a Harness decoy, complete 100 MiB stdout (including invalid UTF-8 bytes), separate stderr, independently computed lengths/SHA-256/tails, the original Session/turn/execution identity and durable receipt admission before the next model request. The model sees an explicit bounded-preview notice and no recommendation to read an inaccessible worker output file. After the driver exits, the fixture captures the live producer processes, closes the Broker, waits for each process to exit and checks that none is alive. The final run confirms Fault cases drop a start response, reject content publication with an injected HTTP 503 at the SQL-backed Store API, and drop the first successful raw-write response. This is a publication-API failure test, not an injected database-engine failure. Each effect marker remains exactly ReproductionBuild the repository and run mvn -q -f packages/sdk-java/runtime-broker/pom.xml -Dnode.executable="$(command -v node)" -Dtest=RuntimeBrokerServiceTest,RuntimeBrokerHttpServerTest,HttpRuntimeTransportTest,InMemoryRepositoryTest,JdbcRepositoryTest clean install
mvn -q -f packages/sdk-java/managed-agent-server/pom.xml -Phosted-workspace-tools -Dnode.executable="$(command -v node)" -Dtest=WorkspaceRuntimeTest,ManagedSessionStoreIntegrationTest verifyThe unchanged default process regression passed on the same product source before the test-only shutdown assertion change; it was not redundantly rerun. To reproduce it, run from npx vitest run --config vitest.hosted.config.tsThe global CLI 0.24.6 baseline rejects the private Hosted startup path before any model request. This is recorded as an unavailable baseline, not a successful Shell test. The existing file-only profile's Shell refusal is also verified in the real fixture. Build, typecheck, bundle and focused lint/format checks passed on the unchanged product source. The post-audit verification clean-built and installed the current Broker, reran the stricter real-process test plus Workspace/Store tests, and passed Checkstyle for both Java modules. Test processes and listeners were checked after completion; none remained. Earlier failures exposed and led to fixes for turn checkpoint identity, stale-owner preparation/admission, overlapping pipe writes and misleading temporary-preview-file guidance. The final test assertions distinguish provider preview text from durable receipt metadata. Limits: H2 is not a real MySQL deployment. Producer shutdown is not a database restart test. Windows, Linux, remote provisioners, real providers, public Artifact APIs and automatic recovery/replay were not validated here. Additional full-diff audits before reviewThree additional open-ended passes covered all 46 changed files and their consumers. The first pass identified a verification gap: requesting Broker shutdown did not prove that worker processes had exited before retained reading. Commit 7a41a5d42 adds the explicit exit barrier. The refreshed integration passes with six producer exits, followed by successful independent output reads. Two subsequent complete passes over the same frozen source found no additional confirmed defects or unresolved suspicions. No product code changed during these additional passes. |
Real-stack verification: #12848 Hosted foreground Shell (head
|
| Part | What ran |
|---|---|
| Database | MySQL 8.4.7 (native), Flyway V1–V13 applied by the PR's jar |
| Server | managed-agent-server jar built from the PR: Session Store, embedded Runtime Broker, local-process workers from the PR bundle |
| Harness | Packaged dist/cli.js serve --profile hosted-harness, owner of the loopback publisher |
| Models | Real qwen3.8-max (DashScope) for the behaviour trials; the repo's fake OpenAI server for deterministic probes |
| Build and tests | build + bundle OK; eslint + prettier clean on all 26 changed TS files; core 44/44; CLI 899/899 (the author's nine files); full shell.test.ts 361/361 (the PR ran 3 selected tests); broker 374 (1 skipped); managed-agent-server 114 + checkstyle |
PR claims vs evidence
| Claim | Result | How |
|---|---|---|
| Shell affects only the selected Workspace, exactly once; file-only Sessions still refuse Shell | ✅ | Driver on MySQL: shell-once.txt is x in all 4 Shell Workspaces, and the decoy is untouched. Real model R1: pwd is the saved Workspace, the file lands there, and the Harness directory with the decoy is untouched |
| Complete 100 MiB output (invalid UTF-8, separate stderr), preview only after a durable receipt | ✅ | Driver on MySQL: 107 content rows, 104,858,219 B |
| Retained output readable after producers stop | ✅, stronger than claimed | Stop Spring, restart mysqld, start a fresh Spring, then run the author's reader: HOSTED_SHELL_RETAINED_OUTPUT_OK, 104,857,735 B, digests and tails match |
| Lost start reply, lost raw-write ACK and SQL publish failure block without repeating the effect | ✅ | Driver on MySQL (fault relays unchanged) |
| Cancellation | ✅ | Driver. Also cancelling mid-stream in a 1 GiB yes: cancelled ~2.3 s later, capture committed, no yes process left, and the next turn works |
| Preview stays under the 64 KiB history limit | ✅ | Worst case, 70 KB of \x01 plus exit 3 (6× JSON escaping): the record is 55,578 B |
| Ordinary Shell truncation unchanged | ✅ | shell.test.ts 361/361 |
| Long commands (the Broker's worker request times out at 30 s) | ✅ | A 45 s command settles once, success |
| Stale-activation fencing | Unit tests pass. Mutant C4 (drop the activation clause in prepare) survives, possibly because other clauses cover it |
F1: the preview keeps only the first 8 KiB, so failures and the exit code never reach the model
boundedShellPreview() takes the first 8 KiB of the tool text. The core Shell tool deliberately keeps head and tail. shell.ts sets keep: 'both' because that preserves the command's "trailing exit/error summary (where shell failures report)". With raw capture, that truncation is bypassed and replaced by the head-only cut.
- Deterministic check. Probe P2 prints 16 KB of output, then
FINAL-42-TAIL, thenexit 2. At head the model sees 8,540 B containing neither the tail norExit Code:. The notice says onlyShell execution: error. - Real model.
build.shprints 600 progress lines, then an ERROR on stderr, thenBUILD FAILED, then exits 3; the tool text is 31,897 B. The preview ends atcompiling module 156 ... ok (cache war. How many timesbuild.shactually ran:
| Arm | Runs per trial | Re-ran |
|---|---|---|
| Hosted, PR head | 2 2 1 2 2 2 | 5/6 |
| Hosted, candidate below | 2 1 1 1 1 1 | 1/6 (that trial already had the answer and re-ran to separate stderr from stdout) |
Ordinary Shell, same bundle, -p |
1 1 1 | 0/3 |
A typical re-run at head: sh build.sh >/tmp/build.stdout 2>/tmp/build.err; echo "EXIT_CODE=$?". With a build this is harmless. With a migration, a deploy or a non-idempotent script, the Hosted profile will now routinely run it twice to find out why it failed.
Candidate (patch): same 8 KiB budget, split as 2 KiB head + a marker + 6 KiB tail, with UTF-8-safe cuts at both ends.
- P2: the tail and the exit code now reach the model.
- The worst-case record shrinks to 54,068 B.
- The nine CLI files pass 901/901 (the patch also adds the auth/close test below).
- The new test fails on the head version.
The tail of the tool text only reaches the end of the output when that output fits in the 64 KiB in-memory buffer (maxBufferedOutputBytes). For larger outputs the exit code still arrives, but the true output tail would have to come from the durable capture. That could be a follow-up.
Candidate boundedShellPreview
const PREVIEW_BYTES = 8 * 1024;
const PREVIEW_HEAD_BYTES = 2 * 1024;
const PREVIEW_GAP =
'\n[... preview truncated; the end of the output follows ...]\n';
export function boundedShellPreview(parts: readonly unknown[]): unknown[] {
const text = parts
.map((part) =>
part && typeof part === 'object' && 'text' in part && typeof part.text === 'string'
? part.text
: '',
)
.join('');
const bytes = Buffer.from(text);
if (bytes.byteLength <= PREVIEW_BYTES) return text ? [{ text }] : [];
// Shell failures and the exit status are reported last: keep both ends.
const head = new TextDecoder().decode(bytes.subarray(0, PREVIEW_HEAD_BYTES), { stream: true });
let start =
bytes.byteLength - (PREVIEW_BYTES - PREVIEW_HEAD_BYTES - Buffer.byteLength(PREVIEW_GAP));
while (start < bytes.byteLength && (bytes[start]! & 0xc0) === 0x80) start++;
return [{ text: head + PREVIEW_GAP + bytes.subarray(start).toString('utf8') }];
}F2 (low): is_background: true kills the whole turn, and core Shell tells the model to use it
When core Shell's sleep N guard blocks a command, it tells the model to use is_background: true or the Monitor tool. Neither exists in this profile. If the model follows that advice, execute() throws before dispatch and the whole turn ends turn_error in 109 ms (probe P7), instead of returning a function error the model could recover from. Over 3 real-model runs the guard fired every time and the model never chose is_background: once it reported the block and stopped, twice it used the # intentional-sleep escape. Suggestion: answer profile-refused arguments with a function error response, or tailor the guard text for Hosted.
Test gaps from mutation (one-line mutants, the author's own suites)
TypeScript: 10/20 killed. Java: 5/10 killed.
- M13 (publisher bearer check deleted) and M11 (the turn never closes its publisher) survive the unit tests and the full six-Workspace driver on MySQL. Live, the PR head is correct: 401 without a token or with a wrong one, and the listener is gone after every turn. Under M13, requests reach the handler (409); under M11, listeners pile up 1 → 2 → 3 across turns. The patch above adds a test that passes on head and kills M13.
- The triage review lists several of these guards as the PR's safety properties. They hold at head, but no test pins them:
- J6: ACK requires a settled execution.
- J7: v3 is Shell-only.
- J4: Store kind whitelist.
- J5: local-process only.
- J9:
inputDigestformat. - M1: raw-write offset check.
- M6/M7: accept/receipt equality.
- C2: seal check on read.
- C6: read bounded by the declared length.
Notes
- N1, storage growth (declared as a follow-up; here are the numbers). One
yescall with a 20 s timeout stored 135,360,328 B ofPUBLISHEDrows. Extrapolated, that is ~0.8 GB per call at the default 120 s timeout and ~4 GB at the 600 s maximum, with no per-call or per-Session cap. A per-call byte cap would be cheap insurance before this profile leaves private use. - N2, throughput. Every 64 KiB raw write calls
assertWritable(), which renews the writer lease (an HTTP call plus a locked UPDATE on the Session head row). A 64 MiB call made 1,097writers:renewrequests against 68tool-results:publish, and ran at 4.4 MiB/s; 300 MiB took 58.9 s. Renewing at most every few seconds (the lease is 60 s) would likely lift this. - N3, resolving the conflict with
c6c76750. In this PR, a Shell result'soutcomeRefcomes from the durable receipt. The oversize fallback inc6c76750replaces the committed parts. If the fallback were applied to a Shell result while the receipt'soutcomeRefis reused, history and the outcome resource would diverge. With the 8 KiB bound, the worst Shell record I could construct is 55.6 KB, so the fallback should never fire for Shell. It is worth asserting that in the resolution. - Not covered: Windows/Linux, remote provisioners, a worker crash mid-command, and live activation replacement.
Evidence (images, rig scripts, per-run logs, mutation JSON, candidate patch): wenshao/qwen-code@9ab1ced6/pr12848.
中文版
真实环境验证:#12848 Hosted 前台 Shell(head 7a41a5d4)
实测提交为 782c1774。之后唯一的新提交 7a41a5d4 只在 HostedWorkspaceToolTurnIT.java 里加了 7 行,生产代码完全一致。本报告是对 triage 评审 的补充:该评审明确说明没有执行 PR 代码,并希望有人实际跑一遍恢复与保留相关的声明。下面所有结论都经过实际运行。
结论
- 在真实 MySQL 8.4.7 上,持久化契约成立。 作者使用的是 H2 的 MySQL 模式。PR 自带的六 Workspace driver 原样通过。保留的 100 MiB 输出在 Spring 与 mysqld 都重启之后逐字节读回一致。每个 Shell 副作用都恰好执行一次。取消、超时、300 MiB 输出以及超过 30 s 的命令都能正常结算。在实测代码中没有发现正确性阻塞问题。
- 目前还不能直接合并,原因有两个,都属于流程问题:
- 它与 base 已经冲突(
CONFLICTING):feat(serve): add gated Hosted Workspace file tool turns #12831 前进到了c6c76750,冲突位于hosted-workspace-tool-turn.ts与hosted-harness-session.ts,恰好是 Shell 结果提交路径。冲突解决后需要重跑一遍,装置已经就绪。 - 由于 base 是堆叠分支,这个 head 上没有跑过任何 CI。
- 它与 base 已经冲突(
- 在真实模型使用这个 profile 之前应该修复:F1。 面向模型的预览只保留前 8 KiB,失败摘要和退出码都会被截掉。用同一个构建脚本,真实模型在 6 次试验中有 5 次 重跑了它;普通 Shell 工具 3 次中 0 次。候选补丁和测试见下文。
- 测试缺口: 删除 publisher 的 bearer 校验(M13),或者让 publisher 永不关闭(M11),都能通过全部 899 个单测,也能通过作者的完整六 Workspace driver。这里同样附了候选测试。
装置
| 部分 | 实际运行内容 |
|---|---|
| 数据库 | MySQL 8.4.7(原生安装),由 PR 构建的 jar 执行 Flyway V1–V13 |
| 服务端 | 由 PR 构建的 managed-agent-server jar:Session Store、内嵌 Runtime Broker,以及使用 PR bundle 的 local-process worker |
| Harness | 打包后的 dist/cli.js serve --profile hosted-harness,是回环 publisher 的 owner |
| 模型 | 行为试验用真实 qwen3.8-max(DashScope);确定性探针用仓库自带的 fake OpenAI server |
| 构建与测试 | build + bundle 通过;26 个改动的 TS 文件 eslint + prettier 干净;core 44/44;CLI 899/899(与作者相同的 9 个文件);完整 shell.test.ts 361/361(PR 只跑了其中 3 个);broker 374(1 个 skip);managed-agent-server 114 + checkstyle |
PR 声明与证据
| 声明 | 结果 | 方式 |
|---|---|---|
| Shell 只作用于选定的 Workspace,且恰好一次;仅文件工具的 Session 仍拒绝 Shell | ✅ | MySQL 上跑 driver:4 个 Shell Workspace 的 shell-once.txt 都是 x,诱饵未被改动。真实模型 R1:pwd 是保存的 Workspace,文件落在那里,放诱饵的 Harness 目录未被改动 |
| 完整的 100 MiB 输出(含非法 UTF-8、独立 stderr),持久回执之后才给出预览 | ✅ | MySQL 上跑 driver:content 107 行,104,858,219 B |
| 生产者退出后仍能读取保留的输出 | ✅,比声明更强 | 停 Spring、重启 mysqld、再起一个新 Spring,然后用作者的 reader 读取:HOSTED_SHELL_RETAINED_OUTPUT_OK,104,857,735 B,摘要与尾部一致 |
| start 回复丢失、raw write 确认丢失、SQL 发布失败时都会阻塞,且不重复副作用 | ✅ | MySQL 上跑 driver(故障中继未修改) |
| 取消 | ✅ | driver 覆盖。另外在 1 GiB yes 输出途中取消:约 2.3 s 后 cancelled,捕获已提交,没有残留 yes 进程,下一轮正常 |
| 预览不超过 64 KiB 历史上限 | ✅ | 最坏情况:70 KB 的 \x01 加 exit 3(JSON 转义 6 倍),记录为 55,578 B |
| 普通 Shell 的截断行为不变 | ✅ | shell.test.ts 361/361 |
| 长命令(Broker 到 worker 的请求 30 s 超时) | ✅ | 45 s 命令只结算一次,结果 success |
| 过期 activation 隔离 | 单测通过。变异体 C4(删掉 prepare 中的 activation 条件)存活,可能因为其他条件已覆盖 |
F1:预览只保留前 8 KiB,失败信息和退出码到不了模型
boundedShellPreview() 只取工具文本的前 8 KiB。核心 Shell 工具是有意同时保留头部和尾部的:shell.ts 设置 keep: 'both',因为这样能保留命令"末尾的退出/错误摘要(shell 失败信息所在位置)"。raw capture 路径绕过了这一截断,改成了只保留头部。
- 确定性检查。 探针 P2 先输出 16 KB,再输出
FINAL-42-TAIL,然后exit 2。在 head 上,模型看到的 8,540 B 里既没有尾部也没有Exit Code:,提示里只有Shell execution: error。 - 真实模型。
build.sh先输出 600 行进度,然后在 stderr 输出 ERROR,再输出BUILD FAILED,最后 exit 3;工具文本共 31,897 B。预览停在compiling module 156 ... ok (cache war。build.sh实际执行的次数:
| 分组 | 每次试验执行次数 | 重跑 |
|---|---|---|
| Hosted,PR head | 2 2 1 2 2 2 | 5/6 |
| Hosted,下方候选补丁 | 2 1 1 1 1 1 | 1/6(那一次已经拿到答案,重跑是为了区分 stderr 与 stdout) |
普通 Shell,同一 bundle,-p |
1 1 1 | 0/3 |
head 上典型的重跑命令是 sh build.sh >/tmp/build.stdout 2>/tmp/build.err; echo "EXIT_CODE=$?"。对构建来说无害;但如果是迁移、部署或非幂等脚本,Hosted profile 会为了弄清失败原因而经常把它执行两次。
候选补丁(patch):预算仍是 8 KiB,拆成 2 KiB 头部 + 标记 + 6 KiB 尾部,两端切点都保证 UTF-8 安全。
- P2:尾部和退出码都能到达模型。
- 最坏情况的记录降到 54,068 B。
- 9 个 CLI 测试文件 901/901 通过(补丁同时加入了下文的鉴权/关闭测试)。
- 新测试在 head 版本上失败。
工具文本的尾部只有在输出能放进 64 KiB 内存缓冲(maxBufferedOutputBytes)时才对应输出的真正末尾。输出更大时,退出码仍然能到达模型,但真正的输出尾部需要从持久捕获中读取,可以作为后续工作。
F2(低):is_background: true 会让整轮失败,而核心 Shell 恰好建议模型这样做
核心 Shell 的 sleep N 守卫拦下命令时,会建议模型改用 is_background: true 或 Monitor 工具,这两者在该 profile 中都不存在。模型如果照做,execute() 会在派发前抛错,整轮在 109 ms 内以 turn_error 结束(探针 P7),而不是返回一个模型可以自行恢复的函数错误。3 次真实模型运行中该守卫每次都触发,模型从未选择 is_background:一次报告被拦后停下,两次改用 # intentional-sleep 逃生方式。建议:对 profile 拒绝的参数返回函数错误,或者为 Hosted 定制守卫文案。
变异测试揭示的测试缺口(单行变异,使用作者自己的测试集)
TS 杀死 10/20,Java 杀死 5/10。
- M13(删掉 publisher 的 bearer 校验) 和 M11(本轮从不关闭 publisher) 既能通过单测,也能通过 MySQL 上的完整六 Workspace driver。真机上 PR head 的行为是正确的:不带 token 或 token 错误都返回 401,每轮结束后监听都会消失。M13 下请求能进入处理函数(返回 409);M11 下监听随轮次累积 1 → 2 → 3。上面的补丁附带了一个测试,在 head 上通过,并能杀死 M13。
- triage 评审把其中若干守卫列为 PR 的安全属性。它们在 head 上成立,但没有任何测试钉住:
- J6:ACK 要求执行已结算。
- J7:v3 只允许 Shell。
- J4:Store 的 kind 白名单。
- J5:只允许 local-process。
- J9:
inputDigest格式。 - M1:raw write 的 offset 校验。
- M6/M7:accept/receipt 一致性。
- C2:读取时的 seal 校验。
- C6:读取受声明长度约束。
备注
- N1,存储增长(已声明为后续工作,这里补上数据)。 一次超时 20 s 的
yes调用存了 135,360,328 B 的PUBLISHED行。外推下来,默认 120 s 超时每次约 0.8 GB,600 s 上限约 4 GB,而且没有任何单次调用或单 Session 的上限。在这个 profile 走出私有使用之前,加一个单次调用字节上限是很便宜的保险。 - N2,吞吐。 每个 64 KiB 的 raw write 都会调用
assertWritable(),续租 writer lease(一次 HTTP 请求加一次 Session head 行的加锁 UPDATE)。64 MiB 的一次调用产生了 1,097 次writers:renew,而tool-results:publish只有 68 次,速度 4.4 MiB/s;300 MiB 用时 58.9 s。若改成每隔几秒最多续租一次(lease 为 60 s),吞吐很可能会提升。 - N3,解决与
c6c76750的冲突。 在本 PR 中,Shell 结果的outcomeRef来自持久回执。c6c76750的超限回退会替换已提交的 parts。如果对 Shell 结果套用该回退、同时又复用回执的outcomeRef,历史与 outcome 资源就会不一致。在 8 KiB 上限下,我能构造出的最坏 Shell 记录为 55.6 KB,回退不应触发;建议在解决冲突时加断言保证这一点。 - 未覆盖: Windows/Linux、远端 provisioner、命令执行中 worker 崩溃,以及真机 activation 替换。
证据(图片、装置脚本、每次运行日志、变异 JSON、候选补丁):wenshao/qwen-code@9ab1ced6/pr12848。
|
@qwen-code /resolve |
|
Qwen Code resolved the merge conflicts and pushed the branch update. Merge resolution: PR #12848 ←
|
|
Addressed the actionable fixes from the real-stack report in a209f3a, on top of the conflict-resolution merge c3fbb0f. The PR is currently mergeable.
Validation on the pushed changes:
The stacked base still excludes the normal PR trigger. I manually started Qwen Code CI with Follow-ups remain explicit: true output tails beyond the 64 KiB buffer, per-call/Session storage quotas (N1), renewal-frequency optimization (N2), and the remaining mutation-coverage inventory (J6/J7/J4/J5/J9/M1/M6/M7/C2/C6 and C4). This correction does not change their production guards or claim to close those test gaps. This round did not repeat the report's real-MySQL restart or real-model repeated-execution trials; the fresh full-stack SQL run used H2. |
a209f3a to
b89b7b3
Compare
|
Please do not rebase or force-push to an active PR as it invalidates existing review comments. Note for future reference, the bots always squash all changes into a single commit automatically as part of the integration. 中文请勿对活跃的 PR 执行 rebase 或 force-push,因为这会使已有的评审评论失效。另外,供日后参考:作为集成流程的一部分,机器人始终会自动将所有改动压缩(squash)为单个提交。 |
|
Retargeted this PR to The new head is |
|
@qwen-code /triage |
qqqys
left a comment
There was a problem hiding this comment.
Independent Critical-only review — head b89b7b38
Not approving, on coverage rather than on a finding: this adds remote foreground Shell execution across four durability boundaries in 46 files and +4846/-513, and I could verify the gate and the turn-level guards myself but not the ~1,600 lines of publisher, capture-session and resource-store code, nor the Java half, inside my budget. For a feature whose subject is executing commands on a saved workspace, the unread part is exactly where a Critical would live, so I am naming the boundary instead of implying coverage I did not have.
No historical blocker
There are no reviews and no inline comments on this PR, so nothing has ever been filed against it and there is no thread to re-verify. Triage stage 2 at this exact head reports "No correctness blockers found at b89b7b38", and says it re-verified the load-bearing invariants on this head rather than carrying them forward after the squash and conflict resolution. The maintainer's real-stack run — native MySQL 8.4.7, the packaged dist/cli.js serve --profile hosted-harness, real qwen3.8-max for the behaviour trials — reports the durability contract holding, every Shell side effect running exactly once, and cancellation, timeouts, 300 MiB of output and >30 s commands all settling cleanly, with "no correctness blocker in the tested code". Its two process blockers are gone: the base conflict is resolved and CI has now run on this head.
What I verified in the code myself
The gate is explicit and narrow. HOSTED_WORKSPACE_SHELL_PROFILE = 'hosted-workspace-shell/1' is a second named profile beside the existing hosted-workspace-files/1, hosted-harness-session.ts accepts exactly those two and rejects anything else, and the shell options are wired only when the profile is the shell one — so default no-tool and file-profile hosts are unchanged, and nothing reaches Shell by default.
The turn refuses ambiguity rather than resolving it optimistically. Invalid Shell arguments, is_background: true included, produce durable function-error responses before acquisition or dispatch, with the message naming what is unavailable and asking for corrected arguments. A mixed batch is refused in full and every other call in it is explicitly reported as not executed, so the model cannot be left guessing which half ran. An uncertain outcome raises HostedToolRecoveryRequiredError rather than being reported as either success or failure.
No complete receipt without durable admission. The turn throws 'Complete Shell output was not admitted.' when the capture was not admitted, and the model-facing text on a truncated preview states the execution status and that complete stdout and stderr are retained in the Session result — it does not present a preview as the output. The over-limit guard is there too: 'Admitted Shell result exceeds the inline Session Store limit.' blocks before history replacement, checkpoint resolution and ACK, so an admitted result too large for the history record cannot be swapped for an omission response while keeping the original outcome reference.
The truncated preview keeps the parts a model needs. The head-plus-gap-marker-plus-tail shape within the 8 KiB budget, cutting only at UTF-8 boundaries, is what the maintainer's F1 measured as necessary — with the old head-only preview a real model re-ran a failed build in 5 of 6 trials against 0 of 3 for the ordinary Shell tool. The marker says "buffered output" rather than implying a >64 KiB process buffer holds the true stream tail, which is the honest claim.
tools/shell.ts changes by three lines: the temp-file truncation is skipped only when a raw capture owns the output, so the existing local path keeps its behaviour.
What I did not read, and why it matters
hosted-shell-publisher.ts (413 lines) and managed-shell-publisher.ts (405) — the loopback publisher, its capability handling, per-stream serialization of overlapping pipe callbacks, and the listener-closure paths; managed-shell-result-session.ts (436) and the 393 lines moved out of local-shell-result-session.ts; resource-tool-result-store.ts (366) — the writer-fenced publication, immutable resource ids and seal verification; and the Java side, RuntimeBrokerHttpServer (+65/-8), RuntimeBrokerService (+74/-8), WorkspaceRuntimeTransport (+97/-1), ManagedSessionStore (+64/-1) and HttpRuntimeTransport (+23/-1). The design's route table assigns an owner and a check set to each of five route families, and I verified none of those checks in code. The capability-discipline claims — canonical loopback only, redirects rejected, capability in memory and never in journal references or model arguments, no SQL writer token crossing into the worker — are the ones I would most want to read rather than accept, precisely because they are the boundary between a gated private profile and an unauthenticated local endpoint.
Two smaller notes on evidence currency, not defects: the executed real-stack results are from the pre-squash head, with this head re-verified statically by triage rather than re-run on a real stack; and the sandboxed verification triggered at 16:10 had not reported when I published.
CI
23 checks pass at this head and none has failed, including Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), the full Java matrix with Runtime Broker and Managed Agent MariaDB, Hosted no-tool processes / MySQL 8.4 and Real daemon E2E / Java 11, both Desktop Shell lanes and the routing checks. review-pr is pending and is not a gate. The maintainer's rig also reports the full shell.test.ts at 361/361 where the PR ran three selected tests.
What closes this
A pass over the publisher and capture pair — capability generation and validation, the loopback and redirect rules, per-stream write/finish serialization, listener closure after completed, model-error and uncertain turns, and what a lost raw-write reply does to the receipt — plus resource-tool-result-store.ts's writer fencing and seal verification, and the five route families' checks against the design's table. With those read, the gate and turn-level behaviour above are already verified and I would expect nothing else to stand in the way.
Verdict: COMMENT — No Critical found, and none has ever been filed here; the gate, the pre-dispatch refusals, the batch semantics, the no-admission-no-receipt rule and the preview shape all verify fail-closed in the code at this head. What blocks an approval is that the publisher, capture-session, resource-store and Java transport layers — the parts that hold the capability and the durability guarantees — are outside what I read in budget, and for a remote Shell path I am not willing to approve those on description alone.
Real-stack re-verification, round 2: #12848 at
|
| Item | Result at b89b7b38 |
|---|---|
| Author's six-Workspace driver on a fresh MySQL DB (all 15 migrations) | ✅ HOSTED_WORKSPACE_TOOLS_OK, 40.4 s, same row counts as round 1 (107 content rows, 104,858,219 B) |
| Stop Spring, restart mysqld, start a fresh Spring, run the reader | ✅ HOSTED_SHELL_RETAINED_OUTPUT_OK |
| Round-1 database (V13, holding the round-1 100 MiB output) opened by the new jar | ✅ V14 + V15 applied on startup; round-1 output still reads back byte-identical |
| F1 tail + exit code (16 KB, then a tail line, then exit 2) | ✅ model sees FINAL-42-TAIL and Exit Code: 2 after the gap marker |
| F1 with output larger than the 64 KiB buffer (200 KB, exit 4) | ✅ Exit Code: 4 arrives; the true stream tail past the buffer does not, as documented ("end of buffered output") |
F1 worst case (70 KB of \x01 + exit 3) |
✅ record is 54,061 B, under the 65,536 B limit |
| F1 with CJK + emoji (900 lines, stderr tail, exit 5) | ✅ tail and exit code present; the preview contains no U+FFFD |
F1, real qwen3.8-max, same build.sh as round 1 |
✅ build.sh ran once in 6/6 trials (round-1 head: re-ran in 5/6). The ERROR line reached the model every time |
F2 is_background: true |
✅ function error, turn_complete, 0 Broker starts; reload and next turn OK |
F2 mixed batch (write_file + background Shell) |
✅ both answered as not executed, 0 starts; the corrected Shell then ran once |
| F2 in the middle of a turn: Shell OK, then refused, then Shell OK | ✅ one turn; inter.txt = r1, r3; 1 acquire, 2 starts, 1 release; reload and next turn OK |
F2 with an unknown directory argument |
✅ function error; the corrected call ran once |
| M13 / M11 live | ✅ 401 without a token and with a wrong token; listener closed after 3 completed turns and after a model-error turn |
About F2 with a real model: in 4/4 trials (2 at this head, 2 at the round-1 head) qwen3.8-max declined to send is_background at all, because the declared schema doesn't have it. So F2 is only reachable through the deterministic probes above, and it is fixed there.
Mutation check on the fixes (the author's nine CLI files, baseline 912/912)
11/12 killed:
- the refusal path disabled (invalid arguments would be dispatched in the foreground)
- the refusal not persisting its tool_result
- the refusal answering only the invalid calls
- the refusal persisting nothing at all
- the N3 guard removed
- the F1 UTF-8 alignment removed, the head removed, the gap marker removed
- M13 bearer check removed; M11 turn close removed; the publisher's own
server.close()removed
M13 and M11 both survived in round 1.
Surviving: deleting the size check on the refusal record. It is effectively unreachable, since refusal messages are short. It could be pinned by a test with many calls, but I wouldn't hold the PR for it.
Against the open items in the latest review
That review notes that the executed real-stack results predated the squash. Everything in this report was executed at b89b7b38.
| Item named in the review | Executed evidence at b89b7b38 |
|---|---|
| Capability generation and validation | Live: no token → 401, wrong 43-char token → 401. Mutant M13 (bearer check deleted) is now killed |
| Loopback and redirect rules | Pinned by unit and Java tests: round-1 mutants M8 (worker URL shape) and J8 (Broker URL shape) were killed. Redirects were not exercised live |
| Per-stream write/finish serialization; what a lost raw-write reply does | Author's driver on MySQL: the raw-reply-loss Workspace blocks, the write is not retried, and the effect runs once. Round-1 mutant M2 (queue removed) was killed |
| Listener closure after completed, model-error and uncertain turns | Live: 3 completed turns, a provider failure after a Shell round, and an uncertain turn (status reply lost) all leave no publisher listener. M11 and server.close() mutants are killed |
| Store writer fencing and seal verification | Java mutants J1 (writer fencing), J2 (immutable resource id) and J3 (digest) were killed in round 1; retained reads after a mysqld restart verify the seal and digests. Round-1 C2 (skipping the seal check on read) still survives, as the author listed |
Non-blocking notes
-
A single lost status-poll reply blocks the Session.
HostedWorkspaceBroker.executepolls every 50 ms and does not retry a failed GET. Dropping one reply blocked the Session (hosted_turn_recovery_required); fail-closed, and the command ran once. The loop is from merged feat(serve): add gated Hosted Workspace file tool turns #12831, but Shell turns poll far longer than file tools: 793 polls for a 45 s command, about 12,000 for a 10-minute one. A few retries of the idempotent status read before declaring the outcome unknown would make long Shell turns much less brittle. -
runtimeError.messageis still cut withBuffer.from(message).subarray(0, 1024).toString('utf8')in the worker'sfinalize. With CJK text this gives 1,026 bytes ending in U+FFFD. The model-facing preview is not affected. This is the same kind of cut F1 fixed; it's a nit. -
The follow-ups from round 1 are unchanged, and the author lists them explicitly: true output tails beyond 64 KiB, storage quotas (N1), lease-renewal frequency (N2), and the older unpinned guards J4–J9, M1, M6/M7, C2, C4, C6.
中文版
真实环境复验(第二轮):#12848 @ b89b7b38
本轮承接第一轮报告和修复说明。修复说明提到本轮没有重做真实 MySQL 重启和真实模型试验,这两项都在这里补上。
head 等价性。 b89b7b38 是把 a209f3ac squash 后落在 main 302e7d88 上的结果。46 个文件的 PR 补丁逐字节一致(比较 git diff c6c76750 a209f3ac 与 git diff 302e7d88 b89b7b38,忽略 index 行),这些文件在 main 上也没有被改动。不过 main 带来了其他改动:V14/V15 迁移、事件回放,以及 hosted-harness-profile.ts、run-qwen-serve.ts 等。所以我在 b89b7b38 上全部重新构建:TS bundle、Broker 374 个测试(1 个 skip)、managed-agent-server 139/139、checkstyle。
结论
第一轮四项可执行问题(F1、F2、M13/M11、N3)均已修复,并在 MySQL 8.4.7 上真机验证;场景可达的地方也用了真实模型。没有阻塞问题。从真实环境验证角度,可以合并。 这个 head 上的 CI 全绿:Qwen Code CI、SDK Java、Serve A/B、tui-parity。此前在 a209f3ac 上手动触发的 Windows 任务失败属于既有问题:我在 GitHub 托管的 windows-2022 runner 上按 CI 的配置重跑,main 302e7d88 与本 head 失败的都是同样的 14 个 CLI 测试和 18 个 core 测试,且都不在本 PR 改动的文件里。
复验内容
| 项目 | b89b7b38 上的结果 |
|---|---|
| 作者的六 Workspace driver,全新 MySQL 库(15 个迁移全部执行) | ✅ HOSTED_WORKSPACE_TOOLS_OK,40.4 s,行数与第一轮一致(content 107 行,104,858,219 B) |
| 停 Spring、重启 mysqld、再起新 Spring,运行 reader | ✅ HOSTED_SHELL_RETAINED_OUTPUT_OK |
| 第一轮的数据库(V13,含第一轮的 100 MiB 输出)交给新 jar 打开 | ✅ 启动时执行 V14 + V15;第一轮输出仍逐字节读回一致 |
| F1 尾部 + 退出码(16 KB 输出、尾行、exit 2) | ✅ 模型在截断标记之后能看到 FINAL-42-TAIL 和 Exit Code: 2 |
| F1,输出超过 64 KiB 缓冲(200 KB,exit 4) | ✅ Exit Code: 4 能到达;缓冲之外的真正流尾不会到达,与文档说明一致("end of buffered output") |
F1 最坏情况(70 KB 的 \x01 + exit 3) |
✅ 记录 54,061 B,低于 65,536 B 上限 |
| F1,中文 + emoji(900 行、stderr 尾、exit 5) | ✅ 尾部和退出码都在,预览中没有 U+FFFD |
F1,真实 qwen3.8-max,与第一轮相同的 build.sh |
✅ build.sh 6/6 次都只执行一次(第一轮 head:6 次中 5 次重跑),ERROR 行每次都到达模型 |
F2 is_background: true |
✅ 返回函数错误,turn_complete,Broker start 次数为 0;重新加载和下一轮正常 |
F2 混合批次(write_file + 后台 Shell) |
✅ 两者都答复为未执行,start 次数为 0;修正后的 Shell 随后执行一次 |
| F2 出现在同一轮中间:Shell 成功、然后被拒、再 Shell 成功 | ✅ 同一轮内完成;inter.txt 为 r1、r3;1 次 acquire、2 次 start、1 次 release;重新加载和下一轮正常 |
F2,未知参数 directory |
✅ 返回函数错误;修正后的调用执行一次 |
| M13 / M11 真机检查 | ✅ 无 token 和错误 token 都返回 401;3 个正常结束的轮次以及一个模型出错的轮次之后,监听都已关闭 |
关于真实模型下的 F2:4 次试验中(本 head 2 次、第一轮 head 2 次),qwen3.8-max 都没有发送 is_background,因为声明的 schema 里没有这个参数。所以 F2 只能通过上面的确定性探针触发,在那里已经修复。
针对修复的变异测试(作者的 9 个 CLI 测试文件,基线 912/912)
杀死 11/12:
- 关闭拒绝路径(无效参数会被当作前台命令派发)
- 拒绝时不持久化 tool_result
- 拒绝时只回复无效的调用
- 拒绝时什么都不持久化
- 删除 N3 守卫
- 删除 F1 的 UTF-8 对齐、去掉头部、去掉截断标记
- 删除 M13 bearer 校验;删除 M11 的本轮关闭调用;删除 publisher 自身的
server.close()
M13 和 M11 在第一轮时都存活。
存活: 删除对拒绝记录的大小检查。实际上不可达,因为拒绝消息都很短。可以用一个大量调用的测试钉住,但不必为此卡 PR。
对照最新评审中列出的待办
该评审提到,此前实际运行的真实环境结果早于 squash。本报告的所有内容都在 b89b7b38 上实际执行。
| 评审列出的项目 | 在 b89b7b38 上实际执行的证据 |
|---|---|
| capability 的生成与校验 | 真机:无 token → 401,43 字符的错误 token → 401。变异 M13(删除 bearer 校验)现已被杀死 |
| 回环地址与重定向规则 | 由单测和 Java 测试钉住:第一轮变异 M8(worker 端 URL 形态)和 J8(Broker 端 URL 形态)均被杀死。重定向未做真机验证 |
| 按流串行的 write/finish;raw write 回复丢失时的行为 | MySQL 上跑作者的 driver:raw 回复丢失的那个 Workspace 被阻塞,写入不重试,副作用只执行一次。第一轮变异 M2(删除队列)被杀死 |
| 正常结束、模型出错、不确定三种轮次之后的监听关闭 | 真机:3 个正常结束的轮次、Shell 之后 provider 出错的轮次、以及不确定轮次(status 回复丢失)之后,都没有残留 publisher 监听。M11 和 server.close() 的变异体均被杀死 |
| Store 的 writer fencing 与 seal 校验 | 第一轮 Java 变异 J1(writer fencing)、J2(resource id 不可变)、J3(摘要)均被杀死;mysqld 重启后的保留读取会校验 seal 和摘要。第一轮的 C2(读取时跳过 seal 校验)仍然存活,作者已列入后续 |
非阻塞备注
-
一次 status 轮询回复丢失就会阻塞 Session。
HostedWorkspaceBroker.execute每 50 ms 轮询一次,GET 失败时不重试。我丢掉一次回复,Session 就被阻塞(hosted_turn_recovery_required);这是 fail-closed,命令也只执行了一次。这个循环来自已合入的 feat(serve): add gated Hosted Workspace file tool turns #12831,但 Shell 轮次的轮询时间远长于文件工具:45 s 的命令约 793 次,10 分钟的命令约 12,000 次。在判定结果未知之前,对这个幂等的 status 读取重试几次,可以让长时间的 Shell 轮次稳健得多。 -
worker 的
finalize里,runtimeError.message仍用Buffer.from(message).subarray(0, 1024).toString('utf8')截断。遇到中文会得到 1,026 字节、并以 U+FFFD 结尾。面向模型的预览不受影响。这和 F1 修掉的是同一类截断,属于小问题。 -
第一轮的后续事项不变,作者已明确列出:超过 64 KiB 的真实输出尾部、存储配额(N1)、续租频率(N2),以及较早未被钉住的守卫 J4–J9、M1、M6/M7、C2、C4、C6。
Evidence (round-2 scripts, per-run logs, mutation JSON, Windows outputs): wenshao/qwen-code@7e4be7d6/pr12848/r2.
|
Thanks for the round-6 real-stack review. I checked the reported paths against the current head and ran the targeted publisher tests plus the real inherited-pipe regression.
No review threads were unresolved or replied to (0/0). I will keep #12848 open for the remaining review and #12900 dependency; it has not been merged. 感谢第六轮真实环境评审。我按当前分支核对了这些路径,并运行了发布器定向测试及真实子进程继承管道的回归测试。
本轮没有未解决的行内评审线程,也没有回复或关闭线程(0/0)。#12848 将继续保持开放,等待后续评审及 #12900 依赖处理;没有合并。 |
|
Synced the merged #12900 and current Validation on macOS: the frozen-lockfile install ran the repository build and bundle successfully; The Workspace lease recovery gap remains tracked in #12904; this merge does not change its fail-closed behavior. 已在 本机 macOS 验证:锁文件安装过程成功执行仓库构建和 bundle, Workspace 租约恢复缺口仍由 #12904 跟踪;本次合并没有改变故障时阻断的行为。 |
|
CI follow-up for
Verified on this commit: build, typecheck, ESLint, and the Hosted Harness session test file (20/20). There were no review threads to reply to or resolve in this batch. |
Real-stack verification, round 7: #12848 at
|
| Command | Durable manifest | Read back through the Store |
|---|---|---|
cd . && sleep 20 >/dev/null 2>&1 & echo ok |
partial / producer_lost; stdout incomplete, 3 B |
3 B ok\n; sha256 = manifest |
cd . && { head -c 3000000 /dev/zero | tr "\000" a; sleep 20; } & sleep 5; echo done |
partial / producer_lost; stdout incomplete, 3,000,005 B |
3,000,005 B: 3,000,000 a + one done; sha256 = manifest |
| a logger that keeps writing for 3 s after the shell exits | partial / producer_lost; stdout incomplete, 36 B |
started + tick-0…tick-3; sha256 = manifest |
printf plain-ok (control) |
complete; sealed |
8 B; sha256 = manifest |
missingRangesrecords an open tail,[{"start": <n>, "end": null}], on both streams of the partial captures.- The four minimal background shapes from round 6 split as before. The two compound forms are now
partial/producer_lost;sleep 20 … & echo okandcd . ; sleep 20 … & echo okstill complete. - As before:
- The turn is still recovery-blocked, with
Complete Shell output was not admitted. - Every new Session in that Workspace, with either profile, still gets 409
workspace_busy. - The control Workspace still serves new Sessions normally.
- The turn is still recovery-blocked, with
The fix is one deleted line: the owner no longer calls failCapture() when a stream finishes with complete: false. That leaves the failed flag on finalize as the only way a worker-side write failure reaches the owner. Four mutants check that path:
| Mutant | Result |
|---|---|
| K1: restore the deleted line | killed by the new test (retains the accepted prefix when a producer holds a pipe past the drain boundary) |
K2: owner finalize ignores the worker's failed flag |
killed (irreversibly blocks capture after a lost raw write acknowledgement) |
K3: worker reports complete: true after a failed write |
survives; equivalent, because finalize still carries failed (K2 + K3 together is killed) |
K4: owner finalize no longer fails a stream that never finished |
survives; equivalent. handleExit always finishes both streams first, and a failed finish sets failed, which K2's line handles |
Regression at a4632e98
| Check | Result |
|---|---|
| Server start, fresh MySQL 8.4.7 database | ✅ 17 migrations: V16 runtime_loss_evidence, V17 managed_session_operation |
HostedWorkspaceToolTurnIT, whole class, mvn -P hosted-workspace-tools verify |
✅ 3/3 in 97.6 s. Six-Workspace driver; FG6a 8 cases, including status; FG6b 7 cases |
FG6b with a Shell call in place of edit |
✅ 6/6 |
| Java unit tests | ✅ Broker 390 (1 skipped), managed-agent-server 155 |
| R1-4 / R1-7 validation, R1-1 grandchild, 8 output shapes, F4 Shell → reload → Shell | ✅ unchanged |
Real qwen3.8-max: Shell, then detach + load, then Shell |
✅ both turns complete; notes.txt = first\nsecond |
CI on a4632e98
- SDK Java: green in all 8 jobs, including Runtime Broker and Managed Agent MariaDB and Hosted no-tool processes / MySQL 8.4, which were red on
20126bae. - Serve A/B and tui-parity: green.
- Qwen Code CI: green, including the web-shell E2E smoke.
Merge reference
- Correctness: I see no blocker. Everything from round 6 that belongs in this PR is done: the partial prefix is kept and labelled honestly, and the V16 collision is fixed on main and merged in.
- What remains is the availability decision from round 6.
- After a backgrounded compound command, the Session and the whole Workspace stay blocked, and no API releases them.
- A real model wrote that command in 3 of 6 trials.
- The author keeps it fail-closed on purpose, documents it in the PR body, and tracks recovery in Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904.
- From the real-stack side: mergeable as a private, fail-closed profile, provided
hosted-workspace-shell/1doesn't serve real-model traffic until Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 lands.
中文版
真实环境验证(第七轮):#12848 @ a4632e98
本轮覆盖作者对第六轮的回复和三个提交:
14b77fee在排空超时时保留 Hosted Shell 的部分输出;0b8ca77b合入 main,其中包括把 D4 迁移改为 V17 的 fix(managed-agent): Move the Session operation migration to V17 #12900;a4632e98为两个测试的等待设上限。
环境:MySQL 8.4.7;Spring jar 和 TS bundle 在 a4632e98 上重新构建;真实本地进程 worker;固定脚本模型加真实 qwen3.8-max。
结论
14b77fee在真实环境中的行为与描述一致。- 后台进程让 Shell 的管道保持打开时,持久 manifest 现在记为
partial/producer_lost,而不是unavailable/storage_failed。 - 已接收的前缀可以通过 Store 逐字节读回,包括一个跨越多页的 3,000,005 字节前缀。
- 变异测试表明,新路径以及它现在所依赖的 worker 失败路径,都有测试钉住。
- 后台进程让 Shell 的管道保持打开时,持久 manifest 现在记为
- 合入 main 修好了第六轮的启动失败。 服务带着 V16 和 V17 正常启动。IT 整类、main 的 FG6a 和 FG6b 门禁、Shell 版 FG6b、Java 测试、rig 场景和真实模型全部通过,SDK Java 恢复全绿。
- 按设计未变:捕获不完整后,Session 和 Workspace 仍然被阻塞。
- 作者保持这个私有 profile 的 fail-closed 行为,并已在 PR 说明中写明这项限制。
- Workspace 租约恢复由 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 跟踪,绝对
file_path导致整轮失败的问题由 Hosted file tools: make invalid absolute file_path a durable refusal #12905 跟踪。
14b77fee:未关闭的管道会保留前缀
每条命令在各自的 Session 中运行。我停掉 Harness,等它的租约过期,再通过 Store HTTP API 用 ResourceToolResultSegmentStore 读回每条流。
| 命令 | 持久 manifest | 通过 Store 读回 |
|---|---|---|
cd . && sleep 20 >/dev/null 2>&1 & echo ok |
partial / producer_lost;stdout incomplete,3 B |
3 B ok\n;sha256 = manifest |
cd . && { head -c 3000000 /dev/zero | tr "\000" a; sleep 20; } & sleep 5; echo done |
partial / producer_lost;stdout incomplete,3,000,005 B |
3,000,005 B:3,000,000 个 a 加一个 done;sha256 = manifest |
| shell 退出后还持续写 3 s 的日志进程 | partial / producer_lost;stdout incomplete,36 B |
started + tick-0…tick-3;sha256 = manifest |
printf plain-ok(对照) |
complete;sealed |
8 B;sha256 = manifest |
- 部分捕获的两条流上,
missingRanges都记录了一个开放的尾部:[{"start": <n>, "end": null}]。 - 第六轮的四种最小后台写法结果分布和之前一样。两种复合写法现在是
partial/producer_lost;sleep 20 … & echo ok和cd . ; sleep 20 … & echo ok仍然正常完成。 - 和之前一样:
- 这一轮仍然被恢复阻塞,报
Complete Shell output was not admitted. - 该 Workspace 中的每个新 Session,无论哪种 profile,仍然得到 409
workspace_busy。 - 对照 Workspace 仍能正常为新 Session 服务。
- 这一轮仍然被恢复阻塞,报
修复只删了一行:某条流以 complete: false 结束时,owner 不再调用 failCapture()。这样一来,finalize 上的 failed 标志就成了 worker 端写入失败传到 owner 的唯一途径。四个变异体检查这条路径:
| 变异体 | 结果 |
|---|---|
| K1:恢复被删掉的那一行 | 被新测试杀死(retains the accepted prefix when a producer holds a pipe past the drain boundary) |
K2:owner 的 finalize 忽略 worker 的 failed 标志 |
被杀死(irreversibly blocks capture after a lost raw write acknowledgement) |
K3:写入失败后 worker 仍报 complete: true |
存活;等价,因为 finalize 仍携带 failed(K2 + K3 同时施加会被杀死) |
K4:owner 的 finalize 不再让从未 finish 的流失败 |
存活;等价。handleExit 总是先 finish 两条流,finish 失败会设置 failed,由 K2 那一行处理 |
a4632e98 上的回归
| 检查 | 结果 |
|---|---|
| 在全新 MySQL 8.4.7 库上启动服务 | ✅ 17 个迁移:V16 runtime_loss_evidence、V17 managed_session_operation |
HostedWorkspaceToolTurnIT 整类,mvn -P hosted-workspace-tools verify |
✅ 3/3,97.6 s。六 Workspace driver;FG6a 8 个用例,含 status;FG6b 7 个用例 |
用 Shell 调用替换 edit 的 FG6b |
✅ 6/6 |
| Java 单元测试 | ✅ Broker 390(1 个跳过),managed-agent-server 155 |
| R1-4 / R1-7 参数校验、R1-1 孙进程、8 种输出形态、F4 Shell → 重新加载 → Shell | ✅ 未变 |
真实 qwen3.8-max:Shell,然后 detach + load,再 Shell |
✅ 两轮都完成;notes.txt = first\nsecond |
a4632e98 的 CI
- SDK Java: 8 个任务全绿,包括
20126bae上变红的 Runtime Broker and Managed Agent MariaDB 和 Hosted no-tool processes / MySQL 8.4。 - Serve A/B、tui-parity: 绿。
- Qwen Code CI: 全绿,包括 web-shell E2E smoke。
合并参考
- 正确性: 我看不到阻塞项。第六轮提出的、属于本 PR 范围的问题都已解决:部分前缀得到保留且标注准确,V16 撞号已在 main 上修复并合进来。
- 剩下的是第六轮提出的可用性问题。
- 放到后台的复合命令执行之后,Session 和整个 Workspace 都保持阻塞,也没有 API 能释放它们。
- 真实模型在 6 次试验中有 3 次写出了这种命令。
- 作者有意保持 fail-closed,已在 PR 说明中写明,并由 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 跟踪恢复。
- 从真实环境的角度看: 可以作为私有、fail-closed 的 profile 合并,前提是在 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 落地之前,
hosted-workspace-shell/1不接真实模型流量。
Evidence (probe scripts, per-run logs, IT summaries, the mutation log): wenshao/qwen-code@c250d377/pr12848/r7.
|
Follow-up to round 7 real-stack verification and the new
On |
Real-stack verification, round 8: #12848 at
|
Check at d650cdd3 |
Result |
|---|---|
| Server start on a fresh database | ✅ 18 migrations: V16 runtime-loss evidence, V17 Session operation, V18 extension record |
Partial capture read back through the Store (cd . && { 3,000,000 × a; sleep 20; } & sleep 5; echo done) |
✅ partial / producer_lost; 3,000,005 B = 3,000,000 a + done; sha256 = manifest |
| Other probes: the 3 B and 36 B prefixes, the complete control, the four minimal background shapes, Workspace hold, validation, R1-1, 8 output shapes, F4 reload, real-model reload | ✅ identical to round 7 |
| FG6b with a Shell call | ✅ 6/6 |
| CI | ✅ Qwen Code CI, SDK Java, Serve A/B and tui-parity. The Hosted process fault gates / MySQL 8.4 job ran HostedWorkspaceToolTurnIT 3/3, HostedHarnessMySqlIT 2/2 and HostedProcessCrashIT with all 6 FG6c cases |
Main's FG6c, re-run with a Shell call
FG6c uses edit on a FIFO so that the tool blocks mid-execution. For the Shell variant I used c=$(cat proof.txt) && [ "$c" = x ] && rm -f proof.txt && printf %s%s "$c" "$c" > proof.txt, which blocks the same way and has the same effect, and left the Java assertions untouched.
- Pass as is:
harness-prepareandspring-kill. worker-stop: passes once the driver accepts one thing. While the worker is stopped, a status poll gets 503managed_runtime_unavailable, because the v3 observation asks the worker (30 s timeout); withedit, the Broker answers from its record. The Harness doesn't retry thatretryable: true503, so the turn blocks right away rather than at the end of its observation window. Either way it ends blocked.worker-kill: the driver asserts a v2 409 code; Shell getsruntime_execution_evidence_unavailableand thenruntime_admission_closed. With those accepted, the Java ledger differs in one row: the binding isLOST, notREADY, because the v3 observation notices the dead worker, where the v2 path only reads the Broker's record.harness-result: the ledger and the journal-event checks pass. Then the assertion "nomanaged-tool-outcomeresource" fails. The Shell path commits arecordToolResulttransaction, holding the outcome and a reference to the capture manifest, before thetool_resultmessage, so a crash at the message commit leaves a committed outcome that theeditpath doesn't have yet. That is more durable state, not less, and the cold load is still blocked.harness-start: I also ran it on the rig, witheditas a control.
Harness SIGKILL while the tool blocks, then the tool finishes |
Shell | edit |
|---|---|---|
| Effect | xx, once |
xx, once |
| Broker execution, 2 s to 120 s later | UNKNOWN, no result |
SETTLED, result stored |
| Capture resources | none | n/a |
| Cold load of the Session | 409 hosted_turn_recovery_required |
same |
| New Session in the same Workspace | 409 workspace_busy |
same |
The capture publisher runs inside the Harness, so when the Harness dies the worker cannot finish the capture, and a command that did complete is recorded with an unknown outcome. That is fail-closed, and nothing replays, but it means #12904-style recovery can't learn the outcome of a Shell call interrupted this way. It is worth stating in that design.
Merge reference
- Unchanged from round 7: correctness shows no blocker, and the remaining decision is availability.
- After an incomplete capture, the Session and the Workspace stay held.
- As agreed in the author's latest comment,
hosted-workspace-shell/1stays private and serves no real-model traffic until Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 provides recovery.
- FG6c adds one input for Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904: a Harness crash during a Shell call leaves a completed command recorded as
UNKNOWN. - From the real-stack side,
d650cdd3is fine to merge on those terms.
中文版
真实环境验证(第八轮):#12848 @ d650cdd3
d650cdd3 把 main e3e7dbcd 合入 a4632e98,带进了:
- H0c Stage H 记录(feat(managed-agent): Commit Stage H records and serve the task list (H0c) #12855),新增迁移 V18;
- FG6c 进程崩溃门禁(test(managed-agent): add FG6c process crash gates #12896);
- codec 位数预算修复(fix(runtime-broker): enforce the codec reader's total-digit budget #12910);
- 以及 main 的其他改动。
涉及本 PR 文件的冲突只有一处,在 ManagedSessionStore(见作者说明)。
环境:MySQL 8.4.7;Spring jar 和 TS bundle 在 d650cdd3 上重新构建;真实本地进程 worker;固定脚本模型加真实 qwen3.8-max。
结论
- 合并在真实环境中是干净的。
- 服务带着 V16、V17、V18 正常启动。Java 测试全部通过:Broker 394(1 个跳过),managed-agent-server 172。
- H0c 新增的只接受
REFERENCED的查找,与本 PR 预先发布的 Shell 资源并存,互不影响。部分捕获和完整捕获仍能通过 Store 逐字节读回。 - 第七轮的每个 rig 场景结果都相同,CI 全绿。
- 把 main 新增的 FG6c 崩溃门禁换成 Shell 调用重跑,所有用例都没有重放,也没有重复的副作用。
- 两个用例原样通过。
- 其余用例与
edit的期望不同,但都能用 Shell/v3 的设计解释。 - 有一处差异值得为 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 记下来:Shell 命令执行期间 Harness 崩溃,已经完成的命令会被记为
UNKNOWN。
- 更正第六、七轮的说法。
- 我写过"FG6b 是手动门禁,没有 workflow 运行它",这是错的。
- CI 的 Hosted 任务运行
-Phosted-harness-mysql,它的 failsafe include 是**/Hosted*IT.java。所以从018f3318起,整个HostedWorkspaceToolTurnIT(含 FG6a 和 FG6b)一直在 CI 中运行,任务日志里能看到FG6B arguments: …。 - CI 没有覆盖的是 FG6b 和 FG6c 的 Shell 变体。
合并部分
- 冲突:
ManagedSessionStore两边都保留了。- PR 的
publishToolResult路径仍以PUBLISHED状态发布 Shell 结果。 - H0c 的
storedResource只接受REFERENCED资源。它唯一的调用方是commit里的 Stage H 扩展记录钩子,所以预先发布的工具结果不会经过它。这与作者的描述一致。
- PR 的
- 真实环境结果:
d650cdd3 上的检查 |
结果 |
|---|---|
| 在全新库上启动服务 | ✅ 18 个迁移:V16 运行时丢失证据、V17 Session 操作、V18 扩展记录 |
通过 Store 读回部分捕获(cd . && { 3,000,000 × a; sleep 20; } & sleep 5; echo done) |
✅ partial / producer_lost;3,000,005 B = 3,000,000 个 a 加 done;sha256 = manifest |
| 其余探针:3 B 与 36 B 前缀、完整捕获对照、四种最小后台写法、Workspace 占用、参数校验、R1-1、8 种输出形态、F4 重新加载、真实模型重新加载 | ✅ 与第七轮完全相同 |
| 换成 Shell 调用的 FG6b | ✅ 6/6 |
| CI | ✅ Qwen Code CI、SDK Java、Serve A/B、tui-parity。Hosted process fault gates / MySQL 8.4 任务跑了 HostedWorkspaceToolTurnIT 3/3、HostedHarnessMySqlIT 2/2,以及含 6 个 FG6c 用例的 HostedProcessCrashIT |
把 main 的 FG6c 换成 Shell 调用重跑
FG6c 让 edit 作用于一个 FIFO,使工具在执行中途阻塞。Shell 变体我用的是 c=$(cat proof.txt) && [ "$c" = x ] && rm -f proof.txt && printf %s%s "$c" "$c" > proof.txt,阻塞方式和副作用都与 edit 相同;Java 断言一行没改。
- 原样通过:
harness-prepare和spring-kill。 worker-stop: 驱动放宽一处后通过。worker 被停住时,状态轮询得到 503managed_runtime_unavailable,因为 v3 的观察要询问 worker(30 s 超时);用edit时 Broker 直接用自己的记录回答。Harness 不会重试这个带retryable: true的 503,所以这一轮立刻阻塞,而不是等到观察窗口结束才阻塞。两种情况最终都是阻塞。worker-kill: 驱动断言的是 v2 的 409 错误码;Shell 先得到runtime_execution_evidence_unavailable,后得到runtime_admission_closed。接受这两个码之后,Java ledger 只差一行:binding 是LOST而不是READY。原因是 v3 的观察会发现 worker 已死,而 v2 路径只读 Broker 自己的记录。harness-result: ledger 检查和 journal 事件检查都通过。接着"不存在managed-tool-outcome资源"这条断言失败。Shell 路径会在tool_result消息之前先提交一笔recordToolResult事务,其中包含 outcome 和对捕获 manifest 的引用;所以在消息提交处崩溃时,会留下一个已提交的 outcome,而edit路径此时还没有。这是更多的持久状态,而不是更少;冷加载仍然被阻塞。harness-start: 我也在 rig 上跑了一遍,用edit做对照。
工具阻塞时 SIGKILL Harness,随后工具完成 |
Shell | edit |
|---|---|---|
| 副作用 | xx,一次 |
xx,一次 |
| 2 s 到 120 s 后的 Broker 执行记录 | UNKNOWN,无结果 |
SETTLED,结果已保存 |
| 捕获资源 | 无 | 不适用 |
| 冷加载该 Session | 409 hosted_turn_recovery_required |
相同 |
| 同一 Workspace 的新 Session | 409 workspace_busy |
相同 |
捕获发布器运行在 Harness 进程里,所以 Harness 死后 worker 无法完成捕获,一条实际已经完成的命令会被记为结果未知。这是 fail-closed 的,也没有任何重放;但这意味着 #12904 那样的恢复,无法得知以这种方式中断的 Shell 调用的结果。值得在那份设计里写明。
合并参考
- 与第七轮相同: 正确性方面没有阻塞项,剩下需要决定的是可用性。
- 捕获不完整之后,Session 和 Workspace 仍然被占住。
- 按作者最新评论中的约定,在 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 提供恢复之前,
hosted-workspace-shell/1保持私有、不接真实模型流量。
- FG6c 为 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 增加了一个输入: Shell 调用期间 Harness 崩溃,会让一条已完成的命令被记为
UNKNOWN。 - 从真实环境的角度看,
d650cdd3可以按上述条件合并。
Evidence (probe scripts, the FG6c Shell driver diffs, per-run logs, IT summaries): wenshao/qwen-code@60e8cf13/pr12848/r8.
|
Follow-up to round 8 real-stack verification and the latest
On |
|
@qwen-code /triage |
Real-stack verification, round 9: #12848 at
|
Check at dfabe7c7 |
Result |
|---|---|
Conflict in HostedWorkspaceToolTurnIT |
✅ Keeps main's cancellationProbe.assertReport branch and this PR's index < 2 guard; the six-Workspace driver's Shell Workspaces at index ≥ 2 don't write after |
HostedWorkspaceToolTurnIT, whole class, MySQL 8.4.7 |
✅ 4/4 in 90.7 s: six Workspaces; FG6a 8 cases; FG6b 7 cases; FG6d prepared, running, status-unavailable, cancel-reply |
| CI (Hosted process fault gates / MySQL 8.4) | ✅ the same 4/4, plus HostedHarnessMySqlIT 2/2 and HostedProcessCrashIT (FG6c, 6 cases) |
| Rig: partial capture read back (3,000,005 B = manifest sha), Workspace hold and control, minimal background shapes, validation, R1-1, 8 output shapes, F4 reload, real-model reload | ✅ identical to round 8 |
FG6d's question for Shell: cancelling a running command
Each command ran in its own Session and Workspace. I cancelled it 2 s after the Broker reported EXECUTING, then waited 35 s for late effects.
| Command | After cancel | 35 s later | Afterwards |
|---|---|---|---|
echo started; sleep 30; echo late > late.txt |
idle in ~1.3 s; turn_complete(cancelled); execution SETTLED / cancelled; capture complete / committed, stdout 8 B |
no late.txt |
new Session in the Workspace works; detach + load 200 |
same, plus (while true; do echo tick >> ticks.txt; sleep 0.2; done) & |
same | ticks.txt stopped growing: the loop died with the group |
same |
cd . && sleep 60 >/dev/null 2>&1 & echo started; sleep 30; … |
same, in ~0.6 s: the pipe-holding subshell is in the group | nothing | same |
perl -MPOSIX -e 'POSIX::setsid(); sleep 5; …write escaped.txt…' >/dev/null 2>&1 & echo started; sleep 30; … |
same | escaped.txt written |
same |
So cancellation is a working exit even for the compound background shape, as long as the foreground command is still running. The Workspace hold from rounds 6 and 7 only happens once the foreground has already exited.
Input for #12904: detached descendants outlive the lease
- Session A runs
perl -MPOSIX -e 'POSIX::setsid(); sleep 6; …append "from-detached-descendant" to shared.txt…' >/dev/null 2>&1 & echo ok. The turn completes normally, the capture iscomplete, and the Workspace is released. - Session B, in the same Workspace, starts right after. It runs
echo from-next-session >> shared.txt; sleep 8; cat shared.txt, and its own output reads:
from-next-session
from-detached-descendant
The descendant left both the pipe (output redirected) and the process group (setsid). Neither the drain boundary nor the group kill sees it, so the capture is complete and the lease is released while it is still running.
- Existing wording: the PR body already says the profile "is not a filesystem sandbox", and the author's round-6 reply says pipe EOF and a group kill don't prove descendants stopped.
- What this adds:
- The same gap exists on the success path, not only on incomplete captures.
- So keeping the lease after an incomplete capture protects against pipe holders, but not against fully detached writers.
- Worth stating in Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 together with round 8's
UNKNOWN-after-Harness-crash case.
Merge reference
- Unchanged: correctness shows no blocker.
- The remaining decision is availability, plus now the scope of the Workspace-lease guarantee for Shell.
- As the author agreed,
hosted-workspace-shell/1stays private and serves no real-model traffic until Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904.
- From the real-stack side,
dfabe7c7is fine to merge on those terms.
中文版
真实环境验证(第九轮):#12848 @ dfabe7c7
dfabe7c7 把 main 92f4d4f6 合入 d650cdd3,带进了:
- FG6d Hosted 取消门禁(test(managed-agent): add FG6d hosted cancellation gates #12924);
- main 针对同一个 Hosted 测试等待的修复(test(cli): give the hosted tool-turn waitFor an explicit timeout (#12911) #12915);
- 与本 PR 无关的 CLI、core 和 VS Code 改动。
本 PR 的生产文件都没有变化,Java 生产代码与 d650cdd3 完全相同。两处冲突都在测试文件里(见作者说明)。
环境:MySQL 8.4.7;沿用第八轮的 Spring jar;TS bundle 在 dfabe7c7 上重新构建;真实本地进程 worker;固定脚本模型加真实 qwen3.8-max。
结论
-
合并是正确的。
HostedWorkspaceToolTurnIT既保留了 main 新增的取消分支,也保留了本 PR 的index < 2判断。- 该类在本地 MySQL 上 4/4 通过(90.7 s),CI 中同样通过,包括 FG6d 的 4 个用例。
- 第八轮的每个 rig 场景结果都相同,CI 全绿。
-
取消一个正在运行的 Hosted Shell 调用是有效的,而且很干净。 FG6d 用的是无法打断的
edit。在真实环境中,被取消的 Shell 调用会:- 在约 1.3 s 内以
cancelled结清,捕获完整; - 杀掉整个进程组,包括后台循环和占着管道的复合命令;
- 释放 Session 和 Workspace。
- 在约 1.3 s 内以
-
为 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 再提供一个输入:对 Shell 而言,Workspace 租约并不保证只有一个写入者。 一个同时脱离管道和进程组的后代(例如用
setsid并重定向输出),会在以下情况下继续写入:- 一轮正常完成之后,此时捕获完整、Workspace 已释放;
- 一次干净的取消之后。
有一次运行中,随后接手同一 Workspace 的第二个 Session,在自己的命令输出里读到了这个后代写入的内容。这与 PR 说明中"该 profile 不是文件系统沙箱"的描述一致;它也说明,捕获不完整时占住租约,覆盖不到这种情况。
合并部分
dfabe7c7 上的检查 |
结果 |
|---|---|
HostedWorkspaceToolTurnIT 的冲突 |
✅ 保留了 main 的 cancellationProbe.assertReport 分支和本 PR 的 index < 2 判断;六 Workspace driver 中索引 ≥ 2 的 Shell Workspace 不会写 after |
HostedWorkspaceToolTurnIT 整类,MySQL 8.4.7 |
✅ 4/4,90.7 s:六 Workspace;FG6a 8 个用例;FG6b 7 个用例;FG6d 的 prepared、running、status-unavailable、cancel-reply |
| CI(Hosted process fault gates / MySQL 8.4) | ✅ 同样 4/4,另有 HostedHarnessMySqlIT 2/2 和 HostedProcessCrashIT(FG6c,6 个用例) |
| rig:部分捕获读回(3,000,005 B = manifest sha)、Workspace 占用与对照、最小后台写法、参数校验、R1-1、8 种输出形态、F4 重新加载、真实模型重新加载 | ✅ 与第八轮完全相同 |
FG6d 留给 Shell 的问题:取消一个正在运行的命令
每条命令都在各自的 Session 和 Workspace 中运行。Broker 报告 EXECUTING 2 s 后取消,然后再等 35 s,看是否有迟到的副作用。
| 命令 | 取消之后 | 35 s 之后 | 随后 |
|---|---|---|---|
echo started; sleep 30; echo late > late.txt |
约 1.3 s 后空闲;turn_complete(cancelled);执行记录 SETTLED / cancelled;捕获 complete / committed,stdout 8 B |
没有 late.txt |
Workspace 中的新 Session 可用;detach + load 返回 200 |
同上,加上 (while true; do echo tick >> ticks.txt; sleep 0.2; done) & |
同上 | ticks.txt 停止增长:循环随进程组一起结束 |
同上 |
cd . && sleep 60 >/dev/null 2>&1 & echo started; sleep 30; … |
同上,约 0.6 s:占着管道的子 shell 也在进程组里 | 无 | 同上 |
perl -MPOSIX -e 'POSIX::setsid(); sleep 5; …写 escaped.txt…' >/dev/null 2>&1 & echo started; sleep 30; … |
同上 | 写出了 escaped.txt |
同上 |
所以只要前台命令还在运行,取消就是一条有效的出路,对复合后台写法也一样。第六、七轮的 Workspace 占用,只在前台命令已经退出之后才会出现。
为 #12904 提供的输入:脱离的后代比租约活得更久
- Session A 运行
perl -MPOSIX -e 'POSIX::setsid(); sleep 6; …向 shared.txt 追加 "from-detached-descendant"…' >/dev/null 2>&1 & echo ok。这一轮正常完成,捕获为complete,Workspace 被释放。 - Session B 紧接着在同一个 Workspace 中启动,运行
echo from-next-session >> shared.txt; sleep 8; cat shared.txt。它自己的输出是:
from-next-session
from-detached-descendant
这个后代同时脱离了管道(输出已重定向)和进程组(setsid)。排空边界和进程组终止都察觉不到它,所以它还在运行时,捕获就已完整、租约就已释放。
- 已有的表述: PR 说明已经写明该 profile "不是文件系统沙箱",作者在第六轮的回复中也指出,管道 EOF 和杀进程组都不能证明后代已经停止。
- 这次补充的是:
- 同样的缺口在成功路径上也存在,并不只出现在捕获不完整的时候。
- 所以,捕获不完整时保留租约,能防住占着管道的进程,但防不住完全脱离的写入者。
- 值得和第八轮"Harness 崩溃后记为
UNKNOWN"那条一起写进 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904。
合并参考
- 未变: 正确性方面没有阻塞项。
- 剩下需要决定的是可用性,现在再加上 Shell 场景下 Workspace 租约保证的范围。
- 按作者的约定,在 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 之前,
hosted-workspace-shell/1保持私有、不接真实模型流量。
- 从真实环境的角度看,
dfabe7c7可以按上述条件合并。
Evidence (probe scripts, per-run logs, the IT summary): wenshao/qwen-code@88dd8cfd/pr12848/r9.
qqqys
left a comment
There was a problem hiding this comment.
Independent Critical-only review — head dfabe7c7c7a0c9955ac9fceb9890f23f6132de09
Verdict: COMMENT. Both Criticals previously filed against this PR are confirmed fixed in code at this head, and I found no new provable Critical in the part of the diff I read. What blocks an approval is coverage, not a finding: this is 46 files and +5193/−516 adding remote foreground Shell execution, and the capture-session, resource-store and Java production layers remained outside what I could read inside this review's budget. I am naming that boundary rather than implying coverage I did not have.
Historical blocking issues — both verified FIXED at this head
The CHANGES_REQUESTED at a4c35e89 carried two sev:C findings, both in packages/cli/src/serve/managed-shell-publisher.ts. Thread replies and the later approval do not establish a fix, so I read the current file.
R1-1 (late write after finish → blocked receipt and hard turn failure on a command that exited 0) — FIXED. RemoteShellCapture now keeps per-stream ended state and a per-stream FIFO queue:
private readonly ended = { stdout: false, stderr: false };private readonly queues = { stdout: Promise.resolve(), stderr: Promise.resolve() };write()andfinish()both chain ontothis.queues[stream], so awriteRPC can no longer be issued after that stream'sfinishRPC.append()returns immediately onif (this.failed || this.ended[stream]) return;— the same guard the in-process sink uses, so a latedataevent is dropped instead of uploaded.end()setsended[stream] = truebefore the RPC and is idempotent.finalize()starts withawait Promise.all(Object.values(this.queues));, so finalization cannot race a still-queued stream.
The reported chain (owner sees entry.ended[stream], calls sink.failCapture(), receipt becomes blocked, turn throws Complete Shell output was not admitted.) is no longer reachable, because the client never emits the out-of-order write.
R1-2 (process reported without setStarted → owner rejects the pair → turn destroyed instead of not_started) — FIXED. finalize now gates the physical result on started and reports not_started, exactly as prescribed:
process:
this.started && physical
? { exitCode: physical.exitCode, signal: physical.signal,
previewBytes: physical.rawOutput.byteLength }
: null,
executionStatus: this.started ? executionStatus : 'not_started',A failed spawn or pre-spawn abort therefore no longer produces the {started: false, process: {...}} pair that hosted-shell-publisher.ts rejects with Unstarted Shell has a physical result. (that owner-side guard is still present, which is correct — the client now never sends the pair it rejects).
The deferred sev:S item on preview derivation is also addressed: shellPreview computes bytes once from the joined text and returns { parts, truncated }, and finalize reads both from that single call, so previewTruncated and the preview it describes can no longer diverge. Non-blocking either way.
What I verified myself at this head
Publisher authentication and binding. hosted-shell-publisher.ts mints randomBytes(32).toString('base64url'), compares the presented token with timingSafeEqual behind an explicit length check, answers 401 on mismatch, and binds this.server.listen(0, '127.0.0.1'). Route body limits are declared (SHELL_PUBLISHER_BODY_LIMIT = 256 KiB, CHUNK_BYTES = 64 KiB, 16 KiB on the managed publisher route) with cache-control: no-store.
Descriptor validation is canonical-loopback-only. parseShellPublisher requires an exact field set (url, token), a token matching /^[A-Za-z0-9_-]{43}$/, and a URL matching /^http:\/\/127\.0\.0\.1:([1-9][0-9]{0,4})\/internal\/hosted-shell-publisher\/v1$/ with the port additionally range-checked. localhost, IPv6 literals, other schemes and any path variation are all rejected before the descriptor is usable.
The gate stays narrow. The Shell profile is a second named profile beside hosted-workspace-files/1; the shell wiring is conditional on it, so default and file-profile hosts are unchanged.
CI. Every check passes at this head — Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), Serve A/B, the full Java matrix, Runtime Broker and Managed Agent MariaDB / Java 21, Hosted process fault gates / MySQL 8.4 / Java 21, Real daemon E2E / Java 11, both Desktop Shell lanes and the TUI parity/no-flicker gates. Only review-pr is pending, which is not a gating check. I saw no failure attributable to this PR.
Gates I could not confirm inside budget
These are the unconfirmed items, listed so the next pass can start from them. None is an allegation:
packages/core/src/managed-runtime/resource-tool-result-store.ts(new, 366 lines) — writer-fenced publication, immutable resource ids and seal verification. The corresponding test file's fixture was separately noted as always returning the ref the store derives, so the publication-receipt check may be unwitnessed.packages/core/src/managed-runtime/managed-shell-result-session.ts(new, 436 lines) against the 393 lines removed fromlocal-shell-result-session.ts— I did not establish that the extraction preserves the existing local behaviour.- The Java production half —
ManagedSessionStore(+68),WorkspaceRuntimeTransport(+97),RuntimeBrokerHttpServer(+72/−10),HttpRuntimeTransport(+23),ManagedSessionStoreController(+10) and the twoToolExecutionRepositoryedits. In particular the ownership scope of the new publish route and its per-kind byte limits, which were flagged as being declared in Java while the client enforces its own. - The
rpchelper's transport options inmanaged-shell-publisher.ts— I read the descriptor validation but not the fetch call itself, so I am not asserting that redirects are rejected at the fetch layer. integration-tests/helpers/hosted-workspace-tool-turn-driver.ts(+343/−11) — the two deferred notes about the store proxy's capture assertions and the session-scoped (not turn-scoped) durable-admission ordering check on a never-reset accumulator.
Per this channel's policy I do not track the remaining deferred Suggestions; they do not gate.
Next step: a pass over items 1–3 above would close this. Items 1 and 3 are the ones I would read first, since they hold the durability fencing and the route-ownership guarantees for a path that executes commands on a saved workspace.
wenshao
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed.
7 Suggestion-level finding(s) this review confirmed are already reported on this PR and are not repeated:
- R2-1 worker publisher registry never evicts its per-turn
sessions/executionsentries — already reported (comment 5863868679, Finding 1) - R2-4 integration-tests/helpers/hosted-workspace-tool-turn-driver.ts:278 durable-admission ordering check is session-scoped on a never-reset accumulator — already reported (review 5333665604 deferral list)
- R2-5 packages/cli/src/serve/hosted-harness-session.test.ts:140 tool-set assertion thrown inside the mocked model — already reported (review 5333665604 deferral list)
- R2-8 packages/cli/src/serve/hosted-workspace-tool-turn.ts:178 admitted Shell keys hand-listed apart from the declared schema — already reported (comment 5863868679, Finding 3)
- R2-12 packages/core/src/managed-runtime/resource-tool-result-store.test.ts:40 publication-receipt guard unwitnessed — already reported (review 5333665604 deferral list)
- R2-18 packages/cli/src/serve/hosted-harness-session.ts:622 publisher-cleanup catch untested — already reported (review 5333665604 deferral list)
- R2-22 packages/cli/src/serve/hosted-harness-session.ts:332 shell wiring covered only through a diverged store mock — already reported (review 5333665604 deferral list)
Unresolved, please confirm:
- [Critical] issue comments 5864594581 (real-stack round 6), 5867309104 (round 7), 5857539063 and 5863868679 (sandboxed verification) — ruled from their Verdict, Previous-finding status and Findings sections plus the open-defect sections, not read end t…
Not reviewed: reverse audit — stopped after the convergence pair (rounds 1 and 2, both reporting: 31 of 46 auditors returned findings) without reaching two consecutive dry rounds; the operator chose to stop the loop rather than run to the plan's 5-round cap, so the audit is not converged.
Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally (the npm integration-tests lane needs the CLI bundle, which this host cannot build: unaccepted Xcode license aborts audio-capture).
Not reviewed: build-and-test — HostedWorkspaceToolTurnIT and the Java failsafe integration lane did not run locally (need -Dnode.executable, a built dist/cli.js, Java 21 and a packaged harness); the Java findings rest on unit-level mvn runs and static traces.
Not explored to full depth (tool budget reached): "agent reverse-audit (round 1)": enumerating every RuntimeResourceHandle construction site to confirm no handle value map can carry a null — that is the only case in which the new WriteNulls…; "agent reverse-audit (round 1)": the JIT static tier ( javac -proc:none + javap -c -p bytecode sizes) for the methods this diff grows — startExecution , prepareExecution , execution , and…; "agent reverse-audit (round 2)": whether the runtime-side publisher route ( managed-runtime-tool-v3-routes.ts / ManagedShellPublisherRegistry.register ) independently validates the descriptor…; "agent reverse-audit (round 2)": whether HttpRuntimeTransport.execute (3-arg) routes a runtimeProtocol: 3 reference to executeV3 , which is the remaining premise of my "fail-closed" conclu…; "agent test-matrix": none — all 23 assigned diff pages were read in full, and every pairing claim above was checked against the worktree at HEAD., and 5 more.
Mechanism health: this round did not close cleanly, so it withholds the incremental anchor — and the round it recovered had no anchor this round could use either — none at all, one with no certifier, one certified by an identity other than the one this round runs under, or one this round's fetch refused or resolved to the head — so the next review re-reads the whole diff unless recovery grafts an earlier own anchor that the round running it can use onto the complete work list this round leaves behind, and keeps doing so until a round's marker carries an anchor again or a graft lands that the round running it can use. (Stated, not acted on — this changes nothing about what the round posts.)
中文说明
仅完成部分审查,审查缺口已披露。
本轮确认的 7 条建议级发现已在 PR 上报告过,不再重复发布(列表见上方英文部分)。
未决,请确认:共 1 条(原文未翻译,列表见上方英文部分)。
未审查(原文为英文):reverse audit — stopped after the convergence pair (rounds 1 and 2, both reporting: 31 of 46 auditors returned findings) without reaching two consecutive dry rounds; the operator chose to stop the loop rather than run to the plan's 5-round cap, so the audit is not converged.
未审查(原文为英文):build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally (the npm integration-tests lane needs the CLI bundle, which this host cannot build: unaccepted Xcode license aborts audio-capture).
未审查(原文为英文):build-and-test — HostedWorkspaceToolTurnIT and the Java failsafe integration lane did not run locally (need -Dnode.executable, a built dist/cli.js, Java 21 and a packaged harness); the Java findings rest on unit-level mvn runs and static traces.
未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 1)":enumerating every RuntimeResourceHandle construction site to confirm no handle value map can carry a null — that is the only case in which the new WriteNulls…;"agent reverse-audit (round 1)":the JIT static tier ( javac -proc:none + javap -c -p bytecode sizes) for the methods this diff grows — startExecution , prepareExecution , execution , and…;"agent reverse-audit (round 2)":whether the runtime-side publisher route ( managed-runtime-tool-v3-routes.ts / ManagedShellPublisherRegistry.register ) independently validates the descriptor…;"agent reverse-audit (round 2)":whether HttpRuntimeTransport.execute (3-arg) routes a runtimeProtocol: 3 reference to executeV3 , which is the remaining premise of my "fail-closed" conclu…;"agent test-matrix":none — all 23 assigned diff pages were read in full, and every pairing claim above was checked against the worktree at HEAD.,另有 5 条。
机制健康:本轮未能干净收尾,因而扣留了增量锚点,而它恢复到的那一轮也没有留下本轮可用的锚点——要么完全没有、要么没有认证者、要么由本轮运行身份之外的身份认证、要么被本轮的获取拒绝或解析为头提交——因此下一次评审将重读整个 diff,除非恢复流程把本轮能使用的更早自有锚点嫁接到本轮留下的完整工作清单上;并会一直如此,直到某一轮的标记重新带上锚点,或落地的嫁接能被运行该轮的评审使用。(仅陈述,不据此行动——这不改变本轮发布的任何内容。)
— qwen3.8-max via Qwen Code /review (v0.24.6)
Round-2 review disposition — all 39 threads handledEvery unresolved round-2 thread was re-verified against current
The verification evidence and provenance for each item is in its thread reply; #13307 groups them by package. The two comments the round-2 review flagged as "Unresolved, please confirm" are closed out as follows: the duplicate-V16 startup failure from real-stack round 6 was fixed on For the record: getting these fixes here surfaced one false start worth naming — a full-suite run on the rebased branch showed a single failure whose identity was lost to output truncation; two subsequent identical full runs passed 381/381 each, so the suites stand green. |









What this PR does
Adds foreground Shell turns to the private Hosted Workspace loop under the explicit
hosted-workspace-shell/1profile. It includes the existing file tools, runs commands in the saved Workspace, retains complete stdout/stderr in the SQL Session Store, and gives the model a bounded preview. Default no-tool and file-only profiles keep their existing behavior.The original invocation and current-turn checkpoint are durable before dispatch. The Session owner verifies the captured bytes and commits a receipt before history, acknowledgement and model continuation. Lost start responses query the original execution; incomplete capture or uncertain effects block continuation. Raw pipe writes are serialized per stream, and activation replacement is fenced during preparation and receipt admission. Captured previews do not advertise inaccessible worker-local files as complete output.
This PR now targets
mainand builds on merged #12831 and O1c #12821.Why it's needed
The Hosted file-tool loop could not use O1c's local Session publisher across a worker process. Its ordinary HTTP resource path stages small resources until a journal commit, which cannot acknowledge durable streaming output. This bridge adds bounded, immutable, writer-fenced output publication and connects the production Broker to the explicit Tool v3 path.
Reviewer Test Plan
How to verify
Evidence (Before & After)
Before: the installed CLI 0.24.6 rejects the private Hosted startup path; this is a startup baseline, not a Shell execution test. The parent file-only profile separately refuses Shell.
After: the packaged Harness, production Java Broker, separate workers and HTTP SQL Store pass a six-Workspace fixture covering file tools, complete 100 MiB Shell output, lost responses, storage failure, cancellation and retained reads after producer shutdown. Independent activation-race probes and the eight default Hosted process tests pass.
Build, typecheck, bundle, formatting, lint and focused tests pass: 44 core tests, three selected Shell truncation regressions, 899 CLI tests, and the focused Java Broker/Store suites. Detailed commands and evidence limits are in the separate E2E report.
Tested on
Environment
Node.js 22.14.0, Java 21, real filesystem/process execution, deterministic loopback OpenAI fixture, Spring HTTP Session Store and H2 in MySQL mode with Flyway migrations. No real model provider was used.
Risk & Scope
Design: English · 简体中文. Both versions cover the same decisions, limits, ownership, acceptance criteria and follow-ups.
Linked Issues
Part of #12380. Builds on merged #12831 and #12821. The duplicate migration fix #12900 is merged and included in this branch. Follow-ups are Workspace recovery #12904, absolute file-path refusal #12905, #12766 and #12670.
中文说明
本 PR 做了什么
通过显式
hosted-workspace-shell/1profile,为私有 Hosted Workspace 循环接入前台 Shell。包含现有文件工具,在保存的 Workspace 中执行命令,将完整 stdout/stderr 保存在 SQL Session Store,并向模型提供有界预览。默认无工具与仅文件工具 profile 保持原行为。派发前持久保存原始调用和当前回合检查点。Session owner 核验捕获字节并提交回执,然后才提交历史、确认结果并继续模型推理。start 响应丢失时查询原执行;捕获不完整或副作用不确定时阻塞后续执行。原始管道写入按流串行,参数准备和回执接纳期间都防止旧 activation 越权。捕获预览不会把不可访问的 worker 本地文件冒充完整输出。
本 PR 现以
main为目标分支,基于已合并的 #12831 和 O1c #12821。为什么需要
Hosted 文件工具循环无法跨 worker 进程使用 O1c 的本地 Session 发布器。其普通 HTTP 资源路径仅暂存小资源,直到 journal 事务才提交,无法确认流式输出已经持久保存。本桥接新增有界、不可变且受 writer fencing 保护的输出发布,并将生产 Broker 接到显式 Tool v3 路径。
评审验证计划
如何验证
前后证据
变更前:已安装的 CLI 0.24.6 拒绝私有 Hosted 启动路径;这是启动基线,不是 Shell 执行测试。父 PR 的仅文件工具 profile 另行验证了拒绝 Shell。
变更后:打包 Harness、生产 Java Broker、独立 worker 和 HTTP SQL Store 通过六 Workspace fixture,覆盖文件工具、完整 100 MiB Shell 输出、响应丢失、存储失败、取消和生产者退出后的保留输出读取。独立 activation 竞态探针和八项默认 Hosted 进程测试通过。
构建、类型检查、打包、格式、lint 与定向测试通过:44 项 core 测试、三项选定的 Shell 截断回归、899 项 CLI 测试,以及定向 Java Broker/Store 测试。详细命令和证据边界见单独的 E2E 报告。
测试平台
环境
Node.js 22.14.0、Java 21、真实文件系统和进程执行、确定性的本地 OpenAI fixture、Spring HTTP Session Store,以及启用 MySQL 模式和 Flyway 迁移的 H2。没有调用真实模型服务。
风险与范围
设计:English · 简体中文。两版决策、限制、所有权、验收标准和后续事项一致。
关联事项
属于 #12380。基于已合并的 #12831 和 #12821。重复迁移修复 #12900 已合入并同步到本分支。后续工作包括 Workspace 恢复 #12904、绝对文件路径的可纠正拒绝 #12905、#12766 和 #12670。