Skip to content

feat(serve): add gated Hosted foreground Shell turns - #12848

Merged
doudouOUC merged 15 commits into
mainfrom
codex/hosted-shell-tool-turn
Sep 28, 2026
Merged

doudouOUC merged 15 commits into
mainfrom
codex/hosted-shell-tool-turn

Conversation

@doudouOUC

@doudouOUC doudouOUC commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Adds foreground Shell turns to the private Hosted Workspace loop under the explicit hosted-workspace-shell/1 profile. It includes the existing file tools, runs commands in the saved Workspace, retains complete stdout/stderr in the SQL Session Store, and gives the model a bounded preview. Default no-tool and file-only profiles keep their existing behavior.

The original invocation and current-turn checkpoint are durable before dispatch. The Session owner verifies the captured bytes and commits a receipt before history, acknowledgement and model continuation. Lost start responses query the original execution; incomplete capture or uncertain effects block continuation. Raw pipe writes are serialized per stream, and activation replacement is fenced during preparation and receipt admission. Captured previews do not advertise inaccessible worker-local files as complete output.

This PR now targets main and builds on merged #12831 and O1c #12821.

Why it's needed

The Hosted file-tool loop could not use O1c's local Session publisher across a worker process. Its ordinary HTTP resource path stages small resources until a journal commit, which cannot acknowledge durable streaming output. This bridge adds bounded, immutable, writer-fenced output publication and connects the production Broker to the explicit Tool v3 path.

Reviewer Test Plan

How to verify

  • Select the Shell profile for a saved Workspace and keep a decoy directory under the Harness. A foreground command must affect the selected Workspace exactly once; default/file-only Sessions must continue to refuse Shell.
  • Produce 100 MiB stdout with invalid UTF-8 bytes, a separate stderr stream and known tails. The model must receive a bounded, explicitly truncated preview only after a durable receipt. After the Harness and workers stop, a fresh reader must reproduce the original stream lengths, hashes, tails and execution identity.
  • Drop a start response, drop a raw-write acknowledgement, and fail SQL content publication. Start-response loss may recover by observing the same invocation; uncertain raw output or failed persistence must block the next model request without repeating the side effect.
  • Cancel an executing command. Confirm physical settlement and capture admission before a cancelled turn completes, then verify a new text turn works. Replace the owner activation during parameter reads and result verification; the old owner must not prepare or commit a receipt.

Evidence (Before & After)

Before: the installed CLI 0.24.6 rejects the private Hosted startup path; this is a startup baseline, not a Shell execution test. The parent file-only profile separately refuses Shell.

After: the packaged Harness, production Java Broker, separate workers and HTTP SQL Store pass a six-Workspace fixture covering file tools, complete 100 MiB Shell output, lost responses, storage failure, cancellation and retained reads after producer shutdown. Independent activation-race probes and the eight default Hosted process tests pass.

Build, typecheck, bundle, formatting, lint and focused tests pass: 44 core tests, three selected Shell truncation regressions, 899 CLI tests, and the focused Java Broker/Store suites. Detailed commands and evidence limits are in the separate E2E report.

Tested on

OS Status
🍏 macOS ✅ tested
🪟 Windows ⚠️ not tested
🐧 Linux ⚠️ not tested

Environment

Node.js 22.14.0, Java 21, real filesystem/process execution, deterministic loopback OpenAI fixture, Spring HTTP Session Store and H2 in MySQL mode with Flyway migrations. No real model provider was used.

Risk & Scope

  • Main risk or tradeoff: this crosses core, worker, Broker and SQL durability boundaries and needs maintainer review. The existing trusted, preapproved Workspace execution profile is not a filesystem sandbox. Output remains retained with the Session, including abandoned publications; per-Session cleanup and storage quotas are follow-up work. A backgrounded compound command can leave incomplete pipe capture and hold the Workspace execution lease indefinitely; this explicit private profile remains fail-closed, and safe lease recovery is tracked in Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 before wider use.
  • Not validated / out of scope: real MySQL, remote provisioners, PTY/background jobs, public Artifact download/UI, object storage and automatic restart/replay of uncertain executions. The first publisher bridge requires the existing same-host process provisioner.
  • Breaking changes / migration notes: no database migration or default capability enablement. Ordinary Shell truncation remains unchanged; the captured path uses its own bounded preview and durable output. Existing Tool v2 execution remains supported.

Design: English · 简体中文. Both versions cover the same decisions, limits, ownership, acceptance criteria and follow-ups.

Linked Issues

Part of #12380. Builds on merged #12831 and #12821. The duplicate migration fix #12900 is merged and included in this branch. Follow-ups are Workspace recovery #12904, absolute file-path refusal #12905, #12766 and #12670.

中文说明

本 PR 做了什么

通过显式 hosted-workspace-shell/1 profile,为私有 Hosted Workspace 循环接入前台 Shell。包含现有文件工具,在保存的 Workspace 中执行命令,将完整 stdout/stderr 保存在 SQL Session Store,并向模型提供有界预览。默认无工具与仅文件工具 profile 保持原行为。

派发前持久保存原始调用和当前回合检查点。Session owner 核验捕获字节并提交回执,然后才提交历史、确认结果并继续模型推理。start 响应丢失时查询原执行;捕获不完整或副作用不确定时阻塞后续执行。原始管道写入按流串行,参数准备和回执接纳期间都防止旧 activation 越权。捕获预览不会把不可访问的 worker 本地文件冒充完整输出。

本 PR 现以 main 为目标分支,基于已合并的 #12831 和 O1c #12821。

为什么需要

Hosted 文件工具循环无法跨 worker 进程使用 O1c 的本地 Session 发布器。其普通 HTTP 资源路径仅暂存小资源,直到 journal 事务才提交,无法确认流式输出已经持久保存。本桥接新增有界、不可变且受 writer fencing 保护的输出发布,并将生产 Broker 接到显式 Tool v3 路径。

评审验证计划

如何验证

  • 对保存的 Workspace 选择 Shell profile,并在 Harness 下保留诱饵目录。前台命令应恰好影响所选 Workspace 一次;默认和仅文件工具 Session 应继续拒绝 Shell。
  • 生成含非法 UTF-8 字节的 100 MiB stdout、独立 stderr 和已知尾部。持久回执提交后,模型才能收到有界且明确标注截断的预览。Harness 和 worker 退出后,新的 reader 应复原原始流长度、摘要、尾部和执行身份。
  • 分别丢弃 start 响应、原始 write 确认,以及让 SQL content 发布失败。start 响应丢失可通过观察同一调用恢复;原始输出不确定或持久化失败时,必须阻止下一次模型请求且不得重复副作用。
  • 取消执行中的命令。确认物理执行结算并接纳捕获后,回合才可结束为取消,然后验证新文本回合可用。在参数读取和结果核验期间替换 owner activation;旧 owner 不得完成准备或提交回执。

前后证据

变更前:已安装的 CLI 0.24.6 拒绝私有 Hosted 启动路径;这是启动基线,不是 Shell 执行测试。父 PR 的仅文件工具 profile 另行验证了拒绝 Shell。

变更后:打包 Harness、生产 Java Broker、独立 worker 和 HTTP SQL Store 通过六 Workspace fixture,覆盖文件工具、完整 100 MiB Shell 输出、响应丢失、存储失败、取消和生产者退出后的保留输出读取。独立 activation 竞态探针和八项默认 Hosted 进程测试通过。

构建、类型检查、打包、格式、lint 与定向测试通过:44 项 core 测试、三项选定的 Shell 截断回归、899 项 CLI 测试,以及定向 Java Broker/Store 测试。详细命令和证据边界见单独的 E2E 报告。

测试平台

OS 状态
🍏 macOS ✅ 已测试
🪟 Windows ⚠️ 未测试
🐧 Linux ⚠️ 未测试

环境

Node.js 22.14.0、Java 21、真实文件系统和进程执行、确定性的本地 OpenAI fixture、Spring HTTP Session Store,以及启用 MySQL 模式和 Flyway 迁移的 H2。没有调用真实模型服务。

风险与范围

  • 主要风险或取舍:跨越 core、worker、Broker 和 SQL 持久化边界,需要维护者评审。现有可信、预批准的 Workspace 执行 profile 不构成文件系统沙箱。输出随 Session 保留,包括遗留发布;Session 清理与存储配额属于后续工作。后台复合命令可能使管道捕获不完整,并无限期占住 Workspace 执行租约;该显式私有 profile 继续在故障时阻断,更大范围使用前的安全租约恢复由 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 跟踪。
  • 未验证或范围外:真实 MySQL、远端 provisioner、PTY/后台任务、公开 Artifact 下载/UI、对象存储,以及不确定执行的自动重启/重放。首个发布器桥接要求使用现有同主机进程 provisioner。
  • 兼容性与迁移:无数据库迁移,不默认启用能力。普通 Shell 截断保持原行为;捕获路径使用自己的有界预览和持久输出。继续支持现有 Tool v2 执行。

设计:English · 简体中文。两版决策、限制、所有权、验收标准和后续事项一致。

关联事项

属于 #12380。基于已合并的 #12831 和 #12821。重复迁移修复 #12900 已合入并同步到本分支。后续工作包括 Workspace 恢复 #12904、绝对文件路径的可纠正拒绝 #12905、#12766 和 #12670。

@doudouOUC

doudouOUC commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator Author

E2E verification report

PASS on macOS, Node.js 22.14.0 and Java 21. The final run uses the packaged CLI, production Java Broker, separate Workspace workers, Spring HTTP Session Store and H2 in MySQL mode with Flyway. The model is a deterministic loopback OpenAI fixture; no real provider is involved.

Group Result
Real six-Workspace integration 1/1, no skips, 37.33 s
Java Workspace routing 12/12, no skips, 0.720 s
SQL Session Store integration 3/3, no skips, 2.260 s
Fresh clean-installed Java Broker focused tests 189/189, no skips
Default Hosted no-tool process regression 8/8, 8.68 s
Core HTTP store, segment store, capture/admission and checkpoint tests 44/44
Selected raw-capture and ordinary Shell truncation regressions 3/3; other 358 Shell tests intentionally excluded by the filter
Hosted/publisher/worker CLI tests 899/899 across nine files
Original activation-replacement reproduction scripts Both verified fixed

The real integration runs file-only and Shell Sessions in six distinct Workspaces. It checks a Harness decoy, complete 100 MiB stdout (including invalid UTF-8 bytes), separate stderr, independently computed lengths/SHA-256/tails, the original Session/turn/execution identity and durable receipt admission before the next model request. The model sees an explicit bounded-preview notice and no recommendation to read an inaccessible worker output file. After the driver exits, the fixture captures the live producer processes, closes the Broker, waits for each process to exit and checks that none is alive. The final run confirms HOSTED_SHELL_PRODUCERS_EXITED: 6 before starting a separate reader process that verifies the retained SQL output.

Fault cases drop a start response, reject content publication with an injected HTTP 503 at the SQL-backed Store API, and drop the first successful raw-write response. This is a publication-API failure test, not an injected database-engine failure. Each effect marker remains exactly x, each execution starts once, failed capture blocks further inference, and raw writes are not retried. Physical cancellation settles and permits a subsequent text turn. The two original race scripts replace the owner during parameter reading and output verification; stale preparation and stale receipt admission both reject.

Reproduction

Build the repository and run npm run typecheck and npm run bundle. With Java 21, Maven and Node 22 available, run from the repository root:

mvn -q -f packages/sdk-java/runtime-broker/pom.xml -Dnode.executable="$(command -v node)" -Dtest=RuntimeBrokerServiceTest,RuntimeBrokerHttpServerTest,HttpRuntimeTransportTest,InMemoryRepositoryTest,JdbcRepositoryTest clean install
mvn -q -f packages/sdk-java/managed-agent-server/pom.xml -Phosted-workspace-tools -Dnode.executable="$(command -v node)" -Dtest=WorkspaceRuntimeTest,ManagedSessionStoreIntegrationTest verify

The unchanged default process regression passed on the same product source before the test-only shutdown assertion change; it was not redundantly rerun. To reproduce it, run from integration-tests:

npx vitest run --config vitest.hosted.config.ts

The global CLI 0.24.6 baseline rejects the private Hosted startup path before any model request. This is recorded as an unavailable baseline, not a successful Shell test. The existing file-only profile's Shell refusal is also verified in the real fixture.

Build, typecheck, bundle and focused lint/format checks passed on the unchanged product source. The post-audit verification clean-built and installed the current Broker, reran the stricter real-process test plus Workspace/Store tests, and passed Checkstyle for both Java modules. Test processes and listeners were checked after completion; none remained. Earlier failures exposed and led to fixes for turn checkpoint identity, stale-owner preparation/admission, overlapping pipe writes and misleading temporary-preview-file guidance. The final test assertions distinguish provider preview text from durable receipt metadata.

Limits: H2 is not a real MySQL deployment. Producer shutdown is not a database restart test. Windows, Linux, remote provisioners, real providers, public Artifact APIs and automatic recovery/replay were not validated here.

Additional full-diff audits before review

Three additional open-ended passes covered all 46 changed files and their consumers. The first pass identified a verification gap: requesting Broker shutdown did not prove that worker processes had exited before retained reading. Commit 7a41a5d42 adds the explicit exit barrier. The refreshed integration passes with six producer exits, followed by successful independent output reads. Two subsequent complete passes over the same frozen source found no additional confirmed defects or unresolved suspicions. No product code changed during these additional passes.

@doudouOUC
doudouOUC marked this pull request as ready for review September 27, 2026 12:32
@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

Real-stack verification: #12848 Hosted foreground Shell (head 7a41a5d4)

I tested 782c1774. The only later commit, 7a41a5d4, adds 7 lines to HostedWorkspaceToolTurnIT.java, so the production code is identical. This complements the triage review. That review deliberately did not execute the PR's code and asked for someone to run the recovery and retention claims. Everything below was actually executed.

Verdict

  • The durability contract holds on real MySQL 8.4.7. The author used H2 in MySQL mode. The PR's own six-Workspace driver passes unchanged. The retained 100 MiB output reads back byte-identical after Spring and mysqld restarts. Every Shell side effect runs exactly once. Cancellation, timeouts, 300 MiB of output and commands longer than 30 s all settle cleanly. I found no correctness blocker in the tested code.
  • It cannot merge as-is, for two process reasons:
    1. It is now CONFLICTING with its base. feat(serve): add gated Hosted Workspace file tool turns #12831 moved to c6c76750, and the conflicts are in hosted-workspace-tool-turn.ts and hosted-harness-session.ts, which is exactly the Shell result-commit path. The resolved code needs a re-run; the rig is ready.
    2. No CI has run on this head, because the base is a stacked branch.
  • Should fix before real models use this profile: F1. The model-facing preview keeps only the first 8 KiB, so the failure summary and the exit code are cut off. With the same build script, the real model re-ran it in 5 of 6 trials; the ordinary Shell tool re-ran it in 0 of 3. A candidate patch plus tests is below.
  • Test gaps: deleting the publisher's bearer check (M13), or never closing the publisher (M11), passes all 899 unit tests and the author's full six-Workspace driver. I also have a candidate test for this.

Rig

Part What ran
Database MySQL 8.4.7 (native), Flyway V1–V13 applied by the PR's jar
Server managed-agent-server jar built from the PR: Session Store, embedded Runtime Broker, local-process workers from the PR bundle
Harness Packaged dist/cli.js serve --profile hosted-harness, owner of the loopback publisher
Models Real qwen3.8-max (DashScope) for the behaviour trials; the repo's fake OpenAI server for deterministic probes
Build and tests build + bundle OK; eslint + prettier clean on all 26 changed TS files; core 44/44; CLI 899/899 (the author's nine files); full shell.test.ts 361/361 (the PR ran 3 selected tests); broker 374 (1 skipped); managed-agent-server 114 + checkstyle

real stack

PR claims vs evidence

Claim Result How
Shell affects only the selected Workspace, exactly once; file-only Sessions still refuse Shell ✅ Driver on MySQL: shell-once.txt is x in all 4 Shell Workspaces, and the decoy is untouched. Real model R1: pwd is the saved Workspace, the file lands there, and the Harness directory with the decoy is untouched
Complete 100 MiB output (invalid UTF-8, separate stderr), preview only after a durable receipt ✅ Driver on MySQL: 107 content rows, 104,858,219 B
Retained output readable after producers stop ✅, stronger than claimed Stop Spring, restart mysqld, start a fresh Spring, then run the author's reader: HOSTED_SHELL_RETAINED_OUTPUT_OK, 104,857,735 B, digests and tails match
Lost start reply, lost raw-write ACK and SQL publish failure block without repeating the effect ✅ Driver on MySQL (fault relays unchanged)
Cancellation ✅ Driver. Also cancelling mid-stream in a 1 GiB yes: cancelled ~2.3 s later, capture committed, no yes process left, and the next turn works
Preview stays under the 64 KiB history limit ✅ Worst case, 70 KB of \x01 plus exit 3 (6× JSON escaping): the record is 55,578 B
Ordinary Shell truncation unchanged ✅ shell.test.ts 361/361
Long commands (the Broker's worker request times out at 30 s) ✅ A 45 s command settles once, success
Stale-activation fencing ⚠️ not re-tested live Unit tests pass. Mutant C4 (drop the activation clause in prepare) survives, possibly because other clauses cover it

probes

F1: the preview keeps only the first 8 KiB, so failures and the exit code never reach the model

boundedShellPreview() takes the first 8 KiB of the tool text. The core Shell tool deliberately keeps head and tail. shell.ts sets keep: 'both' because that preserves the command's "trailing exit/error summary (where shell failures report)". With raw capture, that truncation is bypassed and replaced by the head-only cut.

  • Deterministic check. Probe P2 prints 16 KB of output, then FINAL-42-TAIL, then exit 2. At head the model sees 8,540 B containing neither the tail nor Exit Code:. The notice says only Shell execution: error.
  • Real model. build.sh prints 600 progress lines, then an ERROR on stderr, then BUILD FAILED, then exits 3; the tool text is 31,897 B. The preview ends at compiling module 156 ... ok (cache war. How many times build.sh actually ran:
Arm Runs per trial Re-ran
Hosted, PR head 2 2 1 2 2 2 5/6
Hosted, candidate below 2 1 1 1 1 1 1/6 (that trial already had the answer and re-ran to separate stderr from stdout)
Ordinary Shell, same bundle, -p 1 1 1 0/3

A typical re-run at head: sh build.sh >/tmp/build.stdout 2>/tmp/build.err; echo "EXIT_CODE=$?". With a build this is harmless. With a migration, a deploy or a non-idempotent script, the Hosted profile will now routinely run it twice to find out why it failed.

Candidate (patch): same 8 KiB budget, split as 2 KiB head + a marker + 6 KiB tail, with UTF-8-safe cuts at both ends.

  • P2: the tail and the exit code now reach the model.
  • The worst-case record shrinks to 54,068 B.
  • The nine CLI files pass 901/901 (the patch also adds the auth/close test below).
  • The new test fails on the head version.

The tail of the tool text only reaches the end of the output when that output fits in the 64 KiB in-memory buffer (maxBufferedOutputBytes). For larger outputs the exit code still arrives, but the true output tail would have to come from the durable capture. That could be a follow-up.

preview

Candidate boundedShellPreview
const PREVIEW_BYTES = 8 * 1024;
const PREVIEW_HEAD_BYTES = 2 * 1024;
const PREVIEW_GAP =
  '\n[... preview truncated; the end of the output follows ...]\n';

export function boundedShellPreview(parts: readonly unknown[]): unknown[] {
  const text = parts
    .map((part) =>
      part && typeof part === 'object' && 'text' in part && typeof part.text === 'string'
        ? part.text
        : '',
    )
    .join('');
  const bytes = Buffer.from(text);
  if (bytes.byteLength <= PREVIEW_BYTES) return text ? [{ text }] : [];
  // Shell failures and the exit status are reported last: keep both ends.
  const head = new TextDecoder().decode(bytes.subarray(0, PREVIEW_HEAD_BYTES), { stream: true });
  let start =
    bytes.byteLength - (PREVIEW_BYTES - PREVIEW_HEAD_BYTES - Buffer.byteLength(PREVIEW_GAP));
  while (start < bytes.byteLength && (bytes[start]! & 0xc0) === 0x80) start++;
  return [{ text: head + PREVIEW_GAP + bytes.subarray(start).toString('utf8') }];
}

F2 (low): is_background: true kills the whole turn, and core Shell tells the model to use it

When core Shell's sleep N guard blocks a command, it tells the model to use is_background: true or the Monitor tool. Neither exists in this profile. If the model follows that advice, execute() throws before dispatch and the whole turn ends turn_error in 109 ms (probe P7), instead of returning a function error the model could recover from. Over 3 real-model runs the guard fired every time and the model never chose is_background: once it reported the block and stopped, twice it used the # intentional-sleep escape. Suggestion: answer profile-refused arguments with a function error response, or tailor the guard text for Hosted.

Test gaps from mutation (one-line mutants, the author's own suites)

TypeScript: 10/20 killed. Java: 5/10 killed.

  • M13 (publisher bearer check deleted) and M11 (the turn never closes its publisher) survive the unit tests and the full six-Workspace driver on MySQL. Live, the PR head is correct: 401 without a token or with a wrong one, and the listener is gone after every turn. Under M13, requests reach the handler (409); under M11, listeners pile up 1 → 2 → 3 across turns. The patch above adds a test that passes on head and kills M13.
  • The triage review lists several of these guards as the PR's safety properties. They hold at head, but no test pins them:
    • J6: ACK requires a settled execution.
    • J7: v3 is Shell-only.
    • J4: Store kind whitelist.
    • J5: local-process only.
    • J9: inputDigest format.
    • M1: raw-write offset check.
    • M6/M7: accept/receipt equality.
    • C2: seal check on read.
    • C6: read bounded by the declared length.

mutation

Notes

  • N1, storage growth (declared as a follow-up; here are the numbers). One yes call with a 20 s timeout stored 135,360,328 B of PUBLISHED rows. Extrapolated, that is ~0.8 GB per call at the default 120 s timeout and ~4 GB at the 600 s maximum, with no per-call or per-Session cap. A per-call byte cap would be cheap insurance before this profile leaves private use.
  • N2, throughput. Every 64 KiB raw write calls assertWritable(), which renews the writer lease (an HTTP call plus a locked UPDATE on the Session head row). A 64 MiB call made 1,097 writers:renew requests against 68 tool-results:publish, and ran at 4.4 MiB/s; 300 MiB took 58.9 s. Renewing at most every few seconds (the lease is 60 s) would likely lift this.
  • N3, resolving the conflict with c6c76750. In this PR, a Shell result's outcomeRef comes from the durable receipt. The oversize fallback in c6c76750 replaces the committed parts. If the fallback were applied to a Shell result while the receipt's outcomeRef is reused, history and the outcome resource would diverge. With the 8 KiB bound, the worst Shell record I could construct is 55.6 KB, so the fallback should never fire for Shell. It is worth asserting that in the resolution.
  • Not covered: Windows/Linux, remote provisioners, a worker crash mid-command, and live activation replacement.

Evidence (images, rig scripts, per-run logs, mutation JSON, candidate patch): wenshao/qwen-code@9ab1ced6/pr12848.

中文版

真实环境验证:#12848 Hosted 前台 Shell(head 7a41a5d4)

实测提交为 782c1774。之后唯一的新提交 7a41a5d4 只在 HostedWorkspaceToolTurnIT.java 里加了 7 行,生产代码完全一致。本报告是对 triage 评审 的补充:该评审明确说明没有执行 PR 代码,并希望有人实际跑一遍恢复与保留相关的声明。下面所有结论都经过实际运行。

结论

  • 在真实 MySQL 8.4.7 上,持久化契约成立。 作者使用的是 H2 的 MySQL 模式。PR 自带的六 Workspace driver 原样通过。保留的 100 MiB 输出在 Spring 与 mysqld 都重启之后逐字节读回一致。每个 Shell 副作用都恰好执行一次。取消、超时、300 MiB 输出以及超过 30 s 的命令都能正常结算。在实测代码中没有发现正确性阻塞问题。
  • 目前还不能直接合并,原因有两个,都属于流程问题:
    1. 它与 base 已经冲突(CONFLICTING):feat(serve): add gated Hosted Workspace file tool turns #12831 前进到了 c6c76750,冲突位于 hosted-workspace-tool-turn.ts 与 hosted-harness-session.ts,恰好是 Shell 结果提交路径。冲突解决后需要重跑一遍,装置已经就绪。
    2. 由于 base 是堆叠分支,这个 head 上没有跑过任何 CI。
  • 在真实模型使用这个 profile 之前应该修复:F1。 面向模型的预览只保留前 8 KiB,失败摘要和退出码都会被截掉。用同一个构建脚本,真实模型在 6 次试验中有 5 次 重跑了它;普通 Shell 工具 3 次中 0 次。候选补丁和测试见下文。
  • 测试缺口: 删除 publisher 的 bearer 校验(M13),或者让 publisher 永不关闭(M11),都能通过全部 899 个单测,也能通过作者的完整六 Workspace driver。这里同样附了候选测试。

装置

部分 实际运行内容
数据库 MySQL 8.4.7(原生安装),由 PR 构建的 jar 执行 Flyway V1–V13
服务端 由 PR 构建的 managed-agent-server jar:Session Store、内嵌 Runtime Broker,以及使用 PR bundle 的 local-process worker
Harness 打包后的 dist/cli.js serve --profile hosted-harness,是回环 publisher 的 owner
模型 行为试验用真实 qwen3.8-max(DashScope);确定性探针用仓库自带的 fake OpenAI server
构建与测试 build + bundle 通过;26 个改动的 TS 文件 eslint + prettier 干净;core 44/44;CLI 899/899(与作者相同的 9 个文件);完整 shell.test.ts 361/361(PR 只跑了其中 3 个);broker 374(1 个 skip);managed-agent-server 114 + checkstyle

PR 声明与证据

声明 结果 方式
Shell 只作用于选定的 Workspace,且恰好一次;仅文件工具的 Session 仍拒绝 Shell ✅ MySQL 上跑 driver:4 个 Shell Workspace 的 shell-once.txt 都是 x,诱饵未被改动。真实模型 R1:pwd 是保存的 Workspace,文件落在那里,放诱饵的 Harness 目录未被改动
完整的 100 MiB 输出(含非法 UTF-8、独立 stderr),持久回执之后才给出预览 ✅ MySQL 上跑 driver:content 107 行,104,858,219 B
生产者退出后仍能读取保留的输出 ✅,比声明更强 停 Spring、重启 mysqld、再起一个新 Spring,然后用作者的 reader 读取:HOSTED_SHELL_RETAINED_OUTPUT_OK,104,857,735 B,摘要与尾部一致
start 回复丢失、raw write 确认丢失、SQL 发布失败时都会阻塞,且不重复副作用 ✅ MySQL 上跑 driver(故障中继未修改)
取消 ✅ driver 覆盖。另外在 1 GiB yes 输出途中取消:约 2.3 s 后 cancelled,捕获已提交,没有残留 yes 进程,下一轮正常
预览不超过 64 KiB 历史上限 ✅ 最坏情况:70 KB 的 \x01 加 exit 3(JSON 转义 6 倍),记录为 55,578 B
普通 Shell 的截断行为不变 ✅ shell.test.ts 361/361
长命令(Broker 到 worker 的请求 30 s 超时) ✅ 45 s 命令只结算一次,结果 success
过期 activation 隔离 ⚠️ 未做真机复测 单测通过。变异体 C4(删掉 prepare 中的 activation 条件)存活,可能因为其他条件已覆盖

F1:预览只保留前 8 KiB,失败信息和退出码到不了模型

boundedShellPreview() 只取工具文本的前 8 KiB。核心 Shell 工具是有意同时保留头部和尾部的:shell.ts 设置 keep: 'both',因为这样能保留命令"末尾的退出/错误摘要(shell 失败信息所在位置)"。raw capture 路径绕过了这一截断,改成了只保留头部。

  • 确定性检查。 探针 P2 先输出 16 KB,再输出 FINAL-42-TAIL,然后 exit 2。在 head 上,模型看到的 8,540 B 里既没有尾部也没有 Exit Code:,提示里只有 Shell execution: error。
  • 真实模型。 build.sh 先输出 600 行进度,然后在 stderr 输出 ERROR,再输出 BUILD FAILED,最后 exit 3;工具文本共 31,897 B。预览停在 compiling module 156 ... ok (cache war。build.sh 实际执行的次数:
分组 每次试验执行次数 重跑
Hosted,PR head 2 2 1 2 2 2 5/6
Hosted,下方候选补丁 2 1 1 1 1 1 1/6(那一次已经拿到答案,重跑是为了区分 stderr 与 stdout)
普通 Shell,同一 bundle,-p 1 1 1 0/3

head 上典型的重跑命令是 sh build.sh >/tmp/build.stdout 2>/tmp/build.err; echo "EXIT_CODE=$?"。对构建来说无害;但如果是迁移、部署或非幂等脚本,Hosted profile 会为了弄清失败原因而经常把它执行两次。

候选补丁(patch):预算仍是 8 KiB,拆成 2 KiB 头部 + 标记 + 6 KiB 尾部,两端切点都保证 UTF-8 安全。

  • P2:尾部和退出码都能到达模型。
  • 最坏情况的记录降到 54,068 B。
  • 9 个 CLI 测试文件 901/901 通过(补丁同时加入了下文的鉴权/关闭测试)。
  • 新测试在 head 版本上失败。

工具文本的尾部只有在输出能放进 64 KiB 内存缓冲(maxBufferedOutputBytes)时才对应输出的真正末尾。输出更大时,退出码仍然能到达模型,但真正的输出尾部需要从持久捕获中读取,可以作为后续工作。

F2(低):is_background: true 会让整轮失败,而核心 Shell 恰好建议模型这样做

核心 Shell 的 sleep N 守卫拦下命令时,会建议模型改用 is_background: true 或 Monitor 工具,这两者在该 profile 中都不存在。模型如果照做,execute() 会在派发前抛错,整轮在 109 ms 内以 turn_error 结束(探针 P7),而不是返回一个模型可以自行恢复的函数错误。3 次真实模型运行中该守卫每次都触发,模型从未选择 is_background:一次报告被拦后停下,两次改用 # intentional-sleep 逃生方式。建议:对 profile 拒绝的参数返回函数错误,或者为 Hosted 定制守卫文案。

变异测试揭示的测试缺口(单行变异,使用作者自己的测试集)

TS 杀死 10/20,Java 杀死 5/10。

  • M13(删掉 publisher 的 bearer 校验) 和 M11(本轮从不关闭 publisher) 既能通过单测,也能通过 MySQL 上的完整六 Workspace driver。真机上 PR head 的行为是正确的:不带 token 或 token 错误都返回 401,每轮结束后监听都会消失。M13 下请求能进入处理函数(返回 409);M11 下监听随轮次累积 1 → 2 → 3。上面的补丁附带了一个测试,在 head 上通过,并能杀死 M13。
  • triage 评审把其中若干守卫列为 PR 的安全属性。它们在 head 上成立,但没有任何测试钉住:
    • J6:ACK 要求执行已结算。
    • J7:v3 只允许 Shell。
    • J4:Store 的 kind 白名单。
    • J5:只允许 local-process。
    • J9:inputDigest 格式。
    • M1:raw write 的 offset 校验。
    • M6/M7:accept/receipt 一致性。
    • C2:读取时的 seal 校验。
    • C6:读取受声明长度约束。

备注

  • N1,存储增长(已声明为后续工作,这里补上数据)。 一次超时 20 s 的 yes 调用存了 135,360,328 B 的 PUBLISHED 行。外推下来,默认 120 s 超时每次约 0.8 GB,600 s 上限约 4 GB,而且没有任何单次调用或单 Session 的上限。在这个 profile 走出私有使用之前,加一个单次调用字节上限是很便宜的保险。
  • N2,吞吐。 每个 64 KiB 的 raw write 都会调用 assertWritable(),续租 writer lease(一次 HTTP 请求加一次 Session head 行的加锁 UPDATE)。64 MiB 的一次调用产生了 1,097 次 writers:renew,而 tool-results:publish 只有 68 次,速度 4.4 MiB/s;300 MiB 用时 58.9 s。若改成每隔几秒最多续租一次(lease 为 60 s),吞吐很可能会提升。
  • N3,解决与 c6c76750 的冲突。 在本 PR 中,Shell 结果的 outcomeRef 来自持久回执。c6c76750 的超限回退会替换已提交的 parts。如果对 Shell 结果套用该回退、同时又复用回执的 outcomeRef,历史与 outcome 资源就会不一致。在 8 KiB 上限下,我能构造出的最坏 Shell 记录为 55.6 KB,回退不应触发;建议在解决冲突时加断言保证这一点。
  • 未覆盖: Windows/Linux、远端 provisioner、命令执行中 worker 崩溃,以及真机 activation 替换。

证据(图片、装置脚本、每次运行日志、变异 JSON、候选补丁):wenshao/qwen-code@9ab1ced6/pr12848。

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /resolve

@qwen-code-dev-bot

Copy link
Copy Markdown
Collaborator

Qwen Code resolved the merge conflicts and pushed the branch update.

Merge resolution: PR #12848 ← codex/hosted-tool-turn

Merge c3fbb0f00b (parents 7a41a5d42e, c6c767507f); only conflicted files were edited.

Root cause

Base c6c767507f collided with this PR's 782c17749f in two places: the
HostedWorkspaceToolTurn constructor (base swapped the throwing
validateAssistant for a boolean messageFitsInline(type, …); the PR appended an
optional shell?), and the publication path (base added an oversize guard
rebinding converted to an omission receipt; the PR added modelParts, capture,
receipt?.outcomeRef ?? precedence).

Semantic, not adjacent

Resolved tail (converted now let; receipt literal, contract prefix elided):

let outcome = Buffer.from(JSON.stringify({ executionCallId, ...converted[0] }));
if (outcome.byteLength > maxInlineResourceBytes ||
    !this.messageFitsInline('tool_result', converted, model)) {
  converted = [/* bounded omission receipt: outputOmitted, executionStatus */];
  outcome = Buffer.from(JSON.stringify({ executionCallId, ...converted[0] }));
}
const outcomeRef = receipt?.outcomeRef ??
  (await this.session.resources.publish('managed-tool-outcome', outcome));

Load-bearing

  • messageFitsInline (required) must precede shell? (optional); both construction
    sites rely on it.
  • The guard must stay after the response.response enrichment: it measures the
    committed record, so capture/runtimeError must be on it.
  • receipt?.outcomeRef ?? must stay on the publish, with the rebinding before it:
    Shell reuses the publisher's admitted durable ref, a file-tool turn publishes the
    possibly-replaced bytes. Reversing either double-publishes Shell output or commits
    an oversized record.

Not verified

No build, typecheck, lint or test was run.

The guard now covers Shell too, but no test exercises Shell + oversize:
hosted-workspace-tool-turn.test.ts passes no shell; base's new case uses
hosted-workspace-files/1. Were a Shell tool_result to exceed 64 KiB, the omission
receipt drops capture — the pointer to retained output the PR promises the
model. The PR budgets that preview at 8 KiB UTF-8 so escaped JSON plus metadata fit,
so I treated the branch as unreachable. Non-conflicted
hosted-workspace-tool-turn-driver.ts shares the path, bounded by
boundedShellPreview.

中文说明

合并 c3fbb0f00b(父 7a41a5d42e、c6c767507f)。只改了两个冲突文件。

根因:base 的 c6c767507f 与本 PR 的 782c17749f 在两处相撞:构造函数参数表
(base 把抛异常的 validateAssistant 换成布尔的 messageFitsInline(type, …),PR 追加
可选 shell?),以及已结算结果的发布路径(base 加入把 converted 重绑定为省略回执的
超限保护,PR 加入 modelParts、capture 与 receipt?.outcomeRef ?? 优先级)。

语义冲突:双方改了同一段逻辑,见上方代码块(converted 改为 let)。

关键约束:messageFitsInline(必填)须在 shell?(可选)之前;超限保护须在
response.response 补充字段之后,因为它度量被提交的记录;
receipt?.outcomeRef ?? 须留在发布表达式上且重绑定在其之前——Shell 回合复用 publisher
已准入的持久引用,文件工具发布可能被替换的字节;颠倒任一处都会重复发布 Shell 输出或提交
超限记录。

未能验证:未运行构建、类型检查、lint 或测试。保护现也覆盖 Shell,但无测试覆盖
Shell + 超限
:hosted-workspace-tool-turn.test.ts 不传 shell,base 新增用例用
hosted-workspace-files/1。若 Shell 的 tool_result 超过 64 KiB,省略回执会丢掉
capture,即 PR 承诺给模型的保留输出指针;PR 把预览限定在 8 KiB UTF-8 内以保证不超限,
故判定不可达。未冲突的 hosted-workspace-tool-turn-driver.ts 同路径,由
boundedShellPreview 限界。

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

Addressed the actionable fixes from the real-stack report in a209f3a, on top of the conflict-resolution merge c3fbb0f. The PR is currently mergeable.

  • F1: The same 8 KiB UTF-8 budget now includes up to 2 KiB of head, an explicit gap marker, and the remaining tail. Both cuts preserve UTF-8 characters. Failure summaries and exit codes survive, and the marker says “buffered output” so it does not imply that a >64 KiB process buffer contains the true stream tail.
  • F2: Invalid Shell arguments, including is_background: true, now produce durable function-error responses before acquisition or dispatch. A mixed batch is refused in full, with every other call explicitly reported as not executed; the model can correct the batch in the same turn. Failure to persist the refusal remains recovery-blocked.
  • M13 / M11: Added missing/wrong bearer tests and real listener-closure checks after completed, model-error, and uncertain-execution turns. Removing the bearer guard, the turn's close call, or the publisher's close call now fails the new tests. These were tested mutations, not inferred coverage.
  • N3: An already-admitted Shell result that exceeds the complete history-record limit now blocks before history replacement, checkpoint resolution, and ACK. It cannot substitute an omission response while retaining the original admitted outcome reference. Worst-case JSON-escaped previews retain capture metadata and fit the 64 KiB bound; removing the guard fails its regression test. File-tool omission behavior is unchanged.

Validation on the pushed changes:

  • npm run build, npm run typecheck, npm run bundle, changed-file ESLint/Prettier, and two clean full-diff self-audits passed. Independent static review of these seven changed files reported no findings.
  • CLI: 912/912 tests across nine relevant files; core: 396/396, including the full ordinary Shell suite; default Hosted process tests: 8/8.
  • Java Broker: 374 tests, 1 skipped, no failures; managed-agent-server: 114/114, Checkstyle clean.
  • The original six-Workspace integration test passed against the production Java Broker, separate workers, HTTP SQL Store, and H2 in MySQL mode: 100 MiB capture, lost start/raw replies, SQL publication failure, cancellation, and retained reads after all six producer processes exited.
  • Fresh F1/F2 bundle verification also passed through real Spring/Broker/worker processes and a deterministic model fixture: a 28 KB Chinese/emoji failure output preserved its head, stdout/stderr failure tails and exit 2 with one execution; a background-Shell + write-file batch produced two function errors with zero starts/effects, then its corrected foreground command succeeded once in the same turn. A 70,000-byte control-character probe preserved exit 3; the largest complete stored tool message was 54,381 bytes, below 65,536. All three prompts ended normally.
  • All four targeted mutants (bearer guard, either close call, admitted-result bound) were killed. After restoring production code, all 64 affected tests passed again.

The stacked base still excludes the normal PR trigger. I manually started Qwen Code CI with branch_ref pinned to a209f3a; it is running, not yet a green CI result.

Follow-ups remain explicit: true output tails beyond the 64 KiB buffer, per-call/Session storage quotas (N1), renewal-frequency optimization (N2), and the remaining mutation-coverage inventory (J6/J7/J4/J5/J9/M1/M6/M7/C2/C6 and C4). This correction does not change their production guards or claim to close those test gaps. This round did not repeat the report's real-MySQL restart or real-model repeated-execution trials; the fresh full-stack SQL run used H2.

@doudouOUC doudouOUC reopened this Sep 27, 2026
@doudouOUC
doudouOUC changed the base branch from codex/hosted-tool-turn to main September 27, 2026 15:13
@doudouOUC
doudouOUC force-pushed the codex/hosted-shell-tool-turn branch from a209f3a to b89b7b3 Compare September 27, 2026 15:13
@github-actions

Copy link
Copy Markdown
Contributor

Please do not rebase or force-push to an active PR as it invalidates existing review comments. Note for future reference, the bots always squash all changes into a single commit automatically as part of the integration.

中文

请勿对活跃的 PR 执行 rebase 或 force-push,因为这会使已有的评审评论失效。另外,供日后参考:作为集成流程的一部分,机器人始终会自动将所有改动压缩(squash)为单个提交。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Retargeted this PR to main after #12831 was squash-merged. Deleting the former base branch had automatically closed #12848; it is now reopened, and the temporary branch used to recover it has been removed.

The new head is b89b7b38621a8b2ba3f4a72993097b5fdbd7b352. Its 46-file patch is byte-for-byte identical to the previous head's net change against #12831, and none of those files changed on main in the meantime. Local npm run build, npm run typecheck, npm run bundle, 76 focused CLI tests, 395 focused core tests, and the commit's formatting/lint hooks passed. Maven is unavailable in this local environment, so the Java paths are being verified by the newly triggered main-targeted CI.

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /triage

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head b89b7b38

Not approving, on coverage rather than on a finding: this adds remote foreground Shell execution across four durability boundaries in 46 files and +4846/-513, and I could verify the gate and the turn-level guards myself but not the ~1,600 lines of publisher, capture-session and resource-store code, nor the Java half, inside my budget. For a feature whose subject is executing commands on a saved workspace, the unread part is exactly where a Critical would live, so I am naming the boundary instead of implying coverage I did not have.

No historical blocker

There are no reviews and no inline comments on this PR, so nothing has ever been filed against it and there is no thread to re-verify. Triage stage 2 at this exact head reports "No correctness blockers found at b89b7b38", and says it re-verified the load-bearing invariants on this head rather than carrying them forward after the squash and conflict resolution. The maintainer's real-stack run — native MySQL 8.4.7, the packaged dist/cli.js serve --profile hosted-harness, real qwen3.8-max for the behaviour trials — reports the durability contract holding, every Shell side effect running exactly once, and cancellation, timeouts, 300 MiB of output and >30 s commands all settling cleanly, with "no correctness blocker in the tested code". Its two process blockers are gone: the base conflict is resolved and CI has now run on this head.

What I verified in the code myself

The gate is explicit and narrow. HOSTED_WORKSPACE_SHELL_PROFILE = 'hosted-workspace-shell/1' is a second named profile beside the existing hosted-workspace-files/1, hosted-harness-session.ts accepts exactly those two and rejects anything else, and the shell options are wired only when the profile is the shell one — so default no-tool and file-profile hosts are unchanged, and nothing reaches Shell by default.

The turn refuses ambiguity rather than resolving it optimistically. Invalid Shell arguments, is_background: true included, produce durable function-error responses before acquisition or dispatch, with the message naming what is unavailable and asking for corrected arguments. A mixed batch is refused in full and every other call in it is explicitly reported as not executed, so the model cannot be left guessing which half ran. An uncertain outcome raises HostedToolRecoveryRequiredError rather than being reported as either success or failure.

No complete receipt without durable admission. The turn throws 'Complete Shell output was not admitted.' when the capture was not admitted, and the model-facing text on a truncated preview states the execution status and that complete stdout and stderr are retained in the Session result — it does not present a preview as the output. The over-limit guard is there too: 'Admitted Shell result exceeds the inline Session Store limit.' blocks before history replacement, checkpoint resolution and ACK, so an admitted result too large for the history record cannot be swapped for an omission response while keeping the original outcome reference.

The truncated preview keeps the parts a model needs. The head-plus-gap-marker-plus-tail shape within the 8 KiB budget, cutting only at UTF-8 boundaries, is what the maintainer's F1 measured as necessary — with the old head-only preview a real model re-ran a failed build in 5 of 6 trials against 0 of 3 for the ordinary Shell tool. The marker says "buffered output" rather than implying a >64 KiB process buffer holds the true stream tail, which is the honest claim.

tools/shell.ts changes by three lines: the temp-file truncation is skipped only when a raw capture owns the output, so the existing local path keeps its behaviour.

What I did not read, and why it matters

hosted-shell-publisher.ts (413 lines) and managed-shell-publisher.ts (405) — the loopback publisher, its capability handling, per-stream serialization of overlapping pipe callbacks, and the listener-closure paths; managed-shell-result-session.ts (436) and the 393 lines moved out of local-shell-result-session.ts; resource-tool-result-store.ts (366) — the writer-fenced publication, immutable resource ids and seal verification; and the Java side, RuntimeBrokerHttpServer (+65/-8), RuntimeBrokerService (+74/-8), WorkspaceRuntimeTransport (+97/-1), ManagedSessionStore (+64/-1) and HttpRuntimeTransport (+23/-1). The design's route table assigns an owner and a check set to each of five route families, and I verified none of those checks in code. The capability-discipline claims — canonical loopback only, redirects rejected, capability in memory and never in journal references or model arguments, no SQL writer token crossing into the worker — are the ones I would most want to read rather than accept, precisely because they are the boundary between a gated private profile and an unauthenticated local endpoint.

Two smaller notes on evidence currency, not defects: the executed real-stack results are from the pre-squash head, with this head re-verified statically by triage rather than re-run on a real stack; and the sandboxed verification triggered at 16:10 had not reported when I published.

CI

23 checks pass at this head and none has failed, including Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), the full Java matrix with Runtime Broker and Managed Agent MariaDB, Hosted no-tool processes / MySQL 8.4 and Real daemon E2E / Java 11, both Desktop Shell lanes and the routing checks. review-pr is pending and is not a gate. The maintainer's rig also reports the full shell.test.ts at 361/361 where the PR ran three selected tests.

What closes this

A pass over the publisher and capture pair — capability generation and validation, the loopback and redirect rules, per-stream write/finish serialization, listener closure after completed, model-error and uncertain turns, and what a lost raw-write reply does to the receipt — plus resource-tool-result-store.ts's writer fencing and seal verification, and the five route families' checks against the design's table. With those read, the gate and turn-level behaviour above are already verified and I would expect nothing else to stand in the way.

Verdict: COMMENT — No Critical found, and none has ever been filed here; the gate, the pre-dispatch refusals, the batch semantics, the no-admission-no-receipt rule and the preview shape all verify fail-closed in the code at this head. What blocks an approval is that the publisher, capture-session, resource-store and Java transport layers — the parts that hold the capability and the durability guarantees — are outside what I read in budget, and for a remote Shell path I am not willing to approve those on description alone.

wenshao pushed a commit to wenshao/qwen-code that referenced this pull request Sep 27, 2026
wenshao pushed a commit to wenshao/qwen-code that referenced this pull request Sep 27, 2026
@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

Real-stack re-verification, round 2: #12848 at b89b7b38

This follows up my round-1 report and the fix notes. Those notes said this round did not repeat the real-MySQL restart or the real-model trials, so both are covered here.

Head equivalence. b89b7b38 is a209f3ac squashed onto main 302e7d88. The 46-file PR patch is byte-identical (git diff c6c76750 a209f3ac vs git diff 302e7d88 b89b7b38, ignoring index lines), and none of those files changed on main. Main did bring other changes: V14/V15 migrations, event replay, and edits to hosted-harness-profile.ts and run-qwen-serve.ts. So I rebuilt everything at b89b7b38: TS bundle, Broker 374 tests (1 skipped), managed-agent-server 139/139, checkstyle.

Verdict

All four actionable round-1 items (F1, F2, M13/M11, N3) are fixed, and verified live on MySQL 8.4.7 with a real model where the scenario is reachable. I have no blocking findings. From the real-stack side this is ready to merge. CI on this head is green: Qwen Code CI, SDK Java, Serve A/B and tui-parity. The red Windows job on the earlier manual dispatch at a209f3ac is pre-existing on main. I re-ran it on a GitHub-hosted windows-2022 runner with CI's setup: main 302e7d88 and this head fail the same 14 CLI and 18 core tests, and none of them are in files this PR touches.

round 2

mutation, CI, Windows

What was re-run

Item Result at b89b7b38
Author's six-Workspace driver on a fresh MySQL DB (all 15 migrations) ✅ HOSTED_WORKSPACE_TOOLS_OK, 40.4 s, same row counts as round 1 (107 content rows, 104,858,219 B)
Stop Spring, restart mysqld, start a fresh Spring, run the reader ✅ HOSTED_SHELL_RETAINED_OUTPUT_OK
Round-1 database (V13, holding the round-1 100 MiB output) opened by the new jar ✅ V14 + V15 applied on startup; round-1 output still reads back byte-identical
F1 tail + exit code (16 KB, then a tail line, then exit 2) ✅ model sees FINAL-42-TAIL and Exit Code: 2 after the gap marker
F1 with output larger than the 64 KiB buffer (200 KB, exit 4) ✅ Exit Code: 4 arrives; the true stream tail past the buffer does not, as documented ("end of buffered output")
F1 worst case (70 KB of \x01 + exit 3) ✅ record is 54,061 B, under the 65,536 B limit
F1 with CJK + emoji (900 lines, stderr tail, exit 5) ✅ tail and exit code present; the preview contains no U+FFFD
F1, real qwen3.8-max, same build.sh as round 1 ✅ build.sh ran once in 6/6 trials (round-1 head: re-ran in 5/6). The ERROR line reached the model every time
F2 is_background: true ✅ function error, turn_complete, 0 Broker starts; reload and next turn OK
F2 mixed batch (write_file + background Shell) ✅ both answered as not executed, 0 starts; the corrected Shell then ran once
F2 in the middle of a turn: Shell OK, then refused, then Shell OK ✅ one turn; inter.txt = r1, r3; 1 acquire, 2 starts, 1 release; reload and next turn OK
F2 with an unknown directory argument ✅ function error; the corrected call ran once
M13 / M11 live ✅ 401 without a token and with a wrong token; listener closed after 3 completed turns and after a model-error turn

About F2 with a real model: in 4/4 trials (2 at this head, 2 at the round-1 head) qwen3.8-max declined to send is_background at all, because the declared schema doesn't have it. So F2 is only reachable through the deterministic probes above, and it is fixed there.

Mutation check on the fixes (the author's nine CLI files, baseline 912/912)

11/12 killed:

  • the refusal path disabled (invalid arguments would be dispatched in the foreground)
  • the refusal not persisting its tool_result
  • the refusal answering only the invalid calls
  • the refusal persisting nothing at all
  • the N3 guard removed
  • the F1 UTF-8 alignment removed, the head removed, the gap marker removed
  • M13 bearer check removed; M11 turn close removed; the publisher's own server.close() removed

M13 and M11 both survived in round 1.

Surviving: deleting the size check on the refusal record. It is effectively unreachable, since refusal messages are short. It could be pinned by a test with many calls, but I wouldn't hold the PR for it.

Against the open items in the latest review

That review notes that the executed real-stack results predated the squash. Everything in this report was executed at b89b7b38.

Item named in the review Executed evidence at b89b7b38
Capability generation and validation Live: no token → 401, wrong 43-char token → 401. Mutant M13 (bearer check deleted) is now killed
Loopback and redirect rules Pinned by unit and Java tests: round-1 mutants M8 (worker URL shape) and J8 (Broker URL shape) were killed. Redirects were not exercised live
Per-stream write/finish serialization; what a lost raw-write reply does Author's driver on MySQL: the raw-reply-loss Workspace blocks, the write is not retried, and the effect runs once. Round-1 mutant M2 (queue removed) was killed
Listener closure after completed, model-error and uncertain turns Live: 3 completed turns, a provider failure after a Shell round, and an uncertain turn (status reply lost) all leave no publisher listener. M11 and server.close() mutants are killed
Store writer fencing and seal verification Java mutants J1 (writer fencing), J2 (immutable resource id) and J3 (digest) were killed in round 1; retained reads after a mysqld restart verify the seal and digests. Round-1 C2 (skipping the seal check on read) still survives, as the author listed

Non-blocking notes

  • A single lost status-poll reply blocks the Session. HostedWorkspaceBroker.execute polls every 50 ms and does not retry a failed GET. Dropping one reply blocked the Session (hosted_turn_recovery_required); fail-closed, and the command ran once. The loop is from merged feat(serve): add gated Hosted Workspace file tool turns #12831, but Shell turns poll far longer than file tools: 793 polls for a 45 s command, about 12,000 for a 10-minute one. A few retries of the idempotent status read before declaring the outcome unknown would make long Shell turns much less brittle.

  • runtimeError.message is still cut with Buffer.from(message).subarray(0, 1024).toString('utf8') in the worker's finalize. With CJK text this gives 1,026 bytes ending in U+FFFD. The model-facing preview is not affected. This is the same kind of cut F1 fixed; it's a nit.

  • The follow-ups from round 1 are unchanged, and the author lists them explicitly: true output tails beyond 64 KiB, storage quotas (N1), lease-renewal frequency (N2), and the older unpinned guards J4–J9, M1, M6/M7, C2, C4, C6.

中文版

真实环境复验(第二轮):#12848 @ b89b7b38

本轮承接第一轮报告和修复说明。修复说明提到本轮没有重做真实 MySQL 重启和真实模型试验,这两项都在这里补上。

head 等价性。 b89b7b38 是把 a209f3ac squash 后落在 main 302e7d88 上的结果。46 个文件的 PR 补丁逐字节一致(比较 git diff c6c76750 a209f3ac 与 git diff 302e7d88 b89b7b38,忽略 index 行),这些文件在 main 上也没有被改动。不过 main 带来了其他改动:V14/V15 迁移、事件回放,以及 hosted-harness-profile.ts、run-qwen-serve.ts 等。所以我在 b89b7b38 上全部重新构建:TS bundle、Broker 374 个测试(1 个 skip)、managed-agent-server 139/139、checkstyle。

结论

第一轮四项可执行问题(F1、F2、M13/M11、N3)均已修复,并在 MySQL 8.4.7 上真机验证;场景可达的地方也用了真实模型。没有阻塞问题。从真实环境验证角度,可以合并。 这个 head 上的 CI 全绿:Qwen Code CI、SDK Java、Serve A/B、tui-parity。此前在 a209f3ac 上手动触发的 Windows 任务失败属于既有问题:我在 GitHub 托管的 windows-2022 runner 上按 CI 的配置重跑,main 302e7d88 与本 head 失败的都是同样的 14 个 CLI 测试和 18 个 core 测试,且都不在本 PR 改动的文件里。

复验内容

项目 b89b7b38 上的结果
作者的六 Workspace driver,全新 MySQL 库(15 个迁移全部执行) ✅ HOSTED_WORKSPACE_TOOLS_OK,40.4 s,行数与第一轮一致(content 107 行,104,858,219 B)
停 Spring、重启 mysqld、再起新 Spring,运行 reader ✅ HOSTED_SHELL_RETAINED_OUTPUT_OK
第一轮的数据库(V13,含第一轮的 100 MiB 输出)交给新 jar 打开 ✅ 启动时执行 V14 + V15;第一轮输出仍逐字节读回一致
F1 尾部 + 退出码(16 KB 输出、尾行、exit 2) ✅ 模型在截断标记之后能看到 FINAL-42-TAIL 和 Exit Code: 2
F1,输出超过 64 KiB 缓冲(200 KB,exit 4) ✅ Exit Code: 4 能到达;缓冲之外的真正流尾不会到达,与文档说明一致("end of buffered output")
F1 最坏情况(70 KB 的 \x01 + exit 3) ✅ 记录 54,061 B,低于 65,536 B 上限
F1,中文 + emoji(900 行、stderr 尾、exit 5) ✅ 尾部和退出码都在,预览中没有 U+FFFD
F1,真实 qwen3.8-max,与第一轮相同的 build.sh ✅ build.sh 6/6 次都只执行一次(第一轮 head:6 次中 5 次重跑),ERROR 行每次都到达模型
F2 is_background: true ✅ 返回函数错误,turn_complete,Broker start 次数为 0;重新加载和下一轮正常
F2 混合批次(write_file + 后台 Shell) ✅ 两者都答复为未执行,start 次数为 0;修正后的 Shell 随后执行一次
F2 出现在同一轮中间:Shell 成功、然后被拒、再 Shell 成功 ✅ 同一轮内完成;inter.txt 为 r1、r3;1 次 acquire、2 次 start、1 次 release;重新加载和下一轮正常
F2,未知参数 directory ✅ 返回函数错误;修正后的调用执行一次
M13 / M11 真机检查 ✅ 无 token 和错误 token 都返回 401;3 个正常结束的轮次以及一个模型出错的轮次之后,监听都已关闭

关于真实模型下的 F2:4 次试验中(本 head 2 次、第一轮 head 2 次),qwen3.8-max 都没有发送 is_background,因为声明的 schema 里没有这个参数。所以 F2 只能通过上面的确定性探针触发,在那里已经修复。

针对修复的变异测试(作者的 9 个 CLI 测试文件,基线 912/912)

杀死 11/12:

  • 关闭拒绝路径(无效参数会被当作前台命令派发)
  • 拒绝时不持久化 tool_result
  • 拒绝时只回复无效的调用
  • 拒绝时什么都不持久化
  • 删除 N3 守卫
  • 删除 F1 的 UTF-8 对齐、去掉头部、去掉截断标记
  • 删除 M13 bearer 校验;删除 M11 的本轮关闭调用;删除 publisher 自身的 server.close()

M13 和 M11 在第一轮时都存活。

存活: 删除对拒绝记录的大小检查。实际上不可达,因为拒绝消息都很短。可以用一个大量调用的测试钉住,但不必为此卡 PR。

对照最新评审中列出的待办

该评审提到,此前实际运行的真实环境结果早于 squash。本报告的所有内容都在 b89b7b38 上实际执行。

评审列出的项目 在 b89b7b38 上实际执行的证据
capability 的生成与校验 真机:无 token → 401,43 字符的错误 token → 401。变异 M13(删除 bearer 校验)现已被杀死
回环地址与重定向规则 由单测和 Java 测试钉住:第一轮变异 M8(worker 端 URL 形态)和 J8(Broker 端 URL 形态)均被杀死。重定向未做真机验证
按流串行的 write/finish;raw write 回复丢失时的行为 MySQL 上跑作者的 driver:raw 回复丢失的那个 Workspace 被阻塞,写入不重试,副作用只执行一次。第一轮变异 M2(删除队列)被杀死
正常结束、模型出错、不确定三种轮次之后的监听关闭 真机:3 个正常结束的轮次、Shell 之后 provider 出错的轮次、以及不确定轮次(status 回复丢失)之后,都没有残留 publisher 监听。M11 和 server.close() 的变异体均被杀死
Store 的 writer fencing 与 seal 校验 第一轮 Java 变异 J1(writer fencing)、J2(resource id 不可变)、J3(摘要)均被杀死;mysqld 重启后的保留读取会校验 seal 和摘要。第一轮的 C2(读取时跳过 seal 校验)仍然存活,作者已列入后续

非阻塞备注

  • 一次 status 轮询回复丢失就会阻塞 Session。 HostedWorkspaceBroker.execute 每 50 ms 轮询一次,GET 失败时不重试。我丢掉一次回复,Session 就被阻塞(hosted_turn_recovery_required);这是 fail-closed,命令也只执行了一次。这个循环来自已合入的 feat(serve): add gated Hosted Workspace file tool turns #12831,但 Shell 轮次的轮询时间远长于文件工具:45 s 的命令约 793 次,10 分钟的命令约 12,000 次。在判定结果未知之前,对这个幂等的 status 读取重试几次,可以让长时间的 Shell 轮次稳健得多。

  • worker 的 finalize 里,runtimeError.message 仍用 Buffer.from(message).subarray(0, 1024).toString('utf8') 截断。遇到中文会得到 1,026 字节、并以 U+FFFD 结尾。面向模型的预览不受影响。这和 F1 修掉的是同一类截断,属于小问题。

  • 第一轮的后续事项不变,作者已明确列出:超过 64 KiB 的真实输出尾部、存储配额(N1)、续租频率(N2),以及较早未被钉住的守卫 J4–J9、M1、M6/M7、C2、C4、C6。

Evidence (round-2 scripts, per-run logs, mutation JSON, Windows outputs): wenshao/qwen-code@7e4be7d6/pr12848/r2.

wenshao
wenshao previously approved these changes Sep 27, 2026
qqqys
qqqys previously approved these changes Sep 27, 2026
@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Thanks for the round-6 real-stack review. I checked the reported paths against the current head and ran the targeted publisher tests plus the real inherited-pipe regression.

Finding Action
Two V16 migrations on main Confirmed. #12900 already contains the V17 renumbering and its MySQL/MariaDB Java jobs pass. I did not open a duplicate fix. This branch needs to re-sync after #12900 merges.
Incomplete pipe marked storage_failed, accepted prefix discarded Fixed in 14b77feee: an unclosed pipe no longer calls the storage-failure path. The retained manifest now reports partial/producer_lost, keeps the accepted prefix, and still records a blocked receipt. The lost-write-acknowledgement case remains unavailable/storage_failed and blocked. Hosted publisher tests pass 23/23; the real inherited-pipe test passes 1/1.
Backgrounded compound command holds the Workspace lease Confirmed. I am keeping the fail-closed behavior in this explicit private profile. Killing the original process group and observing pipe EOF does not prove a detached descendant has stopped writing. #12904 records the recovery design and end-to-end acceptance gates; the PR description now calls out this limitation. This remains unresolved in #12848.
Absolute file_path fails the turn Confirmed as a pre-existing file-profile problem; deferred to #12905 with a durable, correctable refusal and accurate argument name.
Idle keep-alive closure and capture: null Your Node 22 measurements do not show the proposed close hang. The null-capture path has both real-stack and TypeScript coverage described in this review, so I made no change here.

No review threads were unresolved or replied to (0/0). I will keep #12848 open for the remaining review and #12900 dependency; it has not been merged.


感谢第六轮真实环境评审。我按当前分支核对了这些路径,并运行了发布器定向测试及真实子进程继承管道的回归测试。

发现 处理
主干存在两个 V16 迁移 已确认。#12900 已将后合入迁移顺延为 V17,MySQL/MariaDB Java CI 通过;我没有重复提交修复。#12900 合入后本分支需同步主干。
不完整管道误标 storage_failed 且丢弃已接收前缀 已在 14b77feee 修复:未关闭的管道不再走存储故障路径,持久 manifest 记录 partial/producer_lost 并保留前缀,回执仍是 blocked。丢失写入确认仍为 unavailable/storage_failed 且阻断。发布器测试 23/23、真实继承管道测试 1/1 通过。
后台复合命令长期占住 Workspace 租约 已确认。显式私有 profile 继续在故障时阻断。杀原进程组并观察到管道 EOF,仍不能证明脱离进程组的后代已停止写入。#12904 记录恢复设计和端到端验收门禁,PR 说明已写明限制;#12848 中此问题尚未解决。
绝对 file_path 导致整轮失败 已确认是原有文件 profile 问题;推迟到 #12905,目标是持久、可纠正的拒绝及准确的参数名称。
空闲 keep-alive 关闭与 capture: null Node 22 实测未显示所担忧的关闭等待;null-capture 路径已有本轮所述真实环境和 TypeScript 覆盖,因此没有改动。

本轮没有未解决的行内评审线程,也没有回复或关闭线程(0/0)。#12848 将继续保持开放,等待后续评审及 #12900 依赖处理;没有合并。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Synced the merged #12900 and current main into this branch at 0b8ca77b7; the merge had no conflicts. The branch now has one runtime-loss V16 migration and the Session-operation migration at V17. The PR body reflects that #12900 is no longer pending.

Validation on macOS: the frozen-lockfile install ran the repository build and bundle successfully; npm run typecheck passed; the Hosted publisher and Workspace tool-turn suites passed 53/53; git diff --check passed. This host has no JDK, so the new Java/MySQL and MariaDB checks must verify the migration in CI. pnpm continued trying to download unrelated platforms' optional binaries after its successful build and bundle, so I stopped that lingering download process. No source or lockfile changes were made by the sync beyond the merge.

The Workspace lease recovery gap remains tracked in #12904; this merge does not change its fail-closed behavior.


已在 0b8ca77b7 将合入的 #12900 和最新 main 同步到本分支,没有冲突。分支现在仅有一个运行时丢失证据 V16 迁移,Session 操作迁移为 V17;PR 说明也已更新。

本机 macOS 验证:锁文件安装过程成功执行仓库构建和 bundle,npm run typecheck 通过,Hosted 发布器和 Workspace 工具回合测试 53/53 通过,git diff --check 通过。本机没有 JDK,迁移的 Java/MySQL 与 MariaDB 验证交给新一轮 CI。pnpm 在构建及 bundle 完成后仍重试下载其他平台的可选二进制包,我停止了这个遗留下载进程。同步没有引入额外的源码或锁文件改动。

Workspace 租约恢复缺口仍由 #12904 跟踪;本次合并没有改变故障时阻断的行为。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

CI follow-up for a4632e984:

Finding Action
Ubuntu Node test job timed out while waiting for three Hosted Harness turn-settlement assertions; cleanup then raced with background writes. Set a 10-second bound on the two affected asynchronous waits (covering all three cases). Production behavior is unchanged.

Verified on this commit: build, typecheck, ESLint, and the Hosted Harness session test file (20/20). There were no review threads to reply to or resolve in this batch.

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

Real-stack verification, round 7: #12848 at a4632e98

This round covers the author's reply to round 6 and three commits:

Stack: MySQL 8.4.7; the Spring jar and TS bundle rebuilt at a4632e98; real local-process workers; the fixture model plus real qwen3.8-max.

Verdict

  • 14b77fee does what it says on the real stack.
    • When a background process keeps the Shell's pipe open, the durable manifest is now partial / producer_lost instead of unavailable / storage_failed.
    • The accepted prefix reads back through the Store byte for byte, including a 3,000,005-byte prefix spread over several pages.
    • Mutation testing shows the new path, and the worker-failure path it now depends on, are both pinned by tests.
  • The main merge fixes the startup failure from round 6. The server starts with V16 and V17. The whole IT class, main's FG6a and FG6b gates, the Shell FG6b variant, the Java suites, the rig scenarios and a real model all pass. SDK Java is green again.
  • Unchanged, by design: the Session and the Workspace stay blocked after an incomplete capture.

round 7

14b77fee: an unclosed pipe keeps its prefix

Each command ran in its own Session. I stopped the Harness, waited for its lease to lapse, and read every stream back through ResourceToolResultSegmentStore over the Store HTTP API.

Command Durable manifest Read back through the Store
cd . && sleep 20 >/dev/null 2>&1 & echo ok partial / producer_lost; stdout incomplete, 3 B 3 B ok\n; sha256 = manifest
cd . && { head -c 3000000 /dev/zero | tr "\000" a; sleep 20; } & sleep 5; echo done partial / producer_lost; stdout incomplete, 3,000,005 B 3,000,005 B: 3,000,000 a + one done; sha256 = manifest
a logger that keeps writing for 3 s after the shell exits partial / producer_lost; stdout incomplete, 36 B started + tick-0…tick-3; sha256 = manifest
printf plain-ok (control) complete; sealed 8 B; sha256 = manifest
  • missingRanges records an open tail, [{"start": <n>, "end": null}], on both streams of the partial captures.
  • The four minimal background shapes from round 6 split as before. The two compound forms are now partial / producer_lost; sleep 20 … & echo ok and cd . ; sleep 20 … & echo ok still complete.
  • As before:
    • The turn is still recovery-blocked, with Complete Shell output was not admitted.
    • Every new Session in that Workspace, with either profile, still gets 409 workspace_busy.
    • The control Workspace still serves new Sessions normally.

The fix is one deleted line: the owner no longer calls failCapture() when a stream finishes with complete: false. That leaves the failed flag on finalize as the only way a worker-side write failure reaches the owner. Four mutants check that path:

Mutant Result
K1: restore the deleted line killed by the new test (retains the accepted prefix when a producer holds a pipe past the drain boundary)
K2: owner finalize ignores the worker's failed flag killed (irreversibly blocks capture after a lost raw write acknowledgement)
K3: worker reports complete: true after a failed write survives; equivalent, because finalize still carries failed (K2 + K3 together is killed)
K4: owner finalize no longer fails a stream that never finished survives; equivalent. handleExit always finishes both streams first, and a failed finish sets failed, which K2's line handles

Regression at a4632e98

Check Result
Server start, fresh MySQL 8.4.7 database ✅ 17 migrations: V16 runtime_loss_evidence, V17 managed_session_operation
HostedWorkspaceToolTurnIT, whole class, mvn -P hosted-workspace-tools verify ✅ 3/3 in 97.6 s. Six-Workspace driver; FG6a 8 cases, including status; FG6b 7 cases
FG6b with a Shell call in place of edit ✅ 6/6
Java unit tests ✅ Broker 390 (1 skipped), managed-agent-server 155
R1-4 / R1-7 validation, R1-1 grandchild, 8 output shapes, F4 Shell → reload → Shell ✅ unchanged
Real qwen3.8-max: Shell, then detach + load, then Shell ✅ both turns complete; notes.txt = first\nsecond

CI on a4632e98

  • SDK Java: green in all 8 jobs, including Runtime Broker and Managed Agent MariaDB and Hosted no-tool processes / MySQL 8.4, which were red on 20126bae.
  • Serve A/B and tui-parity: green.
  • Qwen Code CI: green, including the web-shell E2E smoke.

Merge reference

  • Correctness: I see no blocker. Everything from round 6 that belongs in this PR is done: the partial prefix is kept and labelled honestly, and the V16 collision is fixed on main and merged in.
  • What remains is the availability decision from round 6.
  • From the real-stack side: mergeable as a private, fail-closed profile, provided hosted-workspace-shell/1 doesn't serve real-model traffic until Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 lands.
中文版

真实环境验证(第七轮):#12848 @ a4632e98

本轮覆盖作者对第六轮的回复和三个提交:

环境:MySQL 8.4.7;Spring jar 和 TS bundle 在 a4632e98 上重新构建;真实本地进程 worker;固定脚本模型加真实 qwen3.8-max。

结论

  • 14b77fee 在真实环境中的行为与描述一致。
    • 后台进程让 Shell 的管道保持打开时,持久 manifest 现在记为 partial / producer_lost,而不是 unavailable / storage_failed。
    • 已接收的前缀可以通过 Store 逐字节读回,包括一个跨越多页的 3,000,005 字节前缀。
    • 变异测试表明,新路径以及它现在所依赖的 worker 失败路径,都有测试钉住。
  • 合入 main 修好了第六轮的启动失败。 服务带着 V16 和 V17 正常启动。IT 整类、main 的 FG6a 和 FG6b 门禁、Shell 版 FG6b、Java 测试、rig 场景和真实模型全部通过,SDK Java 恢复全绿。
  • 按设计未变:捕获不完整后,Session 和 Workspace 仍然被阻塞。

14b77fee:未关闭的管道会保留前缀

每条命令在各自的 Session 中运行。我停掉 Harness,等它的租约过期,再通过 Store HTTP API 用 ResourceToolResultSegmentStore 读回每条流。

命令 持久 manifest 通过 Store 读回
cd . && sleep 20 >/dev/null 2>&1 & echo ok partial / producer_lost;stdout incomplete,3 B 3 B ok\n;sha256 = manifest
cd . && { head -c 3000000 /dev/zero | tr "\000" a; sleep 20; } & sleep 5; echo done partial / producer_lost;stdout incomplete,3,000,005 B 3,000,005 B:3,000,000 个 a 加一个 done;sha256 = manifest
shell 退出后还持续写 3 s 的日志进程 partial / producer_lost;stdout incomplete,36 B started + tick-0…tick-3;sha256 = manifest
printf plain-ok(对照) complete;sealed 8 B;sha256 = manifest
  • 部分捕获的两条流上,missingRanges 都记录了一个开放的尾部:[{"start": <n>, "end": null}]。
  • 第六轮的四种最小后台写法结果分布和之前一样。两种复合写法现在是 partial / producer_lost;sleep 20 … & echo ok 和 cd . ; sleep 20 … & echo ok 仍然正常完成。
  • 和之前一样:
    • 这一轮仍然被恢复阻塞,报 Complete Shell output was not admitted.
    • 该 Workspace 中的每个新 Session,无论哪种 profile,仍然得到 409 workspace_busy。
    • 对照 Workspace 仍能正常为新 Session 服务。

修复只删了一行:某条流以 complete: false 结束时,owner 不再调用 failCapture()。这样一来,finalize 上的 failed 标志就成了 worker 端写入失败传到 owner 的唯一途径。四个变异体检查这条路径:

变异体 结果
K1:恢复被删掉的那一行 被新测试杀死(retains the accepted prefix when a producer holds a pipe past the drain boundary)
K2:owner 的 finalize 忽略 worker 的 failed 标志 被杀死(irreversibly blocks capture after a lost raw write acknowledgement)
K3:写入失败后 worker 仍报 complete: true 存活;等价,因为 finalize 仍携带 failed(K2 + K3 同时施加会被杀死)
K4:owner 的 finalize 不再让从未 finish 的流失败 存活;等价。handleExit 总是先 finish 两条流,finish 失败会设置 failed,由 K2 那一行处理

a4632e98 上的回归

检查 结果
在全新 MySQL 8.4.7 库上启动服务 ✅ 17 个迁移:V16 runtime_loss_evidence、V17 managed_session_operation
HostedWorkspaceToolTurnIT 整类,mvn -P hosted-workspace-tools verify ✅ 3/3,97.6 s。六 Workspace driver;FG6a 8 个用例,含 status;FG6b 7 个用例
用 Shell 调用替换 edit 的 FG6b ✅ 6/6
Java 单元测试 ✅ Broker 390(1 个跳过),managed-agent-server 155
R1-4 / R1-7 参数校验、R1-1 孙进程、8 种输出形态、F4 Shell → 重新加载 → Shell ✅ 未变
真实 qwen3.8-max:Shell,然后 detach + load,再 Shell ✅ 两轮都完成;notes.txt = first\nsecond

a4632e98 的 CI

  • SDK Java: 8 个任务全绿,包括 20126bae 上变红的 Runtime Broker and Managed Agent MariaDB 和 Hosted no-tool processes / MySQL 8.4。
  • Serve A/B、tui-parity: 绿。
  • Qwen Code CI: 全绿,包括 web-shell E2E smoke。

合并参考

  • 正确性: 我看不到阻塞项。第六轮提出的、属于本 PR 范围的问题都已解决:部分前缀得到保留且标注准确,V16 撞号已在 main 上修复并合进来。
  • 剩下的是第六轮提出的可用性问题。
  • 从真实环境的角度看: 可以作为私有、fail-closed 的 profile 合并,前提是在 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 落地之前,hosted-workspace-shell/1 不接真实模型流量。

Evidence (probe scripts, per-run logs, IT summaries, the mutation log): wenshao/qwen-code@c250d377/pr12848/r7.

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Follow-up to round 7 real-stack verification and the new main conflict:

Finding Action
main added Stage H transaction resource lookup beside this PR's pre-published Shell result resource path. Merged main in d650cdd3f, retained both paths, and kept Stage H lookup limited to transaction-referenced resources.
Round 7 found no remaining correctness blocker, but incomplete capture can hold a Workspace lease. Agreed: keep hosted-workspace-shell/1 private and do not serve real-model traffic through it until #12904 provides safe recovery.

On d650cdd3f, local build and typecheck passed; related CLI tests passed 50/50 and core tests 11/11. This host has no Java/Maven runtime, so the merged Java integration tests await CI. No review threads required a reply or resolution (0 unresolved).

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

Real-stack verification, round 8: #12848 at d650cdd3

d650cdd3 merges main e3e7dbcd into a4632e98. It brings in:

Only one conflict involves this PR's files, in ManagedSessionStore (see the author's note).

Stack: MySQL 8.4.7; the Spring jar and TS bundle rebuilt at d650cdd3; real local-process workers; the fixture model plus real qwen3.8-max.

Verdict

  • The merge is clean on the real stack.
    • The server starts with V16, V17 and V18. The Java suites pass: Broker 394 (1 skipped) and managed-agent-server 172.
    • H0c's new REFERENCED-only lookup sits next to this PR's pre-published Shell resources without affecting them. Partial and complete captures still read back through the Store byte for byte.
    • Every rig scenario from round 7 gives the same result, and CI is green.
  • Main's new FG6c crash gate, re-run with a Shell call, shows no replay and no duplicate effect in any case.
  • A correction to rounds 6 and 7.
    • I wrote that FG6b is a manual gate that no workflow runs. That is wrong.
    • CI's Hosted job runs -Phosted-harness-mysql, whose failsafe include is **/Hosted*IT.java. So the whole HostedWorkspaceToolTurnIT class, FG6a and FG6b included, has run in CI since 018f3318; the job logs show FG6B arguments: … there.
    • The Shell variants of FG6b and FG6c are the part CI doesn't run.

round 8

The merge

  • The conflict: ManagedSessionStore keeps both sides.
    • The PR's publishToolResult path still publishes Shell results as PUBLISHED resources.
    • H0c's storedResource accepts only REFERENCED resources. Its only caller is the Stage H extension-record hook in commit, so pre-published tool results never go through it. That matches the author's description.
  • On the stack:
Check at d650cdd3 Result
Server start on a fresh database ✅ 18 migrations: V16 runtime-loss evidence, V17 Session operation, V18 extension record
Partial capture read back through the Store (cd . && { 3,000,000 × a; sleep 20; } & sleep 5; echo done) ✅ partial / producer_lost; 3,000,005 B = 3,000,000 a + done; sha256 = manifest
Other probes: the 3 B and 36 B prefixes, the complete control, the four minimal background shapes, Workspace hold, validation, R1-1, 8 output shapes, F4 reload, real-model reload ✅ identical to round 7
FG6b with a Shell call ✅ 6/6
CI ✅ Qwen Code CI, SDK Java, Serve A/B and tui-parity. The Hosted process fault gates / MySQL 8.4 job ran HostedWorkspaceToolTurnIT 3/3, HostedHarnessMySqlIT 2/2 and HostedProcessCrashIT with all 6 FG6c cases

Main's FG6c, re-run with a Shell call

FG6c uses edit on a FIFO so that the tool blocks mid-execution. For the Shell variant I used c=$(cat proof.txt) && [ "$c" = x ] && rm -f proof.txt && printf %s%s "$c" "$c" > proof.txt, which blocks the same way and has the same effect, and left the Java assertions untouched.

  • Pass as is: harness-prepare and spring-kill.
  • worker-stop: passes once the driver accepts one thing. While the worker is stopped, a status poll gets 503 managed_runtime_unavailable, because the v3 observation asks the worker (30 s timeout); with edit, the Broker answers from its record. The Harness doesn't retry that retryable: true 503, so the turn blocks right away rather than at the end of its observation window. Either way it ends blocked.
  • worker-kill: the driver asserts a v2 409 code; Shell gets runtime_execution_evidence_unavailable and then runtime_admission_closed. With those accepted, the Java ledger differs in one row: the binding is LOST, not READY, because the v3 observation notices the dead worker, where the v2 path only reads the Broker's record.
  • harness-result: the ledger and the journal-event checks pass. Then the assertion "no managed-tool-outcome resource" fails. The Shell path commits a recordToolResult transaction, holding the outcome and a reference to the capture manifest, before the tool_result message, so a crash at the message commit leaves a committed outcome that the edit path doesn't have yet. That is more durable state, not less, and the cold load is still blocked.
  • harness-start: I also ran it on the rig, with edit as a control.
Harness SIGKILL while the tool blocks, then the tool finishes Shell edit
Effect xx, once xx, once
Broker execution, 2 s to 120 s later UNKNOWN, no result SETTLED, result stored
Capture resources none n/a
Cold load of the Session 409 hosted_turn_recovery_required same
New Session in the same Workspace 409 workspace_busy same

The capture publisher runs inside the Harness, so when the Harness dies the worker cannot finish the capture, and a command that did complete is recorded with an unknown outcome. That is fail-closed, and nothing replays, but it means #12904-style recovery can't learn the outcome of a Shell call interrupted this way. It is worth stating in that design.

Merge reference

中文版

真实环境验证(第八轮):#12848 @ d650cdd3

d650cdd3 把 main e3e7dbcd 合入 a4632e98,带进了:

涉及本 PR 文件的冲突只有一处,在 ManagedSessionStore(见作者说明)。

环境:MySQL 8.4.7;Spring jar 和 TS bundle 在 d650cdd3 上重新构建;真实本地进程 worker;固定脚本模型加真实 qwen3.8-max。

结论

  • 合并在真实环境中是干净的。
    • 服务带着 V16、V17、V18 正常启动。Java 测试全部通过:Broker 394(1 个跳过),managed-agent-server 172。
    • H0c 新增的只接受 REFERENCED 的查找,与本 PR 预先发布的 Shell 资源并存,互不影响。部分捕获和完整捕获仍能通过 Store 逐字节读回。
    • 第七轮的每个 rig 场景结果都相同,CI 全绿。
  • 把 main 新增的 FG6c 崩溃门禁换成 Shell 调用重跑,所有用例都没有重放,也没有重复的副作用。
  • 更正第六、七轮的说法。
    • 我写过"FG6b 是手动门禁,没有 workflow 运行它",这是错的。
    • CI 的 Hosted 任务运行 -Phosted-harness-mysql,它的 failsafe include 是 **/Hosted*IT.java。所以从 018f3318 起,整个 HostedWorkspaceToolTurnIT(含 FG6a 和 FG6b)一直在 CI 中运行,任务日志里能看到 FG6B arguments: …。
    • CI 没有覆盖的是 FG6b 和 FG6c 的 Shell 变体。

合并部分

  • 冲突: ManagedSessionStore 两边都保留了。
    • PR 的 publishToolResult 路径仍以 PUBLISHED 状态发布 Shell 结果。
    • H0c 的 storedResource 只接受 REFERENCED 资源。它唯一的调用方是 commit 里的 Stage H 扩展记录钩子,所以预先发布的工具结果不会经过它。这与作者的描述一致。
  • 真实环境结果:
d650cdd3 上的检查 结果
在全新库上启动服务 ✅ 18 个迁移:V16 运行时丢失证据、V17 Session 操作、V18 扩展记录
通过 Store 读回部分捕获(cd . && { 3,000,000 × a; sleep 20; } & sleep 5; echo done) ✅ partial / producer_lost;3,000,005 B = 3,000,000 个 a 加 done;sha256 = manifest
其余探针:3 B 与 36 B 前缀、完整捕获对照、四种最小后台写法、Workspace 占用、参数校验、R1-1、8 种输出形态、F4 重新加载、真实模型重新加载 ✅ 与第七轮完全相同
换成 Shell 调用的 FG6b ✅ 6/6
CI ✅ Qwen Code CI、SDK Java、Serve A/B、tui-parity。Hosted process fault gates / MySQL 8.4 任务跑了 HostedWorkspaceToolTurnIT 3/3、HostedHarnessMySqlIT 2/2,以及含 6 个 FG6c 用例的 HostedProcessCrashIT

把 main 的 FG6c 换成 Shell 调用重跑

FG6c 让 edit 作用于一个 FIFO,使工具在执行中途阻塞。Shell 变体我用的是 c=$(cat proof.txt) && [ "$c" = x ] && rm -f proof.txt && printf %s%s "$c" "$c" > proof.txt,阻塞方式和副作用都与 edit 相同;Java 断言一行没改。

  • 原样通过: harness-prepare 和 spring-kill。
  • worker-stop: 驱动放宽一处后通过。worker 被停住时,状态轮询得到 503 managed_runtime_unavailable,因为 v3 的观察要询问 worker(30 s 超时);用 edit 时 Broker 直接用自己的记录回答。Harness 不会重试这个带 retryable: true 的 503,所以这一轮立刻阻塞,而不是等到观察窗口结束才阻塞。两种情况最终都是阻塞。
  • worker-kill: 驱动断言的是 v2 的 409 错误码;Shell 先得到 runtime_execution_evidence_unavailable,后得到 runtime_admission_closed。接受这两个码之后,Java ledger 只差一行:binding 是 LOST 而不是 READY。原因是 v3 的观察会发现 worker 已死,而 v2 路径只读 Broker 自己的记录。
  • harness-result: ledger 检查和 journal 事件检查都通过。接着"不存在 managed-tool-outcome 资源"这条断言失败。Shell 路径会在 tool_result 消息之前先提交一笔 recordToolResult 事务,其中包含 outcome 和对捕获 manifest 的引用;所以在消息提交处崩溃时,会留下一个已提交的 outcome,而 edit 路径此时还没有。这是更多的持久状态,而不是更少;冷加载仍然被阻塞。
  • harness-start: 我也在 rig 上跑了一遍,用 edit 做对照。
工具阻塞时 SIGKILL Harness,随后工具完成 Shell edit
副作用 xx,一次 xx,一次
2 s 到 120 s 后的 Broker 执行记录 UNKNOWN,无结果 SETTLED,结果已保存
捕获资源 无 不适用
冷加载该 Session 409 hosted_turn_recovery_required 相同
同一 Workspace 的新 Session 409 workspace_busy 相同

捕获发布器运行在 Harness 进程里,所以 Harness 死后 worker 无法完成捕获,一条实际已经完成的命令会被记为结果未知。这是 fail-closed 的,也没有任何重放;但这意味着 #12904 那样的恢复,无法得知以这种方式中断的 Shell 调用的结果。值得在那份设计里写明。

合并参考

Evidence (probe scripts, the FG6c Shell driver diffs, per-run logs, IT summaries): wenshao/qwen-code@60e8cf13/pr12848/r8.

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Follow-up to round 8 real-stack verification and the latest main conflict:

Finding Action
Main merged an independent fix for the same Hosted wait and added FG6d cancellation gates. Merged main in dfabe7c7c, retained the already CI-proven 10-second wait, and combined the cancellation assertions with this PR's Shell-specific checks.
A Harness crash can leave a completed Shell command recorded as UNKNOWN. Deferred recovery semantics to #12904. This PR keeps the Session and Workspace fail-closed and does not replay the command. The private Shell profile must not serve real-model traffic before safe recovery is implemented.

On dfabe7c7c, build, typecheck, formatting, and related CLI tests (50/50) passed locally. This host lacks Java/Maven, so FG6d and the merged Java gates await CI. No review threads required a reply or resolution (0 unresolved).

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /triage

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

Real-stack verification, round 9: #12848 at dfabe7c7

dfabe7c7 merges main 92f4d4f6 into d650cdd3. It brings in:

No production file of this PR changed, and Java production code is identical to d650cdd3. The two conflicts are both in test files (see the author's note).

Stack: MySQL 8.4.7; the round-8 Spring jar; the TS bundle rebuilt at dfabe7c7; real local-process workers; the fixture model plus real qwen3.8-max.

Verdict

  • The merge is correct.

    • HostedWorkspaceToolTurnIT keeps main's new cancellation branch and this PR's index < 2 check.
    • The class passes 4/4 locally on MySQL (90.7 s) and in CI, FG6d's 4 cases included.
    • Every rig scenario from round 8 gives the same result, and CI is green.
  • Cancelling a running Hosted Shell works, and cleanly. FG6d uses edit, which can't be interrupted. On the stack, a cancelled Shell call:

    • settles as cancelled in about 1.3 s with a complete capture;
    • kills its process group, including a background loop and a compound command that holds the pipe;
    • frees the Session and the Workspace.
  • One more input for Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904: for Shell, the Workspace lease is not a single-writer guarantee. A descendant that leaves both the pipe and the process group, for example via setsid with its output redirected, keeps writing:

    • after a turn that completed normally, with a complete capture and the Workspace released;
    • after a clean cancellation.

    In one run, a second Session that took the same Workspace read that descendant's write in its own command output. This fits the PR body's note that the profile is not a filesystem sandbox. It also means that holding the lease on an incomplete capture doesn't cover this case.

round 9

The merge

Check at dfabe7c7 Result
Conflict in HostedWorkspaceToolTurnIT ✅ Keeps main's cancellationProbe.assertReport branch and this PR's index < 2 guard; the six-Workspace driver's Shell Workspaces at index ≥ 2 don't write after
HostedWorkspaceToolTurnIT, whole class, MySQL 8.4.7 ✅ 4/4 in 90.7 s: six Workspaces; FG6a 8 cases; FG6b 7 cases; FG6d prepared, running, status-unavailable, cancel-reply
CI (Hosted process fault gates / MySQL 8.4) ✅ the same 4/4, plus HostedHarnessMySqlIT 2/2 and HostedProcessCrashIT (FG6c, 6 cases)
Rig: partial capture read back (3,000,005 B = manifest sha), Workspace hold and control, minimal background shapes, validation, R1-1, 8 output shapes, F4 reload, real-model reload ✅ identical to round 8

FG6d's question for Shell: cancelling a running command

Each command ran in its own Session and Workspace. I cancelled it 2 s after the Broker reported EXECUTING, then waited 35 s for late effects.

Command After cancel 35 s later Afterwards
echo started; sleep 30; echo late > late.txt idle in ~1.3 s; turn_complete(cancelled); execution SETTLED / cancelled; capture complete / committed, stdout 8 B no late.txt new Session in the Workspace works; detach + load 200
same, plus (while true; do echo tick >> ticks.txt; sleep 0.2; done) & same ticks.txt stopped growing: the loop died with the group same
cd . && sleep 60 >/dev/null 2>&1 & echo started; sleep 30; … same, in ~0.6 s: the pipe-holding subshell is in the group nothing same
perl -MPOSIX -e 'POSIX::setsid(); sleep 5; …write escaped.txt…' >/dev/null 2>&1 & echo started; sleep 30; … same escaped.txt written same

So cancellation is a working exit even for the compound background shape, as long as the foreground command is still running. The Workspace hold from rounds 6 and 7 only happens once the foreground has already exited.

Input for #12904: detached descendants outlive the lease

  1. Session A runs perl -MPOSIX -e 'POSIX::setsid(); sleep 6; …append "from-detached-descendant" to shared.txt…' >/dev/null 2>&1 & echo ok. The turn completes normally, the capture is complete, and the Workspace is released.
  2. Session B, in the same Workspace, starts right after. It runs echo from-next-session >> shared.txt; sleep 8; cat shared.txt, and its own output reads:
from-next-session
from-detached-descendant

The descendant left both the pipe (output redirected) and the process group (setsid). Neither the drain boundary nor the group kill sees it, so the capture is complete and the lease is released while it is still running.

  • Existing wording: the PR body already says the profile "is not a filesystem sandbox", and the author's round-6 reply says pipe EOF and a group kill don't prove descendants stopped.
  • What this adds:
    • The same gap exists on the success path, not only on incomplete captures.
    • So keeping the lease after an incomplete capture protects against pipe holders, but not against fully detached writers.
  • Worth stating in Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 together with round 8's UNKNOWN-after-Harness-crash case.

Merge reference

  • Unchanged: correctness shows no blocker.
  • From the real-stack side, dfabe7c7 is fine to merge on those terms.
中文版

真实环境验证(第九轮):#12848 @ dfabe7c7

dfabe7c7 把 main 92f4d4f6 合入 d650cdd3,带进了:

本 PR 的生产文件都没有变化,Java 生产代码与 d650cdd3 完全相同。两处冲突都在测试文件里(见作者说明)。

环境:MySQL 8.4.7;沿用第八轮的 Spring jar;TS bundle 在 dfabe7c7 上重新构建;真实本地进程 worker;固定脚本模型加真实 qwen3.8-max。

结论

  • 合并是正确的。

    • HostedWorkspaceToolTurnIT 既保留了 main 新增的取消分支,也保留了本 PR 的 index < 2 判断。
    • 该类在本地 MySQL 上 4/4 通过(90.7 s),CI 中同样通过,包括 FG6d 的 4 个用例。
    • 第八轮的每个 rig 场景结果都相同,CI 全绿。
  • 取消一个正在运行的 Hosted Shell 调用是有效的,而且很干净。 FG6d 用的是无法打断的 edit。在真实环境中,被取消的 Shell 调用会:

    • 在约 1.3 s 内以 cancelled 结清,捕获完整;
    • 杀掉整个进程组,包括后台循环和占着管道的复合命令;
    • 释放 Session 和 Workspace。
  • 为 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904 再提供一个输入:对 Shell 而言,Workspace 租约并不保证只有一个写入者。 一个同时脱离管道和进程组的后代(例如用 setsid 并重定向输出),会在以下情况下继续写入:

    • 一轮正常完成之后,此时捕获完整、Workspace 已释放;
    • 一次干净的取消之后。

    有一次运行中,随后接手同一 Workspace 的第二个 Session,在自己的命令输出里读到了这个后代写入的内容。这与 PR 说明中"该 profile 不是文件系统沙箱"的描述一致;它也说明,捕获不完整时占住租约,覆盖不到这种情况。

合并部分

dfabe7c7 上的检查 结果
HostedWorkspaceToolTurnIT 的冲突 ✅ 保留了 main 的 cancellationProbe.assertReport 分支和本 PR 的 index < 2 判断;六 Workspace driver 中索引 ≥ 2 的 Shell Workspace 不会写 after
HostedWorkspaceToolTurnIT 整类,MySQL 8.4.7 ✅ 4/4,90.7 s:六 Workspace;FG6a 8 个用例;FG6b 7 个用例;FG6d 的 prepared、running、status-unavailable、cancel-reply
CI(Hosted process fault gates / MySQL 8.4) ✅ 同样 4/4,另有 HostedHarnessMySqlIT 2/2 和 HostedProcessCrashIT(FG6c,6 个用例)
rig:部分捕获读回(3,000,005 B = manifest sha)、Workspace 占用与对照、最小后台写法、参数校验、R1-1、8 种输出形态、F4 重新加载、真实模型重新加载 ✅ 与第八轮完全相同

FG6d 留给 Shell 的问题:取消一个正在运行的命令

每条命令都在各自的 Session 和 Workspace 中运行。Broker 报告 EXECUTING 2 s 后取消,然后再等 35 s,看是否有迟到的副作用。

命令 取消之后 35 s 之后 随后
echo started; sleep 30; echo late > late.txt 约 1.3 s 后空闲;turn_complete(cancelled);执行记录 SETTLED / cancelled;捕获 complete / committed,stdout 8 B 没有 late.txt Workspace 中的新 Session 可用;detach + load 返回 200
同上,加上 (while true; do echo tick >> ticks.txt; sleep 0.2; done) & 同上 ticks.txt 停止增长:循环随进程组一起结束 同上
cd . && sleep 60 >/dev/null 2>&1 & echo started; sleep 30; … 同上,约 0.6 s:占着管道的子 shell 也在进程组里 无 同上
perl -MPOSIX -e 'POSIX::setsid(); sleep 5; …写 escaped.txt…' >/dev/null 2>&1 & echo started; sleep 30; … 同上 写出了 escaped.txt 同上

所以只要前台命令还在运行,取消就是一条有效的出路,对复合后台写法也一样。第六、七轮的 Workspace 占用,只在前台命令已经退出之后才会出现。

为 #12904 提供的输入:脱离的后代比租约活得更久

  1. Session A 运行 perl -MPOSIX -e 'POSIX::setsid(); sleep 6; …向 shared.txt 追加 "from-detached-descendant"…' >/dev/null 2>&1 & echo ok。这一轮正常完成,捕获为 complete,Workspace 被释放。
  2. Session B 紧接着在同一个 Workspace 中启动,运行 echo from-next-session >> shared.txt; sleep 8; cat shared.txt。它自己的输出是:
from-next-session
from-detached-descendant

这个后代同时脱离了管道(输出已重定向)和进程组(setsid)。排空边界和进程组终止都察觉不到它,所以它还在运行时,捕获就已完整、租约就已释放。

  • 已有的表述: PR 说明已经写明该 profile "不是文件系统沙箱",作者在第六轮的回复中也指出,管道 EOF 和杀进程组都不能证明后代已经停止。
  • 这次补充的是:
    • 同样的缺口在成功路径上也存在,并不只出现在捕获不完整的时候。
    • 所以,捕获不完整时保留租约,能防住占着管道的进程,但防不住完全脱离的写入者。
  • 值得和第八轮"Harness 崩溃后记为 UNKNOWN"那条一起写进 Hosted Shell: recover a Workspace lease after incomplete pipe capture #12904。

合并参考

  • 未变: 正确性方面没有阻塞项。
  • 从真实环境的角度看,dfabe7c7 可以按上述条件合并。

Evidence (probe scripts, per-run logs, the IT summary): wenshao/qwen-code@88dd8cfd/pr12848/r9.

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head dfabe7c7c7a0c9955ac9fceb9890f23f6132de09

Verdict: COMMENT. Both Criticals previously filed against this PR are confirmed fixed in code at this head, and I found no new provable Critical in the part of the diff I read. What blocks an approval is coverage, not a finding: this is 46 files and +5193/−516 adding remote foreground Shell execution, and the capture-session, resource-store and Java production layers remained outside what I could read inside this review's budget. I am naming that boundary rather than implying coverage I did not have.

Historical blocking issues — both verified FIXED at this head

The CHANGES_REQUESTED at a4c35e89 carried two sev:C findings, both in packages/cli/src/serve/managed-shell-publisher.ts. Thread replies and the later approval do not establish a fix, so I read the current file.

R1-1 (late write after finish → blocked receipt and hard turn failure on a command that exited 0) — FIXED. RemoteShellCapture now keeps per-stream ended state and a per-stream FIFO queue:

  • private readonly ended = { stdout: false, stderr: false };
  • private readonly queues = { stdout: Promise.resolve(), stderr: Promise.resolve() };
  • write() and finish() both chain onto this.queues[stream], so a write RPC can no longer be issued after that stream's finish RPC.
  • append() returns immediately on if (this.failed || this.ended[stream]) return; — the same guard the in-process sink uses, so a late data event is dropped instead of uploaded.
  • end() sets ended[stream] = true before the RPC and is idempotent.
  • finalize() starts with await Promise.all(Object.values(this.queues));, so finalization cannot race a still-queued stream.

The reported chain (owner sees entry.ended[stream], calls sink.failCapture(), receipt becomes blocked, turn throws Complete Shell output was not admitted.) is no longer reachable, because the client never emits the out-of-order write.

R1-2 (process reported without setStarted → owner rejects the pair → turn destroyed instead of not_started) — FIXED. finalize now gates the physical result on started and reports not_started, exactly as prescribed:

process:
  this.started && physical
    ? { exitCode: physical.exitCode, signal: physical.signal,
        previewBytes: physical.rawOutput.byteLength }
    : null,
executionStatus: this.started ? executionStatus : 'not_started',

A failed spawn or pre-spawn abort therefore no longer produces the {started: false, process: {...}} pair that hosted-shell-publisher.ts rejects with Unstarted Shell has a physical result. (that owner-side guard is still present, which is correct — the client now never sends the pair it rejects).

The deferred sev:S item on preview derivation is also addressed: shellPreview computes bytes once from the joined text and returns { parts, truncated }, and finalize reads both from that single call, so previewTruncated and the preview it describes can no longer diverge. Non-blocking either way.

What I verified myself at this head

Publisher authentication and binding. hosted-shell-publisher.ts mints randomBytes(32).toString('base64url'), compares the presented token with timingSafeEqual behind an explicit length check, answers 401 on mismatch, and binds this.server.listen(0, '127.0.0.1'). Route body limits are declared (SHELL_PUBLISHER_BODY_LIMIT = 256 KiB, CHUNK_BYTES = 64 KiB, 16 KiB on the managed publisher route) with cache-control: no-store.

Descriptor validation is canonical-loopback-only. parseShellPublisher requires an exact field set (url, token), a token matching /^[A-Za-z0-9_-]{43}$/, and a URL matching /^http:\/\/127\.0\.0\.1:([1-9][0-9]{0,4})\/internal\/hosted-shell-publisher\/v1$/ with the port additionally range-checked. localhost, IPv6 literals, other schemes and any path variation are all rejected before the descriptor is usable.

The gate stays narrow. The Shell profile is a second named profile beside hosted-workspace-files/1; the shell wiring is conditional on it, so default and file-profile hosts are unchanged.

CI. Every check passes at this head — Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), Serve A/B, the full Java matrix, Runtime Broker and Managed Agent MariaDB / Java 21, Hosted process fault gates / MySQL 8.4 / Java 21, Real daemon E2E / Java 11, both Desktop Shell lanes and the TUI parity/no-flicker gates. Only review-pr is pending, which is not a gating check. I saw no failure attributable to this PR.

Gates I could not confirm inside budget

These are the unconfirmed items, listed so the next pass can start from them. None is an allegation:

  1. packages/core/src/managed-runtime/resource-tool-result-store.ts (new, 366 lines) — writer-fenced publication, immutable resource ids and seal verification. The corresponding test file's fixture was separately noted as always returning the ref the store derives, so the publication-receipt check may be unwitnessed.
  2. packages/core/src/managed-runtime/managed-shell-result-session.ts (new, 436 lines) against the 393 lines removed from local-shell-result-session.ts — I did not establish that the extraction preserves the existing local behaviour.
  3. The Java production half — ManagedSessionStore (+68), WorkspaceRuntimeTransport (+97), RuntimeBrokerHttpServer (+72/−10), HttpRuntimeTransport (+23), ManagedSessionStoreController (+10) and the two ToolExecutionRepository edits. In particular the ownership scope of the new publish route and its per-kind byte limits, which were flagged as being declared in Java while the client enforces its own.
  4. The rpc helper's transport options in managed-shell-publisher.ts — I read the descriptor validation but not the fetch call itself, so I am not asserting that redirects are rejected at the fetch layer.
  5. integration-tests/helpers/hosted-workspace-tool-turn-driver.ts (+343/−11) — the two deferred notes about the store proxy's capture assertions and the session-scoped (not turn-scoped) durable-admission ordering check on a never-reset accumulator.

Per this channel's policy I do not track the remaining deferred Suggestions; they do not gate.

Next step: a pass over items 1–3 above would close this. Items 1 and 3 are the ones I would read first, since they hold the durability fencing and the route-ownership guarantees for a path that executes commands on a saved workspace.

@doudouOUC
doudouOUC added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit 0ce7b5b Sep 28, 2026
88 of 89 checks passed

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Partially reviewed — gaps disclosed.

7 Suggestion-level finding(s) this review confirmed are already reported on this PR and are not repeated:

  • R2-1 worker publisher registry never evicts its per-turn sessions/executions entries — already reported (comment 5863868679, Finding 1)
  • R2-4 integration-tests/helpers/hosted-workspace-tool-turn-driver.ts:278 durable-admission ordering check is session-scoped on a never-reset accumulator — already reported (review 5333665604 deferral list)
  • R2-5 packages/cli/src/serve/hosted-harness-session.test.ts:140 tool-set assertion thrown inside the mocked model — already reported (review 5333665604 deferral list)
  • R2-8 packages/cli/src/serve/hosted-workspace-tool-turn.ts:178 admitted Shell keys hand-listed apart from the declared schema — already reported (comment 5863868679, Finding 3)
  • R2-12 packages/core/src/managed-runtime/resource-tool-result-store.test.ts:40 publication-receipt guard unwitnessed — already reported (review 5333665604 deferral list)
  • R2-18 packages/cli/src/serve/hosted-harness-session.ts:622 publisher-cleanup catch untested — already reported (review 5333665604 deferral list)
  • R2-22 packages/cli/src/serve/hosted-harness-session.ts:332 shell wiring covered only through a diverged store mock — already reported (review 5333665604 deferral list)

Unresolved, please confirm:

  • [Critical] issue comments 5864594581 (real-stack round 6), 5867309104 (round 7), 5857539063 and 5863868679 (sandboxed verification) — ruled from their Verdict, Previous-finding status and Findings sections plus the open-defect sections, not read end t…

Not reviewed: reverse audit — stopped after the convergence pair (rounds 1 and 2, both reporting: 31 of 46 auditors returned findings) without reaching two consecutive dry rounds; the operator chose to stop the loop rather than run to the plan's 5-round cap, so the audit is not converged.

Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally (the npm integration-tests lane needs the CLI bundle, which this host cannot build: unaccepted Xcode license aborts audio-capture).

Not reviewed: build-and-test — HostedWorkspaceToolTurnIT and the Java failsafe integration lane did not run locally (need -Dnode.executable, a built dist/cli.js, Java 21 and a packaged harness); the Java findings rest on unit-level mvn runs and static traces.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 1)": enumerating every RuntimeResourceHandle construction site to confirm no handle value map can carry a null — that is the only case in which the new WriteNulls…; "agent reverse-audit (round 1)": the JIT static tier ( javac -proc:none + javap -c -p bytecode sizes) for the methods this diff grows — startExecution , prepareExecution , execution , and…; "agent reverse-audit (round 2)": whether the runtime-side publisher route ( managed-runtime-tool-v3-routes.ts / ManagedShellPublisherRegistry.register ) independently validates the descriptor…; "agent reverse-audit (round 2)": whether HttpRuntimeTransport.execute (3-arg) routes a runtimeProtocol: 3 reference to executeV3 , which is the remaining premise of my "fail-closed" conclu…; "agent test-matrix": none — all 23 assigned diff pages were read in full, and every pairing claim above was checked against the worktree at HEAD., and 5 more.

Mechanism health: this round did not close cleanly, so it withholds the incremental anchor — and the round it recovered had no anchor this round could use either — none at all, one with no certifier, one certified by an identity other than the one this round runs under, or one this round's fetch refused or resolved to the head — so the next review re-reads the whole diff unless recovery grafts an earlier own anchor that the round running it can use onto the complete work list this round leaves behind, and keeps doing so until a round's marker carries an anchor again or a graft lands that the round running it can use. (Stated, not acted on — this changes nothing about what the round posts.)

中文说明

仅完成部分审查,审查缺口已披露。

本轮确认的 7 条建议级发现已在 PR 上报告过,不再重复发布(列表见上方英文部分)。

未决,请确认:共 1 条(原文未翻译,列表见上方英文部分)。

未审查(原文为英文):reverse audit — stopped after the convergence pair (rounds 1 and 2, both reporting: 31 of 46 auditors returned findings) without reaching two consecutive dry rounds; the operator chose to stop the loop rather than run to the plan's 5-round cap, so the audit is not converged.

未审查(原文为英文):build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally (the npm integration-tests lane needs the CLI bundle, which this host cannot build: unaccepted Xcode license aborts audio-capture).

未审查(原文为英文):build-and-test — HostedWorkspaceToolTurnIT and the Java failsafe integration lane did not run locally (need -Dnode.executable, a built dist/cli.js, Java 21 and a packaged harness); the Java findings rest on unit-level mvn runs and static traces.

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 1)":enumerating every RuntimeResourceHandle construction site to confirm no handle value map can carry a null — that is the only case in which the new WriteNulls…;"agent reverse-audit (round 1)":the JIT static tier ( javac -proc:none + javap -c -p bytecode sizes) for the methods this diff grows — startExecution , prepareExecution , execution , and…;"agent reverse-audit (round 2)":whether the runtime-side publisher route ( managed-runtime-tool-v3-routes.ts / ManagedShellPublisherRegistry.register ) independently validates the descriptor…;"agent reverse-audit (round 2)":whether HttpRuntimeTransport.execute (3-arg) routes a runtimeProtocol: 3 reference to executeV3 , which is the remaining premise of my "fail-closed" conclu…;"agent test-matrix":none — all 23 assigned diff pages were read in full, and every pairing claim above was checked against the worktree at HEAD.,另有 5 条。

机制健康:本轮未能干净收尾,因而扣留了增量锚点,而它恢复到的那一轮也没有留下本轮可用的锚点——要么完全没有、要么没有认证者、要么由本轮运行身份之外的身份认证、要么被本轮的获取拒绝或解析为头提交——因此下一次评审将重读整个 diff,除非恢复流程把本轮能使用的更早自有锚点嫁接到本轮留下的完整工作清单上;并会一直如此,直到某一轮的标记重新带上锚点,或落地的嫁接能被运行该轮的评审使用。(仅陈述,不据此行动——这不改变本轮发布的任何内容。)

— qwen3.8-max via Qwen Code /review (v0.24.6)

Comment thread packages/cli/src/serve/hosted-harness-session.ts
Comment thread packages/cli/src/serve/hosted-workspace-tool-turn.ts
Comment thread packages/cli/src/serve/managed-shell-publisher.ts
Comment thread integration-tests/helpers/hosted-shell-result-reader.ts
Comment thread integration-tests/helpers/hosted-shell-result-reader.ts
@wenshao

wenshao commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Round-2 review disposition — all 39 threads handled

Every unresolved round-2 thread was re-verified against current main before being answered in place and is now resolved. Disposition:

Finding(s) Outcome
R2-33 (publisher drain gates session.active) Fixed in #13304 — availability is cleared before the unbounded drain; witness test proves hasActivePrompt === false and DELETE → 204 with a never-settling close, and fails without the fix
R2-23 (single-shot acknowledgement escalates to a blocked Session) Fixed in #13304 — one transport-failure replay, mirroring prepare(); a 409 on retry deliberately still propagates (verified: Java records stay settled and the runtime deduplicates identical receipts, so a retried 409 can only be genuine, e.g. acknowledged differently)
R2-35 (worker-side finalize without a failure path) Still stands; deferred to #13307 with analysis. The suggested local patch cannot produce a durable blocked receipt exactly when the owner cannot write — the complete reasoning is in the thread reply and the issue
R2-6 Already fixed by the merged revision itself (exact-shape assertion pins the omit direction)
35 Suggestions + worker registry eviction (round-2 dedup R2-1) Verified still standing on current main (each checked against main at the time of handling); tracked in #13307 with per-item evidence and acceptance direction

The verification evidence and provenance for each item is in its thread reply; #13307 groups them by package. The two comments the round-2 review flagged as "Unresolved, please confirm" are closed out as follows: the duplicate-V16 startup failure from real-stack round 6 was fixed on main by #12900 (V17 renumber) and confirmed green in round 7; the Session/Workspace stay blocked after an incomplete capture behavior is by design fail-closed for this private profile, with Workspace-lease recovery tracked in #12904 (open); the round-1 sandboxed verification's only Major finding was re-measured by the round-2 sandboxed run and does not reproduce (the publisher registry is keyed per turn via promptId).

For the record: getting these fixes here surfaced one false start worth naming — a full-suite run on the rebased branch showed a single failure whose identity was lost to output truncation; two subsequent identical full runs passed 381/381 each, so the suites stand green.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants