Skip to content

feat(managed-agent): Add durable remote Shell result delivery - #12894

Merged
doudouOUC merged 35 commits into
mainfrom
codex/managed-tool-result-o2-design
Sep 30, 2026
Merged

doudouOUC merged 35 commits into
mainfrom
codex/managed-tool-result-o2-design

Conversation

@doudouOUC

@doudouOUC doudouOUC commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Adds the O2 remote result path for foreground Hosted Shell calls: bounded raw stdout and stderr publication, immutable catalog and object storage, fixed-version range reads, Session receipt admission, Broker and independent worker Tool v3 routing, and Hosted recovery. The private Shell profile stays disabled until deployment supplies explicit capacity and storage settings. The existing no-tool, file-only, v2, and local capture paths remain available.

Why it's needed

O1c can preserve complete Shell output and its receipt when the Session owner and worker share a process. Hosted workers need the same original bytes and decision to survive separate processes and host replacement, without rerunning a command whose side effects may already have happened.

Reviewer Test Plan

How to verify

  • With the private Shell profile enabled and a valid original publication grant, run a foreground Shell command that emits more than 64 MiB of raw output (the model-visible preview is separately bounded to 64 KiB). Expect separate, exact stdout and stderr bytes and a fixed manifest whose tail remains readable after reopening the Session.
  • Repeat the original publish, finish, and receipt requests after simulated lost responses. Expect the same object key, reference, decision, and receipt sequence; a changed request or stale claim must be rejected without a second execution.
  • Interrupt after the receipt but before history, checkpoint, or ACK. Expect recovery from the saved receipt, a single model continuation, and an identical ACK to the original Runtime generation. Incomplete capture must block continuation.
  • Verify the old no-tool and file-only Hosted profiles, local Shell capture, and Tool v2 remain functional; an unconfigured remote publisher must refuse v3 before execution.

Evidence (Before & After)

N/A for visual evidence. Local verification passed: full build, typecheck, bundle, scoped source ESLint, formatting, focused core and CLI tests, 21 focused Java publication tests, the runtime Broker suite (374 tests, one skipped), and 11 MySQL 8.4 integration tests on a fresh schema. The 1 GiB incremental H2/file-object-store stress case passed with tail verification. The full Java server suite remains affected by one pre-existing timing-sensitive environment-event test; its isolated class passed 26/26. See the separate acceptance report comment for exact scope and limitations.

Tested on

OS Status
🍏 macOS ✅ local TypeScript, Java, and MySQL verification
🪟 Windows ⚠️ not tested locally
🐧 Linux ✅ independently verified on earlier head 0172d5eca (round 10 report); current head not covered by that run

Environment (optional)

Node.js 22, local Java 21, MySQL 8.4, H2, and a faultable local object-store double. The maintainer independently exercised real OSS and a second host on earlier head 42e2e8b1 (round 14 report); that run does not cover the later expiry-recovery protocol.

Risk & Scope

  • Main risk or tradeoff: This is a large feature change across core infrastructure and Java services. It requires maintainer architectural review, especially for immutable OSS publication, SQL lock ordering, and replay under host replacement.
  • Not validated / out of scope: The complete multi-process fault matrix and measured 100 MiB/1 GiB RSS comparison remain incomplete. Earlier-head real OSS and cross-host evidence is linked above; equivalent evidence for the new expiry-recovery protocol remains required before merge and is tracked in proposal(managed-agent): Recover expired tool publication candidates safely #13019 for maintainer review. Preview/download UI, PTY, background tasks, automatic GC, and running-orphan takeover are outside O2.
  • Breaking changes / migration notes: Three additive SQL migrations define publication ownership, object catalog state and bounded expiry-recovery attempts. This branch adds V20, V21 and V22 without changing earlier migration files; numbering was verified against the merged main ending at V19 at this head. The remote path is off by default and requires an explicitly configured private, never-versioned OSS bucket, publication endpoint, capacity, concurrency, and operation deadlines.
  • Design: English design · 中文设计 · O2a ownership design · O2a 归属设计

Linked Issues

Builds on the merged Hosted file-tool loop in #12831 and the O1a/O1b/O1c groundwork.

中文说明

这个 PR 做了什么

为前台 Hosted Shell 调用增加 O2 远程结果路径:有界发布原始 stdout 和 stderr、不可变 catalog 与对象存储、固定版本范围读取、Session 回执接纳、Broker 与独立 worker 的 Tool v3 路由,以及 Hosted 恢复。私有 Shell profile 在部署方显式提供容量和存储配置前保持关闭。现有无工具、仅文件工具、v2 和本地捕获路径仍可使用。

为什么需要

O1c 能在 Session owner 与 worker 同进程时保留完整 Shell 输出和回执。Hosted worker 分处不同进程,需要让原始字节及接纳决定在进程和宿主替换后仍可恢复,同时避免对已经产生副作用的命令再次执行。

评审验证计划

如何验证

  • 启用私有 Shell profile 并安装有效原始发布授权后,运行原始输出超过 64 MiB 的前台 Shell 命令(模型可见预览另受 64 KiB 上限约束)。预期 stdout 和 stderr 分别按原始字节精确保留,固定 manifest 的尾部在重开 Session 后仍可读取。
  • 模拟响应丢失后重放原发布、finish 和回执请求。预期复用相同对象 key、引用、接纳决定和回执 sequence;变更后的请求或过期 claim 必须被拒绝,且不得再次执行。
  • 在回执提交后、历史记录、checkpoint 或 ACK 前中断。预期从已保存的回执恢复,模型只继续一次,并向原 Runtime generation 重发相同 ACK。不完整捕获必须阻断继续。
  • 验证旧的 Hosted 无工具和仅文件工具 profile、本地 Shell 捕获及 Tool v2 仍可工作;没有配置远程 publisher 时,v3 必须在执行前拒绝。

证据(前后对比)

非可视界面改动,无截图。已通过本地验证:全仓构建、类型检查、bundle、限定源码范围的 ESLint、格式检查、core 与 CLI 定向测试、21 项 Java publication 定向测试、Runtime Broker 测试套件(374 项,跳过 1 项),以及全新 schema 上的 11 项 MySQL 8.4 集成测试。使用 H2 与本地文件对象存储的 1 GiB 增量压力案例通过尾部校验。完整 Java server 套件仍受一个既有时序敏感的环境事件测试影响;该测试所在类独立运行 26/26 通过。精确范围和限制见单独的验收报告评论。

已测试平台

操作系统 状态
🍏 macOS ✅ 本地 TypeScript、Java 和 MySQL 验证
🪟 Windows ⚠️ 本地未测试
🐧 Linux ✅ 独立验证较早的 0172d5eca 提交(第 10 轮报告);该次运行未覆盖当前 head

环境

Node.js 22、本地 Java 21、MySQL 8.4、H2 和可注入故障的本地对象存储替身。维护者已在较早的 42e2e8b1 上独立验证真实 OSS 与第二宿主(第 14 轮报告);该次运行未覆盖后续新增的过期恢复协议。

风险与范围

  • 主要风险或取舍:本次功能改动横跨核心基础设施和 Java 服务,需要维护者审查架构,尤其是不变 OSS 发布、SQL 锁顺序和宿主替换后的重放。
  • 未验证与范围外:完整多进程故障矩阵及 100 MiB 与 1 GiB RSS 对照测量仍未完成。较早 head 的真实 OSS 与跨宿主证据已在上方关联;新增过期恢复协议在合并前仍需对应证据,已由 proposal(managed-agent): Recover expired tool publication candidates safely #13019 跟踪并交由维护者评审。预览/下载界面、PTY、后台任务、自动 GC 和运行中孤儿接管不属于 O2。
  • 兼容和迁移:三个增量 SQL migration 定义发布归属、对象 catalog 状态及有界过期恢复 attempt。本分支新增 V20、V21 与 V22,不修改更早的 migration;当前 head 已与合入且截至 V19 的 main 核对编号。远程路径默认关闭,必须显式配置私有且从未启用版本控制的 OSS bucket、发布入口、容量、并发和操作期限。
  • 设计:English design · 中文设计 · O2a ownership design · O2a 归属设计

关联事项

基于已合并的 Hosted 文件工具循环 #12831,以及 O1a/O1b/O1c 基础能力。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

O2 本地验收报告(2026-09-28,fe4118f2c)

已验证:

  • 全仓 npm run build、npm run typecheck、npm run bundle;限定 packages/core/src 和 packages/cli/src 的 ESLint、变更文件 Prettier check、diff whitespace check 全部通过。普通 npm run lint 会扫描 Maven 生成的 target/site/jacoco 脚本,因此此处使用限定源码的检查。
  • core 定向 58/58、CLI 定向 72/72;包含真实 Shell 子进程的 100 MiB 捕获、慢持久化时的 EOF 排空、未完整捕获阻断、回执后 history 写入失败及 ACK 丢失的恢复。HTTP 契约序列使用受控服务端替身,不等于真实 Java 服务的全部互通验收。
  • Java publication 定向 21/21、Broker 全套 374 项(跳过 1 项)、Java checkstyle 通过。H2 catalog 与临时文件对象存储涵盖固定版本范围读取、seal/finish、故障注入和重开;可选 1 GiB 增量测试以 1024 个 1 MiB 分段写入并验证重开后的尾部字节,未采集完整 RSS 对照。
  • 全新 schema 上的 MySQL 8.4 ManagedAgentMySqlIT 10/10,通过 V17 迁移和既有 Session/Broker 存储回归。该套测试使用固定 fixture ID,复用旧 schema 会发生主键冲突;O2 专属 MySQL 跨实例竞争和完整回执故障矩阵仍待验证。

已知限制:完整 managed-agent-server 套件两次均为 158 项中 1 项失败,落在未修改的 ManagedAgentServerIntegrationTest.ignoresLateEnvironmentResultFromAnOlderTurn;同毫秒环境事件按随机 turn ID 决胜。其所在类独立重跑 26/26 通过,故不把完整套件称为通过。没有可用的专用私有 OSS bucket 和第二宿主,真实 OSS PUT/readback、跨宿主 100 MiB 尾部恢复、完整独立进程故障矩阵及 RSS 对照均未验证。保持 Draft,补齐这些证据与维护者评审后再考虑 Ready。

详细矩阵和测试边界记录在本地 .qwen/e2e-tests/managed-tool-result-o2.md(按仓库约定不跟踪)。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Independent review found three integration issues, fixed in 5f7a2be50:

  • Serialized stdout, stderr, metadata, seal, and finish operations for one publication. Retryable busy responses preserve the original operation ID and bytes.
  • Kept v3 status, cancel, and ACK on the saved original Runtime binding and lease after live Workspace authorization is lost. New grant installation and execution still require current authorization.
  • After a batch reservation failure, stop renewal, cancel prepared Broker calls, and ask the owner to close only confirmed unused reservations. The service releases capacity only with authoritative not-started proof; uncertain effects remain recovery-blocked.

Post-fix verification: full build, typecheck, bundle, scoped ESLint and formatting passed; the three relevant CLI test files passed 75/75. Independent test-engineer reran the v3 routes, publication and Hosted turn tests (62/62) and Java Workspace Runtime plus publication store tests (31/31), including the 100 MiB worker-host reopen case. The new edge tests use mocked HTTP/Broker/worker behavior. Real private OSS and second-host recovery remain unverified, so this PR stays Draft.

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Follow-up verification after merging W0e (d66fdadd2) into this branch; latest head is cdf056fe2.

  • W0e may durably fence a Broker execution as ABANDONED while the O2 publication has independently reached FINISHED. The Broker terminal state remains unchanged. Hosted admission reads the original finished publication only when its full binding matches the original reservation; a mismatched binding remains blocked. Focused tests cover normal, abandoned, and mismatched paths.
  • O2 Flyway migrations now follow W0e V16 as V17/V18. On a fresh local MySQL 8.4.11 schema, ManagedAgentMySqlIT passed 10/10 and all three migrations were applied successfully. Focused Java Broker tests passed 130/130; publication, Workspace, and migration tests passed 36/36.
  • Full repository build, typecheck, and bundle passed. Four relevant CLI test files passed 81/81, including a real worker process reopening a 100 MiB capture and a host-death recovery block. The 1 GiB incremental H2/file-object-store case and 58 focused core tests were reported earlier.

The new W0e fallback test uses controlled HTTP/Broker doubles. A real private OSS bucket and a second host are still unavailable, so this PR remains Draft. No production OSS or cross-host recovery claim is made.

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

The previous head's Hosted MySQL and Managed Agent MariaDB jobs failed for one deterministic merge-base issue: W0e and the later D4 merge both brought a V16 Flyway migration into main. Flyway rejected the duplicate before the relevant tests could run. This is not a flaky database failure.

Head 6b2bc23cb merges current main and preserves W0e as V16, then sequences D4 and O2 as V17/V18/V19. The three renamed SQL migrations are byte-for-byte unchanged. The English/Chinese D4 design, README, and migration-test description now use the corrected number.

Post-fix local evidence: H2 migration/publication tests 24/24; ManagedAgentMySqlIT 11/11 on a fresh MySQL 8.4.11 schema with V16–V19 each recorded successful; full build, typecheck, and bundle passed. New-head GitHub CI, including MariaDB and Hosted processes, is still running. The duplicate-V16 state also exists on current main, so it needs an independent upstream fix while this O2 PR remains Draft.

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

Real-stack verification of #12894 (head 6b2bc23c)

Verdict: not ready to merge. I ran the head on MySQL with the real Broker, worker and Harness processes. No Hosted Shell call completes there. The cause is four integration defects, labelled B1–B4 below: three are still present at 6b2bc23c, and one (B3) was already fixed by 5f7a2be5. None of them shows up in the unit suites, which are green (core 58 and cli 77 locally; CI's H2-backed Java jobs also pass).

I prepared a 3-file candidate patch (+33/−4) for the three open defects. With it applied, the O2 claims I could test hold on the same stack:

  • 100 MiB and 1 GiB outputs are stored byte for byte.
  • A second owner can read fixed-version ranges, including the tail.
  • Lost replies are replayed, and a lost OSS PUT reply ends in FileAlreadyExists followed by a verified readback.
  • Corrupted objects are quarantined and every later read fails closed.
  • After a Harness crash right after the receipt commit, the turn continues exactly once.

Several remaining gaps still leave a Session permanently blocked after a single transient fault; they are listed in section 3.

Setup

  • Builds and database: every build came from the PR worktree. I ran three heads: 5a45f62d, cdf056fe and 6b2bc23c, each on a fresh MySQL 8.4.7 schema through Connector/J 9.7.0.
  • Java server: the packaged Spring jar, with the Session Store, the embedded Runtime Broker and tool-publication.enabled=true.
  • Workers: real dist/cli.js managed-runtime-worker processes, spawned by the Broker.
  • Harness: the packaged dist/cli.js serve --profile hosted-harness, with the hosted-workspace-shell/1 profile.
  • Object store: the PR's own AliyunToolPublicationObjectStore with aliyun-sdk-oss 3.18.4 (SigV4 over TLS). It talks to a local OSS double at https://rig-bucket.oss-cn-hangzhou.aliyuncs.com, reached through a JVM hosts file and a private CA. The double implements GetBucketVersioning, GetBucketAcl, PutObject honouring x-oss-forbid-overwrite, and GetObject, and it can inject faults. It is not a real OSS bucket, and I used no second host.
  • Byte checks: a deterministic generator produces the output. Each stream is rebuilt from the SQL catalog plus the stored objects and compared by SHA-256 against an independent oracle. Range reads go through /range with a second writer, acquired after the Harness detached.
  • Fault injection: a proxy in front of Java carries both the worker's publication ingress and the Harness's Session Store and owner calls.
  • Exception tap: a JDI tap shows the server exceptions. Every publication refusal is returned as 400 invalid_request "The request body is invalid." and nothing is logged, so without the tap the causes are invisible.

1. Blockers at head

head blockers

B1 – every Shell reservation is refused

  • Where: hosted-workspace-tool-turn.ts:401 calls commitAwaitRuntimeBatch(bindings) without a turn identity. The latest checkpoint therefore has identity.turnId and identity.promptId set to null, and ToolPublicationStore.requireCheckpoint rejects that ("Checkpoint does not authorize dispatch").
  • Observed (6/6 across the three heads): POST …/grants returns 400, the command runs 0 times, the Session becomes recoveryBlocked, and the storage lease stays held.
  • Candidate: pass {turnId, promptId} and relabel the identity. This is the same change open PR feat(serve): add gated Hosted foreground Shell turns #12848 makes.

B2 – Broker :start fails on MySQL only

  • Where: ToolPublicationStore.java:120 casts head.get("writer_lease_until") to (Timestamp). Connector/J returns LocalDateTime from queryForMap, while H2 returns Timestamp, so only MySQL hits the ClassCastException.
  • Observed (1/1 on each head once B1 is patched):
    • :start returns 500 runtime_broker_internal_error.
    • The Harness then polls GET /executions/<id> every 100 ms (391 polls in 45 s) for up to 30 min with the prompt still active, and detach returns 409.
    • The same method authorizes every producer operation.
  • Candidate: convert LocalDateTime to Timestamp.

B3 – stdout and stderr publication race

  • What happened: the two streams' seal operations raced into "Publication is busy", which blocked the call even for echo hi.
  • Status: fixed by 5f7a2be5. After that commit, stdout+stderr echo passed 6/6.

B4 – long outputs fail on MySQL after an activation renewal

  • Where: ToolPublicationStore.java:131, which reads journal_tx with a plain SELECT.
  • Observed: a segment POST returns 500 (EmptyResultDataAccessException) 35 ms after a renewActivation commit. The Harness commits one every 20 s. The head row was read FOR UPDATE and shows the new revision, but the later plain SELECT on journal_tx uses the older REPEATABLE READ snapshot and misses it.
  • Frequency: 1 GiB runs were blocked 2/4 times without the fix. With … FOR SHARE on that query, 1 GiB runs were exact 7/7, with 0 occurrences of the exception.
  • Candidate: add FOR SHARE to that query.

Candidate patch, relative to 6b2bc23c: candidate-6b2bc23c.patch. It changes three files (+33/−4). The rest of the jar, worker and Harness is identical to the head: the jar differs only in ToolPublicationStore.class.

2. With the candidate, the O2 claims hold

candidate green

  • 100 MiB stdout + 5 MiB stderr with exit 7: 105 segments, both streams exact. A second owner's reads of head, middle (1.5 MB), tail and zero-length ranges are all exact. Past-end reads and a wrong identity are rejected, and a stale writer gets 409.
  • 1 GiB (+3 MiB stderr): exact 7/7, in 55–93 s. The worker's RSS peaks at 224 MiB, bounded; 16 MiB and 256 MiB runs peak at 199 and 240 MiB. The JVM grows from 373 to 759 MiB during a 1 GiB run under default heap settings.
  • Test plan item 3, crash after the receipt: I let the receipt commit reach the server, withheld its reply, and SIGKILLed the Harness.
    • A new Harness could load after the old writer lease expired (57–59 s).
    • It repaired the history from the saved receipt.
    • The journal holds exactly one tool.receipt, the model continued exactly once, the command ran exactly once, and the ACK to the original generation returned 200.
  • Test plan item 2, lost replies: when the reply to a segment POST or to finish is lost, the worker's status lookup returns the original receipt and no extra PUT happens.
  • OSS faults: a PUT that returns 500 once is retried by the SDK. A PUT whose bytes are stored but whose reply is lost is retried, gets FileAlreadyExists, and the readback is verified. A readback GET that returns 500 once is also retried.
  • Corruption after acceptance: flipping one byte in a stored object quarantines the publication. Every later read fails closed, even for untouched segments and after the bytes are restored.
  • Bucket versioning: reads fail closed with 500 while the bucket is Enabled or Suspended.
  • Quota: a 507 in the middle of a command gives a manifest with quota_exhausted and a preserved executionStatus, and a blocked receipt is committed.
  • Test plan item 4:
    • With the publication service unconfigured, the command runs 0 times.
    • The author's hosted-workspace-tool-turn-driver (file-tool profile) passes unchanged on 6b2bc23c + MySQL.

3. Remaining issues (measured; each one leaves the Session unusable or misleads the model)

fault matrix

  1. One transient ingress failure blocks the Session. One segment POST that is lost before reaching the server (a connection reset), or one 503, leaves the capture failed and the Session blocked. The command has already run, and 2 of its 6 segments are kept.

    • Status for an operation the server never received is 400 "Publication operation is unknown", so the worker cannot tell it apart from a refusal.
    • The worker re-sends only after a 429.
    • The design allows retry within the admitted deadline.
  2. A lost /admissions/prepare reply makes the Session permanently unrecoverable. The publication ends up FINISHED with the admission slot written but no receipt, and the 2 MiB admission allowance stays held.

    • A reload returns 409 hosted_turn_recovery_required, because the load path only resumes from a receipt.
    • The outcome also embeds a fresh messageId UUID and timestamp, so a retry would not match the stored slot.
  3. A lost /receipts/commit reply cannot be recovered in-process. The receipt did commit. detach then returns 503 and load returns 409 hosted_session_already_attached. Only a Harness restart plus lease expiry recovers, as in the crash test above.

  4. An OSS 403 stalls the command for the full operation timeout. The operation stays PENDING until operationTimeout (122 s in my configuration). During that time the Shell's pipe is paused and the command stalls; then the call is blocked.

  5. The finish digest breaks on control characters. Java stores Jackson's re-serialisation of the terminal envelope (\u001F), but the worker hashes its own JSON.stringify output (\u001f). The worker then marks the call unknown, and the ACK returns 409 managed_runtime_identity_conflict. The turn still completes through Broker reconciliation.

    • Trigger: any preview containing control characters (binary output, \b, \x0b, \x1f).
    • 10 of the 45 terminal envelopes in my runs differ this way.
  6. What the model sees from a large Hosted Shell output (image 4, real qwen3.8-max):

    model view
    • What the preview contains: for text output over 64 KiB, only the first 64 KiB. The error at the end on stderr is lost.
    • The misleading note: the preview also tells the model to read_file an absolute worker-local path under ~/.qwen/tmp. That file holds only the head, and Hosted read_file rejects absolute paths with a whole-turn InvalidWorkspaceRelativePathError.
    • Real-model trials (8):
      • 4 ended in turn_error. The recorded calls include read_file("/tmp/build_stderr.txt") and read_file("<worker>/.qwen/tmp/…/run_shell_command_*.output").
      • 3 quoted the error after the model redirected output to /tmp.
      • 1 inferred the error from the script source, saying "I'd need to run it again".
    • The build never ran more than once.
    • This is the same class of problem as feat(serve): add gated Hosted foreground Shell turns #12848 R1 F1, which that PR fixed with a head+tail preview.
  7. Some refusals are safe but end the Session for good. With the publication service unconfigured, the Harness still accepts the Shell profile at /session. Its first Shell call runs nothing, but the Session then needs recovery permanently. A quota overflow under complete_required ends the same way. This follows the design, but it matters before enablement.

  8. Diagnosability:

    • Every refusal in the store becomes an unlogged 400 invalid_request.
    • The Broker maps unexpected exceptions to 500 without logging them.
    • executeV3 swallows a 5xx from :start and polls for 30 min.
    • Range bounds are coerced silently: length: 4294967297 returns 1 byte with 200, and strings and floats are accepted.
  9. Test gap:

    • AliyunToolPublicationObjectStore has no tests. The runs above are the first execution of x-oss-forbid-overwrite, FileAlreadyExists handling and the versioning/ACL checks.
    • B1, B2 and B4 all require a MySQL-backed Shell O2 end-to-end run, like the existing file-tool driver.
    • OSS cost per 1 MiB segment is 1 PUT, 3 full-object GETs and 5 GetBucketVersioning calls; one 1 GiB run makes 3,081 GETs and 5,135 versioning calls.

Merge-order and CI notes

Not verified: a real OSS bucket, a second host, and Windows or Linux.

Evidence, rig and logs are on commit wenshao/qwen-code@41d89c53:

  • harness/: the fake OSS, the fault proxy, the JDI tap and the scenario scripts.
  • results/final-6b2bc23c.log: the complete final pass on the head.
中文版

#12894 真实环境验证(head 6b2bc23c)

结论:目前不能合并。 我在 MySQL 上用真实的 Broker、worker 和 Harness 进程跑了 head,Hosted Shell 调用一次都完成不了。原因是四个集成缺陷(下文记为 B1–B4):其中三个在 6b2bc23c 上仍然存在,另一个(B3)已由 5f7a2be5 修复。单测里一个都看不出来,单测是全绿的(本地 core 58、cli 77;CI 上基于 H2 的 Java 作业也通过)。

我针对仍未修复的三个缺陷准备了一个 3 文件的候选补丁(+33/−4)。打上补丁后,我能测的 O2 主张在同一套环境上都成立:

  • 100 MiB 和 1 GiB 输出逐字节正确落盘;
  • 第二个 owner 可以按固定版本读取范围,包括尾部;
  • 丢失的响应能重放,OSS PUT 响应丢失时走到 FileAlreadyExists 并通过回读校验;
  • 对象损坏后被隔离(quarantine),之后的读取一律失败关闭;
  • 回执提交后 Harness 立即崩溃,回合只续写一次。

但仍有几处缺口,单次瞬时故障就可能让 Session 永久阻断,见第 3 节。

环境

  • 构建与数据库: 全部从 PR worktree 构建。跑了三个 head:5a45f62d、cdf056fe、6b2bc23c,每个 head 用一个全新的 MySQL 8.4.7 schema,驱动为 Connector/J 9.7.0。
  • Java 服务: 打包好的 Spring jar,含 Session Store、内嵌的 Runtime Broker,并设置 tool-publication.enabled=true。
  • worker: 由 Broker 拉起的真实 dist/cli.js managed-runtime-worker 进程。
  • Harness: 打包好的 dist/cli.js serve --profile hosted-harness,使用 hosted-workspace-shell/1 profile。
  • 对象存储: PR 自带的 AliyunToolPublicationObjectStore,配 aliyun-sdk-oss 3.18.4(TLS 上的 SigV4),连接本地 OSS 替身 https://rig-bucket.oss-cn-hangzhou.aliyuncs.com(通过 JVM hosts 文件和私有 CA 指过去)。替身实现了 GetBucketVersioning、GetBucketAcl、遵守 x-oss-forbid-overwrite 的 PutObject 和 GetObject,并且可以注入故障。它不是真实的 OSS bucket,也没有第二台宿主。
  • 字节校验: 输出由确定性的生成器产生。每个流都从 SQL catalog 加已存对象重建,再与独立预言机比对 SHA-256。范围读取走 /range,用的是 Harness detach 之后另行获取的第二个 writer。
  • 故障注入: 在 Java 前面加了一个代理,worker 的发布入口和 Harness 的 Session Store / owner 调用都经过它。
  • 异常探针: 用 JDI 探针查看服务端异常。所有发布拒绝对外都返回 400 invalid_request "The request body is invalid.",而且不打日志,不用探针就看不到原因。

1. head 上的阻断点

B1 – 每次 Shell 预留都被拒

  • 位置: hosted-workspace-tool-turn.ts:401 调用 commitAwaitRuntimeBatch(bindings) 时没有带 turn 标识。于是最新 checkpoint 的 identity.turnId 和 identity.promptId 都是 null,被 ToolPublicationStore.requireCheckpoint 拒绝("Checkpoint does not authorize dispatch")。
  • 实测(三个 head 共 6/6): POST …/grants 返回 400,命令执行 0 次,Session 变成 recoveryBlocked,存储租约一直不释放。
  • 候选修复: 传入 {turnId, promptId} 并重标记 identity,与在飞 PR feat(serve): add gated Hosted foreground Shell turns #12848 的改动相同。

B2 – Broker 的 :start 只在 MySQL 上失败

  • 位置: ToolPublicationStore.java:120 把 head.get("writer_lease_until") 强转为 (Timestamp)。Connector/J 的 queryForMap 返回的是 LocalDateTime,H2 返回的是 Timestamp,所以只有 MySQL 会抛 ClassCastException。
  • 实测(修掉 B1 后,每个 head 1/1):
    • :start 返回 500 runtime_broker_internal_error;
    • 随后 Harness 每 100 ms 轮询一次 GET /executions/<id>(45 s 内 391 次),最长持续 30 分钟,其间 prompt 一直处于活动状态,detach 返回 409;
    • 同一个方法还负责所有 producer 操作的鉴权。
  • 候选修复: 把 LocalDateTime 转成 Timestamp。

B3 – stdout 与 stderr 的发布竞态

  • 现象: 两个流的 seal 操作并发,撞上 "Publication is busy",连 echo hi 都会被阻断。
  • 状态: 已由 5f7a2be5 修复。之后 stdout+stderr 的 echo 6/6 通过。

B4 – 长输出在 MySQL 上遇到 activation 续期即失败

  • 位置: ToolPublicationStore.java:131 用普通 SELECT 读取 journal_tx。
  • 实测: 分段 POST 在某次 renewActivation 提交 35 ms 之后返回 500(EmptyResultDataAccessException)。Harness 每 20 s 提交一次续期。head 行是用 FOR UPDATE 读的,已经能看到新修订号;但随后对 journal_tx 的普通 SELECT 用的是更早的 REPEATABLE READ 快照,看不到这一行。
  • 频率: 不修复时 1 GiB 运行 4 次中 2 次被阻断;在该查询上加 … FOR SHARE 后 7/7 字节一致,该异常出现 0 次。
  • 候选修复: 给该查询加上 FOR SHARE。

候选补丁见上文链接(相对 6b2bc23c),改 3 个文件(+33/−4)。其余的 jar、worker 和 Harness 与 head 完全一致:jar 只有 ToolPublicationStore.class 不同。

2. 打上候选补丁后,O2 主张成立

  • 100 MiB stdout + 5 MiB stderr、exit 7: 105 个分段,两个流都一致。第二个 owner 读取头部、中段(1.5 MB)、尾部和零长度范围全部一致;越界读取和错误身份被拒,陈旧的 writer 得到 409。
  • 1 GiB(另加 3 MiB stderr): 7/7 一致,用时 55–93 s。worker 的 RSS 峰值 224 MiB,是有界的;16 MiB 和 256 MiB 的运行峰值分别为 199 和 240 MiB。1 GiB 运行期间 JVM 在默认堆设置下从 373 MiB 涨到 759 MiB。
  • 测试计划第 3 项,回执后崩溃: 让回执提交到达服务端,但扣住它的响应,然后 SIGKILL Harness。
    • 旧 writer 租约过期后(57–59 s),新 Harness 可以 load;
    • 它根据已保存的回执补写了 history;
    • 日志里只有 1 条 tool.receipt,模型只续写 1 次,命令只执行 1 次,发给原 generation 的 ACK 返回 200。
  • 测试计划第 2 项,丢失响应: 分段 POST 或 finish 的响应丢失时,worker 的状态查询会返回原回执,也不会多出 PUT。
  • OSS 故障: PUT 返回一次 500 时 SDK 会重试。PUT 已落盘但响应丢失时,重试得到 FileAlreadyExists,随后回读校验通过。回读 GET 返回一次 500 时也会重试。
  • 接纳后损坏: 翻转已存对象的一个字节后,publication 被隔离;之后所有读取都失败关闭,包括未受影响的分段,以及把字节恢复原样之后。
  • bucket 版本控制: bucket 处于 Enabled 或 Suspended 时,读取以 500 失败关闭。
  • 容量: 命令执行中途收到 507 时,manifest 记录 quota_exhausted,executionStatus 保留原值,并提交一条 blocked 回执。
  • 测试计划第 4 项:
    • 未配置发布服务时,命令执行 0 次;
    • 作者的 hosted-workspace-tool-turn-driver(文件工具 profile)在 6b2bc23c + MySQL 上原样通过。

3. 仍存在的问题(均为实测;每一项都会让 Session 无法继续使用,或误导模型)

  1. 一次瞬时的入口故障就会阻断 Session。 某个分段 POST 在到达服务端前丢失(连接重置),或者收到一次 503,捕获就会失败、Session 被阻断。此时命令已经执行过,6 个分段只保留了 2 个。
    • 服务端没收到的操作,状态查询返回 400 "Publication operation is unknown",worker 无法把它和"拒绝"区分开;
    • worker 只有在收到 429 时才会重发;
    • 设计允许在准入期限内重试。
  2. /admissions/prepare 的响应丢失后,Session 永久无法恢复。 publication 停在 FINISHED,admission 槽已写入但没有回执,2 MiB 的准入额度一直被占着。
    • reload 返回 409 hosted_turn_recovery_required,因为 load 路径只能从回执恢复;
    • outcome 里嵌着新生成的 messageId(UUID)和时间戳,即使重试也对不上已存的槽。
  3. /receipts/commit 的响应丢失后,同一进程内无法恢复。 回执其实已经提交。之后 detach 返回 503,load 返回 409 hosted_session_already_attached,只有重启 Harness 并等租约过期才能恢复(即上面崩溃测试的做法)。
  4. OSS 返回一次 403,命令会卡满整个操作超时。 该操作一直停在 PENDING,直到 operationTimeout(本环境为 122 s)。这段时间里 Shell 的管道被暂停、命令卡住,之后调用被阻断。
  5. 遇到控制字符时 finish 摘要对不上。 Java 存的是 Jackson 重新序列化后的终态信封(\u001F),worker 却对自己 JSON.stringify 的结果(\u001f)求摘要。于是 worker 把调用记为 unknown,ACK 返回 409 managed_runtime_identity_conflict;回合仍能通过 Broker 对账完成。
    • 触发条件:预览里出现控制字符(二进制输出、\b、\x0b、\x1f);
    • 在我的运行里,45 个终态信封中有 10 个出现这种不一致。
  6. 大输出的 Hosted Shell 调用,模型能看到什么(图 4,真实的 qwen3.8-max):
    • 预览内容: 超过 64 KiB 的文本输出只保留前 64 KiB,stderr 末尾的报错丢失;
    • 误导性提示: 预览还提示模型用 read_file 读取 ~/.qwen/tmp 下的 worker 本机绝对路径。那个文件只有头部,而 Hosted 的 read_file 会拒绝绝对路径,整个回合以 InvalidWorkspaceRelativePathError 失败;
    • 真实模型试验(8 次):
      • 4 次以 turn_error 结束,录到的调用包括 read_file("/tmp/build_stderr.txt") 和 read_file("<worker>/.qwen/tmp/…/run_shell_command_*.output");
      • 3 次在模型把输出重定向到 /tmp 后引用出了报错;
      • 1 次是从脚本源码推断出报错,原话是 "I'd need to run it again";
    • 构建从未执行超过一次;
    • 这与 feat(serve): add gated Hosted foreground Shell turns #12848 R1 F1 属于同一类问题,那个 PR 用头 + 尾预览修复了。
  7. 有些拒绝虽然安全,却让 Session 就此作废。 未配置发布服务时,Harness 仍然在 /session 接受 Shell profile;第一次 Shell 调用不会执行任何东西,但 Session 从此一直需要恢复。在 complete_required 下容量溢出也是同样结局。这符合设计,但启用前需要考虑。
  8. 可诊断性:
    • store 里的每一种拒绝都变成不记日志的 400 invalid_request;
    • Broker 把意外异常映射成 500,也不记日志;
    • executeV3 吞掉 :start 返回的 5xx,然后轮询 30 分钟;
    • 范围参数会被静默转换:length: 4294967297 返回 1 字节并给 200,字符串和小数也会被接受。
  9. 测试缺口:
    • AliyunToolPublicationObjectStore 没有任何测试。上面这些运行是 x-oss-forbid-overwrite、FileAlreadyExists 处理以及版本控制/ACL 检查第一次被真正执行;
    • B1、B2、B4 都需要一个基于 MySQL 的 Shell O2 端到端运行才能暴露,可参照现有的文件工具驱动;
    • OSS 开销:每个 1 MiB 分段 1 次 PUT、3 次整对象 GET、5 次 GetBucketVersioning;一次 1 GiB 运行共 3,081 次 GET、5,135 次版本查询。

合并顺序与 CI

未验证:真实 OSS bucket、第二台宿主、Windows 或 Linux。

证据、测试环境和日志都在上文链接的 commit 下:

  • harness/:OSS 替身、故障代理、JDI 探针和各场景脚本;
  • results/final-6b2bc23c.log:在 head 上完整复测的日志。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Thanks for the real-stack report in this comment. I reproduced B1 with a real local Session and B2/B4 with isolated MySQL 8.4 transactions, then pushed the first blocker batch in 2599ad4cb (after merging current main in 8f7c69e42).

Finding Action
B1 checkpoint turn identity Fixed: the Hosted batch commits its original turn and prompt identity; a later call cannot relabel an unfinished turn.
B2 MySQL lease type Fixed: accept Connector/J's LocalDateTime as well as H2's Timestamp.
B3 stdout/stderr publication race Already fixed in 5f7a2be5.
B4 renewal snapshot race Fixed: read the journal revision with FOR UPDATE after locking the head. FOR SHARE worked in MySQL but failed the existing H2 tests; FOR UPDATE works in both.

Verification for this batch: Core 24/24, CLI 25/25, Java publication store 19/19; repository build, TypeScript typecheck and Java verify passed. An isolated MySQL 8.4 two-transaction reproduction confirmed that FOR UPDATE sees the revision hidden from a plain repeatable-read SELECT. These checks do not replace the full Broker/worker/Harness process run.

The other measured issues in your report remain open, especially lost ingress/admission/receipt replies, finish digest agreement, the misleading Hosted preview, and strict range validation. Real OSS and second-host evidence also remain merge gates. This update does not claim that the PR is ready to merge. The independent migration correction is now on main as #12900; #12848 remains open and its overlap will need reconciliation.

@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Real-stack verification, round 14 — R13-1 fix, real Aliyun OSS and a second host (head af962df7)

Follow-up to round 13 and the author's follow-up af962df7. This round has two parts:

  • af962df7: it changes only the O2 unstarted branch of hosted-workspace-tool-turn.ts (7 lines plus a test). I rebuilt the bundle, checked R13-1 on the real stack and ran a focused regression. The Java side is unchanged, so the jars are the ones from round 13.
  • Real Aliyun OSS and a second host, as the maintainer asked. These ran on 42e2e8b1, before af962df7 was pushed. That commit does not touch the publication, object store or recovery code, so these results carry over.

Verdict:

  • R13-1: fixed. On O2 the model now sees why a Shell call did not start. With real qwen3.8-max, the round-13 task now takes 3 calls in 23–28 s (3/3) instead of 11 blind Shell calls.
  • Real OSS: the O2 path holds against a real bucket. Every object the byte checks downloaded back from the bucket matched. One new issue, R13-2 (medium, needs a bucket misconfiguration): turning on bucket versioning while the service runs stalls each Shell turn for 4 minutes and then blocks the Session.
  • Second host: O2 crash recovery and Session hand-off work between macOS and linux-arm64 in both directions.
  • New, R14-1 (low): the new head/tail cut for long reasons can split an emoji into lone surrogates. Nothing failed on this stack; it is a small fix.
  • Regression on af962df7: clean. The results match round 13.
  • Summary: nothing blocks the merge while the profile stays disabled. R13-2 is worth fixing before the profile runs against a real bucket.

round 14 status

R13-1 fix (af962df7)

R13-1 before and after

Local path O2 at 42e2e8b1 O2 at af962df7
sleep 3; echo slept Blocked: sleep 3 followed by: echo slept. … (291 chars) Runtime Shell did not start. the same 291 chars as the local path
Ownership refused in the install window (R5-1) Workspace execution was refused before dispatch. Runtime Shell did not start. Workspace execution was refused before dispatch.
Real qwen3.8-max, "Wait 3 seconds, then print the current date with the date command. Do it in a single shell command." 2/2, 3 calls 2/2, 11 calls each, 105–114 s 3/3, 3 calls, 23–28 s
  • The real model after the fix. Each run went sleep 3 && date (blocked), then sleep 3 # intentional-sleep: …, then date. Each answer explained that the harness had required two calls.
  • Durability. The reason is stored with the Session. I detached the Session and loaded it on a new Harness process, and the replayed history held the same text.
  • Size limit. A command of 200 KB is refused earlier on both paths, with Hosted tool input exceeds the inline Session Store limit. (turn error, Session not blocked). So reasons of about 5 KB are the realistic way to reach the new branch for reasons over 4096 characters.

R14-1 (low): the head/tail cut can split a surrogate pair

af962df7 keeps errorMessage.slice(0, 2048) and errorMessage.slice(-2048) for reasons over 4096 UTF-16 units. Both cuts count code units, so an emoji or other character outside the BMP that sits on a cut point loses half of its surrogate pair. The Shell tool puts the rest of the command into the reason, so a long command with emoji is enough to trigger it:

Reason (O2, fake model) What the model received
about 4.9K ASCII 4123 units with the marker, no lone surrogates
emoji whose high surrogate sits at index 2047 lone U+D83D at 2047
emoji at index 2046 the pair is kept whole
all emoji lone U+D83D at 2047 and lone U+DE00 at 2075, where the tail starts
  • What happened next: nothing failed on this stack.
    • Every turn completed and no Session was blocked.
    • The Session Store accepted the message, and the replay on a new Harness process was identical.
    • I then loaded the two Sessions with lone surrogates, and the ASCII one as a control, on a Harness backed by real qwen3.8-max. The captured request bodies carried 1 and 2 escaped lone surrogates (0 for the control). DashScope returned 200 and the model explained the block correctly.
  • Why it still matters:
  • Test coverage: no test reaches this branch. The marker text [... error truncated ...] appears only in the source, and the new test uses a 42-character reason.
  • Suggestion: cut on code-point boundaries. cutProviderFitSlot in managed-runtime-provider-protocol.ts:572 already does this: it steps back when a cut would split a pair. Also add a test with a reason over 4096 units that has a surrogate pair at a cut point.

For comparison, the local path gives the model the same all-emoji reason cut to 530 units, ending in U+FFFD. That is main's code, not this PR's.

Real Aliyun OSS (42e2e8b1)

Setup:

  • Bucket: a temporary bucket in cn-hangzhou, private and unversioned, with a 1-day lifecycle rule.
  • Cleanup: all 1,599 objects (1.59 GB), every version and the bucket itself were deleted afterwards.
  • Object store: Spring ran the real AliyunToolPublicationObjectStore with the JDK default trust store.
  • Credentials: they came from a local file and are not in any published file.
  • Checking: after each run that checks output bytes, I downloaded every stored object from the bucket and compared its SHA-256 with the catalog and the generator's oracle.
Scenario Result
echo pass. 2 objects, 21 B. The small resources stay inline in SQL
100 MiB stdout + 5 MiB stderr, exit 7 105 segment objects downloaded and exact; every range read exact; 76 s for the turn (upload at about 12–14 Mbps)
Crash after the receipt commit, a lost segment request, a lost admission reply, control characters plus a crash, reload plus a second turn pass, bytes exact
2 O2 Sessions × 64 MiB at once exact, 0 deadlocks
Preview matrix same results as with the OSS double
Real qwen3.8-max, failing 3,000-line build 2/2, error quoted
Anonymous GET of a segment object 403
One bit flipped in stored segment 50 reads return 400. The publication becomes FENCED with quarantined=1 and segment 50 QUARANTINED. Reads of other segments are refused too
1 GiB stdout stopped at 326 MiB by the default 120 s Shell timeout, because backpressure paces the command to the upload rate. The captured 341,966,848 bytes match the generator's prefix (3341ea68…), and the model is told the command timed out. 300 MiB completed exact

During the 1 GiB run, the ACK returned one 409 managed_runtime_identity_conflict, and the Harness logged that it can be retried. It did not recur with 300 MiB, with a 5 s timeout over 60 MB, or with a 2 s timeout.

R13-2: bucket versioning turned on while running

R13-2

The constructor refuses a versioned or non-private bucket at startup, so this case only covers versioning being turned on later. I enabled versioning on the bucket and started a new O2 Shell turn:

  • The retries: requireUnversioned() refused with "Tool publication bucket cannot enforce immutable objects". The refusals were retried 153, 555, 549, 573 and 426 times in five consecutive minutes, about 2,256 in all. Each refusal first calls GetBucketVersioning on OSS (AliyunToolPublicationObjectStore.java:28-35).
  • How the turn ended: after 241 s the execution outcome became unknown. /finished returned 400, then close_not_started returned 400, and the Session ended recoveryBlocked. The publication stayed OPEN/FINISHING.
  • Existing publications: range reads returned 500 internal_error, and the catalog was unchanged.

Failing closed is correct here. The cost is a four-minute stall, a burst of OSS requests and a blocked Session for every Shell turn until someone fixes the bucket.

Suggestion: treat this refusal as a configuration error rather than a retryable one. The capture would then fail at once with a clear error, or end as a correctable not_started.

Efficiency note: every object PUT and GET also asks OSS for the bucket versioning first. That is one extra round trip per 1 MiB segment.

Second host (42e2e8b1)

Setup: the Harness ran on an Orange Pi 6 Plus (linux-arm64, Node 24.14), built from the same commit with pnpm 11.24.0. Spring, MySQL, the workers and the fake model stayed on the Mac, reached over SSH tunnels.

Scenario (O2 unless noted) Result
Crash after the receipt commit; Mac Harness killed, the Pi Harness recovers load after 48 s. The command ran once, bytes exact
The same, from the Pi to the Mac load after 55 s. The command ran once, bytes exact
Control characters plus a crash after the receipt, Mac to Pi pass
Clean hand-off, Mac → Pi → Mac 3 turns, 3 publications REFERENCED. Each turn saw its own output, and the history held 1, 2, then 3 tool results
Local capture path (#12848), Harness on the Pi the first Shell call returns 409 execution_unknown and the Session is blocked. A hand-off from the Mac to the Pi blocks the same way, and loading it back on the Mac returns 409

The local path's publisher listens on 127.0.0.1 (hosted-shell-publisher.ts:126), so it needs the Harness and the worker on one host. That is main's design, and it fails closed. Cross-host deployments need O2.

Regression on af962df7 (fake OSS)

  • O2:
    • echo writing to stdout and stderr: 2/2.
    • 100 MiB stdout + 5 MiB stderr with exit 7: exact.
    • Crash after the receipt commit: passes.
    • One lost segment request: completes.
    • One lost admission reply: completes.
    • Control characters plus a crash: passes.
    • Reload, then a second Shell turn: passes.
    • ACK request lost, then reload: passes.
    • Receipt commit delayed 12 s: completes.
  • Local path: echo 2/2; 100 MiB + 5 MiB with exit 7; reload, then a second Shell turn.
  • Mixed: 2 O2 Sessions and 2 local Sessions writing 64 MiB each at the same time: 4/4, with 0 deadlocks.
  • Earlier fixes:
    • R9-1: invalid arguments get correctable refusals on both paths.
    • R2-6: a raw ESC completes.
    • R2-5: cancelling before start gives a receipt, and reload works on both paths.
    • Gate 2: an absolute read_file path gets a correctable refusal on both paths.
  • Preview matrix and UTF-8 tail fixtures: the same as round 13 with one exception. The local-path case with 58,000 B stdout and short stderr (the F3 band) showed the truncation marker once instead of twice. That case also did so in round 7. af962df7 does not touch the local path, so I read this as run-to-run variance in F3.
  • Real qwen3.8-max, failing 3,000-line build: O2 2/2 and local 2/2. The build ran once each time, and the error was quoted.
  • Unchanged:
    • Gate 1: load still returns 409 after 150 s.
    • Lone surrogates in arguments still block O2 Sessions.
    • On the local path they do not settle within the probe's 120 s window.

Round 13 already covered the migrations and the bot merge, and af962df7 does not touch them. I did not rerun 1 GiB (it ran on real OSS above), R6-2 load scaling or the author's six-Workspace driver.

Still open (deferred)

Evidence at wenshao/qwen-code@9623cc38:

  • r14/results/s20-*.log and r14/results/s2-r14-real-sleep-o2.json: R13-1 and R14-1, with the provider request summaries in s20-r14real2-o.log;
  • r14/results/s15-r14lsi.log: the R5-1 message;
  • r14/results/r14-af962df7.log: the regression;
  • r13b/results/ro-*.log, s1-ro-m100.log, s17-ro-m100.log, s19-rov.log and ro-bucket-inventory.txt: real OSS. notes.txt explains the 1 GiB prefix check and a wrong label in s19-rov.log;
  • r13b/results/xh-42e2e8b1.log and s18-xh*.log: the second host;
  • r13b/harness/: the scripts, including the OSS helper and the remote Harness patch.
中文版

真实环境验证第 14 轮 — R13-1 修复、真实阿里云 OSS 与第二台宿主机(head af962df7)

接续第 13 轮和作者的跟进提交 af962df7。本轮分两部分:

  • af962df7: 只改了 hosted-workspace-tool-turn.ts 的 O2 未启动分支(7 行,外加一个测试)。我重新打了 bundle,在真实栈上验证 R13-1,并跑了一轮有针对性的回归。Java 侧没有变化,jar 沿用第 13 轮的。
  • 真实阿里云 OSS 和第二台宿主机(按维护者要求):这部分在 af962df7 推送之前、在 42e2e8b1 上跑完。该提交不涉及发布、对象存储和恢复代码,所以这些结论可以沿用。

结论:

  • R13-1: 已修复。O2 下模型现在能看到 Shell 调用未启动的原因。第 13 轮那个任务用真实 qwen3.8-max 重跑,3 次调用、23–28 s 完成(3/3),不再是 11 次盲目的 Shell 调用。
  • 真实 OSS: O2 路径在真实 bucket 上成立,字节校验从 bucket 下载回来的每个对象都一致。新问题 R13-2(中,需要 bucket 配置错误才会触发): 服务运行中打开 bucket 版本控制后,每个 Shell 轮次会卡 4 分钟,然后阻塞 Session。
  • 第二台宿主机: O2 的崩溃恢复和 Session 交接在 macOS 与 linux-arm64 之间双向可用。
  • 新问题 R14-1(低): 新加的长原因首尾截断可能把 emoji 切成孤立代理项。这套栈上没有出故障,修起来很小。
  • af962df7 上的回归: 干净,与第 13 轮一致。
  • 总体: profile 保持关闭时,没有阻塞合并的问题。在 profile 接入真实 bucket 之前,值得先修 R13-2。

R13-1 修复(af962df7)

本地路径 O2 @ 42e2e8b1 O2 @ af962df7
sleep 3; echo slept Blocked: sleep 3 followed by: echo slept. …(291 字符) Runtime Shell did not start. 与本地路径相同的 291 字符
install 窗口内所有权被拒(R5-1) Workspace execution was refused before dispatch. Runtime Shell did not start. Workspace execution was refused before dispatch.
真实 qwen3.8-max,"等 3 秒,然后用 date 命令打印当前日期,用一条 shell 命令完成" 2/2,3 次调用 2/2,每次 11 次调用,105–114 s 3/3,3 次调用,23–28 s
  • 修复后的真实模型: 每次都是先 sleep 3 && date(被拦),再 sleep 3 # intentional-sleep: …,再 date;回答里都说明了运行时要求拆成两次调用。
  • 持久化: 原因随 Session 存储。detach 后在一个新的 Harness 进程上 load,回放的历史文本一致。
  • 长度上限: 200 KB 的命令在两条路径上都会更早被拒,提示 Hosted tool input exceeds the inline Session Store limit.(轮次报错,Session 不阻塞)。所以约 5 KB 的原因是走到新的超长分支(超过 4096 字符)的现实途径。

R14-1(低):首尾截断可能切断代理对

af962df7 对超过 4096 个 UTF-16 单元的原因保留 errorMessage.slice(0, 2048) 和 errorMessage.slice(-2048)。两处都按代码单元计数,所以恰好落在截断点上的 emoji(或其他 BMP 以外的字符)会丢掉代理对的一半。Shell 工具会把命令的剩余部分写进原因,因此一条带 emoji 的长命令就能触发:

原因(O2,假模型) 模型收到的内容
约 4.9K ASCII 4123 个单元,带截断标记,无孤立代理项
emoji 的高位代理落在第 2047 位 第 2047 位孤立 U+D83D
emoji 落在第 2046 位 代理对完整保留
全部是 emoji 第 2047 位孤立 U+D83D,第 2075 位(尾段开头)孤立 U+DE00
  • 后续表现: 这套栈上没有出故障。
    • 每个轮次都完成了,没有 Session 被阻塞。
    • Session Store 接受了这条消息,在新 Harness 进程上回放也一致。
    • 我又把带孤立代理项的两个 Session(外加 ASCII 那个作对照)加载到接真实 qwen3.8-max 的 Harness 上。抓到的请求体里分别带 1 个和 2 个转义的孤立代理项(对照为 0),DashScope 返回 200,模型也正确解释了被拦原因。
  • 为什么仍值得修:
  • 测试覆盖: 没有测试走到这个分支。截断标记 [... error truncated ...] 只出现在源码里,新测试用的原因只有 42 个字符。
  • 建议: 按码点边界截断。managed-runtime-provider-protocol.ts:572 的 cutProviderFitSlot 已经这样做了:截断点会切断代理对时就后退一位。再补一个超过 4096 单元、截断点上带代理对的测试。

作为对照,本地路径对同样的全 emoji 原因截到 530 个单元,末尾是 U+FFFD。那是 main 的代码,不属于本 PR。

真实阿里云 OSS(42e2e8b1)

环境:

  • bucket: cn-hangzhou 的临时 bucket,私有、未开版本控制,带 1 天的生命周期规则。
  • 清理: 结束后删除了全部 1,599 个对象(1.59 GB)、所有版本和 bucket 本身。
  • 对象存储: Spring 使用真实的 AliyunToolPublicationObjectStore,信任库是 JDK 默认的。
  • 凭据: 来自本地文件,没有出现在任何公开文件里。
  • 校验方式: 每次做输出字节校验的运行结束后,把每个存储对象从 bucket 下载回来,按 SHA-256 与目录和生成器的基准比对。
场景 结果
echo 通过。2 个对象,21 B;小资源仍内联在 SQL 里
100 MiB stdout + 5 MiB stderr,exit 7 105 个分段对象下载比对一致;每次范围读都一致;本轮 76 s(上传约 12–14 Mbps)
receipt 提交后崩溃、分段请求丢一次、admission 响应丢一次、控制字符加崩溃、reload 后第二轮 通过,字节一致
2 个 O2 Session 同时各写 64 MiB 一致,0 死锁
预览矩阵 与 OSS 替身的结果相同
真实 qwen3.8-max,3,000 行失败构建 2/2,都引用了报错
匿名 GET 分段对象 403
在第 50 段存储对象里翻转一个比特 读返回 400。发布变为 FENCED、quarantined=1,第 50 段为 QUARANTINED;其他分段的读取也被拒绝
1 GiB stdout 在 326 MiB 处被默认 120 s Shell 超时终止,因为背压把命令节奏压到了上传速率。已捕获的 341,966,848 字节与生成器的前缀一致(3341ea68…),模型被告知命令超时。300 MiB 完整且一致

1 GiB 这次运行中 ACK 返回过一次 409 managed_runtime_identity_conflict,Harness 记录为可重试。用 300 MiB、60 MB 加 5 s 超时、2 s 超时都没有复现。

R13-2:运行中打开 bucket 版本控制

构造函数在启动时就会拒绝已开版本控制或非私有的 bucket,所以本场景只覆盖启动之后才打开版本控制的情况。我在 bucket 上打开版本控制,然后新开一个 O2 Shell 轮次:

  • 重试: requireUnversioned() 以 "Tool publication bucket cannot enforce immutable objects" 拒绝,拒绝后被不断重试,连续五分钟里每分钟分别是 153、555、549、573、426 次,共约 2,256 次。每次拒绝前都先向 OSS 调一次 GetBucketVersioning(AliyunToolPublicationObjectStore.java:28-35)。
  • 轮次如何结束: 241 s 后执行结果变为 unknown。/finished 返回 400,接着 close_not_started 返回 400,Session 最终 recoveryBlocked;发布停在 OPEN/FINISHING。
  • 已有发布: 范围读返回 500 internal_error,目录不变。

这里失败即关闭是对的。代价是在有人修好 bucket 之前,每个 Shell 轮次都要卡 4 分钟、打出一大批 OSS 请求,并阻塞 Session。

建议: 把这种拒绝当作配置错误而不是可重试错误。这样捕获会立即失败并给出清楚的错误,或者以可纠正的 not_started 结束。

效率备注: 每次对象 PUT 和 GET 之前都先向 OSS 查询一次 bucket 版本控制状态,即每个 1 MiB 分段多一次往返。

第二台宿主机(42e2e8b1)

环境: Harness 跑在 Orange Pi 6 Plus 上(linux-arm64,Node 24.14),用同一提交、pnpm 11.24.0 构建。Spring、MySQL、worker 和假模型留在 Mac 上,通过 SSH 隧道访问。

场景(未注明即 O2) 结果
receipt 提交后崩溃:杀掉 Mac 上的 Harness,由 Pi 上的 Harness 恢复 48 s 后 load 成功;命令只跑了一次,字节一致
同上,从 Pi 到 Mac 55 s 后 load 成功;命令只跑了一次,字节一致
控制字符加 receipt 后崩溃,Mac → Pi 通过
正常交接 Mac → Pi → Mac 3 轮,3 个发布都是 REFERENCED;每轮都看到自己的输出,历史里的工具结果依次为 1、2、3 个
本地捕获路径(#12848),Harness 在 Pi 上 第一次 Shell 调用返回 409 execution_unknown,Session 被阻塞。从 Mac 交接到 Pi 也同样阻塞,再回 Mac load 返回 409

本地路径的 publisher 监听在 127.0.0.1(hosted-shell-publisher.ts:126),所以要求 Harness 和 worker 在同一台机器上。这是 main 的设计,并且是失败即关闭。跨宿主机部署需要 O2。

af962df7 上的回归(fake OSS)

  • O2:
    • echo 同时写 stdout 和 stderr:2/2。
    • 100 MiB stdout + 5 MiB stderr、exit 7:一致。
    • receipt 提交后崩溃:通过。
    • 分段请求丢一次:完成。
    • admission 响应丢一次:完成。
    • 控制字符加崩溃:通过。
    • reload 后第二个 Shell 轮次:通过。
    • ACK 请求丢失后 reload:通过。
    • receipt 提交延迟 12 s:完成。
  • 本地路径: echo 2/2;100 MiB + 5 MiB、exit 7;reload 后第二个 Shell 轮次。
  • 混合: 2 个 O2 Session 和 2 个本地 Session 同时各写 64 MiB:4/4,0 死锁。
  • 之前的修复:
    • R9-1:非法参数在两条路径上都得到可纠正的拒绝。
    • R2-6:原始 ESC 能完成。
    • R2-5:启动前取消得到 receipt,两条路径 reload 都正常。
    • Gate 2:绝对路径的 read_file 在两条路径上都得到可纠正的拒绝。
  • 预览矩阵和 UTF-8 尾部夹具: 与第 13 轮一致,只有一个例外:本地路径上 58,000 B stdout 加短 stderr 的用例(F3 区间)截断标记出现 1 次而不是 2 次。这个用例第 7 轮也出现过同样情况。af962df7 不涉及本地路径,所以我判断为 F3 的逐次波动。
  • 真实 qwen3.8-max,3,000 行失败构建: O2 2/2、本地 2/2。每次构建只跑一次,报错都被引用。
  • 未变:
    • Gate 1:load 在 150 s 后仍返回 409。
    • 参数里的孤立代理项仍会阻塞 O2 Session。
    • 在本地路径上它们在探针的 120 s 窗口内不会结束。

迁移和机器人合并第 13 轮已经覆盖,af962df7 也没有涉及。本轮没有重跑 1 GiB(上面在真实 OSS 上跑过)、R6-2 加载规模和作者的六 Workspace 驱动。

仍未解决(已延后)

证据见 wenshao/qwen-code@9623cc38(清单同英文版)。

@doudouOUC

doudouOUC commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

Review fixes and main synchronization

Review item Action Commit
R4-1: expired segment candidate / frozen finish cannot recover Fixed with explicit authenticated, bounded verification recovery; new claim epoch, immutable original operation/bytes/key/ref/envelope/deadline/quota e177e1d
R2-4: definite Workspace refusal permanently blocks resume Fixed shared exact-409 classification and acquisition before load attachment; later load resumes the original durable continuation e177e1d
Main conflicts after D6a approvals Merged main and retained approval before v3 prepare/publication reservation, refusal ordinal identity and O2 recovery 0027f7a

Final verification on 0027f7a7f: build, typecheck and bundle pass; independent CLI 165/165; core harness 31/31; Java focused unit/schema/controller 31/31; real MySQL 8.4.11 recovery races 3/3; Checkstyle zero violations. The review-fix delta and merge resolutions completed two consecutive clean open-ended self-audit passes.

The MySQL races use two store instances and an immutable object-store fixture. They prove original slot/key/ref/digest/deadline/quota preservation and refusal of late old-epoch installation for segment and terminal variants. This is not real OSS or cross-host evidence; those acceptance requirements remain tracked in #13019 and the O2 rollout verification. Recovery does not reset FINISHING or replace invalid/corrupt results, and does not re-execute Shell.

中文:已采纳并修复两条 Critical,同时解决与主线审批接线的冲突。最终构建、类型检查、bundle、CLI/core/Java 定向回归和真实 MySQL 恢复测试通过;完成两轮无新问题的自审。真实 OSS/跨宿主证据仍需补齐,未把测试替身计作该类验收。

New round-14 feedback

I checked round 14 against this head. Both Suggestions are recorded in #12986 under the repository's review-round cap:

Item Action Reason / next verification
R13-2: enabling OSS bucket versioning after startup causes repeated configuration refusals Deferred to #12986 Default-off profile still fails closed; classify configuration drift distinctly and stop retries promptly before enabling deployment. Preserve the physical Shell outcome and original object identity.
R14-1: long reason head/tail truncation splits surrogate pairs Deferred to #12986 No failure on the reviewed stack; add code-point-safe boundaries and bounded durable-history regression separately.

The reviewer's older-head real OSS/two-host tests are useful evidence; they do not verify the newly added expiry recovery protocol. That remaining evidence stays in #13019. Both original inline Critical threads are now resolved (2/2), with zero unresolved review threads. The PR is mergeable; new-head CI and renewed maintainer review are pending. One Qwen Autofix routing run was cancelled before any steps/logs, consistent with its per-PR cancel-in-progress routing; no code failure was observed and no rerun was requested.

中文:新增两项非阻塞建议已在 #12986 记录后续验收,原两条 Critical 已修复并解决。PR 无冲突,等待新 head 的 CI 和维护者复审。

Triage re-run follow-up

[codex] The current-head static re-run independently confirms R2-4 and R4-1 are fixed. The PR description is corrected in English and Chinese to V20/V21/V22 and links the current design. It distinguishes the earlier-head real OSS/two-host report from the new recovery protocol's remaining evidence in #13019.

New / repeated Suggestion Decision Tracking
Publication response size guard releases its reader without cancelling the remaining body Confirmed by source; defer under the review-round cap #12986
Recovery IT class name routes it to MariaDB CI rather than the Hosted MySQL CI job Confirmed; local MySQL 8.4.11 3/3 remains valid, but is not MySQL CI evidence #12986

Current live GraphQL returns all 43 review threads, hasNextPage=false, with zero unresolved; both R5-1 thread IDs are resolved. This supersedes the re-run's stale thread counts. There are no new blocking code findings. The standing CHANGES_REQUESTED remains for maintainer reassessment; the sandboxed verification job is still running. A successful recovery acquisition clears uncertainty; ambiguous/lost acquisition remains recovery-required and is not claimed to release uncertain Workspace ownership.

中文:本轮独立核验确认两条 Critical 已修复;双语 PR 描述及设计链接已更新,非阻塞的响应体清理与 CI 数据库路由已记入 #12986。实时 API 确认 43/43 线程已解决。仍等待维护者复审及当前沙箱验证,不将旧 head 的 OSS 证据计作新增恢复协议验收。

@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /triage

@doudouOUC
doudouOUC requested a review from chiga0 September 30, 2026 10:44
wenshao pushed a commit to wenshao/qwen-code that referenced this pull request Sep 30, 2026
wenshao pushed a commit to wenshao/qwen-code that referenced this pull request Sep 30, 2026
@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Real-stack verification, round 15 — R4-1 and R2-4 fixes, with real Aliyun OSS and a second host (head 0027f7a7)

Follow-up to round 14 and the author's review fixes and main sync:

  • e177e1da: the R4-1 fix adds recovery for expired publication operations, and the R2-4 fix classifies Workspace refusals when a Session resumes.
  • 0027f7a7: a main merge that brings in the D6a approvals.

The author noted that R4-1 still lacked real-OSS and cross-host evidence, so this round covers both. I rebuilt the bundle and jar for 0027f7a7, and also built af962df7 (the head before these commits) for A/B comparisons.

Verdict:

  • Migrations: V22 applies on a fresh schema, on main's V19 schema and on the round-14 V21 schema.
  • R4-1: fixed only when the producer loses the reply. The new /recover path works on the fake OSS and on real OSS: bytes stay exact and quota is unchanged.
    • New, R15-1 (high for enabling O2): when an operation runs past its deadline while it is still in flight, the server answers 400 invalid_request. The producer treats that as final and never calls /recover, so the Session is blocked exactly as before the fix.
    • The blocked Session also keeps the Workspace. Other Sessions there get 409 workspace_busy; I still saw this 80 minutes later.
    • On real OSS this happens with no injected fault once a stream seal or the finish takes longer than the operation timeout, because both re-read the whole capture.
  • R2-4: a resume whose Workspace acquisition fails no longer leaves a blocked Session attached; the next load resumes the turn. The exact R2-4 trigger (a definite 409 workspace_busy during a resume) cannot be reached on this stack, so the unit tests remain the evidence for that mapping.
  • Approvals merged into O2: they work on O2. A denied or expired call runs nothing, reserves nothing and makes no Broker call.
  • Second host (linux-arm64): all pass.
  • Pre-existing, found this round: after a Broker restart, a Session whose last Shell call was not acknowledged cannot resume, on either path. Its stale Workspace lease also blocks other Sessions. This predates e177e1da.
  • Regression: clean. The results match round 14.
  • Summary: nothing blocks the merge while the profile stays disabled. R15-1 needs fixing before O2 is enabled, because R4-1 is closed only for lost replies.

round 15 status

R15-1: an operation that outlives its deadline in flight still wedges the capture

R4-1 matrix

Setup:

  • Fake OSS with operation-timeout 6s and claim-timeout 3s.
  • A fault proxy arms a fault in the object store when a named publication request passes. For example, "the object PUT for segments/stdout/2 takes 9 s".
  • "Blocked" means the Session ends recoveryBlocked.
Case before (af962df7) after (0027f7a7) candidate (below)
F2: segment PUT takes 9 s once blocked blocked, 400 recovered, exact
F3: one finish read-back GET takes 9 s FINISHING, blocked FINISHING, blocked recovered, exact
F4 / F5: the same 8 times in a row blocked blocked 3 recoveries, then blocked
F6: PUT takes 4.5 s (only the 3 s claim lapses) blocked blocked blocked
F7: F2, and the producer loses the reply blocked recovered, exact recovered, exact
F8: F3, and the producer loses the reply blocked recovered, FINISHED, exact recovered, exact

Relation to #13019: its acceptance scenario is the one where the first PUT response is lost, the deadline expires and the command is not repeated. That now passes on the real stack, including real OSS (F7/F8 here and RF7/RF8 below). R15-1 is the other case, where the response does arrive, as a 400.

Why:

  • Server: the in-flight checks throw IllegalArgumentException("… claim expired"), which becomes 400 invalid_request. They are in install() (ToolPublicationDataStore.java:1127), installFinish() (:660), the heartbeat (:952) and checkScanClaim() (:837); the JDI tap shows install:1125. Only the start-of-request checks were changed to requireUnexpired (409 managed_tool_publication_operation_expired).
  • Producer: remote-shell-result-publication.ts:251-259 rethrows any 4xx other than that code unless the observed status is SUCCEEDED. The status at that point is already EXPIRED, but the recovery loop at :272 is never reached.
  • Test coverage: the new Java test holds the first write, moves the deadline back and checks a second store instance. It never asserts what the held request itself answers.

Real Aliyun OSS:

  • Bucket: a temporary private bucket, deleted afterwards (367 objects, 372 MB).
  • Fault injection: the OSS SDK ignores https.proxyHost. So the JVM resolved the OSS host names to 127.0.0.1, and a TCP forwarder relayed the TLS bytes unchanged to the real OSS address, holding one chunk for 9 s when armed.
Case Result (after)
RF2: segment PUT body held 9 s 400, blocked
RF7: RF2, and the producer loses the reply recovered. The delayed first PUT had already created the object before the retry; the bytes read back exact
RF3: finish read-back body held 9 s 400, FINISHING, blocked
RF8: RF3, and the producer loses the reply recovered, FINISHED, bytes exact
RT1: 100 MiB + 5 MiB at 120 s / 30 s, no fault complete. The finish took 15.4 s; segment requests had a median of 255 ms and a slowest of 5.2 s
RT2: the same at 10 s / 5 s, no fault the stdout seal re-read 100 MiB in 10.5 s, got 400, and the Session was blocked with an incomplete capture
RT3: RT2 with the candidate the seal took 8.0 s. The finish then got 400 after about 10.3 s four times (3 recoveries) and stayed FINISHING, blocked

What the real-OSS runs show:

  • The seal and the finish both re-read the whole stream from OSS.
  • From this Mac, 110 MB took 15 s. At the rig's 120 s timeout, a capture of roughly 800 MiB would outlive the finish at that bandwidth, and the rig's execution limit is 2 GiB. In-region bandwidth is higher, so in production the likelier trigger is a stall longer than the remaining deadline.
  • Recovery restarts that read from the beginning. So when the read-back itself is longer than the timeout, recovery cannot finish it (RT3).

Candidate (tested, not proposed as the whole fix):

  • The change: one condition in the producer, so that an observed EXPIRED after a 4xx goes to the recovery loop:

                 failure.code !== 'managed_tool_publication_operation_expired' &&
    -            status['state'] !== 'SUCCEEDED'
    +            status['state'] !== 'SUCCEEDED' &&
    +            status['state'] !== 'EXPIRED'
  • What it fixes: single stalls (F2, F3), while staying bounded (F4, F5).

  • What it does not fix:

    • A lapse of the claim alone (F6). That needs a retryable code from the server when the deadline has not passed. I did not let the producer retry RETRYABLE after a 4xx, because a deterministic refusal would then loop until the 30-minute client deadline.
    • A seal or finish whose read-back is longer than the timeout (RT3). That needs either verification progress that survives a recovery, or a verification window sized to the bytes.
  • Suggestion: return 409 managed_tool_publication_operation_expired from the in-flight checks once the deadline has passed, and a distinct retryable code when only the claim lapsed. Also add a test that asserts what the held request answers.

R2-4: resume classification

R2-4, Broker restart, approvals

Setup: Harness A is killed after the receipt commit. Harness B loads the Session once A's writer grant has lapsed (65 s).

Case before after
The Broker's reply to the resume's acquire is lost load 200, then recoveryBlocked; prompts get 409 hosted_turn_recovery_required, and a third Harness is needed load 503 managed_session_open_failed with the Session not attached; the load 5 s later returns 200 and the turn completes. The command ran once
The same, with B on the second host — 503, then 200, completes
The Workspace lease is handed to another holder first 200, completes 200, completes
  • Why the busy case does not trigger: the Broker re-acquires the existing, still-live runtime session without calling WorkspaceExecutionStore.claim(). That is why the tampered lease is never checked, and why a definite 409 workspace_busy during a resume cannot be produced here.
  • Releasing A's runtime session by hand only leads to 404 runtime_session_not_found on the recovery ACK (next section).
  • Load latency: load now waits for the resume's acquisition before answering.

Pre-existing: Broker restart with an unacknowledged execution

I restarted Spring (the Broker) between A's crash and B's load:

  • O2 before: load returns 200, the recovery ACK gets 404 runtime_session_not_found, and the Session is blocked.
  • O2 after: every load takes about 30 s and returns 503 (9 loads in 4 minutes).
  • Local capture path: load returns 409 hosted_turn_recovery_required every time.
  • The Workspace: in all three, the lease still named the dead runtime session 12 minutes later, and a new Session in that Workspace got 409 workspace_busy.

This is not caused by e177e1da. On O2, though, the receipt is already committed, so a definite 404 on the recovery ACK could count as settled rather than blocking the Session.

Approvals merged into O2 (0027f7a7)

With approvalMode: default, each run_shell_command asks first:

  • allow: the command runs once, and the publication is REFERENCED.
  • deny: the model sees "The Session owner denied this tool call, so it was not run." There is no Broker call and no publication row.
  • Two Shell calls in one reply, allow then deny: the first runs, and the second is refused against its own call id. This holds on O2 and on the local path.
  • A 3 s timeout with no answer: the call is refused as expired and not run.
  • Harness killed while an approval is pending: load returns 409 hosted_turn_recovery_required. The same happens on the local path, so this is main's D6a behavior.

Second host (0027f7a7 on the Orange Pi)

The Harness ran on linux-arm64 with Node 24; everything else stayed on the Mac:

  • Crash after the receipt commit, Mac → Pi and Pi → Mac: load after 56–57 s. The command ran once and the bytes were exact.
  • Clean hand-off Mac → Pi → Mac: 3 publications REFERENCED, and the history held 1, 2, then 3 tool results.
  • Lost acquire reply during a resume, with B on the Pi: 503, then 200, and the turn completes.
  • Approvals (allow; allow + deny): correct.

Regression on 0027f7a7 (fake OSS, fresh schema)

  • O2:
    • echo writing to stdout and stderr: 2/2.
    • 100 MiB stdout + 5 MiB stderr with exit 7: exact.
    • Crash after the receipt commit: passes.
    • One lost segment request: completes.
    • One lost admission reply: completes.
    • Control characters plus a crash: passes.
    • Reload, then a second Shell turn: passes.
    • ACK request lost, then reload: passes.
    • Receipt commit delayed 12 s: completes.
  • Local path: echo 2/2; 100 MiB + 5 MiB with exit 7; reload, then a second Shell turn.
  • Mixed: 2 O2 Sessions and 2 local Sessions writing 64 MiB each at the same time: 4/4, with 0 deadlocks.
  • Earlier fixes: R9-1, R2-6, R2-5 and Gate 2 behave as in round 14.
  • Preview matrix and UTF-8 tail fixtures: the same as round 13. The F3-band case that varied in round 14 showed two markers again.
  • Real qwen3.8-max, failing 3,000-line build: O2 2/2 and local 2/2. The build ran once each time, and the error was quoted.
  • Unchanged:
    • Gate 1: load still returns 409 after 150 s.
    • Lone surrogates in arguments still block O2 Sessions.

Not rerun: 1 GiB, R6-2 load scaling and the author's six-Workspace driver. The unstarted-reason code (R13-1, R14-1) is unchanged since round 14.

Merge coordination

Other in-flight PRs add Flyway migrations with the same numbers:

Git does not report these clashes; Flyway fails at startup.

Still open (deferred)

Evidence at wenshao/qwen-code@5871c7d5:

  • results/notes.txt maps every file to its run;
  • results/r15-r41*.log, jdi-claim-expired-excerpt.txt and candidate-client-change.diff: R15-1 on the fake OSS;
  • results/ro15*.log and ro15-bucket-inventory.txt: real OSS;
  • results/r15-r24*.log: R2-4 and the Broker restarts;
  • results/s22-*.log: approvals;
  • results/xh-0027f7a7.log: the second host;
  • results/flyway-checks-r15.txt: migrations;
  • results/r15-0027f7a7.log: the regression;
  • harness/: the scripts, including the fault proxy, the fake OSS and the TCP forwarder.
中文版

真实环境验证第 15 轮 — R4-1 与 R2-4 修复,含真实阿里云 OSS 与第二台宿主机(head 0027f7a7)

接续第 14 轮和作者的评审修复与 main 同步:

  • e177e1da: R4-1 的修复为过期的发布操作增加了恢复;R2-4 的修复对 Session 续跑时的 Workspace 拒绝做了分类。
  • 0027f7a7: 一次 main 合并,带入了 D6a 审批。

作者说明 R4-1 还缺真实 OSS 和跨宿主证据,本轮两者都覆盖了。我为 0027f7a7 重新打了 bundle 和 jar,同时构建了 af962df7(这两个提交之前的 head)用于 A/B 对照。

结论:

  • 迁移: V22 在全新库、main 的 V19 库和第 14 轮的 V21 库上都能应用。
  • R4-1:只有生产端丢失应答时才修好了。 新的 /recover 路径在 fake OSS 和真实 OSS 上都能用:字节一致,配额不变。
    • 新问题 R15-1(对启用 O2 而言属于高): 操作在执行过程中超过期限时,服务端返回 400 invalid_request。生产端把它当成最终结果,从不调用 /recover,Session 与修复前一样被阻塞。
    • 被阻塞的 Session 还会一直占着 Workspace,那里的其他 Session 会得到 409 workspace_busy;80 分钟后我仍看到这种情况。
    • 在真实 OSS 上,只要流 seal 或 finish 超过操作超时,不注入任何故障也会出现,因为两者都要把整段 capture 重新读一遍。
  • R2-4: Workspace 获取失败的续跑不再挂上一个被阻塞的 Session,下一次 load 就能续完这一轮。R2-4 的确切触发条件(续跑时明确的 409 workspace_busy)在这套栈上无法触发,这个映射仍以单元测试为证据。
  • 审批合入 O2: 在 O2 上工作正常。被拒绝或过期的调用什么都不执行、不预留、也不调用 Broker。
  • 第二台宿主机(linux-arm64): 全部通过。
  • 本轮发现的既有问题: Broker 重启后,最后一次 Shell 调用未确认的 Session 在两条路径上都无法续跑;它残留的 Workspace lease 还会阻塞其他 Session。这在 e177e1da 之前就存在。
  • 回归: 干净,与第 14 轮一致。
  • 总体: profile 保持关闭时,没有阻塞合并的问题。启用 O2 之前需要修复 R15-1,因为 R4-1 只在丢失应答的情况下被解决。

R15-1:执行中超过期限的操作仍会卡死 capture

环境:

  • fake OSS,operation-timeout 6s、claim-timeout 3s。
  • 故障代理在指定的发布请求经过时,向对象存储布置一个故障,例如"segments/stdout/2 的对象 PUT 耗时 9 s"。
  • "阻塞"指 Session 最终处于 recoveryBlocked。
场景 修复前(af962df7) 修复后(0027f7a7) 候选改动(见下)
F2:分段 PUT 一次耗时 9 s 阻塞 阻塞,400 恢复,一致
F3:finish 读回的一次 GET 耗时 9 s FINISHING,阻塞 FINISHING,阻塞 恢复,一致
F4 / F5:同样的故障连续 8 次 阻塞 阻塞 3 次恢复后阻塞
F6:PUT 耗时 4.5 s(只有 3 s 的 claim 过期) 阻塞 阻塞 阻塞
F7:F2 且生产端丢失应答 阻塞 恢复,一致 恢复,一致
F8:F3 且生产端丢失应答 阻塞 恢复,FINISHED,一致 恢复,一致

与 #13019 的关系: 它的验收场景是第一个 PUT 的应答丢失、期限过期、命令不被重复执行。这个场景现在在真实栈上通过了,包括真实 OSS(这里的 F7/F8 和下面的 RF7/RF8)。R15-1 是另一种情况:应答确实回来了,只是一个 400。

原因:

  • 服务端: 执行中的检查抛出 IllegalArgumentException("… claim expired"),对外变成 400 invalid_request。这些检查位于 install()(ToolPublicationDataStore.java:1127)、installFinish()(:660)、heartbeat(:952)和 checkScanClaim()(:837);JDI 探针显示抛出点在 install:1125。只有请求开始时的检查改成了 requireUnexpired(409 managed_tool_publication_operation_expired)。
  • 生产端: remote-shell-result-publication.ts:251-259 对这个码以外的 4xx 一律重新抛出,除非查到的状态是 SUCCEEDED。此时状态其实已经是 EXPIRED,但 :272 的恢复循环永远走不到。
  • 测试覆盖: 新的 Java 测试扣住第一次写入、把期限改到过去,然后检查第二个 store 实例;它没有断言被扣住的那个请求本身返回什么。

真实阿里云 OSS:

  • bucket: 临时私有 bucket,结束后已删除(367 个对象、372 MB)。
  • 故障注入: OSS SDK 不读取 https.proxyHost。所以让 JVM 把 OSS 域名解析到 127.0.0.1,由一个 TCP 转发器把 TLS 字节原样转给真实 OSS 地址,布置后扣住一个数据块 9 s。
场景 结果(修复后)
RF2:分段 PUT 请求体被扣 9 s 400,阻塞
RF7:RF2 且生产端丢失应答 恢复。重试之前,被延迟的第一次 PUT 已经创建了对象;读回的字节一致
RF3:finish 读回的响应体被扣 9 s 400,FINISHING,阻塞
RF8:RF3 且生产端丢失应答 恢复,FINISHED,字节一致
RT1:100 MiB + 5 MiB,120 s / 30 s,无故障 完成。finish 用时 15.4 s;分段请求中位数 255 ms,最慢 5.2 s
RT2:同上但 10 s / 5 s,无故障 stdout seal 用 10.5 s 重读 100 MiB,得到 400,Session 因 capture 不完整被阻塞
RT3:RT2 加候选改动 seal 用时 8.0 s。之后 finish 四次都在约 10.3 s 后得到 400(3 次恢复),停在 FINISHING,阻塞

真实 OSS 的结果说明:

  • seal 和 finish 都会把整条流从 OSS 重新读一遍。
  • 从这台 Mac 读 110 MB 用了 15 s。按装置的 120 s 超时和这个带宽,约 800 MiB 的 capture 就会让 finish 超过期限,而装置配置的执行上限是 2 GiB。同区域带宽更高,所以生产环境里更可能的触发条件是一次停顿超过了剩余期限。
  • 恢复会从头开始读,所以当读回本身就比超时长时,恢复也完成不了(RT3)。

候选改动(已实测,但不是完整修复):

  • 改动: 生产端的一个条件,让 4xx 之后查到的 EXPIRED 进入恢复循环。
  • 能修好的: 单次停顿(F2、F3),并且有上限(F4、F5)。
  • 修不好的:
    • 只有 claim 过期(F6)。这需要服务端在期限未到时返回一个可重试的码。我没有让生产端在 4xx 之后重试 RETRYABLE,因为遇到确定性的拒绝时会一直循环到客户端 30 分钟的总期限。
    • 读回时间长于超时的 seal 或 finish(RT3)。这需要让校验进度在恢复后保留下来,或者让校验窗口随字节数调整。
  • 建议: 期限已过时,执行中的检查返回 409 managed_tool_publication_operation_expired;只有 claim 过期时返回另一个可重试的码。再补一个断言"被扣住的请求本身返回什么"的测试。

R2-4:续跑时的分类

环境: Harness A 在 receipt 提交后被杀;A 的写者授权过期后(65 s),Harness B 加载这个 Session。

场景 修复前 修复后
Broker 对续跑 acquire 的应答丢失 load 200,随后 recoveryBlocked;prompt 得到 409 hosted_turn_recovery_required,需要第三个 Harness load 返回 503 managed_session_open_failed,Session 未挂上;5 s 后的 load 返回 200,轮次完成。命令只执行了一次
同上,B 在第二台宿主机上 — 503,然后 200,完成
先把 Workspace lease 交给另一个持有者 200,完成 200,完成
  • 为什么触发不了 busy: Broker 重新获取仍然存活的 runtime session 时,不会调用 WorkspaceExecutionStore.claim()。所以被篡改的 lease 从未被检查,续跑时也就产生不了明确的 409 workspace_busy。
  • 手动释放 A 的 runtime session: 只会让恢复 ACK 得到 404 runtime_session_not_found(见下一节)。
  • load 延迟: load 现在要等续跑的 acquire 完成后才返回。

既有问题:Broker 重启时有未确认的执行

我在 A 崩溃和 B load 之间重启了 Spring(即 Broker):

  • O2 修复前: load 返回 200,恢复 ACK 得到 404 runtime_session_not_found,Session 被阻塞。
  • O2 修复后: 每次 load 约 30 s 后返回 503(4 分钟内 9 次)。
  • 本地捕获路径: load 每次都返回 409 hosted_turn_recovery_required。
  • Workspace: 三种情况下 lease 12 分钟后仍指向已死的 runtime session,同一 Workspace 的新 Session 得到 409 workspace_busy。

这不是 e177e1da 引起的。不过在 O2 上 receipt 已经提交了,所以恢复 ACK 收到明确的 404 时,可以当作已结算,而不是阻塞 Session。

审批合入 O2(0027f7a7)

在 approvalMode: default 下,每次 run_shell_command 都会先询问:

  • 允许: 命令执行一次,发布为 REFERENCED。
  • 拒绝: 模型看到 "The Session owner denied this tool call, so it was not run."。没有 Broker 调用,也没有发布行。
  • 一条回复里两个 Shell 调用,先允许后拒绝: 第一个执行,第二个按自己的 call id 被拒。O2 和本地路径都成立。
  • 3 s 超时无人应答: 该调用以过期为由被拒,不执行。
  • 审批待定时杀掉 Harness: load 返回 409 hosted_turn_recovery_required。本地路径也一样,所以这是 main 的 D6a 行为。

第二台宿主机(Orange Pi 上的 0027f7a7)

Harness 跑在 linux-arm64、Node 24 上,其余都在 Mac 上:

  • receipt 提交后崩溃,Mac → Pi 与 Pi → Mac: 56–57 s 后 load 成功。命令只执行了一次,字节一致。
  • 正常交接 Mac → Pi → Mac: 3 个发布为 REFERENCED,历史中的工具结果依次为 1、2、3 个。
  • 续跑时 acquire 应答丢失,B 在 Pi 上: 503,然后 200,轮次完成。
  • 审批(允许;允许 + 拒绝): 正确。

0027f7a7 上的回归(fake OSS,全新库)

  • O2:
    • echo 同时写 stdout 和 stderr:2/2。
    • 100 MiB stdout + 5 MiB stderr、exit 7:一致。
    • receipt 提交后崩溃:通过。
    • 分段请求丢一次:完成。
    • admission 应答丢一次:完成。
    • 控制字符加崩溃:通过。
    • reload 后第二个 Shell 轮次:通过。
    • ACK 请求丢失后 reload:通过。
    • receipt 提交延迟 12 s:完成。
  • 本地路径: echo 2/2;100 MiB + 5 MiB、exit 7;reload 后第二个 Shell 轮次。
  • 混合: 2 个 O2 Session 与 2 个本地 Session 同时各写 64 MiB:4/4,0 死锁。
  • 之前的修复: R9-1、R2-6、R2-5 和 Gate 2 与第 14 轮表现相同。
  • 预览矩阵和 UTF-8 尾部夹具: 与第 13 轮相同。第 14 轮波动过的 F3 区间用例,这次截断标记又是两个。
  • 真实 qwen3.8-max,3,000 行失败构建: O2 2/2、本地 2/2。每次构建只运行一次,报错都被引用。
  • 未变:
    • Gate 1:load 在 150 s 后仍返回 409。
    • 参数里的孤立代理项仍会阻塞 O2 Session。

未重跑:1 GiB、R6-2 加载规模和作者的六 Workspace 驱动。未启动原因相关代码(R13-1、R14-1)自第 14 轮以来没有变化。

合并协调

其他在飞 PR 加了编号相同的 Flyway 迁移:

Git 不会报告这种冲突;Flyway 会在启动时失败。

仍未解决(已延后)

证据见 wenshao/qwen-code@5871c7d5(清单同英文版,results/notes.txt 列出了每个文件对应的运行)。

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head 0027f7a7f4d85fc6387fa0fa074fd649ea44cf2f

Verdict: COMMENT. Two Criticals from the maintainer's outstanding CHANGES_REQUESTED have never been adjudicated at this head and I could not verify them within one review budget; a third I did verify, and it is fixed. The diff is 71 files and +13,909/−321, so the current-head scan this channel's Approve path requires is not reachable in the time available. Under the rule that any historical blocking issue which still stands or cannot be confirmed forbids an Approve, this is a COMMENT — not a judgement that the PR is unsound, and I file no new Critical of my own.

The head has moved six times since that review landed (e13b415b → af423eb7 → 299a6fa8 → 42e2e8b1 → af962df7 → e177e1da → 0027f7a7), and no automated round has reviewed anything after 42e2e8b1. My own earlier review at e13b415b is dismissed and is not carried forward.

R2-4 — the second acquire() site missing the definite-refusal contract: FIXED

The finding was that the recovery/open path's acquire() did not carry the contract the class's own execute path implements and its tests pin, so a definite 409 workspace_busy / workspace_unavailable from the broker was indistinguishable from an ambiguous failure: the caller's blanket catch turned a live, actionable workspace conflict into 503 managed_session_open_failed.

At this head both halves of the prescribed fix are present in packages/cli/src/serve/hosted-harness-session.ts. Line 466, on the acquisition path, rethrows the refusal unwrapped instead of wrapping it:

if (isRetryableWorkspaceAcquisition(cause)) throw cause;

and the route's catch, at line 852, now classifies before falling through to the 503:

if (isRetryableWorkspaceAcquisition(cause)) {
  error(res, 409, cause.code);
} else if (cause instanceof ManagedSessionAlreadyExistsError) {
  error(res, 409, 'managed_session_already_exists');
} else if (cause instanceof ManagedSessionNotFoundError) {
  error(res, 404, 'managed_session_not_found');
} else {
  error(res, 503, 'managed_session_open_failed');
}

So a definite refusal reaches the caller as a 409 carrying its own code, and only genuinely ambiguous failures answer 503. The predicate is imported from hosted-workspace-tool-turn.ts — the module that already implemented this contract on the execute path — so the two sites now share one classifier rather than each spelling the condition. Honest limit on this confirmation: I verified the wiring at both call sites and the shared import, but did not read the predicate's body, so I am not ruling on the exact set of codes it matches.

Unconfirmed at this head: R5-1 and R4-1

Both are severity C in the round-5 ledger at e13b415b, and nothing has ruled on them since. Their full text is in that review's inline threads; what the ledger titles name is:

  • R5-1 — WorkspaceRuntimeTransport.java:145: the new grant-mode entry point lets a provably pre-dispatch ownership refusal take a path it should not. This is the newest of the three and sits on the Java transport that this PR's whole delivery mechanism runs through.
  • R4-1 — ToolPublicationDataStore.java:549: beginFinish moves the publication to producer_phase = 'FINISHING' before the step that the finding objects to, i.e. a phase transition committed ahead of the fact it is meant to follow.

Both are in files this PR adds or heavily changes, and both are the shape where a green suite says nothing: a phase written too early, or a refusal classified on the wrong branch, is exactly what tests written against the intended behaviour pass over. I had no budget to trace either call chain, so they stand as unconfirmed, which under this channel's rules is treated the same as still standing.

Gate not completed: the current-head scan

I read the acquisition and close paths of hosted-harness-session.ts and the review record. I did not scan this diff. For scale, and so the gap is concrete rather than rhetorical: 71 files, +13,909/−321, spanning the durable publication store and its SQL, the Java transport and broker service, the TypeScript runtime and session layers, the OpenAPI contract and its regenerated client, the record contracts and their fixtures, and the design docs in both languages. Nothing in that list was read here beyond the one path above.

Two things in it deserve naming because a green suite does not bound them. Any new migration in this diff should be checked for its release ordering against the migrations still pending on main, since an out-of-order migration is invisible to every test in the PR. And the OpenAPI change should be checked the way the last contract PR was: read x-qwen-implementation-status on each touched operation at this head, so a route the server does not serve cannot be certified as one — planned routes cannot mis-certify shipped behaviour, and the regenerated-client diff shows the real blast radius.

Other reviewers

@chiga0 posted review rounds at af423eb7 and 42e2e8b1, both three heads behind this one, so their findings are likewise unadjudicated here. @wenshao's CHANGES_REQUESTED at e13b415b is the latest maintainer review state on record and has not been superseded by an approval.

CI

Green at this head: Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), Serve A/B, web-shell E2E Smoke, TUI parity snapshots, OpenTUI no-flicker gate, Hosted process fault gates / MySQL 8.4 / Java 21, Runtime Broker and Managed Agent MariaDB / Java 21, Real daemon E2E / Java 11 and the full Java matrix all pass. No review-pr check is reported against this head, so no automated round is pending that would rule on R5-1 and R4-1 by itself.

Next step: the shortest path to an Approve is a ruling on R5-1 and R4-1 at the head that will actually merge — either a fix, or a re-review round that traces both call chains and records them fixed. R2-4 can be closed as fixed on the evidence above. Given this PR has been through five maintainer rounds, it is also worth applying the repository's own guidance about not letting review rounds balloon it: beyond the two Criticals, remaining Suggestions are better deferred to a follow-up issue than folded into another round here.

wenshao pushed a commit to wenshao/qwen-code that referenced this pull request Sep 30, 2026
@doudouOUC
doudouOUC added this pull request to the merge queue Sep 30, 2026
Merged via the queue into main with commit 3b18cfe Sep 30, 2026
157 of 159 checks passed
@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Follow-up to the Critical-only review at 0027f7a7: the items it left unconfirmed

The PR has since been approved at this head, by @wenshao (11:20Z) and by @qqqys (11:36Z). For the record, here is how each item that review could not confirm stands on the real stack at 0027f7a7:

Item Status Evidence
R5-1: ownership refused before dispatch Fixed Rerun at this head: ownership is lost in the publications:install window, and executeV3 settles not_started. The model gets "Workspace execution was refused before dispatch.", the receipt is committed, and a reload returns 200 with the Session not blocked. The Broker → worker ledger has no /v3/execute. The Java transport is unchanged since 42e2e8b1, and the fix had an A/B in round 12.
R4-1: an expired operation wedges the capture Fixed only for lost replies When the producer loses the reply, /recover works on the fake OSS and on real OSS: bytes exact, quota unchanged, command run once. That is also #13019's acceptance scenario. An operation that expires while in flight still answers 400 invalid_request ("claim expired"), and the producer treats that as final. The segment stays CANDIDATE or the publication stays FINISHING, and the Session is blocked. This is R15-1 in round 15.
R2-4: resume classification Fixed as wired isRetryableWorkspaceAcquisition (hosted-workspace-tool-turn.ts:63) matches only a 409 workspace_busy or workspace_unavailable. On the real stack, a lost acquire reply during a resume now gives 503 and then 200 on the retried load, and the turn completes (Mac and second host). A definite 409 during a resume could not be produced: the Broker re-acquires a live runtime session without WorkspaceExecutionStore.claim(), so the unit tests remain the evidence for that branch.
Migration ordering No clash with main V20–V22 apply on main's V19 schema. Other in-flight PRs do collide (#13037, #13084, #13087, #13088, #12946); they are listed in round 15.

Two corrections to that review:

  • Maintainer review state: by the time the review was posted, @wenshao had already approved 0027f7a7 (11:20:12Z), so the CHANGES_REQUESTED at e13b415b was no longer the latest.
  • OpenAPI: this PR changes no OpenAPI document and no generated client. Its contract files are the internal managed-tool-publication-v1 schema and fixtures, so the x-qwen-implementation-status check does not apply.

Also, before e177e1da the R2-4 failure was a 200 load followed by a blocked Session (prompts got 409 hosted_turn_recovery_required). The 503 managed_session_open_failed answer is new in that commit.

Still open: R15-1 does not block this merge while the profile stays disabled, but it should be fixed before O2 is enabled. It could be tracked in #13019 or in a new issue. Round 15 has the fault matrix, the real-OSS runs and a tested one-line client change, which covers single stalls but not claim-only lapses or read-backs longer than the timeout.

R5-1 rerun logs: s15-r15lsi.log (O2) and s15-r15lsl.log (the local path has no install window, so the command runs normally there).

中文版

关于 0027f7a7 上的 Critical 专项评审:它未能确认的各项

此后 PR 已在这个 head 上获得批准:@wenshao(11:20Z)和 @qqqys(11:36Z)。以下记录那条评审未能确认的各项,在 0027f7a7 的真实栈上的状态:

项 状态 证据
R5-1: 派发前所有权被拒 已修复 在这个 head 上重跑:在 publications:install 窗口内丢失所有权,executeV3 结算为 not_started。模型收到 "Workspace execution was refused before dispatch.",receipt 已提交,reload 返回 200,Session 未被阻塞。Broker → worker 的账本里没有 /v3/execute。Java transport 自 42e2e8b1 起未变,修复在第 12 轮做过 A/B。
R4-1: 过期的操作卡死 capture 只修好了应答丢失的情况 生产端丢失应答时,/recover 在 fake OSS 和真实 OSS 上都能用:字节一致,配额不变,命令只执行一次。这也是 #13019 的验收场景。操作在执行过程中过期时,仍返回 400 invalid_request("claim expired"),生产端把它当成最终结果。分段停在 CANDIDATE 或发布停在 FINISHING,Session 被阻塞。这就是第 15 轮的 R15-1。
R2-4: 续跑时的分类 接线已修复 isRetryableWorkspaceAcquisition(hosted-workspace-tool-turn.ts:63)只匹配 409 的 workspace_busy 或 workspace_unavailable。在真实栈上,续跑时 acquire 应答丢失,现在会先得到 503,重试 load 后得到 200,轮次完成(Mac 和第二台宿主机都成立)。续跑时明确的 409 无法产生:Broker 重新获取仍然存活的 runtime session 时不会调用 WorkspaceExecutionStore.claim(),所以那一支仍以单元测试为证据。
迁移顺序 与 main 不冲突 V20–V22 能在 main 的 V19 库上应用。其他在飞 PR 之间确有冲突(#13037、#13084、#13087、#13088、#12946),已列在第 15 轮报告中。

对那条评审的两处更正:

  • 维护者评审状态: 那条评审发出时,@wenshao 已经批准了 0027f7a7(11:20:12Z),所以 e13b415b 上的 CHANGES_REQUESTED 已不是最新状态。
  • OpenAPI: 这个 PR 没有改动任何 OpenAPI 文档,也没有生成客户端。它的契约文件是内部的 managed-tool-publication-v1 schema 和 fixtures,所以 x-qwen-implementation-status 的检查不适用。

另外,e177e1da 之前 R2-4 的表现是 load 返回 200,随后 Session 被阻塞(prompt 得到 409 hosted_turn_recovery_required)。503 managed_session_open_failed 是那个提交新加的。

仍未解决: profile 保持关闭时,R15-1 不阻塞本次合并,但在启用 O2 之前应当修复。可以放在 #13019 或新开 issue 跟踪。第 15 轮有故障矩阵、真实 OSS 的运行结果,以及一个实测过的一行客户端改动;这个改动能处理单次停顿,但处理不了只有 claim 过期、以及读回时间长于超时的情况。

R5-1 重跑日志:s15-r15lsi.log(O2)和 s15-r15lsl.log(本地路径没有 install 窗口,所以命令正常执行)。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants