Skip to content

feat(managed-agent): H3 background Shell and Monitor runtime - #13265

Merged
wenshao merged 131 commits into
mainfrom
docs/h3-shell-monitor-design
Oct 6, 2026
Merged

wenshao merged 131 commits into
mainfrom
docs/h3-shell-monitor-design

Conversation

@wenshao

@wenshao wenshao commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Implements slice H3 — background Shell and Monitor on the Managed path. The bilingual design documents (English · 中文) landed first in this PR; the code follows them incrementally and is updated here as each batch lands.

The design: H3 is the first Stage H slice that produces user-visible tasks, so it covers both the two Legacy capabilities and the task events route that the earlier contracts left planned. Background shells become durable session records on the extension-record pipeline, each process runs in a dedicated Linux cgroup v2 unit as its restart-stable identity, logs stream into one growing output Artifact per task with pipe-level backpressure, the runtime hold rides the execution ledger, and session close terminates and drains in order. Monitors run on the already-frozen monitor_run record with debounce-bounded observations and watermark-deduplicated notifications, and only purely observational targets may be rebuilt after Runtime loss. Recovery either re-attaches the original process by its cgroup unit or blocks accurately; nothing auto-reruns, and cross-boot attach is explicitly not claimed.

Landed so far, each batch with its own review-audit and both-language test suites green:

  • The managed-child_run record body, schema version 1 (kind: "shell"), in TypeScript and Java: a closed eleven-key body for one background Shell — fixed identity and start-call pin, a start receipt that appears exactly once and is reused by every re-attach (a changed receipt is refused as a rerun shape), stopRequested set-once-never-clears driving the shared draining runtime projection, one growing output reference, closed stop reasons, exit code/signal proven exactly when the Shell exits, and terminal-state execution-line linkage. The shared fixture file pins the body shapes and chains, replayed in both languages; the projection registry gains the domain with task kind background_shell on both sides.

  • Authority end-to-end and the hosted orchestration funnel: one background Shell committing its full revision life through the Session authority (cold-reopen rebuild, skip-step refusals, replayed commands answered with the original receipt), and a HostedChildRunSession that funnels admit → dispatch-started → attach → advance → settle through the authority with command idempotence and deep-equal skip.

  • The isolated-capture base: hook-command-cgroup create/attach/terminate on Linux cgroup v2 with accurate refusal where no delegated root exists (Linux probes refuse before any side effect), and managed-child-run-supervisor with TERM → drain → kill semantics; plus the locked capture (LocalShellStreamCapture) that publishes bounded segments/pages as output arrives, revises its manifest per page and seals at exit, with the ended-stream revision published only after the seal decision.

  • The detached result family: the managed-tool-result/1 envelope gains captureStatus: "detached" with a blocked decision, threaded through the schema, TS parse, the Java projector, worker settle, ack and both replay validators — every journal tool.receipt lands in the result store, so a background start settles row 1 with a handle and must not corrupt replay.

  • Worker background admission: the hosted tool turn's two background refusals flip to a three-state admission (requested / ill-formed / admitted), start journals entry-first, an eight-per-session quota, env merging through the allowlist, and precise not-started refusals.

  • The maintenance route /internal/managed-runtime/v3/shells: shell-status and shell-terminate answered only from physical registry knowledge (never guessed), with the view's operationId pinned to the target process identity on both writer and wire validation.

  • The broker's two-row ledger: the start invocation settles row 1 with the handle; a second :process row admits at start in PREPARED so hasActiveBy* and session release naturally count it, claim/fence never touch it, and it settles only on physical evidence via the new settlePrepared repository verb; pollV3Result and ack gain the detached branches, and the v3 guard opens for Shell hosts.

  • The hosted publisher's background exit leg: a single publisher serves both capture families per session (a second session-level publisher would conflict at the broker registry); admission checks the child-run record instead of a model checkpoint; open() publishes revision 1 at prepare, every awaited write advances outputRef to the newest manifest, finalize feeds the same evidence object to the final manifest and settleExited, and accept returns a detached-blocked receipt with no second journal receipt. Manifests publish through the authority's journal-backed resource store so the domain closure can read them.

  • The monitor_run resource closure (enablement precondition): both stores now read the four monitor refs at commit time (commandRef, startReceiptRef, outputRef, lastObservationRef), both fixture rigs were rebuilt from fictional identities to real published placeholders memoized on the cited identity, and each side carries a dedicated witness that a run citing a resource outside its commit is refused.

  • Reader compatibility (revised after the round-5 cross-version matrix): every new Session keeps minimumReader: managed-session/1 — monitor_run has parsed in every deployed reader since feat(managed-agent): Define the managed-extension-record/1 contract (H0b) #12837 (v0.24.7), so a v2 stamp would only make rollback and mixed-version rollouts lose access to Sessions without protecting anything. The stamp policy is recorded in the design: it rises only when a change genuinely breaks an older reader mid-scan, and only one release after a tolerant reader ships.

  • The Monitor physical side: the cgroup watch executor (supervisor unit as execution identity, Legacy partial-line cap, final-line flush, exit never missed even under a fast watch), the monitor registry, the is_monitor branch of the v3 executor (four-watch quota, background+monitoring markers, detached settle), and the maintenance route family /internal/managed-runtime/v3/monitors (monitor-status/monitor-stop) on the same unknown-never-exited discipline.

  • The turn's Monitor admission arm: a three-state gate mirroring the Shell arm, the monitor tool declaration, record-first admit → dispatch → attach with H6 journal-before-attach ordering, replay-safe accepts, and Session-scoped observation loops.

  • Observation fan-out: monitor captures carry recordDomain 'monitor_run', per-capture observers, and the Legacy line semantics into the hosted loop; monitor finalize feeds the same evidence object to the record settle.

  • The notification wake: the Legacy <task-notification> envelope builder (same tags, same truncate/strip/escape pipeline), the two-pass pending-input intake read straight from the journal (consumed exactly when a turn settles under its turnId), the envelope on the loop's notification input, and the embedded wake scheduler — a monitor notification runs as an ordinary text turn while the Session idles, queues in the journal while a turn runs, parks accurately while anything is blocked, and settles cancelled model-free on the close path so no wedged notification holds a Session unopenable. Parked-turn scans in reopen/takeover skip monitor inputs — the pump owns them exclusively.

  • The task-events lane: listSessionTaskEvents and queryWebShellTaskEvents flip planned→partial on both surfaces behind a bounded per-task SQL journal (V40): one row per committed task event written in the Session commit transaction with per-task sequences allocated in commit order, durable retention floor under the Artifact visibility barrier, output cursors with overlap/gap refusals, backlog bound refusing instead of silently discarding; §6.1 demonstrations run as store-level contract traffic; feat(managed-agent): Track the task contract gaps deferred from the H0a review #12847's C15/C16 instances close alongside.

  • Round-5 physical folds: monitor output is byte-true (raw-chunk capture channel, StringDecoder on the observation leg, exit never missed); both funnels' advanceOutput refuses back-to-older revisions before anything commits; the broker drains a busy background Shell at release through the worker's drain-before-refuse route (busy for non-background work still refuses with zero worker hops); the output pause re-asserts behind every flushed chunk so a writer outliving the launcher cannot balloon the queue; and the shell maintenance validator refuses a view answering another operation.

  • The full-PR audit round (this PR's own gate, and the round this description entry records): every slice handed to a read-only auditor plus a cross-cutting pass over every seam a slice cannot see alone (TS↔Java wire contracts, env-guard claims, migration uniqueness, discriminative-marker survival across merges, drive-by legitimacy), with every finding verified against the exact reviewed commit before being fixed or rejected with evidence. Fixed and mutation-verified: a bounded journal read that left monitor wakes undelivered and unsettled at close once a Session passed the default page; a full task journal propagating its refusal into the extension-record commit (fatally wedging every later revision of that task; output producers were the latent amplifier); a replayed output advance refusing through the very no-op path a redelivered delivery holds; the wake pump misreading an accurately-blocked settle as a consumption bug; and the maintenance protocol's never-set error field removed. Main's fix(managed-agent): stop virtual-thread carrier pinning in the Hosted Harness stream and the Runtime Broker guards #13388 carrier discipline is carried onto every H8 site since. Everything rejected carried its own evidence: the beginControl/endControl imbalance claim (safeStage captures sync throws, whenComplete is downstream), the loop terminal-state mapping (verbs belong to the H4–H5 and rebuild-driver slices), the activation-gate proposal (it would rebuild the wedge round 5 asked us to remove), the recordDomain charge, the "journal row first" ordering claim (record rows always precede journal rows), the artifact-barrier and config-switch claims, and every item found pre-existing on main.

Verification culture so far: two review rounds closed end-to-end (34 threads, then 43, every one replied and resolved), three real-stack Linux probe rounds on containers and bare metal (rounds 4 and 5 in the comments below — round 5's findings are folded into the landed list above), and every commit batch passes a multi-round self-audit before landing.

Still to land in this PR, in order: the H1–H5 physical findings from the real-stack rounds (H1–H4 and H7b stay in the round-3 candidate patch, which is the maintainer's to land), live outputRef advancement with the output producer slice, the rebuild driver, monitor_run and child_run submission enablement once the candidate patch and the exact-head Linux evidence exist, and the Linux physical acceptance of the full chain at exact head.

Why it's needed

The tracker's prerequisites for H3 (workspace restart/reclamation, durable local-process provisioning, durable remote output, the output lifecycle) have all merged — the last was the durable/reboot-recovery defaults flip — so the slice can start. Its acceptance gate in the reference design (real process owner, log Artifacts, runtime hold, stop/drain, restart attach or accurate blocking) and the obligations the H0a/H0b/H0c contracts deferred to H3 (output segmentation, retention-floor mechanics, the artifact_refs cap, reader compatibility, resource closure, notification wake delivery) need a reviewed design, then the code that meets the gate.

Reviewer Test Plan

How to verify

  • Confirm the English and Chinese design versions stay complete and synchronized, including the decisions this PR has since fixed: the two-row ledger, and the reader-compatibility policy that replaced the earlier reader-gating decision.
  • Record side: cd packages/core && npx vitest run src/managed-runtime/ covers the fixture contracts, the authority chains, reader compatibility and the resource closure; the shared Stage H golden is replayed by Java ManagedSessionStoreIntegrationTest. On the Java side run the managed-agent-server and runtime-broker unit suites (ManagedChildRunRecordContractTest, ManagedExtensionRecordStoreTest, RuntimeBrokerServiceTest).
  • Workers and hosted orchestration: cd packages/cli && npx vitest run src/serve/ for the shell route, background registry, publisher, tool turn and worker suites.
  • Confirm the gates that stay closed: child_run and monitor_run are still absent from the enabled-domain list, so no Session can commit them through production configuration; enabling arrives behind the items listed above, with submission enablement as the only remaining domain gate.
  • For the physical findings: the round-3/round-4 real-stack comments carry probe scripts, kernels and the candidate patch; H1–H4 reproduce on bare metal exactly as reported there.

Evidence (Before & After)

N/A (no user-visible surface yet; every path stays disabled until the physical findings above settle)

Tested on

OS Status
🍏 macOS ✅
🐧 Linux ⚠️

Environment (optional)

macOS unit runs (vitest, and Maven with JDK 21 for the Java suites including the broker's MySQL ITs). Linux real-stack probes ran on bare metal (kernel 6.6 cgroup v2; a 5.10 host to pin accurate refusal) in the round-4 comment's setup, with H1–H4 reproduced and a candidate patch recorded; the final Linux acceptance runs at exact merge head per the design.

Risk & Scope

  • Main risks being managed before enablement: the H7 broker↔worker exited-agreement gap (a natural exit currently answers unknown, and observeBackgroundProcess has no production caller yet), the G5 capture backpressure, and the H1–H5 physical findings — all itemized, and both domains stay disabled until they settle.
  • Reader compatibility: new Sessions stay at minimumReader: managed-session/1; the stamp policy (revised after the round-5 cross-version matrix showed the earlier v2 gate would only harm rollback and mixed rollouts) is recorded in the design docs.
  • Breaking changes / migration notes: none for consumers of the current release lines; the reader bump only affects binaries old enough to predate H3 reading brand-new Sessions. Public task cancel, send_input, detach, cross-boot attach, and the macOS process-group profile stay out of scope, each assigned to a named follow-up.

Linked Issues

Part of #12380 and #12827.

中文说明

本 PR 内容

实现 H3 切片(Managed 路径上的后台 Shell 与 Monitor)。中英双语设计文档(English · 中文)在本 PR 中先行落地,代码在其后的提交中增量跟进,本说明随每一批落地同步更新。

设计部分:H3 是第一个产生用户可见任务的 H 阶段切片,因此同时覆盖两项 Legacy 能力与此前契约留在 planned 的任务事件路由。后台 Shell 成为扩展记录管线上的持久 Session 记录;每个进程运行在专用的 Linux cgroup v2 unit 下,作为其跨重启的稳定身份;日志以管道级背压流入每任务一个、持续增长的输出 Artifact;Runtime hold 挂在 execution 账本上;Session 关闭按序 terminate 并 drain。Monitor 在已冻结的 monitor_run 记录上运行,观测按去抖限界、通知按水位去重;Runtime 丢失后仅纯观测目标允许重建。恢复要么按 cgroup unit 重新 attach 原进程,要么准确阻塞;绝不自动重跑,且明确不主张跨 boot attach。

已落地的批次,每批都带各自的评审自审计、双语言测试套件绿:

  • managed-child_run 记录体,schema 版本 1(kind: "shell"),TypeScript 与 Java 双实现:一个后台 Shell 对应一条封闭十一键记录——身份与启动调用 pin 固定;start receipt 只出现一次且每次 re-attach 复用(换了 receipt 按重跑形态拒绝);stopRequested 置位后不再清除,驱动共享的 draining 运行时投影;输出引用单一且随 manifest 增长;stop reason 封闭;退出码/信号恰在退出时被证明;终态状态与 execution 线联动。共享 fixture 钉住记录形态与链条,双语言回放;两侧投影注册表获得该域,task kind 为 background_shell。

  • authority 端到端与托管编排漏斗:一个后台 Shell 经 Session authority 提交其完整 revision 生命周期(冷重开重建一致、跳步拒绝零提交、重复命令返回原回执),以及把 admit → dispatch-started → attach → advance → settle 全部汇入 authority 的 HostedChildRunSession(命令幂等、内容一致即跳过)。

  • 隔离捕获基座:hook-command-cgroup 在 Linux cgroup v2 上的 create/attach/terminate,无委派根时在准入处准确拒绝(Linux 探针验证拒绝先于任何副作用);managed-child-run-supervisor 的 TERM → 排空 → kill 语义;以及锁定捕获(LocalShellStreamCapture)——有界分段/分页来了就发布、每页修订 manifest、退出时封口,ended-stream 修订只在封口决定之后发布。

  • detached 结果家族:managed-tool-result/1 envelope 增加 captureStatus: "detached" 与 blocked decision,贯通 schema、TS 解析、Java projector、worker 结算、ack 与两个回放校验器——每笔 journal tool.receipt 都会落进结果存储,后台启动的 row 1 带 handle 结算时绝不能腐化回放。

  • worker 后台准入:托管 tool turn 的两处后台拒绝翻转为三态准入(requested / ill-formed / admitted),启动先记 journal 再执行,每 session 8 个配额,env 经 allowlist 合并,未启动拒绝精确归类。

  • 维护路由 /internal/managed-runtime/v3/shells:shell-status 与 shell-terminate 只凭物理注册表知识回答(绝不猜终止),视图的 operationId 在 writer 与线上校验两侧都钉在目标进程身份上。

  • Broker 两行账本:启动调用带 handle 结算 row 1;第二条 :process 行在启动时以 PREPARED 准入,使 hasActiveBy* 与 session 释放自然计入,claim/fence 不触碰,只有物理证据经新仓库动词 settlePrepared 才结算;pollV3Result 与 ack 获得 detached 分支,v3 守卫对 Shell 宿主打开。

  • 托管 publisher 的后台退出腿:每 session 单 publisher 服务两个 capture 家族;准入校验 child_run 记录而非模型检查点;prepare 即发出 open() 首版,每次被 await 的写推进 outputRef 到最新 manifest;finalize 把同一 evidence 对象喂给最终 manifest 与 settleExited;accept 返回 detached-blocked receipt,journal 里不再有第二笔 receipt。manifest 经 authority 的 journal-backed 资源存储发布,使域闭包能读到。

  • monitor_run 资源闭包(启用前置):两侧 store 在提交时读取 monitor 的四条引用(commandRef、startReceiptRef、outputRef、lastObservationRef);两套 fixture rig 从虚构身份重建为真实发布的占位资源(按被引身份记忆化);各有一条专门见证钉住——点名了提交之外资源的 run 被拒绝。

  • 读者兼容(依据第 5 轮跨版本矩阵修订):每个新 Session 保持 minimumReader: managed-session/1——monitor_run 自 feat(managed-agent): Define the managed-extension-record/1 contract (H0b) #12837(v0.24.7 起可用)在每个已部署 reader 上都可解析,因此 v2 盖章只会让回滚与混版本部署丢掉 Session,而保护不了任何东西。盖章策略已写入设计:仅当某次改动真的让旧 reader 在扫描中途失败时才抬升,且只有在能读取它的 reader 先发布一个版本之后才写入。

  • Monitor 物理侧:cgroup watch 执行器(supervisor unit 即执行身份、Legacy partial-line 上限、末行冲刷、快速 watch 下也绝不漏掉结束)、monitor 注册表、v3 执行器的 is_monitor 分支(每 Session 四表配额、background+monitoring 标记、detached 结算),以及同一「unknown 绝不等于已退出」纪律下的维护路由族 /internal/managed-runtime/v3/monitors(monitor-status/monitor-stop)。

  • turn 的 Monitor 准入臂:与 Shell 臂对称的三态闸门、monitor 工具声明、按 H6 journal-before-attach 次序的先记录 admit → dispatch → attach、可重放的 accept,以及 Session 级观测环。

  • 观测扇出:monitor capture 携带 recordDomain 'monitor_run'、按 capture 的 observer,以及 Legacy 行语义汇入 hosted 观测环;monitor finalize 把同一个 evidence 对象喂给记录结算。

  • 通知 wake:Legacy <task-notification> 信封构造器(同标签、同 truncate/strip/escape 管线),只从 journal 直读的两段式待决输入入口(恰在同 turnId 的回合结算时消费),观测环通知输入上的信封,以及内嵌 wake 调度器——Monitor 通知在 Session 空闲时作为普通文本回合运行、回合忙碌期间在 journal 排队、任何阻塞时准确保持待决,并在关闭路径上不经模型按取消结算,使任何楔住的通知都不会把 Session 卡在不可重开。重开/接管路径的 parked-turn 扫描跳过 monitor 输入——由泵独占其消费。

  • 任务事件车道:listSessionTaskEvents 与 queryWebShellTaskEvents 在两个表面翻转 planned→partial,后端是有界的每任务 SQL 流水账(V40):每条已提交任务事件一行、写在 Session 提交事务内、按提交顺序分配序号;Artifact 可见性屏障下的持久保留 floor、带重叠/缺口拒绝的输出游标、越界即拒绝的积压上界;§6.1 演示以存储级契约流量运行;feat(managed-agent): Track the task contract gaps deferred from the H0a review #12847 的 C15/C16 实例一并关闭。

  • 第 5 轮物理折叠:monitor 输出字节级保真(raw-chunk 捕获通道 + 观测腿的 StringDecoder + 结束绝不漏接);两个漏斗的 advanceOutput 在提交前拒绝退回旧 revision;Broker 经 worker 的「先排空后拒绝」路由在释放时排空忙碌的后台 Shell(非后台忙碌仍零跳拒绝);输出暂停在每个冲入 chunk 后重断言,使寿命超过 launcher 的写者无法吹爆队列;shell 维护校验器拒绝回答别处的视图。
    验证文化:两轮评审逐条闭环(先是 34 线程,再是 43 线程,全部回复并 resolve),两轮 Linux 真实栈探针(容器与裸机,第 4 轮见下方评论),每个提交批次落地前过多轮自审计。

后续增量按序:真实栈轮次遗留的物理发现(H1–H4 与 H7b 留在维护者自己的第 3 轮候选补丁)、与输出生产者一起落地的运行中 outputRef 推进、rebuild driver、候选补丁与精确头 Linux 证据齐备后的 monitor_run/child_run 提交启用、rebase 时推进 main #13304 的回合级关闭断言,以及完整链路在精确头点上的 Linux 物理验收。

为什么需要

跟踪 issue 为 H3 列出的前提(workspace 重启与回收、durable 本地进程 provisioning、持久远端输出、输出生命周期)已全部合入——最后一项是 durable/重启恢复默认值翻转——因此该切片可以开工。参考设计中的验收门槛(真实进程 owner、日志 Artifact、Runtime hold、stop/drain、重启 attach 或准确阻塞)以及 H0a/H0b/H0c 契约挂给 H3 的义务(输出分段、保留 floor 机制、artifact_refs 上限、读者兼容、资源闭包、通知 wake 投递)需要先有一份经过评审的设计,再有满足门槛的代码。

评审验证要点

如何验证

  • 确认中英文设计版本保持完整同步,包括本 PR 后来定案的内容:两行账本,以及替代早前门禁决定的读者兼容策略。
  • 记录侧:cd packages/core && npx vitest run src/managed-runtime/ 覆盖 fixture 契约、authority 链、读者兼容与资源闭包;共享 Stage H golden 由 Java ManagedSessionStoreIntegrationTest 回放。Java 侧跑 managed-agent-server 与 runtime-broker 单测(ManagedChildRunRecordContractTest、ManagedExtensionRecordStoreTest、RuntimeBrokerServiceTest)。
  • worker 与托管编排:cd packages/cli && npx vitest run src/serve/ 覆盖 shell 路由、后台注册表、publisher、tool turn 与各 worker 套件。
  • 确认仍然关闭的关口:child_run 与 monitor_run 都不在启用域列表中,生产配置下任何 Session 都无法提交它们;启用在上述事项解决之后到来,唯一的域闸门就是启用清单。
  • 物理发现:第 3/4 轮真实栈评论带有探针脚本、内核与候选补丁;H1–H4 在裸机上与报告完全一致地复现。

现场证据(前后对比)

N/A(尚无用户可见入口;每条路径在上述物理问题解决前保持关闭)

已测平台

平台 状态
🍏 macOS ✅
🐧 Linux ⚠️

环境(可选)

macOS 单测(vitest;Java 套件用 JDK 21 的 Maven,含 broker 的 MySQL IT)。Linux 真实栈探针按第 4 轮评论的环境在裸机上运行(6.6 内核 cgroup v2;5.10 主机钉准确拒绝),H1–H4 已复现并记录了候选补丁;最终 Linux 验收按设计在精确合并头点上执行。

风险与范围

  • 启用前正在管理的主要风险:H7 Broker↔worker 退出一致缺口(自然退出目前回答 unknown,observeBackgroundProcess 尚无生产调用方)、G5 捕获背压、H1–H5 物理发现——全部逐条列出,两个域在这些问题解决前保持关闭。
  • 读者兼容:新 Session 保持 minimumReader: managed-session/1;盖章策略(第 5 轮跨版本矩阵证明早前 v2 门禁只会伤害回滚与混跑后已修订)已记录在设计文档。
  • 破坏性变更/迁移说明:对当前发布线的消费者无影响;reader 抬升只影响早于 H3 的二进制读取全新 Session。公开任务的 cancel、send_input、detach、跨 boot attach、macOS 进程组 profile 均不在范围内,各自有指名的后续项。

@wenshao wenshao changed the title docs(managed-agent): design H3 background Shell and Monitor runtime feat(managed-agent): H3 background Shell and Monitor runtime Oct 3, 2026
@wenshao

wenshao commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

CI failure attribution — Hosted process fault gates / MySQL 8.4 / Java 21 is the intermittent tracked in #13255, not this PR.

The single failed check on this PR is that lane (job 111180098809). Its only failure is HostedWorkspaceToolTurnIT.packagedHarnessUsesSavedWorkspacesThroughRealBrokerWorkerAndSqlStore: two turns failed with Hosted Workspace profile refused a tool call, leaving the Session unsettled; the driver's final cold load then answered 409 hosted_turn_recovery_required and the wrapper failed expected: 0 but was: 1. The daemon log also carries the #13255 signature — POST /files/rewind ... durationMs=1042 status=409 (the issue records ≈1016).

Evidence chain, each independently checkable:

  1. The tested path is not in this diff. git diff --name-only 1a4de7486a..817383760e lists 11 files: two design docs, seven record-contract/authority/test files under packages/core/src/managed-runtime/, and two store validators plus one contract test in packages/sdk-java/managed-agent-server. It does not touch hosted-workspace-tool-turn.ts, integration-tests/helpers/hosted-workspace-tool-turn-driver.ts, HostedWorkspaceToolTurnIT.java, the tool declarations, or the session store. The refusal guards (unknown tool, duplicate callId, truncated or incomplete arguments) read only the fake-model stream, and this IT's Session never commits a Stage H record, so the disabled child_run body this PR registers is inert on the path that failed.
  2. Sibling lane: not applicable — grep -c HostedWorkspaceToolTurnIT is 0 in the Runtime Broker and Managed Agent MariaDB / Java 21 log; this class runs only in the Hosted lane.
  3. main is green: the SDK Java workflow is 9/9 successful on main today (latest completed run 37098468609 at 5130c1a734), consistent with an intermittent window rather than a lane outage.
  4. Cross-branch, same signature, same hour: fix/managed-agent-quality-hardening (run 37115978280) and fix/13183-runtime-broker-hardening (run 37118858802) failed the same lane, same method, with the identical full chain — two refused a tool call turns + hosted_turn_recovery_required + expected: 0 but was: 1. Earlier instances are already filed in flaky(ci): HostedWorkspaceToolTurnIT.packagedHarnessUsesSavedWorkspacesThroughRealBrokerWorkerAndSqlStore intermittently fails with 409 on POST /files/rewind #13255: run 37078701187 on fix(cli): bound managed function-hook module evaluation and keep hold-fenced owners recoverable #13243 and run 37086759282 on feat/managed-agent-broker-auth with the same rewind-409, plus a different method of the same class failing on the nightly (run 37069499549).
  5. Local repro: not attempted here — this IT needs MySQL 8.4 and the packaged harness, and flaky(ci): HostedWorkspaceToolTurnIT.packagedHarnessUsesSavedWorkspacesThroughRealBrokerWorkerAndSqlStore intermittently fails with 409 on POST /files/rewind #13255 likewise records no local reproduction as of filing.
  6. Reachability: nothing in this PR can alter the model stream, the tool declarations, or the rewind gate on this path.

I am rerunning the failed job now for a second sample, with the call declared upfront: green ⇒ transient confirmed; red at the same point ⇒ a deterministic component in this window, and attribution still belongs to the lane/IT family tracked in #13255, not to this PR. The rerun outcome goes under this comment.

@wenshao

wenshao commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification — #13265 at 817383760e

Scope. This covers what has landed so far: the H3 design, plus the managed-child_run v1 record body in TS and Java, registered but disabled. The PR body says more increments are coming in this PR, so treat this as a checkpoint report.

Verdict. The landed slice is inert in production and consistent across TS and Java:

  • The shipped build refuses child_run and writes nothing.
  • Shipped artifacts change only by the new body and its registry entry.
  • With the domain enabled in-process, three background-Shell chains commit through the real HTTP store into Spring + MariaDB with 0 view mismatches.

Five items need attention before enablement. Two of them are defects:

  • F1: child_run refs sit outside both reference-closure checks.
  • F2: a Java NPE answers HTTP 500 instead of 409.

Two are design decisions for the maintainers:

  • F3: the re-attach receipt rule.
  • F4: the body fields the design lists but the body lacks.

One is test strength:

  • F5: 15–16 conditions are not pinned by the shared fixtures.

None of them is reachable while child_run stays out of MANAGED_SESSION_ENABLED_DOMAINS.

What was run

Check Result
PR's focused TS suites (child-run, authority.child-run, extension record/projection, hook/MCP record, authority.extension) 955/955
All of packages/core/src/managed-runtime 1993/1995: 2 timeouts in hook-scale, a file this PR does not touch. Run alone, that file fails 1 of 6 on both head and base at the same host load.
Java managed-agent-server clean verify checkstyle:check with MariaDB ITs Unit tests 553 passed (1 skipped), Checkstyle 0. ITs 51/52: ManagedAgentMySqlIT.admitsHookExecutionsWithoutReadingTheirHistoryOnMySql fails alone on both head and base on this host; it passes in CI's MariaDB lane.
S1 production gate (head and base builds, no override) Head: domain child_run is registered but not enabled for submission. Base: has no Stage H record body. Journal unchanged on both.
Artifact diff, base → head core/dist gains managed-child-run-record.js plus one registry entry. MANAGED_SESSION_ENABLED_DOMAINS is identical. Server jar: 570/576 entries have identical CRCs; the 6 that changed are all ManagedExtensionProjection* / ManagedExtensionRecords*.
S2: three chains (life to exit, stop request, Runtime loss + re-attach), real HTTP store → Spring + MariaDB 15 commits, 0 mismatches between the authority view, the Java row and GET /v1/agents/sessions/{id}/tasks (kind background_shell). 15 task.updated events. Refusals leave the journal unchanged. A cold reopen by a second writer gives identical views.
TS/Java differential, 200,000 candidates through the built TS dist and the compiled Java classes 0 disagreements except F2 (1,706) and the known NFC gap from #12837 (290)
Mutation testing of both validators against the PR suites TS 18/41 killed, 7 equivalent; Java 17/38 killed, 6 equivalent (F5)
Trial merge: PR + main 576689d073 + in-flight #13217 7f4d6aa299 (touches ManagedExtensionRecords.java) Textually clean. Java 595 unit tests, Checkstyle 0, TS 955 pass. S2 on the merged jar: 15/15, 0 mismatches. managed-runtime/ is byte-identical to head.

real stack

F1 — child_run references sit outside both closure checks (fix before enablement)

The design (§Resource closure) says the store commits a body "only with every resource the body names", that the writer "checks each reference before publishing the body", and that "No resource kind is exempted". The code covers neither side for child_run:

  • Writer side. verifyExtensionResources (managed-session-authority.ts:1890) reads refs only for mcp_configuration, mcp_operation, hook_registration and hook_execution.
    • A child_run with a never-published or another Session's commandRef is published anyway. The store answers 409 "resource missing", and the Session's writer stops: the next valid commit is refused locally (S3 cases A and B).
    • The same mistake in mcp_operation is refused before publishing, and the writer keeps working (case C).
  • Server side. The nested-ref check in applyRevision (ManagedExtensionRecordStore.java:344) lists the same four domains.
    • When a writer leaves the body's refs out of the commit's resource list, a 4-revision child_run chain naming three nonexistent resources commits (200 ×4) and projects as background_shell running. Reading any of the three refs then fails with "does not exist", and a cold reopen accepts the chain (case D).
    • The same omission for mcp_operation gets a 409 (case E).
  • Test. The authority suite passes only because neither check exists: RECEIPT_1, MANIFEST_1/2 and args-shell-1 (managed-session-authority.child-run.test.ts:132-166) are never published.
  • monitor_run. It has the same gap, carried over from H0c, and H3 enables it too.

closure

F2 — Java NPE on a non-string exitSignal (one-line fix)

  • Cause. ManagedExtensionRecords.java:688 calls EXIT_SIGNAL.matcher(exitSignal.textValue()). textValue() is null for 9, true, {} or [], and accepts() (line 785) catches only InvalidRecordException, so isChildRunStart and isChildRunSuccessor throw as well.
  • Effect on the real server.
    • exitSignal: 9 or true → POST /transactions:commit 500, retried 3× by the client, with an ERROR stack trace in the server log.
    • "term" → a clean 409.
  • Reach. TS refuses these bodies before sending, so only a writer that bypasses the TS validator can trigger it.
  • Fix. exitSignal.isTextual() && …, plus a fixture with a numeric exitSignal. The Java contract test asserts InvalidRecordException, so that fixture would have caught it.

npe

F3 — Re-attach must mint a new start receipt (decision needed)

After recovery_blocked/runtime_lost/outcome_unknown at generation 1, there are two ways to record the re-attach:

  • running_attached under generation 2 with the same startReceiptRef is refused, by TS and by the real Java store (409 "cannot follow its revision 4").
  • A new receipt is accepted.

This is the Monitor rebuild rule carried over to the Shell (managed-child-run-record.ts:324-330). For a Monitor, a rebuild starts a fresh watch, so a new receipt fits. The design (§Recovery) says otherwise for a Shell: it is never rebuilt, only re-attached, "the recorded command digest and start receipt match". §Records also says a receipt "changes only with" a new binding; it does not say "must change".

As written, the contract accepts the shape of a re-run and refuses the shape of a true re-attach. The refusal is also unpinned: there is no reattach-same-receipt fixture, so T39/J39 survive. Either:

  • allow the same receipt on re-attach (and consider refusing a new one), or
  • keep the rule, document the new receipt as an attach receipt, and rename the rebuild-new-receipt fixture.

reattach

The same card shows S5, which confirms the design's server-first rule. An H3 writer against the pre-PR jar commits 3 revisions with no task row. After swapping in this PR's jar on the same database, the next valid revision gets a 409 "first revision … must open its run", the writer stops, and an unrelated new Shell is refused afterwards. This matches §Reader gating and H0c open question 7 exactly.

F4 — Design §Records lists fields the closed body does not have (decision needed)

The design's Shell record carries:

  • workspace generation, cwdRef and the tool profile's definition revision in the pin;
  • processId;
  • a stop-request marker that projects draining.

The closed ten-key body has none of these:

  • run.definition must be null, so the public task never shows definition_revision.
  • runtimeState() (managed-extension-projection.ts:226) cannot produce draining, yet acceptance items 3 and 8 rely on stop and drain.

Both language versions say the same. Either trim §Records to the body as landed, or add at least the stop-request field to the v1 body now. It is cheap while the domain is disabled; after enablement it means a new body version.

F5 — Shared fixtures leave 15–16 conditions unpinned

Each mutant was replayed against the 200k candidates:

  • TS: 16 surviving mutants have an input the head refuses and the mutant accepts.
  • Java: the same conditions give 15; T17 (fractional exitCode) has no Java twin.
  • The rest are equivalent.

Several have a fixture named for them that also breaks a second rule, so another rule refuses it first. For example, stop-mismatch and process-failed-no-receipt put handler_unavailable on a failed run, and quota-mismatch keeps an intent execution. TS and Java agree on all of these today, but the fixture file is the cross-language contract. The smallest isolating input for each is on the card and in results/mutation/minimal-examples.txt.

mutation

Merge reference

Evidence (images, harness, raw logs): wenshao/qwen-code@6ee5112c pr13265/. Environment: macOS arm64, JDK 21.0.12, Node 24.18.1, MariaDB 10.11.18 in docker. child_run was enabled only inside the rig process, the way the PR's own authority test does.

中文说明

真实栈验证 — #13265 @ 817383760e

范围。 本报告覆盖目前已落地的部分:H3 设计,以及 TS/Java 双实现的 managed-child_run v1 记录体(已注册、未启用)。PR 正文说后续增量还会进入本 PR,因此这是一份阶段性报告。

结论。 已落地部分在生产上是惰性的,TS 与 Java 一致:

  • 发布构建拒绝 child_run,零写入。
  • 产物差异只有新记录体及其注册项。
  • 在进程内启用该域后,三条后台 Shell 链经真实 HTTP store 写入 Spring + MariaDB,视图零不一致。

启用前有五项需要处理。其中两项是缺陷:

  • F1: child_run 的引用不在两侧的引用闭包检查内。
  • F2: Java 的 NPE 返回 HTTP 500 而不是 409。

两项是需要维护者拍板的设计决定:

  • F3: re-attach 的 receipt 规则。
  • F4: 设计列出、但记录体没有的字段。

一项是测试强度:

  • F5: 共享 fixture 有 15–16 个条件没有被钉住。

只要 child_run 不在 MANAGED_SESSION_ENABLED_DOMAINS 中,以上各项都不可达。

运行了什么

检查 结果
PR 的 TS 聚焦套件 955/955
managed-runtime 全目录 1993/1995:2 个超时都在本 PR 未改动的 hook-scale 文件中;单独运行该文件,在同等负载下 head 与 base 都是 6 个失败 1 个。
Java clean verify checkstyle:check(含 MariaDB IT) 单测 553 通过(1 跳过),Checkstyle 0。IT 51/52:admitsHookExecutionsWithoutReadingTheirHistoryOnMySql 在本机 head 与 base 单跑都失败,CI 的 MariaDB lane 通过。
S1 生产门禁 head 报 "registered but not enabled",base 报 "no Stage H record body",两侧 journal 均不变。
产物差异 core/dist 只多了新模块和一个注册项,启用域清单逐字相同。jar 中 570/576 个条目 CRC 相同,变化的 6 个全部在 ManagedExtensionProjection* / ManagedExtensionRecords*。
S2:三条链(到退出、stop 请求、Runtime 丢失后 re-attach) 15 次提交;authority 视图、Java 行与公开 tasks API 之间零不一致;15 条 task.updated;被拒的提交不改 journal;第二个 writer 冷重开后视图一致。
20 万候选的 TS/Java 差分 除 F2(1,706)与 #12837 已知的 NFC 差异(290)外零分歧
变异测试 TS 41 个中杀死 18、等价 7;Java 38 个中杀死 17、等价 6(见 F5)
与 main 576689d073 及在飞 #13217 7f4d6aa299 的试合并 无冲突;Java 595 单测与 Checkstyle、TS 955 全过;合并 jar 上 S2 15/15、零不一致。

F1 — child_run 引用不在两侧闭包检查内(启用前修)

设计 §Resource closure 要求:store 只在 body 引用的资源齐备时才提交;writer 在发布前先检查每个引用;"No resource kind is exempted"。代码在两侧都没有覆盖 child_run:

  • writer 侧。 verifyExtensionResources(managed-session-authority.ts:1890)只检查 MCP/Hook 四个域。
    • child_run 引用了未发布的、或其他 Session 的 commandRef 时照样发布。store 回 409 "resource missing",随后整个 Session 的 writer 停写,下一次合法提交在本地就被拒(S3 A/B)。
    • 同样的错误放在 mcp_operation 上,会在发布前被拒,writer 继续可用(C)。
  • server 侧。 applyRevision 中的嵌套引用检查(ManagedExtensionRecordStore.java:344)同样只列这四个域。
    • writer 若不把 body 的引用放进提交的资源清单,一条引用了三个不存在资源的四段链会被接受(200 ×4),并投影为 background_shell running;之后读取这三个引用都报 "does not exist",冷重开也照样接受(D)。
    • mcp_operation 做同样的省略会得到 409(E)。
  • 测试。 authority 测试使用从未发布的引用(child-run.test.ts:132-166),之所以能通过,正是因为两侧都没有这项检查。
  • monitor_run。 存在同样的缺口(H0c 起既有),而 H3 也会启用它。

F2 — 非字符串 exitSignal 触发 Java NPE(一行修复)

  • 原因。 ManagedExtensionRecords.java:688 调用 EXIT_SIGNAL.matcher(exitSignal.textValue())。对 9、true、{}、[],textValue() 为 null;而 accepts()(785 行)只捕获 InvalidRecordException,所以 isChildRunStart 与 isChildRunSuccessor 也会抛出。
  • 真实服务端表现。
    • exitSignal: 9 或 true → POST /transactions:commit 返回 500,客户端重试 3 次,服务端日志打印 ERROR 堆栈。
    • "term" → 干净的 409。
  • 可达性。 TS 在发送前就会拒绝这类 body,只有绕过 TS 校验器的 writer 才能触发。
  • 修法。 加上 exitSignal.isTextual() && …,并补一个数字型 exitSignal 的 fixture。Java 契约测试断言的是 InvalidRecordException,有这个 fixture 就能抓到。

F3 — re-attach 必须换新 start receipt(需决定)

在 generation 1 处于 recovery_blocked/runtime_lost/outcome_unknown 之后,记录 re-attach 有两种写法:

  • 在 generation 2 下以 running_attached 携带同一个 receipt:TS 与真实 Java store 都拒绝(409 "cannot follow its revision 4")。
  • 携带新 receipt:被接受。

这条规则来自 Monitor 的 rebuild(managed-child-run-record.ts:324-330)。Monitor 重建会启动新的 watch,换新 receipt 合理。设计 §Recovery 对 Shell 的说法不同:Shell 从不重建、只 re-attach,要求"记录的命令 digest 与 start receipt 相符";§Records 也只说 receipt "仅随"新 binding 改变,没有说"必须"改变。

按现在的写法,契约接受的是重跑的形态,拒绝的却是真正 re-attach 的形态。而且这条拒绝没有 fixture 钉住(T39/J39 存活)。二选一:

  • 允许 re-attach 沿用原 receipt(并考虑拒绝新 receipt);或
  • 保留现规则,在设计中把新 receipt 写明为 attach receipt,并给 rebuild-new-receipt 这个 fixture 改名。

同一张图中的 S5 印证了设计的"server 先部署"规则:H3 writer 对着 PR 之前的 jar 提交 3 个 revision,Java 没有任务行;同库换成本 PR 的 jar 后,下一个合法 revision 得到 409 "first revision … must open its run",writer 停写,之后连一个无关的新 Shell 也被拒绝。这与 §Reader gating 及 H0c 开放问题 7 的描述完全一致。

F4 — 设计 §Records 列出了封闭记录体没有的字段(需决定)

设计中的 Shell 记录包含:

  • pin 中的 workspace generation、cwdRef 与工具 profile 的 definition revision;
  • processId;
  • 投影为 draining 的 stop 请求标记。

封闭的十键记录体一个都没有:

  • run.definition 必须为 null,因此公开任务永远没有 definition_revision。
  • runtimeState()(managed-extension-projection.ts:226)无法产生 draining,而验收第 3、8 条依赖 stop 与 drain。

中英文设计说法一致。二选一:要么把 §Records 收敛到已落地的记录体;要么趁未启用,至少先把 stop 请求字段加进 v1 记录体——启用之后再加就要升记录体版本。

F5 — 共享 fixture 有 15–16 个条件没被钉住

把每个存活变异体放到 20 万候选上回放:

  • TS:16 个存活变异体存在 head 拒绝、变异体接受的输入。
  • Java:同一批条件得到 15 个;T17(小数 exitCode)没有 Java 对应。
  • 其余为等价变异体。

其中好几个条件虽然有同名 fixture,但这些 fixture 同时违反了另一条规则,被那条规则先拒掉了。例如 stop-mismatch 与 process-failed-no-receipt 在 failed 状态下带了 handler_unavailable,quota-mismatch 的 execution 停留在 intent。目前 TS 与 Java 在这些条件上全部一致,但 fixture 文件就是跨语言契约本身。每个条件的最小隔离输入见图与 results/mutation/minimal-examples.txt。

合并参考

证据(图片、装置、原始日志):wenshao/qwen-code@6ee5112c pr13265/。环境:macOS arm64,JDK 21.0.12,Node 24.18.1,docker 中的 MariaDB 10.11.18。child_run 仅在装置进程内启用,做法与 PR 自带的 authority 测试相同。

@wenshao

wenshao commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Rerun outcome on run 37115037789: success — the Hosted lane went green with no code change, exactly the call declared above (transient confirmed, attribution stays with #13255). The PR is now fully green on 8173837; a new code commit follows with the F1/F2 fixes and F3/F4 decisions from the verification below.

@wenshao

wenshao commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed in 007289fee9 (pushed). Findings accounting: F1 and F2 were real defects on my side, both fixed; F5's worst offenders are now isolated in the fixtures; F3 and F4 are decided, with the design updated in both languages; one item is deliberately carried rather than done here.

F1 — both closure checks now cover child_run. Writer side: verifyExtensionResources reads commandRef, startReceiptRef and outputRef. Server side: the nested-ref check in applyRevision lists the same three for child_run. The authority suite now publishes real managed-tool-args / managed-runtime-receipt / managed-tool-result-manifest bytes and builds records from the returned refs (the fabricated refs were exactly the blind spot you named), and it gained a refusal case: a record whose commandRef was never published is rejected with nothing journaled. monitor_run is deliberately not in this commit. Its gap is identical but older (H0c), and closing it means republishing the refs of the shared H0c fixture chains — the lifecycle bodies in managed-session-authority.extension.test.ts and the monitorChainCases are built on 'a'*64-digest refs that cannot be materialized by LocalManagedSessionResourceStore, so the fixture chains need a ref-mapping step or real-content materialization first. Rather than stretch this commit, I carried it as an explicit before-enablement item (H3's enablement slice owns monitor_run anyway). Flagged so you can overrule.

F2 — guarded and pinned. exitSignal now requires isTextual() before the pattern, so 9/true/{}/[] get InvalidRecordException → 409 instead of an NPE → 500, and isChildRunStart/isChildRunSuccessor no longer throw. New shared case exit-signal-numeric replays on both sides.

F3 — decided: re-attach keeps the same receipt. You were right that accepting only the new-receipt shape blesses a rerun and refuses a true re-attach. The Monitor rebuild rule no longer leaks into the Shell: startReceiptRef is set-once period — a re-attach under a later generation must keep the original receipt (the process it proves never restarted), and a changed receipt is refused as the shape of a rerun. The fixture formerly named rebuild-new-receipt is now reattach-same-receipt (valid) and reattach-changed-receipt (invalid), and the design's Records section states the divergence from the Monitor rule explicitly, EN and ZH.

F4 — decided: trim the design, and add the stop marker to the v1 body now. processId, cwdRef, the workspace generation key and the tool-profile definitionRevision key were design-review residue — the body now carries exactly commandRef, and the cgroup unit name derived from the execution identity is the restart-stable identity (both sentences now in the doc). draining is real: the body gained an eleventh key, stopRequested — set once, never cleared; stop_requested requires it; a start requires it absent; and while set with an unsettled execution, the task projects draining on both languages' projections, with h3-draining-attached and h3-draining-provisioning pinned in the shared projection fixtures (the older five body domains pass false, so their rows are byte-identical).

F5 — the named survivors are now single-rule fixtures. stop-mismatch, process-failed-no-receipt, quota-mismatch and nonquota-mismatch no longer trip a second rule first (the handler_unavailable-under-failed and the intent-at-terminal shapes are gone), and the four new shape/flag cases (cancelled-without-request, start-with-stop-request, stop-request-type, exit-signal-numeric) pin rules that had no isolating input. Remaining survivors beyond these I have not individually re-pinned in this pass.

Verification after the change. TS: 964 focused tests green (child-run body, authority chains incl. the new closure/draining/refusal cases, all sibling contract suites). Java: 43 across the contract/store/projection suites — including ManagedChildRunRecordContractTest replaying the updated shared fixtures with the F2 and F3 cases on the real classes, and ManagedExtensionRecordStoreTest covering the new closure branch. Rerun of the Hosted lane finished green as declared in the CI-evidence comment above.

The design docs were updated in lockstep (EN+ZH): trimmed pin fields, the re-attach receipt rule, the stopRequested/draining marker. The S5 deployment-ordering observation matches the Reader-gating section unchanged. Open to a look at the new head when convenient.

@wenshao

wenshao commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

On the three non-blocking points from the review:

  1. Platform coverage — agreed this is the product-level call, and the PR states it in the open: managed background Shell and Monitor admit only where delegated cgroup v2 exists (Linux), every other platform gets a recorded, accurate refusal rather than a weaker story. That mirrors H2's merged precedent deliberately; the weaker macOS process-group profile is a named follow-up, not inherited silence. Requesting the explicit maintainer nod at merge time rather than treating it as settled by inheritance — the surface where someone downstream first notices is admission failure, so the choice should be consciously signed off.
  2. Normative references in the personal fork — fair exposure, and your framing (convention followed, exposure newly load-bearing for acceptance gates) is the right one. Vendoring the referenced sections is a separate conversation that touches all 38 sibling docs, so this PR keeps the established pinning; suggest it be tracked as its own follow-up rather than resolved here ad hoc.
  3. Density — acknowledged. The compressed paragraphs that are load-bearing (one execution per process, receipt set-once, closure, draining) are the ones the implementation PRs will unpack, and the verification round above already forced two of them (F3/F4) into the open; expect future increment PRs to cite sections rather than restate them.

wenshao pushed a commit to wenshao/qwen-code that referenced this pull request Oct 3, 2026
@wenshao

wenshao commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification, round 2 — #13265 at 9c1437ddd6

Scope. This round covers what changed since my round-1 report (head 817383760e):

  • the round-1 fixes (007289fee9);
  • the new background-Shell pieces: the named-unit cgroup supervisor (0c1425926e), the v3 executor admission and worker registry (cddd231653, 9c1437ddd6), and the open-ended stream capture (72c5e2cd8d, 5cf54e8a67);
  • two merges of main.

It used two environments:

  • macOS: the same real stack as round 1 — the PR's Spring jar on MariaDB, driven by the PR's built TS authority over the HTTP store.
  • Linux: a privileged container with real cgroup v2 (kernel 6.8, a delegated root, cgroupns private). The PR's own supervisor tests only use a fake unit, so this is where the cgroup code actually runs.

Verdict.

  • All four round-1 defects and decisions are fixed and hold on the real stack. F5 is partly done.
  • Not ready to merge: CI is red for two reasons caused by this PR (B1).
  • The new supervisor, registry and stream capture can't be reached in production yet; nothing calls them. On real Linux, though, they have five defects that should be fixed before they are wired in (G1–G5). Three of them, plus one of the CI failures, are fixed by a 4-edit candidate patch.

What was run

Check Result
Java managed-agent-server clean verify checkstyle:check, MariaDB ITs 553 unit (1 skipped), 52/52 ITs, Checkstyle 0, SpotBugs 0 (the bot's review could not run SpotBugs)
TS: child-run, authority, extension contracts, supervisor, stream capture, cgroup (core); background shell, tool v3, worker (cli) 976 + 818 pass on macOS
macOS real stack S1–S4 (S5 not rerun: the reader-gating rules did not change) S1: shipped build still refuses child_run, nothing written. S2: 27 commits, 0 mismatches between the authority, the Java row and the public tasks API; cold reopen equal. S3/S4: below.
TS/Java differential, 200,000 candidates (now including stopRequested) TS crashes 0, Java crashes 0 (was 1,706). 272 disagreements, all the known NFC case from #12837.
TS/Java task-projection differential 8,410/8,410 records identical, 66 of them draining
Linux, real cgroup v2: supervisor + registry (L1–L8), stream capture (L9, L10) see G1–G5
Mutation testing, survivors classified by replaying the 200k candidates TS 49 → 27 killed / 6 equivalent / 16 unpinned; Java 26 → 10 killed / 16 unpinned
Trial merges: main 5ddfacc9d4, #13217 b81978cef2 Textually clean. The #13217 merge passes 613 Java unit tests with Checkstyle 0.

Round-1 findings: fixed

round-1 fixes

  • F1 (reference closure): fixed on both sides.
    • A ghost or foreign commandRef is now refused before publishing, and the writer keeps working. In round 1 it got a 409 and then the writer stopped.
    • A ghost startReceiptRef or outputRef is each refused at the revision that names it.
    • With the local checks off, Java answers 409 "resource missing" with 0 rows. In round 1 it accepted the chain: 200 ×4.
    • monitor_run is deliberately left out, as the author flagged. Whether that's acceptable is the maintainers' call.
  • F2 (Java NPE): fixed. A numeric or boolean exitSignal now gets a single 409 with no retries, and there is no NPE in the server log.
  • F3 (re-attach receipt): fixed as decided. A new receipt under generation 2 is refused and the same receipt is accepted. This holds through the authority and through Java alone.
  • F4 (fields): fixed as decided. The design is trimmed, and stopRequested projects draining both while attached and while provisioning, in Java rows and the public API. Clearing the flag is refused, and so is a start that already carries it.
  • F5 (fixtures): partly.
    • stop-mismatch and quota-mismatch now isolate their rules, so T21/J21 and T27/J27 die.
    • 14 TS and 13 Java record-level conditions are still unpinned (card 5), including the new T54/J54: no fixture says a stop request cannot be cleared.
  • Bot thread R1-19, confirmed: an outputRef that points back at an older manifest is still accepted (S2 shell-6).

B1 — CI is red for two PR-caused reasons (blocker)

CI red

Both failures are in Test (ubuntu-latest, Node 22.x), job 111248462682. I reproduced both locally.

  1. process-env-guard.test.ts. The new backgroundEnv() (managed-runtime-tool-executor.ts:1115-1126) reads process.env[key] with a computed key that isn't in the guard's allowlist. It fails on head and passes 3/3 on the base.
  2. hook-command-cgroup.test.ts › "refuses to create or attach without a delegated Linux root".
    • attach() calls resolveRoot() outside any try. On Linux, a missing root therefore throws a raw ENOENT from realpathSync instead of HookCommandIsolationUnavailableError.
    • On macOS the platform check throws first, so the test passes there. In the container: attach('/tmp/not-a-cgroup') → ENOENT, while create() on the same root gives the isolation error.

G1–G4 — the supervisor on real Linux cgroup v2 (fix before wiring)

supervisor on Linux

What works:

  • Membership survives setsid: all 6 tree members are in the unit.
  • stdout and stderr are captured.
  • When every member obeys TERM, terminate() returns SIGTERM evidence and removes the unit.

The defects:

  • G1 — attach() always returns undefined. hook-command-cgroup.ts:104 checks unitName.includes(''), an empty string, which matches every name. It was meant to be the NUL '\0'. A fresh supervisor cannot re-attach a live, populated unit, so re-attach by unit name, which is the design's recovery primitive, never works. Only the refusal path is tested.
  • G2 — terminate() races cgroup.kill. hook-command-cgroup.ts:178 writes cgroup.kill and returns without waiting. managed-child-run-supervisor.ts:78 then checks empty() immediately. With one TERM-ignoring member, the call answers null ("not proven"), skips remove(), and leaves the unit directory behind. The unit reads empty 0 ms later, and a second call returns SIGTERM.
  • G3 — the registry ends a Shell on the launcher's exit, not on the unit.
    • complete() resolves once the launcher exits and the pipes reach EOF. It never checks the unit and never removes it.
    • Natural exit while a setsid daemon with detached stdio lives on: the hold is released and success is published while the daemon still runs in a populated unit (L3). The design says "exit of the root alone is never proof".
    • Every natural exit (L8), and every registry terminate where G2 answered null (L7), leaves a unit directory behind.
    • The supervisor's process map is never cleared: size 11 after one probe run.
  • G4 — signal deaths are reported as exit code 1. The launcher (hook-command-cgroup.ts:45) exits with code ?? 1, so kill -KILL $$ and kill -SEGV $$ come back as exitCode 1, exitSignal null. child_run.exitSignal can never carry them. The launcher is shared with H2 hooks, but H3 now records this value as exit evidence.
  • Minor: a second start with a unit name already in use answers "requires a delegated Linux cgroup v2 directory" (L6). qwen-bg-<callId> with non-[A-Za-z0-9._-] characters folded to - can make two call IDs collide.

The candidate patch makes four edits to hook-command-cgroup.ts: the NUL check, mapping attach() root errors to the isolation error, waiting after cgroup.kill, and re-raising signals in the launcher.

  • Compiled from source and run in the container: L2, L5, L7 and L4 are fixed, and the missing root gives the isolation error. This fixes G1, G2, G4 and the second CI failure.
  • On macOS: the hooks and supervisor suites pass 1165 tests and tsc is clean.
  • G3 needs a registry-level decision: wait for the unit to empty (or cap and block) before releasing the hold, and remove the unit afterwards.

G5 — stream capture: no backpressure, and no live visibility (fix before wiring)

stream capture

Setup: a real process under the supervisor, its onOutput feeding LocalShellStreamCapture.write, and a segment store persisting at 20 MiB/s.

  • No pipe-level backpressure. The supervisor's onOutput ignores the promise write() returns, so chunks queue in memory.
    • As wired, peak RSS tracks the output size: 128 MiB → 163, 256 → 288, 512 → 522 MiB.
    • With the pipe paused until each write resolves, it stays at 67–69 MiB.
    • Both arms capture every byte with the correct digest.
    • Design §Logs says "the producer pauses at pipe level … so memory stays bounded" (acceptance 2: 1 GiB with flat memory).
  • Output is invisible until exit. The manifest is refreshed only when a 512-segment page fills (512 MiB per stream) or at finalize.
    • A process writing 20 lines over 2 s: a reader of the current manifest sees revision 1 with 0 bytes the whole time, and all 370 bytes only at exit.
    • This contradicts the design's and the module header's "a reader never waits for process exit".
    • The unit tests use a small segmentsPerPage, so the production cadence is never exercised.

Production reachability

The new path is inert today:

  • managed-context-worker.ts builds ManagedToolExecutor without a supervisor.
  • Every upstream gate still refuses is_background: true: hosted-workspace-tool-turn.ts:745, Java ToolPublicationContract, Broker v3 (RuntimeBrokerService.java:547), and managed-runtime-session-worker.ts:841.

So G1–G5 are before-wiring items, not production regressions. I did not drive a Hosted background turn, since nothing can reach the new path from one.

Test strength

mutation

  • New rules are pinned: set-once receipt, the stopRequested rules, both draining projections, the authority passing the stop request, and the writer-side closure for child_run.
  • Unpinned:
    • J57: the store passing stopRequested into the projection.
    • J58/J59: the server-side closure for child_run and for outputRef.
    • T59/T60: the writer-side closure for outputRef and startReceiptRef; only a missing commandRef is tested.
  • The real stack confirms each of these behaviours today, but no test in CI would catch a regression.
  • The reply above says ManagedExtensionRecordStoreTest covers the new closure branch. That file is unchanged in this PR and has no child_run case.

Also noted

  • The PR body is stale. It still says "closed ten-key body" and "39 shapes and 19 chains", and lists the worker admission and the supervisor as still to come.
  • The bot's CHANGES_REQUESTED review at 817383760e still has 34 open threads with no replies. Its Critical R1-1 is the F2 NPE, now fixed. R1-19 is confirmed (above).
  • At posting, the Hosted lane and the web-shell smoke are still pending.

Merge reference

  • Blocking: B1, the two CI failures. The candidate patch covers the second.
  • Before the supervisor, registry or stream capture is wired:
    • G1, G2, G4: the candidate patch.
    • G3: registry end-of-life on unit evidence, and remove units.
    • G5: backpressure, and manifest refresh at a reader-useful cadence.
    • Tests for J57–J59 and T59/T60, and the T54/J54 fixture.
  • Maintainers' call: deferring monitor_run closure to its enablement slice.

Evidence (images, harness for both platforms, raw logs, candidate patch): wenshao/qwen-code@fae4d82b pr13265/r2/.

中文说明

真实栈验证第 2 轮 — #13265 @ 9c1437ddd6

范围。 本轮覆盖第 1 轮报告(head 817383760e)之后的全部变化:

  • 第 1 轮问题的修复(007289fee9);
  • 新增的后台 Shell 组件:具名 unit 的 cgroup supervisor(0c1425926e),v3 executor 准入与 worker 注册表(cddd231653、9c1437ddd6),以及开放式流捕获(72c5e2cd8d、5cf54e8a67);
  • 两次合入 main。

使用了两个环境:

  • macOS: 与第 1 轮相同的真实栈——PR 的 Spring jar 跑在 MariaDB 上,由 PR 构建出的 TS authority 经 HTTP store 驱动。
  • Linux: 带真实 cgroup v2 的特权容器(内核 6.8,委派根,私有 cgroupns)。PR 自带的 supervisor 测试只用假 unit,cgroup 代码只有在这里才真正运行。

结论。

  • 第 1 轮的四项缺陷与决定全部已修复,并在真实栈上成立;F5 部分完成。
  • 目前不能合入:CI 有两处由本 PR 引起的失败(B1)。
  • 新的 supervisor、注册表与流捕获在生产中尚不可达,没有任何地方调用它们;但在真实 Linux 上它们有五项缺陷,接入前应修复(G1–G5)。其中三项加上一处 CI 失败,可由一个包含 4 处改动的候选补丁修复。

运行了什么

检查 结果
Java managed-agent-server 的 clean verify checkstyle:check(含 MariaDB IT) 单测 553(跳过 1),IT 52/52,Checkstyle 0,SpotBugs 0(机器人评审没能跑 SpotBugs)
TS:core(child-run、authority、extension 契约、supervisor、流捕获、cgroup)与 cli(后台 shell、tool v3、worker) macOS 上 976 + 818 全部通过
macOS 真实栈 S1–S4(S5 未重跑:reader gating 规则没有变化) S1:发布构建仍拒绝 child_run,零写入。S2:27 次提交,authority、Java 行与公开 tasks API 之间零不一致,冷重开一致。S3/S4 见下文。
20 万候选的 TS/Java 差分(现已包含 stopRequested) TS 崩溃 0,Java 崩溃 0(原为 1,706)。272 处分歧,全部是 #12837 已知的 NFC 情形。
TS/Java 任务投影差分 8,410/8,410 条记录一致,其中 66 条为 draining
Linux 真实 cgroup v2:supervisor 与注册表(L1–L8),流捕获(L9、L10) 见 G1–G5
变异测试(存活者用 20 万候选回放分类) TS 49 个:杀死 27 / 等价 6 / 未钉住 16;Java 26 个:杀死 10 / 未钉住 16
试合并:main 5ddfacc9d4、#13217 b81978cef2 无冲突。与 #13217 的合并上 613 个 Java 单测通过,Checkstyle 0。

第 1 轮问题:已修复

  • F1(引用闭包):两侧均已修复。
    • 不存在的或来自其他 Session 的 commandRef 现在会在发布前被拒,writer 继续可用。第 1 轮是先得到 409,随后 writer 停写。
    • 不存在的 startReceiptRef 或 outputRef 各自在引用它的那个 revision 被拒。
    • 关掉本地检查后,Java 返回 409 "resource missing",零行写入。第 1 轮它接受了整条链:200 ×4。
    • monitor_run 按作者说明暂不纳入;是否接受由维护者决定。
  • F2(Java NPE):已修复。 数字或布尔类型的 exitSignal 现在只得到一次 409,没有重试,服务端日志也没有 NPE。
  • F3(re-attach receipt):按决定修复。 generation 2 下换新 receipt 被拒,沿用原 receipt 被接受;经 authority 和仅由 Java 判定时都如此。
  • F4(字段):按决定修复。 设计已收敛;stopRequested 在已 attach 和 provisioning 阶段都会投影为 draining,Java 行与公开 API 一致。清除该标记会被拒,带着该标记开始也会被拒。
  • F5(fixture):部分完成。
    • stop-mismatch、quota-mismatch 现在只违反各自那一条规则,T21/J21、T27/J27 被杀死。
    • 仍有 14 个 TS 条件和 13 个 Java 记录级条件未被钉住(见图 5),其中包括新的 T54/J54:没有 fixture 规定 stop 请求不可清除。
  • 机器人 thread R1-19 成立: 指回旧 manifest 的 outputRef 仍被接受(S2 shell-6)。

B1 — CI 有两处由本 PR 引起的失败(阻塞)

两处失败都在 Test (ubuntu-latest, Node 22.x)(job 111248462682),均已在本地复现。

  1. process-env-guard.test.ts。 新增的 backgroundEnv()(managed-runtime-tool-executor.ts:1115-1126)以计算键读取 process.env[key],不在守卫的白名单中。head 失败,base 3/3 通过。
  2. hook-command-cgroup.test.ts 中 "refuses to create or attach without a delegated Linux root"。
    • attach() 在 try 之外调用 resolveRoot(),因此在 Linux 上,根目录不存在时会抛出 realpathSync 的原始 ENOENT,而不是 HookCommandIsolationUnavailableError。
    • macOS 上平台检查先抛出,所以测试能通过。容器内:attach('/tmp/not-a-cgroup') → ENOENT,而同一根目录的 create() 给出隔离错误。

G1–G4 — supervisor 在真实 Linux cgroup v2 上(接入前修)

正常的部分:

  • 成员关系跨 setsid 保持:进程树的 6 个成员都在 unit 内。
  • stdout 与 stderr 都被捕获。
  • 所有成员都响应 TERM 时,terminate() 返回 SIGTERM 证据并删除 unit。

缺陷:

  • G1 — attach() 永远返回 undefined。 hook-command-cgroup.ts:104 检查的是 unitName.includes(''),空串,对任何名字都成立;本意应是 NUL 字符 '\0'。新的 supervisor 无法重新 attach 一个仍有成员的 unit,因此设计中的恢复原语——按 unit 名重新 attach——完全不可用。测试只覆盖了拒绝路径。
  • G2 — terminate() 与 cgroup.kill 存在竞态。 hook-command-cgroup.ts:178 写入 cgroup.kill 后不等待就返回;managed-child-run-supervisor.ts:78 随即检查 empty()。只要有一个成员忽略 TERM,调用就返回 null("未证明"),跳过 remove(),留下 unit 目录。0 ms 后再读,unit 已经为空;第二次调用返回 SIGTERM。
  • G3 — 注册表以 launcher 退出而非 unit 为准结束一个 Shell。
    • complete() 在 launcher 退出且管道 EOF 后就 resolve,从不检查 unit,也从不删除它。
    • 自然退出时若有一个已脱离 stdio 的 setsid 守护进程仍在运行:hold 已释放并发布了 success,而该守护进程仍在一个有成员的 unit 里(L3)。设计写的是 "exit of the root alone is never proof"。
    • 每次自然退出(L8),以及 G2 返回 null 时的每次注册表 terminate(L7),都会留下 unit 目录。
    • supervisor 的进程表从不清理:一次探针运行后 size 为 11。
  • G4 — 信号致死被记成退出码 1。 launcher(hook-command-cgroup.ts:45)以 code ?? 1 退出,于是 kill -KILL $$ 与 kill -SEGV $$ 都被报告为 exitCode 1, exitSignal null,child_run.exitSignal 永远记录不到它们。该 launcher 与 H2 hooks 共用,但 H3 现在把这个值当作退出证据记录。
  • 次要: 用已在使用的 unit 名再次启动,报的是"需要委派的 Linux cgroup v2 目录"(L6);而 qwen-bg-<callId> 会把 [A-Za-z0-9._-] 之外的字符折叠成 -,两个不同的 callId 可能撞名。

候选补丁对 hook-command-cgroup.ts 做 4 处修改:NUL 检查,把 attach() 的根目录错误映射为隔离错误,cgroup.kill 后等待,launcher 重新抛出信号。

  • 从源码编译后在容器中运行:L2、L5、L7、L4 均修复,根目录不存在时给出隔离错误。也就是修复了 G1、G2、G4 以及第二处 CI 失败。
  • macOS 上 hooks 与 supervisor 套件 1165 个测试通过,tsc 无错误。
  • G3 需要注册表层面的决定:释放 hold 前先等 unit 清空(或设上限并阻塞),之后删除 unit。

G5 — 流捕获:没有背压,也没有实时可见性(接入前修)

设置:supervisor 下的真实进程,其 onOutput 喂给 LocalShellStreamCapture.write,segment store 以 20 MiB/s 持久化。

  • 没有管道级背压。 supervisor 的 onOutput 忽略了 write() 返回的 promise,数据块堆积在内存中。
    • 按现有接法,峰值 RSS 随输出量增长:128 MiB → 163,256 → 288,512 → 522 MiB。
    • 每次 write 完成前暂停管道,则稳定在 67–69 MiB。
    • 两种接法都完整捕获了全部字节,digest 正确。
    • 设计 §Logs 写的是 "the producer pauses at pipe level … so memory stays bounded"(验收 2:1 GiB 时内存平稳)。
  • 输出在退出前不可见。 只有当一个 512 segment 的页写满(每个流 512 MiB)或 finalize 时,manifest 才会刷新。
    • 一个 2 秒内写 20 行的进程:读取当前 manifest 的一方全程只看到 revision 1、0 字节,370 字节直到退出才出现。
    • 这与设计和模块头注释中的 "a reader never waits for process exit" 相矛盾。
    • 单测使用了很小的 segmentsPerPage,因此从未覆盖生产节奏。

生产可达性

新路径目前是惰性的:

  • managed-context-worker.ts 构造 ManagedToolExecutor 时没有传 supervisor。
  • 所有上游关口仍拒绝 is_background: true:hosted-workspace-tool-turn.ts:745、Java ToolPublicationContract、Broker v3(RuntimeBrokerService.java:547)、managed-runtime-session-worker.ts:841。

因此 G1–G5 属于接入前事项,不是生产回归。我没有驱动 Hosted 后台 turn,因为从 Hosted turn 无法到达新路径。

测试强度

  • 新规则已被钉住: receipt 只设一次、stopRequested 相关规则、两处 draining 投影、authority 传递 stop 请求,以及 child_run 的 writer 侧闭包。
  • 未被钉住:
    • J57: store 把 stopRequested 传入投影。
    • J58/J59: child_run 及其 outputRef 的服务端闭包。
    • T59/T60: outputRef 与 startReceiptRef 的 writer 侧闭包;只测了缺失的 commandRef。
  • 这些行为今天都已由真实栈确认,但 CI 中没有任何测试能发现它们的回归。
  • 上文回复说 ManagedExtensionRecordStoreTest 覆盖了新的闭包分支;该文件在本 PR 中没有改动,也没有 child_run 用例。

另外

  • PR 正文已过期:仍写着 "closed ten-key body"、"39 shapes and 19 chains",并把 worker 准入与 supervisor 列为尚未到来。
  • 机器人在 817383760e 上的 CHANGES_REQUESTED 评审仍有 34 个 thread 未处理、无回复。其 Critical R1-1 就是 F2 的 NPE,现已修复;R1-19 成立(见上文)。
  • 发帖时 Hosted lane 与 web-shell smoke 仍在运行中。

合并参考

  • 阻塞: B1,两处 CI 失败;第二处由候选补丁覆盖。
  • 在接入 supervisor、注册表或流捕获之前:
    • G1、G2、G4:用候选补丁修复。
    • G3:注册表按 unit 证据结束一个 Shell,并删除 unit。
    • G5:加背压,并让 manifest 以对读者有用的节奏刷新。
    • 为 J57–J59 与 T59/T60 补测试,补 T54/J54 的 fixture。
  • 维护者决定: 是否接受把 monitor_run 的闭包推迟到其启用切片。

证据(图片、两个平台的装置、原始日志、候选补丁):wenshao/qwen-code@fae4d82b pr13265/r2/。

@wenshao

wenshao commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

R1 round closed — 34/34 threads replied and resolved at head 4fb9f0a9c8.

Self-accounting, by disposition:

  • 28 threads addressed with changes in 4fb9f0a9c8 (plus 007289fee9 for the already-landed R1-1 NPE guard): contract body and fixtures (subsumed settle/set-once clauses deleted on both languages, stop-reason union derived from CHILD_RUN_STOP_REASONS, isolated single-rule negative cases, owner-scope/reopen-settled/version/negative/fractional witnesses added, ref kinds renamed to managed-tool-args/managed-runtime-receipt, keys+fixedKeys+stopReasons pinned and replayed both sides), store tests (round-trip deep-equal + deep-frozen, refusal-reason regexes, Java commit-path chain through the real journal including closure refusal), CHILD_RUN_KIND deleted, and both language files revised across the supervision/closure/logs/close/notifications/recovery/acceptance sections flagged.
  • 6 threads answered in place with the obligation recorded where it belongs — each reply names the slice that carries it: runtime_lost reason cleared on the next non-gap revision (③ monitor builder), per-kind projection for H4's multi-kind child_run, exit-evidence ↔ final-manifest outcome alignment (② publisher), outputRef advance-only (② publisher; parse-time refs are content-addressed), dispatch-time startedAt pinned by H0c's frozen projection (not this slice to change), and waiting oscillation as a consequence of the shared no-self-loop line rules.

Verification at head: TS managed-runtime + child-host suites 858 passed / 5 files; Java store/contract/projection suites 31 passed / 4 files including commitsAndProjectsAChildRunChain against real MariaDB; pre-commit hooks clean. The one finding quoted against code that no longer exists in the diff (settle clause) was verified against the exact reviewed commit before disposition, per convention.

On the review's own disclosure — received, and we are not treating the uncovered rounds as covered: the reverse-audit stopped at the round cap without convergence, so further findings from a sixth round land as a normal R2 on this PR; test efficacy was never measured by the review, but the suites named above were re-run at this head and are the load-bearing ones for the R1 fixes; the SpotBugs skip under the workflow-pinned Maven is a CI environment item, tracked separately from this PR's gate.

@wenshao

wenshao commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

复审固定在 6c264cb5071992ab4047b6ba45e53656f260cbf5,对照上轮 c370c5582bd9c68052e5b6e646dd4b769d116b60,覆盖新增两个提交、五个 CLI 文件。上轮两项缺陷的对应行为已修复,本次增量未发现新增阻塞问题;历史遗留项仍须保留,不能据此对整个 PR 作 APPROVE。

上轮问题 当前 head 的独立复验
后台 Shell 使用错误 Session 环境键 启动改用 tools.sessionId。上轮原三例保持断言和顺序,全部通过:真实 Config 自然保留前一 Session 全局目录、显式清空全局目录、原始 ID 正对照。复合键场景下,当前 Session 的前台与后台均取得同一复合 ID 和自身注册项目目录,不再串用前一 Session 的目录。
拒绝释放后将 Monitor 误报为已停止 close 根据 Session 的 blocked 状态选择结算语义,loop 写入 runtime_lost。原公共路由案例中,普通 turn 的 release 被拒后,DELETE 仍为 204,但持久化 Monitor 已保持 recovery_blocked / runtime_lost / outcome_unknown,stopReason=null,不再生成 settled/stop_requested。正常释放对照仍正确结算。

关闭验证的边界需要说清:原负向脚本保留了“DELETE 不应为 204”的旧断言,因此原三例仍是 2 passed / 1 failed,剩余失败仅为该 HTTP 状态断言;“不得写成 settled”的原断言及新增持久化状态检查均通过。当前修复选择关闭本地 attachment 并准确保存未确认终态,这个 204 本身不再作为缺陷报告。追加真实 public load 返回 409 hosted_turn_recovery_required(尚有未结算 input),未额外 acquire/release 或调用模型,记录保持相同阻塞状态。这里没有证明恢复完成、自动 rebuild 或物理进程停止。 双语设计仍明确把 rebuild driver 列为待实现内容。

检查关闭顺序时还排除了 MCP/live Monitor 组合:Shell publication 携带 mcpServers 被公共入口拒绝,MCP profile 即使带 captureBytes 也不能准入 Monitor。因此没有用私有 Session 状态构造不可达组合来新增问题。MCP 追加控制初次错误地等待返回值,实际准入会抛出拒绝;保留初次日志,捕获并重抛同一错误后单例复跑通过,prepare/execute/Monitor 记录均为 0。这次夹具修正没有改动原关闭负向断言。

Linux 验证:CLI 增量 build、bundle、全工作区 typecheck,五个变更 TS 文件的 ESLint/Prettier、增量 diff --check 均通过。仓库定向测试 298/298(Harness 207;background Shell、executor、Monitor loop/session 合计 91)。本轮未重复运行未改动的 capture、publisher 重驱矩阵及 Java SQL/物理验收,不把上轮结果冒充本轮运行。全局 qwen 与当前构建 CLI 的实验 worker 入口均 exit 1 / not implemented,因此继续采用内部脚本;环境验证停在真实执行服务/supervisor 启动参数边界,关闭验证使用真实 HTTP、authority/journal/resource 和受控 Broker/model,未启动物理 watch。

历史项按当前源码重新核对后仍保留:Remote Monitor terminate 仍立即返回;launcher 仍把 signal 退出转为普通 code;普通 turn.finish 的 release 路径仍在;R2-8 的 same-boot loss evidence/reclaim 与 R1-27 的 runtime_lost 投影语义仍待 owner 裁定。本轮状态修复不自动关闭这些事项。测试中的 /tmp/p21-walk.txt 调试写入已删除;既有 harness 内部 assertion 被应用捕获的问题仍留在测试清理跟进项,不扩大本轮修复范围。

测试前后均核对 10715 个受版本控制文件,所有实现源码与精确 head 一致,仅有已知的 git archive Windows 安装脚本换行差异。未修改生产源码;原始与修正版脚本、失败诊断和日志已保留,临时测试及资源已清理。发评前 head 未变;CI 快照 22 success、27 skipped、4 in progress,仍在运行的是 Web Shell visuals、Ubuntu Node 22 测试、MySQL hosted process fault gates 和 review-pr,未称全部通过。

@wenshao
wenshao added this pull request to the merge queue Oct 6, 2026
@chiga0

chiga0 commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Scope: targeted spot-check, not a full-depth review. Given 101 changed files and 23 908 additions, only the highest-risk checkpoints were audited directly; the rest of the diff is marked as unreviewed below.

No blocking findings.
Approval blockers: none.


What was checked

Safety gate — domains still disabled (verified).
MANAGED_SESSION_ENABLED_DOMAINS at head contains only goal_state, session_metadata, file_history, session_source, mcp_configuration, mcp_operation, hook_registration, hook_execution. Neither child_run nor monitor_run is in that list. No Session can commit either domain through production configuration.
Source: packages/core/src/managed-runtime/managed-session-records.ts

Collision guard at load time (verified).
managed-extension-projection.ts throws synchronously on import if any body domain overlaps with an envelope domain, catching the collision before any Session opens.

Prior-round critical findings — spot-checked at head 6c264cb5:

ID Claim Head verdict
R2-3 Page flush inside end() publishes manifests that certify false Fixed: end() now publishes the page only after the seal decision
R2-5 Settle path has no failure branch Fixed: finalize() throws "Finalized Shell capture has no physical outcome." when outcome is missing
R2-6 parseChildRun accepts settled run with wrong execution state Fixed: guard if (run.state === "settled" && execution !== "settled") fail(...) present
R1-1 Java null TextNode on matcher Plausible fix; exitSignal guarded by isTextual() before textValue(); cannot execution-confirm

Cross-check against existing reviews: qqqys approved at this exact head. The qwen-code-ci-bot CHANGES_REQUESTED predates multiple fix commits; individual critical findings appear addressed per spot-checks above.


Unreviewed dimensions

  • Bulk of the diff not reviewed: Java broker, cgroup implementation, monitor registry, observation fan-out, notification wake pump, task-events SQL journal (V40), all test files
  • R2-7 (test spawns real processes) not re-checked at head
  • R1-1 Java null TextNode: assessed but not execution-confirmed
  • Concurrency / state-machine correctness of full broker two-row ledger not audited
  • Execution rungs not run (working tree unavailable)

The production safety argument depends on the domain-gate (verified above) holding until remaining physical findings settle.

Reviewed with AI assistance.

@pomelo-nwu pomelo-nwu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approve merging the current disabled H3 slice at 6c264cb5071992ab4047b6ba45e53656f260cbf5 (base 69d5db2ff2424da01ac6f14e4c484773aae7204c). No blocking finding in the scope reviewed.

I rechecked all five historical Criticals against current source: the Java exit-signal type guard, seal-before-manifest ordering, degraded finalization on storage refusal, terminal execution/receipt invariants, and the Windows guard for POSIX process tests. I also checked Session ownership and unknown-outcome handling on the maintenance routes, the broker's process ledger/release sweep, detached receipt recovery, output lineage/concurrency, monitor wake intake, and public task-event authorization/cursors. The latest fixes select the registered composite Session environment and preserve an unproven Monitor end after a refused release.

Local verification on macOS / Node 26.7.0: build passed; 934 Core tests and 1,107 CLI tests passed across 24 targeted files, plus all three selected Harness close regressions passed. The remaining 204 Harness cases were excluded by the explicit title filter. The built production admission guard accepts all eight enabled domains and refuses both child_run and monitor_run. No source changes were made. Workspace and integration typechecks also passed.

Both Qwen Code CI and SDK Java CI passed for this exact head. Java execution evidence here comes from CI; Linux physical acceptance was not rerun locally. This is a targeted review of the listed areas; a full-depth audit of every changed file was not performed.

The production H3 admission gates remain closed. Keep the recorded physical recovery/reclaim, signal attribution, remote Monitor stop and rebuild/acceptance work outstanding before enablement. Existing test flakiness/isolation follow-ups also remain tracked.

Reviewed with Codex assistance.

@wenshao
wenshao removed this pull request from the merge queue due to a manual request Oct 6, 2026
@wenshao
wenshao added this pull request to the merge queue Oct 6, 2026
Merged via the queue into main with commit b2c95e0 Oct 6, 2026
110 of 111 checks passed
wenshao added a commit that referenced this pull request Oct 6, 2026
The merged task-journal chain (#13265) took V45 meanwhile, so the
renumber hops again; V46 sits past the new top.
wenshao added a commit that referenced this pull request Oct 6, 2026
Conflicts resolved:
- managed-agent-public-api.openapi.json: the task-journal chain (#13265)
  took v1.31.0 with its Stage H3 section; the W2 cwd contract now ships
  v1.32.0 (all three changelog sentences kept: qwenSignature v1.30, task
  events v1.31, cwd change v1.32), and the two (v1.31) feature references
  are re-anchored to (v1.32).
- web-shell generated client regenerated from the resolved contract.
The cwd migration is V46 after the task journal took V45.
qwen-code-dev-bot pushed a commit to chiga0/qwen-code that referenced this pull request Oct 6, 2026
…(W2) (QwenLM#13247)

* feat(managed-agent): let creators change a bound Session's directory (W2)

Implement the W2 slice of QwenLM#12380: a durable, idempotent cwd change for a
Workspace-bound Session under the workspace-files opt-in.

* Admission replays by idempotency key before any revision or busy check,
  is creator-only (unreadable 404 / readable non-creator 403
  session_operation_forbidden, matching the merged bound lifecycle), and
  refuses with typed 400/409 codes in a pinned order.
* Settlement probes the target against the administrator mounts (canonical
  path, file identity, storage guard, the exact directory rule a tool turn
  enforces) and commits one transaction: binding update with a revision
  CAS, operation completion or a terminal failure_code that never retries,
  and one session.context.changed event per completed change.
* The next tool turn installs the committed binding on a fresh Runtime
  Session, so no worker or Harness protocol change is needed.
* Contract v1.29.0 -> v1.30.0 flips both routes and their schemas to
  implemented; Flyway V34 adds the operation's cwd columns.

* fix(managed-agent): read additive operation columns tolerantly on upgrade paths

The V34 cwd columns are additive, but the operation mapper read them
unconditionally, so any read against a schema pinned before V34 - exactly
what the MariaDB retention upgrade IT constructs - failed with bad SQL
grammar. Map them through a metadata existence check instead (absent
becomes null, which no operation row can carry anyway), and pin the
invariant with an H2 twin of the upgrade test.

* fix(managed-agent): refuse cwd probing before the deployment gate and replay

Review round 1 on the W2 slice found real leaks before cosmetic items; the
critical one was ordering: the idempotency replay and the deployment flag
check ran before the creator check, so a revoked grant could replay-read a
live operation and an ungranted caller could map Sessions and the flag from
the refusal codes. The creator check now runs immediately after the binding
gate, matching the sibling lifecycle admission's invisible-caller
discipline; the contract descriptions and the design admission table
enumerate the wired order truthfully (including the 401 step).

* failCwdChangeOperation reports whether the terminal write landed, so a
  lease-lost refusal is logged as contested instead of as a terminal
  refusal that never happened; settlement timestamps now anchor on
  database time like the rest of the lifecycle queue.
* The legacy getPublic/getWebShell operation reads are gone; the
  kind-aware pair is the only projection.
* Test coverage closes the gaps the same review raised: probe-per-conjunct
  failure fences, recovery-scan visibility for cwd rows, a null-binding
  settlement guard, the added-column upgrade path reading a real row, the
  WebShell failed projection carrying failureCode, and the busy barrier
  refusing a lifecycle admission from the other side.

* fix(managed-agent): settle the W2 busy fences on fewer, verified facts

Review round 2 measured four Critical gaps against the merged W2 slice;
every one of them changed production order or a guard, not a doc.

* The idempotency contract outranks the deployment gate: replay now runs
  before the workspace-files check, so a lost 202 resolves through the
  original operation even when execution has since been disabled, while
  the actor invisibility check still precedes both.
* The settlement probe verifies the candidate binding instead of the
  current one — a change away from a destroyed current directory stays
  reachable, which is the escape the feature exists for — and the shared
  directory rule carries the worker's own access(R_OK|X_OK) checks so
  nothing passes a probe that the next turn would refuse as an untyped
  wedge.
* The bound later-Turn barrier counts context-changing operation rows
  only, so a stuck display mutation or an in-flight ACTION_RESPONSE can
  never wedge the Session's only re-acquire path; the wider barrier stays
  where the operation machinery owns delivery.
* The admission service enforces 401-before-key itself, the refusal log
  carries the broker code, class and message, and the cwd refusals use
  error.getCode() instead of a hard-coded literal.

Tests and documents catch up: pipeline and citations rewritten to the
wired truth, the hosted IT runs a real initial tool Turn for the later
Turn to escape with fail-fast diagnostics, and the contract texts and
generated types name unsupported_feature, the deleted-404 and the
over-128 classification divergence where they belong.

* fix(managed-agent): gate cwd changes on undecided Actions and fix the probe's transient class

Review round 3 measured one Critical and a set of acceptance-level gaps
against the merged W2 slice; each is answered in code, not prose.

* A requested permission Action holds the Session busy at both cwd
  admission and the settlement re-check, the same state the sibling
  lifecycle refuses with 409 turn_active - committing under an
  answerable approval would certify a context certainty that does not
  exist.
* A binding detached between the coordinator read and the commit now
  answers workspace_unavailable, not context_revision_conflict.
* Momentary probe I/O failures keep their cause and retry through the
  delivery machine (unavailableTransient + isRetryable consulted at
  settlement) instead of certifying a verification that never happened;
  structural refusals remain the typed terminal workspace_unavailable.

Test hardening from the same round: action-blocked admission/settlement
with answered-or-expired release, settlement busy pinned on a second
open operation (the previously untested disjunct), the two permission
conjuncts discriminated mode by mode with a uid-aware gate, the shared
rule proven to fence acquire() before installContext or ownership.claim,
the guarded probe pinned for the negative half (no current-binding call
slips through), key-form refusals probed on both surfaces including the
129-char divergence, schema-shape validation of the cwd operation DTOs
against the reviewed OpenAPI schemas including the failed => failure_code
conditional, the reclaim tail asserting the settled row is re-read after
the second dispatch, and the sentinel that makes the later-Turn pivot
individually discriminating. Docs and contract texts match the shipped
truth.

* fix(managed-agent): renumber the cwd migration to V36

V35 was claimed by V35__managed_session_tool_profile on main while this
branch was under review; this keeps `V36__managed_cwd_operation.sql` as
the version Flyway can place without colliding, together with the comments
that cite it and the merged-origin shape already verified green.

* fix(managed-agent): classify probe failures by the verdict, not the origin

Review round 4 measured the transient/terminal split at origin type and
found it backfiring the other way: a permanently gone mount root is a
NoSuchFileException - an IOException - which mapped to the retryable arm
and wedged the Session with unbounded retried scans. Classification is now
by the meaning: a vanished mount root or target directory is the
structural verdict the terminal `workspace_unavailable` exists to name
(NoSuchFileException → unavailable()), while only genuine I/O blips
retry (ESTALE/EIO, `unavailableTransient`).

* A requested permission Action now has its own admission witness - 409
  session_context_busy asserted with no open operation and no active Turn,
  so the hasDecidableAction conjunct cannot hide behind the ||
  short-circuit.
* The settlement's status re-check is pinned: a Session that closed
  between claim and commit fails typed with the revision and directory
  untouched.
* The sealed-acquire test now also proves no claim is recorded after the
  refusal (`holder_key IS NOT NULL` count 0), so the ordering the name
  claims is tested.
* The published key-form sentence counts both divergences (over-length
  and blank) and four more probes pin them on both surfaces; the cwd
  instance suite also validates the cwd operations against the serving
  oneOf unions, and its Javadoc stops claiming exclusivity.
* The coordinator's retry prose says unbounded-with-capped-delay now,
  matching the delivery machine.

* fix(managed-agent): renumber the cwd migration to V40

Main landed the session-journal migration chain meanwhile with its own
V36-V39 additions, so the earlier V36 number collided; V40 sits past
the new top without touching the journal sequence.

* test(managed-agent): pin the probe's failure split with witnesses

The round-4 fixes put the machinery in but under-pinned it in two places,
spotted during this round's self-audit:

* The resolver's terminal refusals now assert the retryable flag, and a
  new case deletes the target and then the mount root to prove a
  NoSuchFileException verdict is terminal, not retried (the Critical's
  missing regression witness).
* The delivery machine's isRetryable() consult gets its own witness: a
  retryable probe refusal leaves the operation RUNNING with no failure
  code and re-settles completed after the backoff, so dropping the
  consult turns the suite red.
* The instance suite's Javadoc no longer claims exclusivity over the
  record/schema drift harness - the javadoc reword the round-4 commit
  message announced but never included.

* fix(managed-agent): bound the cwd probe retry and classify by verdict

The round-5 re-report of the transient/terminal finding was still
standing in three places:

* Nothing bounded the retry a retryable classification triggers: a
  permanent mount fault whose shape is not ENOENT (ESTALE/EIO, a late
  EACCES) re-armed the operation on every scan forever, and the open
  CWD_CHANGE row then wedged the Session behind session_context_busy /
  session_operation_active with only database surgery as release. The
  settle's retryable arm now consults CWD_CHANGE_ATTEMPT_BUDGET (8)
  before rethrowing: outliving the budget settles with the typed
  terminal failure, the deliverable set drains, and both admission
  barriers re-open. The budget lives only in the cwd branch — the
  shared retryOperation semantics are untouched — and is greater than
  one by contract.
* The mirror direction classified momentary as permanent twice: the
  Files.is* predicates answer false rather than throwing, and the
  storage-guard leg folded every I/O failure into the terminal verdict.
  requireDirectory now probes with throwing calls (readAttributes
  NOFOLLOW, toRealPath identity, checkAccess READ|EXECUTE), widening
  the terminal arm to NoSuchFileException | AccessDeniedException;
  the guard gained a probe-only entry (WorkspaceStorageGuard.verifyProbe)
  whose I/O legs keep the cause and classify transient while structural
  refusals stay terminal, exposed as verifyMountForProbe — the shared
  acquire path is untouched byte-for-byte.
* The retryable side had no witness either: a sticky-refusal budget
  test (red when the budget is removed) and a deterministic ENAMETOOLONG
  probe (red when the IOException arm collapses) pin both directions,
  and WorkspaceRuntimeTest.assertUnavailable asserts the terminal flag
  at every call site.

Also this round: the sealed-acquire vacuous lease count is replaced by
the non-vacuous rival-claim witness (red under the claim-before-
validate mutation); the deferral and retry logs now carry the trailing
throwable so the preserved cause actually reaches the log; the
H2-twin attribution in both upgrade-test Javadoc and the store comment
is corrected to name this test as the only missing-column-branch
coverage; both design twins record the bounded retry and probe-only
guard entry. A distinct wire code for the retryable classification is
deferred on QwenLM#12380's tracker under the round-5+ convergence posture.

The earlier local hosted-IT red during this round was environmental:
a sibling worktree's build overwrote the fenced local qwencode alpha
jar with a stricter client; A/B worktree, wire trace and jar-content
fingerprint proved the tree was innocent, and the IT is green again on
the disinfected fence.

* fix(managed-agent): renumber the cwd migration to V45

Main landed the session-creator record and the Kubernetes CSI chain
meanwhile with their own V40-V44 migrations, so the earlier V40
collided; V45 sits past the new top.

* fix(managed-agent): keep the terminal verdict on the acquire path

The round-6 re-report was right on both surfaces of the third
narrowing. ENOTDIR and ELOOP surface as a bare FileSystemException on
some JDKs, so the shared rule's residual IOException arm still
classified them as retryable; and because the shared rule had also
become the throwing-call shape on the acquire path, a persisted
plain-file/sub cwd answered retryable there too - the regression this
diff tags - turning an immediate typed failure into transientFailure
retries or a RECOVERY_BLOCKED park.

Split along the duality the guard already has: the acquire path keeps
its long-standing boolean-predicate requireDirectory and
verifyMountIntact with every I/O anomaly terminal, while the probe
twins requireDirectoryForProbe/verifyMountIntactForProbe do the
classification by verdict - the residual IOException arm walks the
ancestor chain (hasStructuralAncestor): a plain-file or a
dangling/looping symlink ancestor is structural, whatever remains
(ESTALE/EIO, ENAMETOOLONG) retries through the 8-attempt budget.

Witnesses per the round's acceptance, both mutation-proven:
assertProbeRefused(plain-file/lib) and (loop-a/lib) beside the nested
suite (terminal code and !isRetryable), and a WorkspaceRuntimeTest
acquire case on a plain-file/sub session (assertUnavailable's
!isRetryable() tripwire plus rival claim). Fix-removed mutant puts
both red; restored, both green.

* ci(sdk-java): double the hosted MySQL verify step's budget to 20m

The step has run at the 12-minute edge for weeks - the previous green
run on 12d6c41 finished in 11m04s with 684 module tests - and the
upstream merge wave (CSI chain, journal chain, H0c) lifted the suite
to 990 tests plus the hosted-mysql failsafes, so the step now times
out at 12 minutes while maven completes BUILD SUCCESS five minutes
later, killing the lane with everything green. Bubble the production
workload, not the test outcome.

* fix(managed-agent): renumber the cwd migration to V46

The merged task-journal chain (QwenLM#13265) took V45 meanwhile, so the
renumber hops again; V46 sits past the new top.

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Downgraded from Request changes to Comment: self-PR; CI failing: review-pr. Partially reviewed — gaps disclosed.

Not reviewed: reverse audit — rounds 1 and 2 (the 3B convergence pair) BOTH reported findings, so the loop never reached two consecutive dry rounds; it was stopped after the pair rather than run to the plan cap of 5.

Not reviewed: reverse audit of the 59 chunks outside the Critical-bearing territories — scoped out for budget: only the 25 chunks holding a Critical-bearing file received rounds 1 and 2.

Not reviewed: build-and-test — "Integration Tests (CLI, No Sandbox)" was skipped in CI and its suite did not run locally; it is the only job that collects integration-tests/helpers/hosted-workspace-tool-turn-driver.ts, which this PR changes.

Not reviewed: build-and-test — "Test (windows-latest, Node 22.x)" and "Test (macos-latest, Node 22.x)" were skipped in CI, so the win32 guard fixed for R2-7 is unverified on the platform it targets.

Not reviewed: test efficacy — harnessValidated: null; the positive control never ran because no probe file was green in the unmutated baseline, so 0 mutants and 0 hunk probes were measured and no coverage conclusion rests on the probe.

Not reviewed: build-and-test — Agent 7 was killed by the workflow timeout before delivering its return; its install/build/test facts were recovered from the reports and logs it wrote on disk (install EXIT=0, build EXIT=0, 934 core changed-tests passed, cli changed-tests 2 failed | 1338 passed), but its own narrative analysis is partial.

Not reviewed: 55 reverse-audit Suggestions and 3 low-confidence ones were never verified this round — terminal-only, not posted.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 1)": 无(计划内核查均已完成;Finding 4 中 remote prepare 端点半发布行为的未决点已按其要求计入 Confidence: low,非预算截断)。; "agent reverse-audit (round 1)": 未能确认 worker/Broker 侧的 publisher install 是否在同一 descriptor 重复注册时抬升 bindingGeneration ——若抬升,则回合 N+1 的 registerPublisher (:1091/:1112,每回合无条件重跑)会使回合 N 仍在运行的 backgr…; chunk 47: 未确认 hosted ack 腿( hosted-workspace-tool-turn.ts:2122/2378/2724 、 hosted-harness-session.ts:1134 ,均在其他 chunk)是否会对 capture: null 的 not_started 结果发送 receipt,因此…; chunk 47: 未审查 packages/cli/src/serve/managed-shell-routes.ts 的 HTTP 层测试覆盖是否存在于 managed-context-worker.test.ts 等其他测试文件中(仅确认了同名 managed-shell-routes.test.ts 不在 HEAD 中…; "agent invariant-c (packages/cli/src/serve/hosted-shell-publ…": 无(本次走查在预算内完成,未截断任何检查项)。, and 25 more.

Not reviewed: "agent verify", "agent verify (round 2)" — pointed at diff lines it never opened: it made tool calls, but none of them read the diff.

⚠️ 82 finding(s) still carried the — [unverified] tag when the loop ended — the verifier never ruled on them, and they are not confirmed.

Deferred under the convergence posture (round 3, not a blocker) — recorded, not requested in this round:

  • docs/design/2026-10-03-managed-shell-monitor-runtime.md:106 — [review] Both design documents name the task-event journal's migration as V41 , but the migration this diff ships is V45__managed_session_task_journal.sql , and V41 is an unr…
  • packages/cli/src/serve/hosted-monitor-wake-turn.ts:74 — [review] The pump's guard that blocks (rather than re-drives) a wake turn which already ran has no test in createMonitorWakeRunTurn .
  • packages/cli/src/serve/managed-workspace-activation.ts:104 — [review] The new deactivate path that drains background work before rechecking hasActiveSession is not exercised end-to-end; the only activation-route release test ( managed-con…
  • packages/core/src/managed-runtime/managed-shell-protocol.ts:27 — [review] 这个共享 helper 的注释宣称"hosted turn 与 worker 都由它派生同一个 unit",但本 PR 同时在另外两处保留了逐字相同的内联派生,helper 实际是第三份副本而非唯一来源。
  • docs/design/2026-10-03-managed-shell-monitor-runtime.md:58 — [review] 文档声称本切片替换掉了**两处**后台 Shell 拒绝门,但在被审提交上只有 executeV3 那一处被替换; ManagedToolExecutor 的 v2 拒绝门原样保留在两个文件里,其中一个 PR 完全没有触碰(中文版 docs/design/2026-10-03-managed-shell-monitor-runti…
  • docs/design/2026-10-03-managed-shell-monitor-runtime.zh-CN.md:106 — [review] 任务事件账的 Flyway 版本号已过期:文中称其为「本 PR 的 V41 表」,但本 PR head 上该表实际是 V45__managed_session_task_journal.sql ,而 V41 在本树上是一张无关的 CSI 表( V41__workspace_csi_reservation.sql )…
  • integration-tests/helpers/hosted-workspace-tool-turn-driver.ts:252 — [test] integration-tests/helpers/hosted-workspace-tool-turn-driver.ts is outside every npm workspace, so the project's unit test command never collects it. The test-effi…
  • packages/cli/src/serve/hosted-child-run-session.test.ts:35 — [review] enablement.childRun 是一个永不翻转的死开关:整个测试文件里它只在 vi.hoisted 声明处被设为 true (第 21 行)、在 mock 里被读(第 35 行)、在 afterEach 里被重置为 true (第 98 行),没有任何用例把它设为 false ,所以 mock 的行为与 if …
  • packages/cli/src/serve/hosted-child-run-session.test.ts:408 — [review] 这条断言声称钉住"终态冻结"( isChildRunSuccessor 的 total freeze),但实际让它变绿的拒绝发生在更早的记录解析层,而四选一的宽松正则恰好把解析错误也接住了——冻结被删掉,本用例依然通过。
  • packages/cli/src/serve/hosted-child-run-session.ts:223 — [review] settleFailed advertises 'quota_exceeded' as a settle reason, but the method can never produce a valid quota_exceeded record — it is a dead, unimplementable option.
  • packages/cli/src/serve/hosted-harness-session.test.ts:975 — [review] 这个新测试与紧邻的上一个测试(822-974 行)在约 153/156 行中只有约 25 行不同,其余逐字重复——包括 broker 的 warm / acquire / registerPublisher / prepareV3 / executeV3 / acknowledgeV3 mock、整段 readMonitorRun …
  • packages/cli/src/serve/hosted-harness-session.test.ts:1322 — [review] 四个 it.each 分支只断言 /load 返回 200、 DELETE 返回 204,没有任何断言检查"恢复出来的会话确实带着那条 settled detached 记录及其历史 revision 谱系",而这正是测试名与注释("its pending history revisions are the ledger")所声…
  • packages/cli/src/serve/hosted-harness-session.ts:525 — [review] settleCancelledHookTurn 现在无条件做一次全量 journal 投影,但 projected 只在存在 monitor 输入时才被 wakeHasPriorAttempt 读到。
  • packages/cli/src/serve/hosted-harness-session.ts:923 — [review] 新增的 detached 扫描把每一条 child_run / monitor_run 记录修订都从 resource store 串行重读并重新解析,而 authority 早已在内存里持有每条记录的最新修订。
  • packages/cli/src/serve/hosted-harness-session.ts:992 — [review] 新增的 detached lineage 校验有四个全新结果分支,但整个仓库没有任何测试触及它们。
  • packages/cli/src/serve/hosted-harness-session.ts:1908 — [review] busy 期间唤醒泵每 500 ms 重新扫描整个已提交前缀,并重新读取+解析同一条通知资源;扫描结果在日志未推进时是完全相同的。
  • packages/cli/src/serve/hosted-harness-session.ts:2110 — [review] 这一段内联的 pending-input 计算与同文件已有的 unsettledInputsThrough(session, throughSequence) (第 322–335 行)逐字重复,本次改动必须在两处并行打上 !isMonitorInput(event) 补丁。
  • packages/cli/src/serve/hosted-harness-session.ts:3819 — [review] 任一 loop.stop() 抛出就会中断后续全部 teardown(其余 loop、publisher close、mcp close、monitor 通知 settle、 managed.close() 、 sessions.delete ),且失败的那个 loop 在重试时被静默跳过。
  • packages/cli/src/serve/hosted-monitor-loop.test.ts:438 — [review] 两处"终止后再来的行被丢弃"的断言不可能失败: settle() / stop() 已经 disarm() 掉 window 与 idle 定时器, advance(5_000) 期间 pending 表为空、不会跑任何 handler,因此 HostedMonitorLoop.onLine 的 ended 守卫没有任何测试见证(…
  • packages/cli/src/serve/hosted-monitor-session.test.ts:553 — [review] The only test of the rebuild gate passes os.tmpdir() for every case, so nothing pins that the cwd argument is actually forwarded — the directory-aware half of the gate…
  • …and 75 more (see the run report)

Convergence: round 3 posted 20 inline comment(s), 20 of them reported for the first time. Findings keep coming back to the same files: packages/core/src/managed-runtime/managed-child-run-supervisor.ts (findings in round 2; 3 more now); packages/cli/src/serve/managed-runtime-tool-executor.ts (findings in round 2; 2 more now); docs/design/2026-10-03-managed-shell-monitor-runtime.md (findings in round 2; 1 more now). (Evidence: the previous round's work list was truncated to fit the marker, so the rounds named above may be an undercount, and a new finding written under an earlier round's id cannot be told from a re-post over a partial list, so the new-finding count may be understated; the previous round was recovered from a marker this account did not post, so those rounds may not be this account's own.) A cluster that keeps producing siblings usually means the fixes are treating instances of a shared root cause — triaging that cause before the next round, or splitting an independent cluster into its own pull request, tends to end the loop faster than fixing them one at a time. (Observation only — nothing was withheld from this review because of this observation.)

Mechanism health: this round did not close cleanly, so it withholds the incremental anchor — and the round it recovered had no anchor this round could use either — none at all, one with no certifier, one certified by an identity other than the one this round runs under, or one this round's fetch refused or resolved to the head — so the next review re-reads the whole diff unless recovery grafts an earlier own anchor that the round running it can use onto the complete work list this round leaves behind, and keeps doing so until a round's marker carries an anchor again or a graft lands that the round running it can use. (Stated, not acted on — this changes nothing about what the round posts.)

中文说明

⚠️ 已从请求修改降级为评论:self-PR; CI failing: review-pr。 仅完成部分审查,审查缺口已披露。

未审查(原文为英文):reverse audit — rounds 1 and 2 (the 3B convergence pair) BOTH reported findings, so the loop never reached two consecutive dry rounds; it was stopped after the pair rather than run to the plan cap of 5.

未审查(原文为英文):reverse audit of the 59 chunks outside the Critical-bearing territories — scoped out for budget: only the 25 chunks holding a Critical-bearing file received rounds 1 and 2.

未审查(原文为英文):build-and-test — "Integration Tests (CLI, No Sandbox)" was skipped in CI and its suite did not run locally; it is the only job that collects integration-tests/helpers/hosted-workspace-tool-turn-driver.ts, which this PR changes.

未审查(原文为英文):build-and-test — "Test (windows-latest, Node 22.x)" and "Test (macos-latest, Node 22.x)" were skipped in CI, so the win32 guard fixed for R2-7 is unverified on the platform it targets.

未审查(原文为英文):test efficacy — harnessValidated: null; the positive control never ran because no probe file was green in the unmutated baseline, so 0 mutants and 0 hunk probes were measured and no coverage conclusion rests on the probe.

未审查(原文为英文):build-and-test — Agent 7 was killed by the workflow timeout before delivering its return; its install/build/test facts were recovered from the reports and logs it wrote on disk (install EXIT=0, build EXIT=0, 934 core changed-tests passed, cli changed-tests 2 failed | 1338 passed), but its own narrative analysis is partial.

未审查(原文为英文):55 reverse-audit Suggestions and 3 low-confidence ones were never verified this round — terminal-only, not posted.

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 1)":无(计划内核查均已完成;Finding 4 中 remote prepare 端点半发布行为的未决点已按其要求计入 Confidence: low,非预算截断)。;"agent reverse-audit (round 1)":未能确认 worker/Broker 侧的 publisher install 是否在同一 descriptor 重复注册时抬升 bindingGeneration ——若抬升,则回合 N+1 的 registerPublisher (:1091/:1112,每回合无条件重跑)会使回合 N 仍在运行的 backgr…;chunk 47:未确认 hosted ack 腿( hosted-workspace-tool-turn.ts:2122/2378/2724 、 hosted-harness-session.ts:1134 ,均在其他 chunk)是否会对 capture: null 的 not_started 结果发送 receipt,因此…;chunk 47:未审查 packages/cli/src/serve/managed-shell-routes.ts 的 HTTP 层测试覆盖是否存在于 managed-context-worker.test.ts 等其他测试文件中(仅确认了同名 managed-shell-routes.test.ts 不在 HEAD 中…;"agent invariant-c (packages/cli/src/serve/hosted-shell-publ…":无(本次走查在预算内完成,未截断任何检查项)。,另有 25 条。

未审查:"agent verify"、"agent verify (round 2)"——启动 prompt 为它指定了 diff 中的行,但它从未打开:有工具调用,却没有一次读取 diff。

⚠️ 循环结束时仍有 82 条发现带着 — [unverified] 标记——验证者从未对它们作出裁决,它们不算已确认。

收敛姿态下延后(第 3 轮,非阻断)——已记录,本轮不要求修改:共 95 条(原文未翻译,列表见上方英文部分)。

收敛情况:第 3 轮发布了 20 条行内评论,其中 20 条是首次提出。发现反复回到同一批文件:packages/core/src/managed-runtime/managed-child-run-supervisor.ts(第 2 轮已出过发现,本轮又有 3 条);packages/cli/src/serve/managed-runtime-tool-executor.ts(第 2 轮已出过发现,本轮又有 2 条);docs/design/2026-10-03-managed-shell-monitor-runtime.md(第 2 轮已出过发现,本轮又有 1 条)。(证据说明:上一轮的工作清单为放进标记而被截断,上述轮次可能少计;在不完整的清单上,冠以早先轮次 id 的新发现无法与重发区分,首次提出的条数可能少计;上一轮的数据来自并非本账号发布的标记,上述轮次可能不属于本账号。)一个不断再生兄弟发现的簇,通常意味着逐条修复只在处理同一根因的实例——先定位并处理该根因,或把独立的簇拆成单独的 PR,通常比逐条修复更快结束循环。(仅为观察——本轮评审未因此扣留任何内容。)

机制健康:本轮未能干净收尾,因而扣留了增量锚点,而它恢复到的那一轮也没有留下本轮可用的锚点——要么完全没有、要么没有认证者、要么由本轮运行身份之外的身份认证、要么被本轮的获取拒绝或解析为头提交——因此下一次评审将重读整个 diff,除非恢复流程把本轮能使用的更早自有锚点嫁接到本轮留下的完整工作清单上;并会一直如此,直到某一轮的标记重新带上锚点,或落地的嫁接能被运行该轮的评审使用。(仅陈述,不据此行动——这不改变本轮发布的任何内容。)

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment on lines +639 to +642
const truncated = background.remainder.slice(0, 4096) + '...';
background.remainder = '';
if (observer) observer.onLine(truncated);
else background.remainder = truncated;

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-1: [fails-closed] [new-surface] fanMonitorLines's if (!observer) return; (line 624) sits before the 4096-byte partial-line cap, so the else background.remainder = truncated; arm at line 642 is dead code — observer is necessarily truthy past line 624. The consequence is that while no observer is registered, background.remainder accumulates the watch's entire stdout with no bound at all; the cap only ever applies once an observer exists.

Reachability at this commit: child_run / monitor_run are not in MANAGED_SESSION_ENABLED_DOMAINS (packages/core/src/managed-runtime/managed-session-records.ts:123-133), and ToolPublicationContract.requirePayload still refuses a background or monitor payload, so this path cannot be entered in production yet. The mechanism is verified; it is a pre-enablement obligation rather than a live regression.

Witness:

not run — settled by reading the code at head 6c264cb5 and the verifier grep/probe trail quoted in its shard report

Suggested fix: 二选一。要么删掉这条死分支,并在注释里写明「attach 前的累积是有意不设上限的」;要么把上限移到早退之前,让两种情况都受约束: ts background.remainder += background.decoder.write(chunk); const observer = background.observer; if (observer) { let at = background.remainder.indexOf('\n'); while (at >= 0) { const line = background.remainder.slice(0, at); background.remainder = background.remainder.slice(at + 1); if (line.length > 0) observer.onLine(line); at = background.remainder.indexOf('\n'); } } if (background.remainder.length > 4096) { const truncated = background.remainder.slice(0, 4096) + '...'; background.remainder = observer ? '' : truncated; if (observer) observer.onLine(truncated); } (两个现有用例都仍为绿:'one\ntwo\nthree' 远低于上限;5000 个 x 的场景在写入时先截成 4096+'...',attach 重放时再截一次得到同样的 'x'.repeat(4096) + '...'。)

The fix must not violate an existing fact: hosted-shell-publisher.background.test.ts:1553 'replays the lines a monitor wrote before its observer registered' 断言 attach 前的行必须按序完整回放(['one','two'] → ['one','two','three-and-a-half','four']);:1615 'force-emits a truncated observation at the partial-line ceiling...' 断言 attach 时回放的是 ['x'.repeat(4096) + '...'] 且随后的 'yyyyy\n' 仍单独成一条观察。新界只允许作用于"无观察者"分支,不得改动这两条已固定的语义。 Acceptance criterion: packages/cli/src/serve/hosted-shell-publisher.background.test.ts 需新增一例:setMonitorObserver 之前先 write('stdout', Buffer.from('x'.repeat(5000))) 再 write('stdout', Buffer.from('tail\n')),然后注册观察者,断言回放出的唯一观察是 'x'.repeat(4096) + '...'。未加上述设界时该断言为红——现状下 remainder 会累积成 x*5000 + 'tail\n',attach 时行分割先命中 \n,onLine 收到的是整条 'x'.repeat(5000) + 'tail'(5004 字符,未截断)。现有两例(:1553、:1615)在此修复下均仍绿,因此它们不能充当 witness。 Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

fanMonitorLines 的 if (!observer) return;(第 624 行)位于 4096 字节 partial-line 上限之前,使第 642 行那条「无 observer 时把 remainder 截断保存」的分支永远不可达;结果是 observer 未注册期间 background.remainder 无任何上界地累积该 watch 的全部 stdout。

触发场景: Monitor watch 的 observer 只由 HostedMonitorRemoteExecutor.start → setMonitorObserver 注册(hosted-monitor-remote-executor.ts:30),而它只在 acceptMonitor 走到 resumeMonitorWatch(hosted-workspace-tool-turn.ts:2002)时才被调用——即在 attach(1977)与 settleAttached(1994)之后。LocalShellStreamResultSession.prepare 明确允许 run.execution === 'dispatch_started'(local-shell-stream-result-session.ts:124-131),所以 worker 在 attach 之前就能持续 POST write。此时每个 write 都在第 623 行把解码文本追加进 remainder,然后在第 624 行返回,第 634 行的上限永远不执行:一个 tail -f 之类持续输出的 watch,在 accept 路径抛错/重试而迟迟没有注册 observer 的期间(loop 不运行意味着 idleTimeoutMs 与 maxEvents 也都不生效),会把整段 stdout 在 CLI 进程里再存一份无上界的字符串——而这些字节 durable stream 已经存过了。第 642 行正是唯一会在该状态下给缓冲区设界的代码,它因第 624 行的提前返回而成为死代码(observer 是 const,第 624 行之后 TS 已把它收窄为必然真值)。现有测试 hosted-shell-publisher.background.test.ts:1586('force-emits a truncated observation at the partial-line ceiling')只覆盖 5000

本提交上的可达性:child_run / monitor_run 不在 MANAGED_SESSION_ENABLED_DOMAINS(managed-session-records.ts:123-133)中,且 ToolPublicationContract.requirePayload 仍拒绝 background / monitor 载荷,因此该路径在生产上尚不可进入。机制已验证;它属于「启用切片前必须修」的义务,而非当前可触发的回归。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment on lines +1474 to +1476
...(request.monitoring
? { background: true, monitoring: true }
: {}),

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-2: This guard newly routes admitted monitor calls into publisher.register, but the registration's reference.argsDigest is request.inputDigest! (line 1464) and inputDigest is only ever computed for run_shell_command (inputDigest: isShell ? managedToolDigest(input) : undefined, line 985) — so every Monitor capture is registered with argsDigest: undefined, hidden from the type checker by the !.

The runtime worker prepares a watch's capture by posting the reference it was dispatched with (this.capturePublisher.prepare({ reference, capture: { ...capture, background: true, monitoring: true } }), managed-runtime-tool-executor.ts:1187-1190). The publisher then compares the registered request against the posted one — managedToolDigest(entry.request) !== managedToolDigest(request) → 'Unregistered Shell capture.' (hosted-shell-publisher.ts:269-274). undefined cannot equal the digest string the runtime's reference carries (managed-runtime-tool-executor.ts:768 requires reference.argsDigest.replace(/^sha256:/, '') to equal the input digest), and canonicalJson throws ManagedToolProtocolError on an undefined member in the first place (managed-tool-protocol.ts:196-198). Even if that comparison passed, LocalShellStreamResultSession.prepare copies reference.argsDigest into identity.invocationDigest (local-shell-stream-result-session.ts:149), and the capture's first manifest snapshot spreads that identity straight into parseToolResultManifest (local-shell-stream-capture.ts:464-467), whose invocationDigest: id(...) = assertManagedSessionStableId rejects a non-st

Merged finding (same root cause, reported independently by two agents): the same undefined also reaches local-shell-stream-result-session.ts:149, which writes invocationDigest: reference.argsDigest into the capture identity, and managed-shell-publisher.ts:376 then fails identity.invocationDigest !== request.reference.argsDigest; and the await_runtime binding at :1441-1443 degrades to request.digest.slice(7) (the payload digest) instead of managedToolDigest(input), so original-receipt-checkpoint.ts:280's dispatchItem.inputDigest === binding.reference.argsDigest.slice(7) never holds for a monitor either.

Reachability at this commit: child_run / monitor_run are not in MANAGED_SESSION_ENABLED_DOMAINS (packages/core/src/managed-runtime/managed-session-records.ts:123-133), and ToolPublicationContract.requirePayload still refuses a background or monitor payload, so this path cannot be entered in production yet. The mechanism is verified; it is a pre-enablement obligation rather than a live regression.

Witness:

not run — settled by reading the code at head 6c264cb5 and the verifier grep/probe trail quoted in its shard report

Suggested fix: Give admitted monitor requests a real input digest so the registered reference follows the shell convention: at line 985 use inputDigest: isShell || monitoring ? managedToolDigest(input) : undefined (hoisting the call.name === 'monitor' && monitorAdmitted test above the returned literal). That also makes the await_runtime binding's inputDigest for a monitor (line 1440-1443, request.inputDigest ?? request.digest.slice(7)) equal request.argsDigest.slice(7) — the same value the shell path binds.

The fix must not violate an existing fact: 摘要比较是精确的 canonical-JSON 相等,且 Runtime 侧另有格式约束——managedToolDigest(entry.request) !== managedToolDigest(request) → throw new Error('Unregistered Shell capture.')(packages/cli/src/serve/hosted-shell-publisher.ts:272-274);!/^(?:sha256:)?[0-9a-f]{64}$/.test(body['argsDigest'])(packages/cli/src/serve/managed-runtime-tool-v3-routes.ts:54);if (reference.argsDigest.replace(/^sha256:/, '') !== inputDigest) {(packages/cli/src/serve/managed-runtime-tool-executor.ts:768)。因此新摘要必须是归一化后 monitor input({ ...args, is_monitor: true })的 managedToolDigest,前缀形式必须与 Runtime 投递给 publisher 的 reference.ar Acceptance criterion: 目前无测试覆盖。应在 packages/cli/src/serve/hosted-workspace-tool-turn.test.ts的 monitor arm(rig 里lane.publisher为undefined,turn 会 new 出真实 HostedShellPublisher)断言注册进 publisher 的 capture reference.argsDigest === managedToolDigest({ command: 'tail -f build.log', is_monitor: true });或把 hosted-shell-publisher.background.test.ts里手写的argsDigest: 'sha256:' + 'a'.repeat(64)换成由 turn 的字段映射产出后跑一次 prepare。变异验证:把inputDigest改回isShell ? …,该测试必须红(ManagedToolProtocolError/Unregistered Shell capture.`)。 Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

本 hunk 新把已准入的 monitor 调用引进 publisher.register,但注册用的 reference.argsDigest 取自 request.inputDigest!(第 1464 行),而 inputDigest 只在 isShell 为真时计算(第 985 行),对 monitor 恒为 undefined。Runtime 侧 prepare 会对该 request 求 managedToolDigest,其 canonical JSON 访问器遇到 undefined 直接抛错,于是每次 Monitor watch 都在 capture prepare 上失败(409 hosted_shell_publication_failed);同一个 undefined 还让 await_runtime binding 退化到 payload 摘要而非输入摘要。

本提交上的可达性:child_run / monitor_run 不在 MANAGED_SESSION_ENABLED_DOMAINS(managed-session-records.ts:123-133)中,且 ToolPublicationContract.requirePayload 仍拒绝 background / monitor 载荷,因此该路径在生产上尚不可进入。机制已验证;它属于「启用切片前必须修」的义务,而非当前可触发的回归。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment on lines +2129 to +2131
throw new Error(
`Monitor watch acknowledgement failed: ${cause instanceof Error ? cause.message : String(cause)}`,
);

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-3: [fails-closed] [new-surface] acceptMonitor escalates any retryable acknowledgeV3 transport failure (429, 5xx, timeout) into a thrown fatal error, after the Session receipt, history and await resolution have all been committed. Its two sibling paths in this same PR — acceptBackgroundShell (:2385) and acceptShell (:2731) — only write to stderr, and the method's own comment calls it "the mirror of acceptBackgroundShell". The throw marks the whole Session blocked, so a transient ACK failure on an already-settled turn wedges the Session rather than being retried.

Reachability at this commit: child_run / monitor_run are not in MANAGED_SESSION_ENABLED_DOMAINS (packages/core/src/managed-runtime/managed-session-records.ts:123-133), and ToolPublicationContract.requirePayload still refuses a background or monitor payload, so this path cannot be entered in production yet. The mechanism is verified; it is a pre-enablement obligation rather than a live regression.

Witness:

`✓ Hosted Harness no-tool session > wires the wake pump with the shared recovery predicate (M3b) 11366ms`(一次 11.4s 的 server + POST /load + DELETE 集成测试,唯一断言是 `expect(wakeDeps.last.needsRecovery).toBe(monitorWakeNeedsRecovery)`);`grep -n wakeDeps` → 4 hits,全部服务于该断言。

Suggested fix: 与两个同类路径一致:把该 catch 改为 writeStderrLineSafe('qwen serve: Monitor watch v3 ACK can be retried after its Session receipt: ' + String(cause)); 并 return converted;(若确实要保留更严的语义,则至少按 ManagedSessionStoreHttpError.status === 429 || status >= 500、TypeError、DOMException(AbortError|TimeoutError) 分类,只对确定性冲突抛出,并带上 { cause })。

The fix must not violate an existing fact: 兄弟路径已把该规则写成代码内的显式前提 —— hosted-workspace-tool-turn.ts:2385-2387:'qwen serve: Background Shell v3 ACK can be retried after its Session receipt: ';致命化的后果由 hosted-harness-session.ts:2504(if (admitted) session.blocked = true;)与 hosted-harness-session.ts:3221(session.blocked = true;)固定。修复不得改变 ACK 之前的提交顺序(receipt → commit('tool_result') → resolveAwaitRuntime → ACK),因为"收据先落、ACK 可重试"正是该规则成立的前提。 Acceptance criterion: packages/cli/src/serve/hosted-workspace-tool-turn.test.ts 中 "hosted Monitor admission arm" 下新增一条用例:broker.acknowledgeV3.mockRejectedValueOnce(new Error('ACK transport down')) 后执行 DETACHED 的 monitor accept,断言 execute() 仍 resolve 出 executionStatus: 'success' 的 parts、tool.receipt 只有 1 条、rig.options.monitorLoops?.has('monitor-execution') 为 true,且不 reject HostedToolRecoveryRequiredError。现成的同形先例是同文件 548 行对 acceptShell 的 ACK 拒绝重放断言(expect(replayed).toEqual(result))。把上面的 log 改回 throw,该用例必须变红。 Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

acceptMonitor 在 Session 收据、history、await 解析全部提交完成之后,把一次可重试的 v3 ACK 传输失败升级为致命异常,与同一 PR 中两个兄弟路径(acceptBackgroundShell:2385、acceptShell:2731 都只写 stderr)以及本方法自述的 "the mirror of acceptBackgroundShell" 契约相矛盾,结果是整个 Session 被置为 blocked。

触发场景: 模型调用 monitor,Runtime 以 captureStatus: 'detached' 正常返回,acceptMonitor 依次完成 attach → settleAttached → resumeMonitorWatch → journal tool.receipt → commit('tool_result') → resolveAwaitRuntime;此时 broker.acknowledgeV3 因一次瞬时故障被拒(hosted-workspace-broker.ts:485-503:HTTP 传输失败,或响应里 acknowledged !== true / status.state !== 'settled' 都会 throw,例如 worker 短暂 503)。异常沿 execute() 的 catch(本文件 1882 行)变成 HostedToolRecoveryRequiredError,该 catch 还会对本批次所有 reserved id 发 broker.cancel、对未删除的 shellBinding 发注定失败的 close_not_started(因为 watch 确实已启动,state !== 'NOT_STARTED',只留下一行 "close was not confirmed")。上层两个生产处理器都据此把会话停摆:hosted-harness-session.ts:2504 if (admitted) session.blocked = true; 与 hosted-harness-session.ts:3221 session.blocked = true;。于是一次可重试的 ACK 抖动,让一个 durable 状态已经完整、watch 正在运行、history 已经写入的 Session 被标记为 recovery blocked,本轮 responses 也永远不会返回给模型;而同样的抖动发生在后台 Shel

本提交上的可达性:child_run / monitor_run 不在 MANAGED_SESSION_ENABLED_DOMAINS(managed-session-records.ts:123-133)中,且 ToolPublicationContract.requirePayload 仍拒绝 background / monitor 载荷,因此该路径在生产上尚不可进入。机制已验证;它属于「启用切片前必须修」的义务,而非当前可触发的回归。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment on lines +967 to +972
for (;;) {
await Promise.allSettled(
[...this.captures.values()].map(
(entry) => entry.background?.redriveInFlight,
),
);

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-4: [certifies-falsely] [new-surface] drain()'s collection loop inspects only redriveInFlight and never background.redrive — the armed-but-not-yet-fired backoff timer. It therefore breaks immediately when a retry is queued but none is in flight, and the clearTimeout at line 981 then silently discards that queued retry. This contradicts the invariant stated in the comment directly above it ("a re-drive already on its way lands inside the drain: no close answers ahead of it"), so a close can answer while a refused finalize is still owed its retry.

Reachability at this commit: child_run / monitor_run are not in MANAGED_SESSION_ENABLED_DOMAINS (packages/core/src/managed-runtime/managed-session-records.ts:123-133), and ToolPublicationContract.requirePayload still refuses a background or monitor payload, so this path cannot be entered in production yet. The mechanism is verified; it is a pre-enablement obligation rather than a live regression.

Witness:

not run — settled by reading the code at head 6c264cb5 and the verifier grep/probe trail quoted in its shard report

Suggested fix: 让 armed 状态也进入 drain 的等待集合:scheduleRedrive 在 arm 定时器的同时把一个 redriveArmed?: Promise<void> 挂到 background 上(在该次 attempt 的整条 .then/.catch/.finally 链结算时 resolve),drain 的每轮快照改为 Promise.allSettled 收集 redriveInFlight 与 redriveArmed 两者,break 条件改为两者全空;只有在"既无飞行也无排队"之后才执行 981 行的 clearTimeout(此时它已退化为纯兜底)。

The fix must not violate an existing fact: hosted-shell-publisher.ts:95-96 const MAX_REDRIVE_ATTEMPTS = 3; / const REDRIVE_BACKOFF_MS = 250; —— 让 drain 等待 armed 定时器会把 close 的额外阻塞上界设为 (MAX_REDRIVE_ATTEMPTS - 1) * REDRIVE_BACKOFF_MS = 500 ms,修复不得引入无界等待(例如轮询直到 settle 成功),必须沿用这两个既有常数作为唯一界限。 Acceptance criterion: packages/cli/src/serve/hosted-shell-publisher.background.test.ts 需新增一例:attached 记录的 finalize 被拒且 attempt 1(0 ms)也失败(照 P2-3 用 advanceOutput 的 mockImplementation 计数抛错即可),在 250 ms 退避尚未 fire、且无任何 in-flight attempt 时调用 publisher.close(),断言 close 完成后 parseChildRun(session.authority.extensionRecord('child_run', id).record).stopReason === 'exited'。未修复时为红:close 立即返回,记录仍为 stopReason: null。P2-3(:1009)与 'drains an in-flight settle'(:920 附近,断言 answered === false 直到 releaseGate())必须同时保持绿。 Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

drain() 的收集循环只看 redriveInFlight,完全不看 background.redrive(已 armed、尚未 fire 的退避定时器),所以它在"有重试正在排队但没有重试正在飞行"时立刻 break,紧接着 981 行 clearTimeout 把这条排队中的重试静默丢弃——违反紧邻上方注释自己声明的不变量"a re-drive already on its way lands inside the drain: no close answers ahead of it either"。

触发场景: handle 的 finalize catch(:405-411)在记录已 attach(owner?.startReceiptRef 为真)时 arm 重试;scheduleRedrive 的 delay 是 (attempt - 1) * REDRIVE_BACKOFF_MS。时序:finalize 被瞬时 store 5xx 拒绝 → attempt 1(0 ms)在 drain 开始前就 fire 并失败 → catch 把 redriveAttempts = 2 并 arm attempt 2(250 ms),此时 redriveInFlight 已被 finally 清成 undefined、redrive 非空 → 会话在这 250 ms 内 close(worker 按注释所述"finalizes once and drops the hold",不会再有客户端重试)→ server.close() 与 allSettled(operations) 都在同一 tick 内解决 → for(;;) 首轮 every(... === undefined) 成立 → break → clearTimeout 干掉 attempt 2/3 → store.close()。结果:该 child_run / monitor_run 记录永远停在未 settle 状态(stopReason === null、run.execution 仍是 running_attached),而 worker 已经退出、客户端已收到 409,再没有任何一方会驱动这次 settle——正是这段重试机制存在的唯一理由。作者在 hosted-shell-publisher.background.test.ts:1009 的 P2-3 用例注释里已把该后果写明:"answering after A alone strands B's record on a Session whose own teardown

本提交上的可达性:child_run / monitor_run 不在 MANAGED_SESSION_ENABLED_DOMAINS(managed-session-records.ts:123-133)中,且 ToolPublicationContract.requirePayload 仍拒绝 background / monitor 载荷,因此该路径在生产上尚不可进入。机制已验证;它属于「启用切片前必须修」的义务,而非当前可触发的回归。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment on lines +51 to +53
executable: shell.executable,
args: [...shell.argsPrefix, text],
env: sanitizeChildEnv(),

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-5: [certifies-falsely] [new-surface] The managed Monitor watch spawns its child with the worker process's environment (sanitizeChildEnv() = process.env minus four internal secrets). It neither injects this Session's shell-context variables nor isolates the process-global identity slot, so a model-chosen command receives another Session's identity variables. This diverges from both sibling paths added in the same PR — the background Shell path resolves backgroundEnv(tools.sessionId) and the foreground path runs inside sessionIdContext.run(sessionId, ...).

Reachability at this commit: child_run / monitor_run are not in MANAGED_SESSION_ENABLED_DOMAINS (packages/core/src/managed-runtime/managed-session-records.ts:123-133), and ToolPublicationContract.requirePayload still refuses a background or monitor payload, so this path cannot be entered in production yet. The mechanism is verified; it is a pre-enablement obligation rather than a live regression.

Witness:

` same rig, same worker process, same supervisor, one session registered as `runtime-session-1` with project dir `/proj/runtime-session-1`; the process-global slot pre-set to a *different* session (`QWEN_CODE_SESSION_ID=session-A-global`, `QWEN_CODE_PROJECT_DIR=/proj/session-A-global`) to model `con

Suggested fix: 让 watch 的 env 由调用方按 Session 解析后传入,而不是在 watcher 内自造。给 MonitorWatchExecutor.start 的 identity 增加一个 env?: NodeJS.ProcessEnv(类型定义在 packages/cli/src/serve/hosted-monitor-loop.ts),watcher 改成 env: identity?.env ?? sanitizeChildEnv(),;executeV3Monitor 在 managed-runtime-tool-executor.ts:1245-1250 的调用里传 env: backgroundEnv(tools.sessionId)(backgroundEnv 已存在于同文件 1810-1831 行),使其真正成为其文档注释所称的 "the mirror of the background Shell admission"。

The fix must not violate an existing fact: packages/cli/src/serve/managed-runtime-tool-executor.ts:1810-1817 —— function backgroundEnv(sessionId: string): NodeJS.ProcessEnv { const env: NodeJS.ProcessEnv = { …sessionIdContext.run(sessionId, getShellContextEnvVars) }; … },其上方注释(1812-1815 行):"the worker serves many, no session context runs on the v3 path, and the process-global slot only ever reflects the first session created in this process";另见 packages/core/src/services/shellContextEnv.ts:18-21:"only the first Config ever claims the process-global env slot (sessionEnvClaimed in config.ts), so process.env alone would report a s Acceptance criterion: packages/cli/src/serve/managed-monitor-watcher.test.ts 的 'spawns under the identity unit through the supervisor'(第 75-96 行)目前用 toMatchObject 断言 supervisor.spec.given 的 unitName/executable/args/cwd,没有断言 env。补一条断言:identity 带 env 时 supervisor.spec.given.env 必须逐字等于该 env(含 QWEN_CODE_SESSION_ID 为本次调用的 Session),且不含未传入的 ambient 值;把修复回退成 sanitizeChildEnv() 后该断言必须变红。另需在 managed-runtime-tool-executor.test.ts 的 monitor 准入用例里断言传给 watcher 的 identity env 由 tools.sessionId 解析而来。 Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

受管 Monitor watch 用 worker 进程的环境变量(sanitizeChildEnv() = process.env 去掉 4 个内部密钥)拉起子进程,既没有注入本 Session 的 shell 上下文变量,又会把进程全局槽里别的 Session 的身份变量交给这条模型选定的命令 —— 与同一 PR 里 Background Shell 路径(backgroundEnv(tools.sessionId))和前台路径(sessionIdContext.run(sessionId, …))都不一致。

触发场景: 一个 managed runtime worker 同时服务多个 Session(managed-context-worker.ts:333 的 resolver 按 reference.sessionId 解析,backgroundEnv 自己的注释也写着 "the worker serves many")。Config 只会为进程内第一个 Session 抢占全局槽(packages/core/src/config/config.ts:3292-3294:process.env['QWEN_CODE_SESSION_ID'] = this.sessionId; sessionEnvClaimed = true;),而 sanitizeChildEnv() 直接复制 process.env、完全不查 ALS。于是 Session B 的 Monitor watch(例如 tail -f build.log 或任何会 shell out 到 qwen … 的监视脚本)拿到的 QWEN_CODE_SESSION_ID / QWEN_CODE_PROJECT_DIR / QWEN_CODE_MODEL 是 Session A 的(若 worker 内从未创建 Config,则干脆全部缺失)。具体错误结果:watch 子进程的 tracing/audit 归因到错误的 Session;按 QWEN_CODE_PROJECT_DIR 去找本 Session harness 记录的子进程会去查另一个 Session 的目录(shellContextEnv.ts:127-133 明确说明该值"keyed on the session's launch cwd … 无法在下游重算")。Background Shell 路径正是为此在修复轮里改成了按调用 Session 解析(commit b732391382 "resolve the background environment for the tool s

本提交上的可达性:child_run / monitor_run 不在 MANAGED_SESSION_ENABLED_DOMAINS(managed-session-records.ts:123-133)中,且 ToolPublicationContract.requirePayload 仍拒绝 background / monitor 载荷,因此该路径在生产上尚不可进入。机制已验证;它属于「启用切片前必须修」的义务,而非当前可触发的回归。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment on lines +3864 to +3865
settleUnstartedBackgroundSiblings(
unknown.getExecutionCallId(), result);

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-17: [fails-closed] [new-surface] The sibling :process row's settle is placed after the parent invocation has already persisted its resolve, with no atomicity between the two repository writes. A crash or kill in that window leaves the parent SETTLED and this code permanently unreachable — after a restart absorbRuntimeStatus's CAS loop skips the parent because it no longer needsReconciliation(), so the PREPARED :process row becomes residue nothing can heal, and that Session stays busy forever.

Reachability at this commit: child_run / monitor_run are not in MANAGED_SESSION_ENABLED_DOMAINS (packages/core/src/managed-runtime/managed-session-records.ts:123-133), and ToolPublicationContract.requirePayload still refuses a background or monitor payload, so this path cannot be entered in production yet. The mechanism is verified; it is a pre-enablement obligation rather than a live regression.

Witness:

not run — settled by reading the code at head 6c264cb5 and the verifier grep/probe trail quoted in its shard report

Suggested fix: 在 absorbRuntimeStatus 中把兄弟 settle 提到 resolveUnknown/resolveUnsettled 之前(同一轮 CAS 迭代内、sameIdentity 校验之后)执行;settleUnstartedBackgroundProcess 对已终态的行是幂等 no-op,所以 CAS 重试重复调用无害。同一形状的调用点在 3401 行(dispatch poll 分支,属另一 chunk),应一并对齐。

The fix must not violate an existing fact: 提前调用不得放宽证据门槛:settleUnstartedBackgroundSiblings 以 "not_started".equals(result.get("executionStatus")) 为唯一入口(RuntimeBrokerService.java:1769),且 settlePrepared 的两个实现都要求 current.getState() == PREPARED 且版本一致(JdbcToolExecutionRepository.java:352-357、InMemoryToolExecutionRepository.java:238-243),所以先写兄弟行不会误 settle 一个此后被 dispatch 的行;接口契约本身也写明「Requires the immutable identity, the current version and state PREPARED」(ToolExecutionRepository.java:44-49)。 Acceptance criterion: RuntimeBrokerServiceTest:在 DelegatingToolExecutionRepository 里让 settlePrepared 持续返回 null(模拟兄弟行写不进去),走 reconcile 吸收一个 not_started 状态,断言 父行仍是 UNKNOWN(而不是 SETTLED),并断言随后一次 reconcile 能把父行与兄弟行一起收敛。按现有顺序,父行会先变 SETTLED、重试返回 ALREADY_SETTLED、兄弟行留在 PREPARED,断言变红。现有的 settlesTheProcessRowWhenTheRuntimeProvesNoStartAndReleasesCleanly(1230 行)只覆盖两次写都成功的happy path。 Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

兄弟行(:process)的 settle 被放在父调用 已持久化 resolve 之后,两次写之间没有原子性;一旦在这中间崩溃/进程被杀,父行已 SETTLED,而这段代码此后 永远不可达,PREPARED 的 :process 行成为无法自愈的残留,使该 Session 永久 busy。

触发场景: resolveUnknown/resolveUnsettled 返回非空即父行已提交为 SETTLED,紧接着才调 settleUnstartedBackgroundSiblings。若 Broker 在这两条仓储写之间死掉:重启后 absorbRuntimeStatus 的 CAS 循环在 3848 行 if (!current.needsReconciliation() || ...) return notUnknown(current, runtimeState); 处提前返回 ALREADY_SETTLED,根本走不到新增的这一行,因此不存在「重试再补写兄弟行」的路径(findUnsettled 也只返回 EXECUTING/CANCEL_REQUESTED/UNKNOWN,不含 PREPARED)。唯一剩下的修复者是 release sweep:settleUnprovenBackgroundRows → observeProcessRow/controlProcessRow 去问物理 owner,但一个从未启动的 unit 在 runtime 侧只会得到 {state:"unknown"}(packages/cli/src/serve/managed-shell-runtime.ts:130 的兜底返回),controlProcessRow 对非 exited 的回答原样返回 process,于是该行保持 PREPARED,releaseSessionAfterSweep 的 backgroundBusy 恒为真,release 一直回 runtime_session_busy——正是这段代码注释里写的「round 6」要消灭的那个 wedge,只是被挪进了一个崩溃窗口。把顺序换成「先 settle 兄弟行、再 resolve 父行」后,同样的崩溃会留下父行 UNKNOWN + 兄弟行 SETTLED,而 UNKNOWN 是可重入的:下一次 reconcile 重新吸收同一个 `not_st

本提交上的可达性:child_run / monitor_run 不在 MANAGED_SESSION_ENABLED_DOMAINS(managed-session-records.ts:123-133)中,且 ToolPublicationContract.requirePayload 仍拒绝 background / monitor 载荷,因此该路径在生产上尚不可进入。机制已验证;它属于「启用切片前必须修」的义务,而非当前可触发的回归。

— qwen3.8-max via Qwen Code /review (v0.25.0)

}
});

it('settles a notification that landed beyond the default event page', async () => {

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-18: The test this PR adds at line 650 (settlePendingMonitorInputs > settles a notification that landed beyond the default event page) takes 16.0-16.5s of wall clock, which exceeds the testTimeout this repository configures for every non-ECS runner. packages/cli/vitest.config.ts:232 sets testTimeout: process.env['RUNNER_NAME']?.startsWith('ecs-qwen-') ? 60_000 : 15_000, so the test is red under the plain npx vitest run src/serve/hosted-monitor-wake.test.ts that AGENTS.md prescribes for contributors, and green only on the ECS CI hosts (60s ceiling).

A contributor follows AGENTS.md and runs cd packages/cli && npx vitest run src/serve/hosted-monitor-wake.test.ts on a normal runner. Vitest kills the test at 15000ms: Error: Test timed out in 15000ms. at src/serve/hosted-monitor-wake.test.ts:650, the file reports 1 failed | 17 skipped (18), exit code 1. The suite is red on every machine that is not an ecs-qwen- host, and the ~1s margin over the ceiling means it is intermittent even there under load. The same test passes at 16023ms when the ceiling is raised, so the assertion logic is sound and only the runtime is out of bounds.

Witness:

$ cd packages/cli && npx vitest run src/serve/hosted-monitor-wake.test.ts   # repo default testTimeout 15000
 × settlePendingMonitorInputs > settles a notification that landed beyond the default event page 15054ms
   → Test timed out in 15000ms.
 ❯ src/serve/hosted-monitor-wake.test.ts:650:3
 Test Files  1 failed (1)
      Tests  1 failed | 17 skipped (18)
EXIT_A=1

$ same single test with testTimeout=300s
 ✓ settlePendingMonitorInputs > settles a notification that landed beyond the default event page 16023ms
 Test Files  1 passed (1)
EXIT_C=0

$ whole file with testTimeout=300s
 ✓ src/serve/hosted-monitor-wake.test.ts (18 tests) 21085ms
EXIT_D=0

packages/cli/vitest.config.ts:232
    testTimeout: process.env['RUNNER_NAME']?.startsWith('ecs-qwen-')
      ? 60_000
      : 15_000,

Suggested fix: Either remove the real-time wait that makes the test cost 16s (drive the paging path with fake timers or by stubbing the bounded read, so it settles in milliseconds), or pin an explicit per-test ceiling on this one case, e.g. it('settles a notification that landed beyond the default event page', { timeout: 60_000 }, async () => { … }). Prefer the first: a 16s unit test also slows every local run of this file.

The fix must not violate an existing fact: packages/cli/vitest.config.ts:232-234 sets testTimeout: process.env['RUNNER_NAME']?.startsWith('ecs-qwen-') ? 60_000 : 15_000 — the 15s local ceiling is deliberate and shared by every other cli test, so raising it globally to accommodate this one case would weaken the guard for all of them. Acceptance criterion: cd packages/cli && npx vitest run src/serve/hosted-monitor-wake.test.ts on a runner whose RUNNER_NAME does not start with ecs-qwen- must exit 0. Today it exits 1 with Test timed out in 15000ms. Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

本 PR 在第 650 行新增的测试 settlePendingMonitorInputs > settles a notification that landed beyond the default event page 实测墙钟耗时 16.0–16.5 秒,超过本仓库为所有非 ECS runner 配置的 testTimeout。packages/cli/vitest.config.ts:232 写的是 process.env['RUNNER_NAME']?.startsWith('ecs-qwen-') ? 60_000 : 15_000,因此在 AGENTS.md 规定的 npx vitest run 下该测试必红,只有 ECS CI(60 秒上限)能过。断言逻辑本身没问题——把上限提到 300 秒后它以 16023ms 通过——问题在于超出默认上限约 1 秒,因此在不同机器上还会间歇性失败。

— qwen3.8-max via Qwen Code /review (v0.25.0)

// observation loop buffers (with the Legacy partial-line cap), and reports
// the watch's physical end exactly once — natural exit versus mid-run
// failure. The durable capture keeps both streams byte-exact; only stdout
// carries observation lines — the Legacy watch's own discipline. See

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-19: [fails-closed] [new-surface] The managed Monitor wake channel is deaf to stderr, and the comment that justifies it states a false fact about the Legacy tool it claims to inherit from. managed-monitor-watcher.ts:56-57 returns early on stream !== 'stdout', and the hosted fan-out gate hosted-shell-publisher.ts:349-352 requires stream === 'stdout' before fanMonitorLines. The bytes are not lost — onChunk still feeds the durable capture — but the observation channel, which is the only reason a Monitor exists, never sees fd 2. The stated premise ("only stdout carries observation lines — the Legacy watch's own discipline", echoed at hosted-monitor-loop.ts:49-50) is contradicted by Legacy's own wiring.

A model calls monitor on cargo build, npm test (jest writes its report to stderr), make, or any command whose diagnostics go to fd 2. The watch produces zero observations while the child is actively printing, so the loop's idle arm fires and the record settles idle_timeout with nothing observed — the model is told the command went silent and never sees the error text. Because MONITOR_DEBOUNCE_FLOOR_MS and the idle default (300000 ms) are the only timers, a stderr-only command yields an empty, wrongly-reasoned settle.

Reachability at this commit: child_run / monitor_run are not in MANAGED_SESSION_ENABLED_DOMAINS (packages/core/src/managed-runtime/managed-session-records.ts:123-133), and ToolPublicationContract.requirePayload still refuses a background or monitor payload, so this path cannot be entered in production yet. The mechanism is verified; it is a pre-enablement obligation rather than a live regression.

Witness:

PROBE RA11 observation lines = ["Compiling demo v01.0"] <- the two stderr lines are absent; PROBE RA11 onExit calls = 1. Legacy parity premise quoted false at head: packages/core/src/tools/monitor.ts:601-602 wires BOTH child?.stdout and child?.stderr into processLines, and :468 registry.emitEvent carries no stream distinction.

Suggested fix: Fan both streams into the observation path (drop the stream !== 'stdout' early return and the && stream === 'stdout' gate), matching Legacy's shared processLines/throttledEmit ordering. If stdout-only is the intended decision, record it in both language versions of the design doc and correct the false Legacy claim in managed-monitor-watcher.ts:22-23 and hosted-monitor-loop.ts:49-50.

The fix must not violate an existing fact: fanMonitorLines's partial-line cap is if (background.remainder.length > 4096) (hosted-shell-publisher.ts:646) while Legacy keeps separate stdoutBuf/stderrBuf under PARTIAL_LINE_BUFFER_CAP (packages/core/src/tools/monitor.ts:433-434,480) — merging both fds into one remainder shares that single 4096-byte budget across the two streams. Acceptance criterion: packages/cli/src/serve/hosted-shell-publisher.background.test.ts — drive a monitor_run capture with a write op on stream: 'stderr' and assert the registered observer's onLine received that line. Today every monitor fan test writes stdout only, so removing the gate changes no existing assertion. Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

受管 Monitor 的唤醒通道对 stderr 完全失聪,而为它辩护的注释陈述了一个关于 Legacy 工具的假事实。managed-monitor-watcher.ts:56-57 在 stream !== 'stdout' 时提前返回,hosted 侧的分发门 hosted-shell-publisher.ts:349-352 也要求 stream === 'stdout' 才进 fanMonitorLines。字节没有丢(onChunk 仍写入耐久 capture),但观测通道——Monitor 存在的唯一理由——永远看不到 fd 2。注释所称「only stdout carries observation lines — the Legacy watch's own discipline」被 Legacy 自己的接线推翻:monitor.ts:601-602 把 stdout 与 stderr 都接进 processLines,:468 的 registry.emitEvent 不区分流。

本提交上的可达性:child_run / monitor_run 不在 MANAGED_SESSION_ENABLED_DOMAINS(managed-session-records.ts:123-133)中,且 ToolPublicationContract.requirePayload 仍拒绝 background / monitor 载荷,因此该路径在生产上尚不可进入。机制已验证;它属于「启用切片前必须修」的义务,而非当前可触发的回归。

— qwen3.8-max via Qwen Code /review (v0.25.0)

Comment on lines +99 to +101
if (this.settled) return this.evidence;
if (!this.unit.empty()) return null;
this.unit.remove();

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-20: [fails-closed] [new-surface] ManagedChildRunProcess.terminate() answers null on every escalated stop. HookCommandCgroup.terminate ends with if (!(await this.waitForEmpty(graceMs))) this.kill();, and kill() is a bare writeFileSync(cgroup.kill, '1') that returns immediately. The await in the supervisor then resumes on a microtask, so if (!this.unit.empty()) return null; reads cgroup.events in the same tick as the kill write — before the worker's own exit handler can reap the child and flip populated to 0. The measured gap is 1-2 ms.

A background Shell ignores SIGTERM and is escalated to cgroup.kill. terminate() returns null, so managed-background-shell-registry.ts:124-126 answers {evidence: null} and keeps the entry; managed-shell-runtime.ts:105-108 reports state: 'unknown'; managed-workspace-activation.ts:105-111 re-checks hasActiveSession and answers 409 managed_activation_conflict; and the Broker's controlProcessRow keeps the row so releaseSessionAfterSweep answers runtime_session_busy. The HEAD commit's own policy — "park an unconfirmed stop on runtime_lost instead of claiming it stopped" — is therefore fed a false negative on every escalated stop, and the fail-closed machinery parks a process that was provably killed.

Reachability at this commit: child_run / monitor_run are not in MANAGED_SESSION_ENABLED_DOMAINS (packages/core/src/managed-runtime/managed-session-records.ts:123-133), and ToolPublicationContract.requirePayload still refuses a background or monitor payload, so this path cannot be entered in production yet. The mechanism is verified; it is a pre-enablement obligation rather than a live regression.

Witness:

INTACT PR: +322ms kill() -> SIGKILL delivered / +322ms empty()#8 -> false (supervisor.ts:100, same millisecond) / +322ms terminate() ANSWERED null / +324ms worker reaped child (code=null signal=SIGKILL) -> populated 0 / +425ms second terminate() ANSWERED {"exitCode":null,"exitSignal":"SIGKILL"}. WITH THE FIX: +389ms terminate() ANSWERED {"exitCode":null,"exitSignal":"SIGKILL"} on the first call. The probe flips.

Suggested fix: Wait for emptiness after the escalation before answering, e.g. if (!(await this.unit.waitForEmpty(2_000, () => this.settled))) return null; in the escalated branch, so the read happens after the reap rather than in the kill's own tick. A unit that genuinely will not empty still times out and still returns null.

The fix must not violate an existing fact: The file header rule at managed-child-run-supervisor.ts:19-20 — "Exit is claimed only with evidence; a unit that cannot be proven empty keeps the hold instead." The bounded wait must still answer null on timeout and must never synthesize a ChildRunExitEvidence. Acceptance criterion: packages/core/src/managed-runtime/managed-child-run-supervisor.test.ts — an escalated-stop case asserting terminate() resolves with {exitCode: null, exitSignal: 'SIGKILL'} rather than null. Measured cost of the fix: the existing keeps the process when emptiness cannot be proven case goes from ~0 ms to 2056 ms because it now pays the bounded wait; suite green with the fix (12/12 + 7/7). Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

ManagedChildRunProcess.terminate() 在每一次升级停止上都回答 null。HookCommandCgroup.terminate 以 if (!(await this.waitForEmpty(graceMs))) this.kill(); 结尾,而 kill() 只是一次立即返回的 writeFileSync(cgroup.kill, '1')。supervisor 里的 await 随后在微任务上恢复,所以 if (!this.unit.empty()) return null; 与写 kill 处在同一个 tick 内读 cgroup.events——此时 worker 自己的 exit 处理器还没回收子进程、populated 还没翻成 0。实测这个间隔是 1–2 毫秒。后果是每次升级停止都给出假阴性,而 HEAD 提交自己的策略(「park an unconfirmed stop on runtime_lost instead of claiming it stopped」)正被这个假阴性喂错。

本提交上的可达性:child_run / monitor_run 不在 MANAGED_SESSION_ENABLED_DOMAINS(managed-session-records.ts:123-133)中,且 ToolPublicationContract.requirePayload 仍拒绝 background / monitor 载荷,因此该路径在生产上尚不可进入。机制已验证;它属于「启用切片前必须修」的义务,而非当前可触发的回归。

— qwen3.8-max via Qwen Code /review (v0.25.0)

child.on('exit', (code, signal) => {
this.exitEvidence = {
exitCode: code,
exitSignal: typeof signal === 'string' ? signal : null,

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R3-21: [certifies-falsely] [new-surface] A background Shell's persisted exit evidence can never carry the command's signal. The HookCommandCgroup launcher ends with child.on('exit', (code) => process.exit(code ?? 1)); — it drops the signal and exits 1. This PR makes that launcher the root process of a supervised background Shell (managed-child-run-supervisor.ts, new) and persists its exit facts as the Shell's evidence (managed-child-run-record.ts, new), so exitSignal is only ever reachable from the terminate path, where the launcher itself is signalled.

A background command is OOM-killed (SIGKILL) or segfaults (SIGSEGV). The record commits exitCode: 1 (or 139, which is the outer shell's 128+signal encoding, not the command's exit code at all) with exitSignal: null, under the invariant "Child run exitCode or exitSignal is proven exactly when it exits" — for a command that never exited. managed-background-shell-registry.ts:216 persists signal: signalNumber(null) === null and :232-236 seals the capture with 'Background Shell exited nonzero.' instead of 'Background Shell terminated with SIGKILL.' The signal is unrecoverable from the durable record, so an operator triaging a killed background job reads a plain nonzero exit.

Reachability at this commit: child_run / monitor_run are not in MANAGED_SESSION_ENABLED_DOMAINS (packages/core/src/managed-runtime/managed-session-records.ts:123-133), and ToolPublicationContract.requirePayload still refuses a background or monitor payload, so this path cannot be entered in production yet. The mechanism is verified; it is a pre-enablement obligation rather than a live regression.

Witness:

RA14 CONTROL (exit 3) OBSERVED {"exitCode":3,"exitSignal":null} / RA14 SIGKILL OBSERVED {"exitCode":1,"exitSignal":null} / RA14 SIGSEGV OBSERVED {"exitCode":139,"exitSignal":null} / RA14 SUPERVISOR (via ManagedChildRunSupervisor.start) OBSERVED {"exitCode":1,"exitSignal":null}. WITH THE FIX (launcher forwards the signal): SIGKILL -> {"exitCode":null,"exitSignal":"SIGKILL"}, CONTROL unchanged. The control arm is what makes this evidence rather than a harness artefact.

Suggested fix: Have the launcher forward the signal — child.on('exit', (code, signal) => { if (signal) process.kill(process.pid, signal); else process.exit(code ?? 1); }) — or write the inner child's code/signal to fd 3 before exiting and have the supervisor prefer that evidence over the launcher's own exit facts.

The fix must not violate an existing fact: managed-child-run-record.ts:226 — fail('Child run exitCode must be an integer from 0 to 255.'). Forwarding the signal makes exitCode null, which is exactly why the field is nullable; the fix must not push a 128+signal value into it. Acceptance criterion: packages/core/src/managed-runtime/managed-child-run-supervisor.test.ts — a case whose command is sh -c 'kill -SEGV $$', asserting proc.evidence answers exitSignal: 'SIGSEGV'. It is red today because the launcher's flattening is invisible to the supervisor. Please prove it by mutation — remove the fix, run that test, and confirm it goes red.

中文说明

后台 Shell 持久化的退出证据永远带不上被监督命令的信号。HookCommandCgroup 的 launcher 以 child.on('exit', (code) => process.exit(code ?? 1)); 结尾——它丢掉信号并以 1 退出。本 PR 让这个 launcher 成为后台 Shell 的 root 进程(新增 managed-child-run-supervisor.ts),并把它的退出事实作为 Shell 的证据持久化(新增 managed-child-run-record.ts),于是 exitSignal 只在 terminate 路径上可达(那里被信号杀死的是 launcher 自己)。命令被 OOM kill 或段错误时,记录会在「exitCode 或 exitSignal 恰在其退出时被证明」这条不变量下提交 exitCode: 1(或 139,那是外层 shell 的 128+signal 编码,根本不是命令的退出码)而 exitSignal: null,信号信息不可恢复。压平那一行本身是既有代码(在 hunk 中以上下文出现,不是 + 行);本 PR 新增的是让它「新近变错」的消费者。

本提交上的可达性:child_run / monitor_run 不在 MANAGED_SESSION_ENABLED_DOMAINS(managed-session-records.ts:123-133)中,且 ToolPublicationContract.requirePayload 仍拒绝 background / monitor 载荷,因此该路径在生产上尚不可进入。机制已验证;它属于「启用切片前必须修」的义务,而非当前可触发的回归。

— qwen3.8-max via Qwen Code /review (v0.25.0)

qwen-code-dev-bot added a commit that referenced this pull request Oct 6, 2026
Main's H3 background Shell and Monitor runtime (#13265, b2c95e0) and
this branch's owed file-history retirement (ee3183b) both inserted
cleanup steps at the same point in the Session close sequence, right
after the lease release. Keep both: the retirement runs first because it
is guarded by its own try/catch, while the observation-loop stop is not,
so a throw there must not strand the retirement and leave the durable
marker outliving the Session.
qwen-code-dev-bot pushed a commit that referenced this pull request Oct 6, 2026
Conflict resolution:

- managed-runtime-tool-executor.ts: take main's realpath/ownership
  containment (glob admission, #13166/#13265/#13291), which subsumes this
  branch's lexical admitsDirectory check; the branch's two containment
  regression specs now assert main's 'Session working directory' message.
- use-managed-session.ts: keep this branch's per-answer-authority signals
  model (the merged panel renders its stoppedReason contract); main's
  errorOwnership writers model covered the same ground and is superseded.
- use-managed-session.test.tsx: union both suites; main's error-ownership
  specs are retimed onto this branch's jittered failure ladder (rung-1
  3000..5999ms), and three specs pinning the replaced slot machine are
  dropped (open-but-silent establishment clearing, fixed reveal priority,
  persistent stall slot) with surviving coverage recorded in the round's
  test-weakening.json.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants