Skip to content

feat(managed-agent): Define the managed-extension-record/1 contract (H0b) - #12837

Merged
wenshao merged 1 commit into
mainfrom
feat/managed-extension-record-contract
Sep 27, 2026
Merged

wenshao merged 1 commit into
mainfrom
feat/managed-extension-record-contract

Conversation

@wenshao

@wenshao wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

What this PR does

This PR defines slice H0b of #12827, the shared record contract of stage H in the Managed Agent proposal #12380: the managed-extension-record/1 records that MCP, Hooks, background Shell and Monitor, child agents, Channels and automation will all use on the Managed path. It is a contract only. Nothing commits these records yet; H0c commits them with the Session authority and rebuilds the task list from them.

The contract has six parts:

  • Operation grant. OperationGrant lets the owner of an accepted operation finish its planned phases without a model activation. It names the Session, operation, domain, revision, owner, Workspace generation, lease and a scope: the committed domain record that holds the plan (its kind must match the domain) and the phases it admits. It has no activation fields, so it can never pass for an activation. A Runtime gate may replace a grant only with a renewal that extends the lease or with a later revision that does not go back to an older Workspace generation.
  • Three state lines. A logical run (reserved … recovery_blocked), its physical execution (intent … outcome_unknown, corrupt) and the delivery or acceptance of its result, each with its allowed single steps. Delivery has two targets: a Channel (sending, partial, delivered) or a Session (accepting, accepted, consumed). The lines are tied together, so an unknown or corrupt execution keeps the run recovery_blocked, a run ends only with an execution proven to have ended, and a Session accepts a result only after the run ended.
  • Revisions. A later revision moves each line by one allowed step at most, so every step, and the evidence it rests on, is committed before the next one acts. Identities and the definition pin are set once, and the pin and the Runtime binding are recorded by the dispatch. A binding changes only when an unknown execution is attached again under a later generation. A run that ended changes only its delivery.
  • Reasons and identities. Five recovery reasons (outcome_unknown, execution_corrupt, runtime_lost, dispatch_unknown, handler_unavailable) and six quota reasons, each allowed only in the states it explains. A running or waiting run may carry a recovery reason, which is how a degraded run is recorded. executionCallId or effectId, dispatchId and deliveryId identify the physical call, the cross-Session dispatch and the external delivery.
  • Definition pin. A run uses one immutable definition revision, identified by ID, revision and digest, and a revision never names two digests.
  • monitor_run. The record body covers section 10 of the reference design: the start call's arguments, limits, debounce window, start receipt, observation and notification watermarks, stop reason and output manifest. The domain joins the closed v1 index, as the storage design lists it among the 33 v1 names, but stays disabled for submission.

One shared JSON schema and fixture file pins the contract, with 520 cases, each invalid case aimed at one rule. A Python implementation written from the design document, independent of both languages and kept outside the repository, labeled every case, and the generator stops when the two disagree. A TypeScript module in core and a Java validator in managed-agent-server, the control plane that reads these records to build the task projection, replay every case, check every pair of states on every line, and pin the constants. The bilingual design document records the decisions, rules and open questions.

Why it's needed

On the Managed path a Runtime can be reclaimed and a Harness replaced. The Legacy implementations of these capabilities recover from in-memory registries, PIDs and local files, and none of those survives. #12827 makes every asynchronous capability a durable resource plus a trigger intent, with stable identities and receipts. If each capability invented its own states, grant, reasons and identity rules, Java, qwen and the Runtime would read "unknown" and "settled" differently, and no single task projection could merge them. As with the attestation, Tool v2, managed-context/1 and managed-tool-result/1 contracts, the shared shapes and fixtures are settled first, so that H0c and H1–H6 build against a fixed, reviewed target.

Reviewer Test Plan

How to verify

  • Run the focused Vitest files in core.
    • The contract suite validates the fixtures against the schema in strict mode, replays every grant, pin, run, monitor, grant replacement, pin pair and revision case, checks every pair of states on every line, and checks that the schema and the module disagree only on the listed rules that JSON Schema cannot state.
    • The records suite checks that the domain index now has 33 names; the contract suite checks that monitor_run is still refused for submission.
  • Run the managed-agent-server contract test and Checkstyle on JDK 21. The module needs qwencode and runtime-broker installed first, as its README says.
  • Expected: every case passes in both languages, and no existing test changes.
(cd packages/core && npx vitest run src/managed-runtime/managed-extension-record.test.ts src/managed-runtime/managed-session-records.test.ts)
(cd packages/sdk-java/managed-agent-server && mvn test -Dtest=ManagedExtensionRecordContractTest && mvn checkstyle:check)

Evidence (Before & After)

N/A: there is no user-visible change. Local results, on the branch rebased onto main at 88d881491b:

  • TypeScript: ESLint, Prettier and tsc are clean for core. The contract suite passes 530 tests and the records suite 80; all 23 test files under src/managed-runtime pass (1443 tests).
  • Java: managed-agent-server passes 111 tests, 7 of them in the new contract test, and Checkstyle is clean.
  • The Python reference, the TypeScript module and the Java class agree on all 520 cases.
  • Mutation sweep, TypeScript: each failure, each early return false and each equality, logical and relational operator of the module was mutated in turn; 218 of 219 mutants fail a test. The survivor turns > into >= right after the equal-revision branch has returned, so it is equivalent.
  • Mutation sweep, Java: each require and each early return false of the validator was disabled in turn; 45 of 45 mutants fail a test.

Tested on

OS Status
🍏 macOS ⚠️
🪟 Windows ⚠️
🐧 Linux ✅

Environment (optional)

Unit tests only. Linux: Node 22 and JDK 21. macOS and Windows were not run.

Risk & Scope

  • Main risk or tradeoff: the contract fixes shapes and rules before any authority commits them, so H0c or an H slice may find a gap; any change then goes through the shared fixtures and the design document, where both languages see it. The fixture file is 612 KB because every case carries a whole record; it is generated, so review the design document, the schema and the tests rather than the file itself. The choices that most need the proposal owner's review:
    • adding monitor_run to the closed v1 index rather than a v2 index (question 1 of feat(managed-agent): Stage H extension runtime, starting with H0 shared records and the task contract #12827);
    • one run block embedded by every record, rather than a separate shared domain;
    • one step per line per revision, rather than any reachable state;
    • the five recovery and six quota reasons, derived from sections 3 and 12 of the reference design, which name the situations but no codes;
    • the grant scope as the committed plan record plus its admitted phases; each domain registers its phases in its own slice;
    • watch_failed as a Monitor stop reason, for the Legacy failures after start (non-zero exit, signal, binary output, stream error).
  • Not validated / out of scope: no authority, store, Harness, Broker, worker, route or CI workflow changes. The outbox, wake intents, grant issuance and the task projection are H0c; the public task API is H0a; the other domain bodies and their phases are H1–H6. The macOS and Windows lanes were not run locally.
  • Breaking changes / migration notes: none. domain.committed now parses monitor_run, while submission still refuses it; when H3 enables the domain, a Session that holds such records must keep readers from before this change out, for example by raising its minimumReader.

Design document: English · 中文

Linked Issues

Refs #12827 (H0b; H0a and H0c follow)
Refs #12380

中文说明

这个 PR 做了什么

本 PR 定义 #12827 的 H0b 切片,即 Managed Agent 提案 #12380 中 H 阶段的共用记录契约:managed-extension-record/1。MCP、Hooks、后台 Shell 与 Monitor、子 Agent、Channels 和自动化在 Managed 路径上都将使用这些记录。本 PR 只定义契约,尚无组件提交这些记录;H0c 会随 Session authority 提交它们,并据此重建任务列表。

契约包含六部分:

  • 操作授权。 OperationGrant 让已受理操作的 owner 在没有模型 activation 的情况下完成计划中的 phase。它写明 Session、操作、domain、修订号、owner、Workspace generation、租约和一个范围:保存计划的已提交 domain 记录(其 kind 必须与 domain 对应)以及它准入的 phase。它没有 activation 字段,因此永远不能冒充 activation。Runtime 门禁只能用延长租约的续租,或 Workspace generation 不回退的更晚修订来替换一个 grant。
  • 三条状态线。 逻辑运行(reserved … recovery_blocked)、其物理执行(intent … outcome_unknown、corrupt)以及结果的交付或接收,各有允许的单步迁移。交付有两个目标:Channel(sending、partial、delivered)或 Session(accepting、accepted、consumed)。三条线相互关联:未知或损坏的执行使运行保持 recovery_blocked,运行只有在执行被证明结束时才能结束,Session 只有在运行结束后才能接收结果。
  • 修订。 后一次修订的每条线最多前进一个允许的单步,因此每一步及其依据的证据都在下一步行动前已经提交。身份和定义固定只设置一次,定义固定和 Runtime 绑定随派发记录。绑定只有在未知的执行于更晚的 generation 下重新 attach 时才会改变。已结束的运行只能改变其交付。
  • 原因与身份。 五个恢复原因(outcome_unknown、execution_corrupt、runtime_lost、dispatch_unknown、handler_unavailable)和六个配额原因,每个只允许出现在它能解释的状态中。运行中或等待中的运行可以携带恢复原因,降级运行就是这样记录的。executionCallId 或 effectId、dispatchId 和 deliveryId 分别标识物理调用、跨 Session 派发和外部交付。
  • 定义固定。 运行使用一个不可变的定义修订,以 ID、修订号和 digest 标识;一个修订永远不会对应两个 digest。
  • monitor_run。 记录正文覆盖参考设计第 10 节:启动调用的参数、限额、去抖窗口、启动回执、观测与通知水位、停止原因和输出 manifest。该 domain 加入封闭的 v1 索引(存储设计把它列在 33 个 v1 名称中),但仍不开放提交。

一份共享的 JSON schema 和 fixture 文件固定整个契约,共 520 个用例,每个无效用例都针对一条规则。一个依据设计文档编写、独立于两种语言且放在仓库之外的 Python 实现为每个用例标注了结论,二者不一致时生成器停止。core 中的 TypeScript 模块和 managed-agent-server(读取这些记录以构建任务投影的控制面)中的 Java 校验器回放每个用例,检查每条线上的每一对状态,并固定常量。双语设计文档记录了各项决策、规则和开放问题。

为什么需要

在 Managed 路径上,Runtime 可能被回收,Harness 也可能被替换。这些能力的 Legacy 实现依靠内存注册表、PID 和本地文件恢复,这些都保不住。#12827 让每项异步能力都成为"持久资源加触发意图",具备稳定身份和回执。如果每项能力各自定义状态、授权、原因和身份规则,Java、qwen 与 Runtime 对"unknown"和"settled"的理解就会不同,也无法用统一的任务投影把它们合并。与 attestation、Tool v2、managed-context/1 和 managed-tool-result/1 契约一样,先定下共用结构和 fixtures,让 H0c 与 H1–H6 面向一个固定且经过评审的目标来构建。

评审测试计划

如何验证

  • 在 core 中运行相关的 Vitest 文件。
    • 契约测试以严格模式按 schema 校验 fixtures,回放每个 grant、定义固定、运行、monitor、grant 替换、定义固定对和修订用例,检查每条线上的每一对状态,并检查 schema 与模块只在列出的、JSON Schema 无法表达的规则上结论不同。
    • 记录测试检查 domain 索引现在有 33 个名称;契约测试检查 monitor_run 提交时仍被拒绝。
  • 在 JDK 21 上运行 managed-agent-server 的契约测试和 Checkstyle。按该模块 README,需先安装 qwencode 与 runtime-broker。
  • 预期:每个用例在两种语言中都通过,已有测试无变化。
(cd packages/core && npx vitest run src/managed-runtime/managed-extension-record.test.ts src/managed-runtime/managed-session-records.test.ts)
(cd packages/sdk-java/managed-agent-server && mvn test -Dtest=ManagedExtensionRecordContractTest && mvn checkstyle:check)

证据(前后对比)

不适用:没有用户可见的变化。以下结果基于 rebase 到 main(88d881491b)之后的分支:

  • TypeScript:core 的 ESLint、Prettier 和 tsc 均无问题。契约测试通过 530 项,记录测试通过 80 项;src/managed-runtime 下全部 23 个测试文件通过(1443 项)。
  • Java:managed-agent-server 通过 111 项测试,其中 7 项来自新的契约测试;Checkstyle 无问题。
  • Python 参考实现、TypeScript 模块和 Java 类在全部 520 个用例上结论一致。
  • TypeScript 变异检查:依次变异模块中的每个失败分支、每个提前 return false 以及每个相等、逻辑和关系运算符;219 个变异体中 218 个使测试失败。幸存者把相等修订分支返回之后的 > 改为 >=,是等价变异。
  • Java 变异检查:依次禁用校验器中的每个 require 和每个提前 return false;45 个变异体全部使测试失败。

测试平台

OS 状态
🍏 macOS ⚠️
🪟 Windows ⚠️
🐧 Linux ✅

环境(可选)

仅单元测试。Linux:Node 22 与 JDK 21。macOS 与 Windows 未运行。

风险与范围

  • 主要风险或取舍:契约在任何 authority 提交这些记录之前就固定了结构与规则,H0c 或某个 H 切片可能发现缺口;届时的任何修改都经由共享 fixtures 和设计文档进行,两种语言都能看到。fixture 文件为 612 KB,因为每个用例都携带完整记录;它是生成的,请评审设计文档、schema 和测试,而不是文件本身。最需要提案负责人评审的选择:
    • 把 monitor_run 加入封闭的 v1 索引,而不是 v2 索引(feat(managed-agent): Stage H extension runtime, starting with H0 shared records and the task contract #12827 的问题 1);
    • 每条记录内嵌一个运行块,而不是单独的共用 domain;
    • 每次修订每条线只走一步,而不是接受任何可达状态;
    • 五个恢复原因和六个配额原因:参考设计第 3、12 节列出了情形但没有给出代码,这些代码由此推导;
    • grant 的范围为已提交的计划记录加其准入的 phase;各 domain 在各自切片中登记 phase;
    • 新增 Monitor 停止原因 watch_failed,对应 Legacy 在启动后的失败(非零退出码、信号、二进制输出、流错误)。
  • 未验证/不在范围内:authority、存储、Harness、Broker、worker、路由和 CI workflow 均无改动。outbox、唤醒意图、grant 签发和任务投影属于 H0c;公开任务 API 属于 H0a;其他 domain 的正文及其 phase 属于 H1–H6。macOS 与 Windows 未在本地运行。
  • 破坏性变更/迁移说明:无。domain.committed 现在可以解析 monitor_run,但提交仍被拒绝;H3 开放该 domain 时,含有这类记录的 Session 必须把本变更之前的 reader 挡在外面,例如提高其 minimumReader。

设计文档:English · 中文

关联 Issue

Refs #12827(H0b;H0a 与 H0c 随后)
Refs #12380

…H0b)

Add the records that the Stage H capabilities share, before anything
commits them: the OperationGrant and when a Runtime gate may replace one;
the logical run, physical execution and delivery lines with their single
steps and the rules that tie them together, as one run block that every
Stage H record embeds; recovery and quota reasons; stable identities and
the Runtime binding; the definitionRevision pin; and the monitor_run
record body with its revision rule.

One shared schema and fixture file pins the contract. A TypeScript module
in core and a Java validator in managed-agent-server replay every case,
and both check every pair of states on every line. monitor_run joins the
closed v1 domain index but stays disabled for submission.

Refs #12827
@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Audit rounds before this PR

Each round froze the diff and gave it to two fresh auditors with no other context: one read the diff undirected, the other started from the design and tried to break the TypeScript module, the Java class and the schema with real probes, fuzzing and mutations. Rounds 1 and 2 fixed every valid finding; from round 3 only a Critical finding would have changed code.

Round 1 — 1 Critical, fixed

  • Critical: a Monitor that started and then failed could not be recorded. Legacy fails a monitor after start on a non-zero exit, a signal, binary output or a stream error, but failed allowed only start_failed and quota_exceeded. Added watch_failed, which requires a start receipt.
  • Seven rules had no fixture of their own, because each negative case also broke another rule. Added targeted cases.
  • A rebuilt Monitor could keep the lost generation's start receipt; a new Runtime binding now requires a new one.
  • After a run ended, only a Session delivery could be planned; the first Channel delivery's deliveryId is now allowed too.
  • A Monitor could be attached without a Runtime binding, and a binding could first appear after dispatch; both are refused now.
  • Observations could grow while the watch was lost or had ended; they now grow only when either revision is running_attached.
  • Java refused 1.0 where TypeScript, Ajv and networknt accept it; Java now accepts a number equal to an integer, as JSON Schema does.
  • Java compared records with JsonNode.equals, so IntNode(4) did not equal LongNode(4); numbers are now compared by value, with a regression test.
  • Docs: a corrupt execution has no exit (open question 4), CANCEL_REQUESTED, the minimumReader note for enabling monitor_run, the Unicode-version note for NFC, the ECMA-262 pattern note. One valid grant per domain pins the schema's domain list.

Round 2 — 1 Critical, fixed

  • Critical: Java threw NumberFormatException on a number past the double range (Jackson reads 1e400 as infinity), so successor checks threw instead of returning false. Non-finite numbers are refused now; four fixtures carry the raw 1e400.
  • The definition pin could first appear after dispatch; like the binding, it is now recorded by the dispatch.
  • TypeScript accepted a sparse phase array; parsing now visits holes, with a unit test.
  • A settled Monitor needs a start receipt; a run blocked on runtime_lost cannot be running_attached.
  • Two more fixture gaps closed.
  • Docs: the monitor_run key set is the content inside the authority's envelope (open question 3); a rebuild's intent and dispatch live in its maintenance phase's ledger; the Broker ledger is the same kind of line but not step for step; renewals repeat phases in order, and an identical grant is idempotent at the gate.

Round 3 — no Critical; deferred

  • Doc accuracy:
    • The Current state section says a cancelled PREPARED call settles as not_started. The Broker settles it as cancelled, as it does a DISPATCHING call cancelled before sending. So H0c must tell "never sent" from "sent, then cancelled" when it maps Broker records.
    • The acceptance criterion "every existing record behavior is unchanged" should say that readers now parse monitor_run in domain.committed while submission stays refused.
  • Fixture gaps under finer mutants. Both languages refuse every counterexample below, but these rules have no case of their own:
    • a recovery reason on reserved, admitted, settled or cancelled;
    • a generation change under the same binding ID;
    • a Runtime change while settling from outcome_unknown;
    • a Monitor start receipt at intent with a binding;
    • TypeScript phases given as a string;
    • Java id() control-character bounds (U+001F, DEL, U+009F) and a lone low surrogate.
  • Design questions for H0c/H3:
    • monitor_run has no notification-policy field beyond debounceMs, though the reference design mentions a notification policy.
    • Revocation is gate state outside the replacement predicate.
    • A lost watch later proven settled can no longer be rebuilt, because rebuild starts only from outcome_unknown.

Across the three rounds, the auditors' differential fuzzing over several million generated records and revision pairs found no case where TypeScript and Java disagree. The schema differs from them only on the rules the design lists as unstatable.

中文说明

提交本 PR 之前的审计轮次

每一轮都冻结 diff,交给两个没有其他上下文的新审计者:一个无方向地通读 diff,另一个从设计出发,用真实探针、模糊测试和变异尝试攻破 TypeScript 模块、Java 类和 schema。第 1、2 轮修复了所有有效发现;从第 3 轮起只有 Critical 才会改动代码。

第 1 轮:1 个 Critical,已修复

  • Critical: 启动后才失败的 Monitor 无法记录。Legacy 会在启动后因非零退出码、信号、二进制输出或流错误而让 monitor 失败,但 failed 只允许 start_failed 和 quota_exceeded。新增 watch_failed,要求有启动回执。
  • 七条规则没有专属 fixture,因为对应反例同时违反了另一条规则。已补充针对性用例。
  • 重建后的 Monitor 可以保留旧代的启动回执;现在新的 Runtime 绑定必须带来新回执。
  • 运行结束后只能规划 Session 交付;现在也允许首个 Channel 交付的 deliveryId。
  • Monitor 可以在没有 Runtime 绑定时 attach,绑定也可以在派发之后才首次出现;现在两者都被拒绝。
  • watch 丢失或结束后观测仍可增长;现在只有当前后任一修订为 running_attached 时才能增长。
  • Java 拒绝 1.0,而 TypeScript、Ajv 和 networknt 接受;现在 Java 与 JSON Schema 一致,接受等于整数的数字。
  • Java 用 JsonNode.equals 比较记录,IntNode(4) 不等于 LongNode(4);现在按数值比较,并有回归测试。
  • 文档:corrupt 执行没有出口(开放问题 4)、CANCEL_REQUESTED、开放 monitor_run 时的 minimumReader 说明、NFC 的 Unicode 版本说明、ECMA-262 正则说明。每个 domain 各一个有效 grant,固定 schema 的 domain 列表。

第 2 轮:1 个 Critical,已修复

  • Critical: Java 遇到超出 double 范围的数字(Jackson 把 1e400 读成无穷大)时抛出 NumberFormatException,后继判定因此抛异常而不是返回 false。现在拒绝非有限数;四个 fixture 带有原始的 1e400。
  • 定义固定可以在派发之后才首次出现;现在它与绑定一样随派发记录。
  • TypeScript 接受稀疏的 phase 数组;现在解析会访问空位,并有单元测试。
  • settled 的 Monitor 需要启动回执;因 runtime_lost 阻塞的运行不能是 running_attached。
  • 另补两个 fixture 缺口。
  • 文档:monitor_run 的键集合是 authority 封装之内的内容(开放问题 3);重建的 intent 与派发记录在其维护 phase 的账本中;Broker 账本是同类状态线但并非逐步对应;续租按原顺序重复 phase,完全相同的 grant 在门禁上幂等。

第 3 轮:无 Critical;延后项

  • 文档准确性:
    • "现状"一节说被取消的 PREPARED 调用以 not_started settled。Broker 实际以 cancelled 结算,发送前被取消的 DISPATCHING 调用也一样。因此 H0c 映射 Broker 记录时必须区分"从未发送"和"已发送后取消"。
    • 验收标准"现有记录行为不变"应改为:reader 现在能在 domain.committed 中解析 monitor_run,而提交仍被拒绝。
  • 更细粒度变异下的 fixture 缺口。 下列反例两种语言都会拒绝,但这些规则没有专属用例:
    • reserved、admitted、settled 或 cancelled 上的恢复原因;
    • 同一 binding ID 下的 generation 变化;
    • 从 outcome_unknown settle 时更换 Runtime;
    • 在 intent 且有绑定时带启动回执的 Monitor;
    • TypeScript 中以字符串给出的 phases;
    • Java id() 的控制字符边界(U+001F、DEL、U+009F)和孤立的低代理项。
  • 留给 H0c/H3 的设计问题:
    • 参考设计提到通知策略,但 monitor_run 除 debounceMs 外没有通知策略字段。
    • 撤销是替换判定之外的门禁状态。
    • 丢失的 watch 若之后被证明已 settled,就不能再重建,因为重建只能从 outcome_unknown 开始。

三轮中,审计者对数百万个生成的记录和修订对做了差分模糊测试,没有发现 TypeScript 与 Java 结论不同的情况。schema 与两者的差异只出现在设计列为无法表达的规则上。

wenshao added a commit that referenced this pull request Sep 27, 2026
…filtering

A task whose read_output capability was withdrawn and restored would
have its output events filtered out while the cursor moved past them,
and a task without the capability kept output events that neither the
events route nor the Artifacts returned. read_output now does not change
during a task's life, and a task without it produces no output events,
so the events route never filters out an event that exists and the
no-loss guarantees hold for every task.

With nothing filtered, next_cursor is simply the cursor of the last
returned event. The capability text also says that only cancel needs an
authorization check, since reading output needs only read access, and
the design note points H0b at #12837, which answers issue question 1.
@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Local verification on macOS: built artifacts, head 18d8313904

Verdict: nothing found that blocks merging. What ships changes by one token, "monitor_run", in the domain index. On 1,201,560 cases the TypeScript module and the Java validator gave the same verdict every time, except 371 instances of the Unicode-version gap that the design doc already documents. The schema was never stricter than the module. I'd fix the two doc sentences in F1 before merging. The 9 fixture edges in F2 can be added here or before H0c. Q1 is a question for H0c.

Environment: macOS 26 arm64, Node 24.18.1, pnpm 11.24.0, JDK 21.0.12, Maven 3.9.16. I built both trees, base 88d881491b and PR 18d8313904, with npm run build && npm run bundle and mvn package. The A/B comparisons, the probes and the differential run against those built artifacts (packages/core/dist, the jar or target/classes), not against source under vitest.

Findings

# Kind Finding
F1 Doc, fix before merge Two sentences are wrong in both languages. Round 3 and the triage bot flagged both statically; I ran both. (a) The Broker's real InMemoryToolExecutionRepository settles a cancelled PREPARED call as SETTLED with executionStatus=cancelled, not not_started. JdbcToolExecutionRepository:241 writes the same value. (b) The real LocalManagedSessionAuthority on base refuses a domain.committed event for monitor_run ("payload.domain must be one of …"), and this PR parses it. So the acceptance criterion "every existing record, authority and store behavior is unchanged" needs the exception "…except that readers now parse monitor_run in domain.committed; submitting it is still refused."
F2 Test coverage, non-blocking I wrote 9 targeted mutants for the gaps round 3 deferred: T1–T5 in TypeScript, J1–J4 in Java. All 9 pass the PR's own suites (530/530 and 7/7). All 9 are killed by my differential run or by a directed pair, so both languages behave correctly today; the shared fixtures just don't pin these edges. t3.jsonl in the harness is a ready-made case for T3, the one that random mutation missed.
Q1 Question for H0c The contract has no rule for a record's first revision. The parse functions judge one record and the successor functions judge a pair, so revision 1 can be any valid state, such as a run born with execution running_attached (the canonical monitorRun fixture is one) or a Monitor born settled with max_events. "An execution first appears as intent" therefore constrains revisions, not creation. Should H0c add an initial-revision predicate (for example: state reserved/admitted, execution null/intent, delivery null/planned), or is creating a record mid-life intended, for instance to adopt a Legacy monitor? Either answer works; one sentence in the doc would settle it before H0c commits records.
N1 Informational The documented NFC gap is concretely U+105D2 U+0307. Unicode 16 made that pair compose to U+105C9. JDK 21 (Unicode 15) accepts it as NFC; Node 24.18.1 (Unicode 17) refuses it. The TS verdict also changes within Node 22: 22.3.0 (Unicode 15.1) accepts, like JDK 21, and 22.23.2 (Unicode 17) refuses. So a TS writer and a TS reader on different Node 22 minors can disagree too. The gap comes from the Managed Session stable-ID rule rather than from this PR, and "identifiers must stay within Unicode 15" is not enforced by either validator. The doc could name the concrete pair and the Node-minor dependence.
N2 Nit The TS lease bounds are literals, but core already exports the same bounds as HTTP_MANAGED_SESSION_STORE_CONTRACT.minimumLeaseDurationMs/maximumLeaseDurationMs (1000/300000) in http-managed-session-store.ts, the TS counterpart of Java's ManagedSessionStoreModels. The triage bot's "there is no shared TS constant to reuse" is therefore not accurate. Optional.

1. What ships: one token

What ships

  • CLI bundle. I compared both dist/ trees (1,319 files each) with esbuild content-hash names normalized. They differ in two files: one chunk gains "monitor_run", and one carries the git-commit stamp. managed-extension-record is not in the bundle.
  • Jar. All 361 base entries have identical CRC-32 in the PR jar; the only additions are ManagedExtensionRecords.class and its InvalidRecordException. Nothing in src/main references them.
  • Real authority, from core dist, on both trees. Committing session_metadata (the control) works. Committing monitor_run is refused with "registered but not enabled" and the journal does not grow, identically on base and PR. Only the reader side differs (F1b): a pre-H0b reader refuses such a Session, so H3 must raise minimumReader, as the design doc says.

2. The PR's suites on macOS, and an independent TS ↔ Java differential

Suites and differential

  • The PR's platform table marks macOS as not run. Here: 530 + 80 tests, all 23 src/managed-runtime files (1,443 tests), tsc, ESLint and Prettier are clean. mvn verify passes 111 tests, 7 of them new, with 0 Checkstyle violations. These numbers match the PR description.
  • Differential: 3 seeds × 400,520 cases. Each seed is all 520 fixtures plus structured random grants, pins, runs, monitors and near-neighbour revision pairs. Number spellings such as 1.0, -0 and 1e400, and hostile ids (lone surrogates, C0/DEL/C1 characters, 512/513 bytes, NFD), reach each parser as raw text.
    • Both drivers reproduce all 520 fixture labels.
    • Neither language threw anything other than a contract refusal, and Jackson read every line.
    • The only disagreements are the 371 N1 cases.
  • Schema vs module (Ajv 2020, strict): 601,191 single records. The schema was never stricter than the module. It was looser 5,998 times, all within the doc's list of rules JSON Schema cannot state.
  • Of the 139,982 valid records in the corpus, every run, monitor and pin is accepted as its own successor, and every grant is refused as its own successor ("an identical grant replaces nothing"). Revisions that differ only in spelling (key order, 1000.0, 1e3, -0, duplicate keys) get the same verdict in both languages, 16/16.

3. The round-3 fixture gaps, measured (F2)

Mutants

4. A Monitor's life, including a Runtime-loss rebuild, in both languages

Monitor lifecycle

Both languages accept all 11 legal revisions: admission, intent, dispatch, attach, observation, Runtime loss, rebuild under generation 2 with a new receipt, max_events, and the watermark catching up after the end. Both refuse all 9 shortcuts, including reusing the generation-1 receipt on rebuild, observing while the watch is lost, and a lost watch "settling" under a new binding.

5. Doc facts, the pinned references, and CI coverage

Doc facts and CI

  • At doudouOUC/code_agent@6891216:
    • Storage design §3.1 lists the same 33 names in the same order as MANAGED_SESSION_DOMAINS, with monitor_run after memory_job.
    • The control protocol's OperationGrant field list is exactly the PR's 9 closed keys.
    • Every item extension-runtime §10 lists for monitor_run has a field.
  • The EN and ZH design docs have identical code spans, identical numbers and 19/19 headings.
  • CI on this head:
    • The TS suites ran only in Test (ubuntu, Node 22).
    • The Java contract test ran only in the two ubuntu database jobs, "Runtime Broker and Managed Agent MariaDB / Java 21" and "Hosted no-tool processes / MySQL 8.4 / Java 21".
    • The plain Java 11/17/21, macOS and Windows jobs don't run it, and the Node macOS/Windows jobs were skipped.
    • So this macOS run is the first for both languages on that platform.

Not covered

  • Windows was not run locally.
  • There is no end-to-end path to drive: nothing imports the module until H0c. "Real environment" here means the shipped bundle, the jar, the real Session authority and the real Broker ledger.
  • I did not weigh in on the six architectural choices the PR routes to the proposal owner.

Harness (generator, both drivers, mutants, probes) and raw results: assets-pr12837@f079755/pr12837. harness/README.md explains how to rerun each piece.

中文版

macOS 本地验证:用构建产物验证,head 18d8313904

结论:没有发现阻止合并的问题。 实际发布的内容只多了一个 token,即 domain 索引里的 "monitor_run",。在 1,201,560 个用例上,TypeScript 模块与 Java 校验器每次结论都相同,唯一的例外是设计文档已经写明的 Unicode 版本差异,共 371 例。schema 从未比模块更严格。建议合并前改掉 F1 的两句文档;F2 的 9 个 fixture 边界可以在本 PR 补,也可以在 H0c 之前补;Q1 是留给 H0c 的问题。

环境:macOS 26 arm64、Node 24.18.1、pnpm 11.24.0、JDK 21.0.12、Maven 3.9.16。base 88d881491b 和 PR 18d8313904 两棵树都用 npm run build && npm run bundle 和 mvn package 构建。A/B 对比、各个探针和差分都跑在这些构建产物上(packages/core/dist、jar 或 target/classes),不是在 vitest 下跑源码。

发现

# 类型 内容
F1 文档,建议合并前修 有两句话在中英文里都不准确。第 3 轮审计和 triage bot 都做过静态指认,我实际跑了这两处。(a) Broker 真实的 InMemoryToolExecutionRepository 把被取消的 PREPARED 调用结算为 SETTLED、executionStatus=cancelled,而不是 not_started;JdbcToolExecutionRepository:241 写入的值相同。(b) base 上真实的 LocalManagedSessionAuthority 会拒绝 monitor_run 的 domain.committed 事件("payload.domain must be one of …"),本 PR 能解析它。因此验收标准"现有记录、authority 和 store 行为不变"要加一个例外:"……但 reader 现在能在 domain.committed 中解析 monitor_run,提交仍被拒绝。"
F2 测试覆盖,不阻塞 针对第 3 轮延后的缺口,我写了 9 个定向变异体:TypeScript 5 个(T1–T5),Java 4 个(J1–J4)。9 个全部通过 PR 自带的测试(530/530 和 7/7)。9 个都能被我的差分或一个定向用例对杀死,说明两种语言现在的行为是对的,只是共享 fixtures 没有固定这些边界。随机变异没有触达 T3,harness 里的 t3.jsonl 就是现成的 T3 用例。
Q1 留给 H0c 的问题 契约没有规定一条记录的首个修订。解析函数判断单条记录,后继函数判断一对记录,所以修订 1 可以是任何合法状态,例如 execution 一出生就是 running_attached 的 run(规范 monitorRun fixture 就是这样),或者一出生就以 max_events settled 的 Monitor。"execution 首次出现必须是 intent"因此只约束修订,不约束创建。H0c 是否要加一个首修订判定(例如 state 为 reserved/admitted,execution 为 null/intent,delivery 为 null/planned),还是有意允许中途创建记录,比如接管 Legacy monitor?两种答案都可以,但最好在 H0c 开始提交记录之前,在文档里用一句话写明。
N1 信息 文档里提到的 NFC 差异,具体就是 U+105D2 U+0307:Unicode 16 让这对字符组合成 U+105C9。JDK 21(Unicode 15)认为它是 NFC 并接受,Node 24.18.1(Unicode 17)拒绝。在 Node 22 内部,TS 的结论也会变:22.3.0(Unicode 15.1)与 JDK 21 一样接受,22.23.2(Unicode 17)拒绝。所以同为 TS,写入方和读取方跑在不同的 Node 22 小版本上也可能结论不同。这个差异来自 Managed Session 稳定 ID 规则,不是本 PR 引入的;"标识符须限于 Unicode 15 已分配字符"这条约束,两个校验器都没有强制。文档可以写明这个具体字符对,以及结论随 Node 小版本变化这一点。
N2 小建议 TS 的租约上下界写成了字面量,但 core 已经在 http-managed-session-store.ts 导出了同样的上下界 HTTP_MANAGED_SESSION_STORE_CONTRACT.minimumLeaseDurationMs/maximumLeaseDurationMs(1000/300000),它对应 Java 的 ManagedSessionStoreModels。所以 triage bot"没有可复用的共享 TS 常量"的说法不准确。可选。

1. 实际发布的内容:只多一个 token(图 01)

  • CLI bundle: 我对比了两边的 dist/(各 1,319 个文件),先把 esbuild 的内容哈希文件名归一化。两边只有两个文件不同:一个 chunk 多了 "monitor_run",,另一个是 git commit 戳。managed-extension-record 不在 bundle 里。
  • jar: base 的全部 361 个条目在 PR 的 jar 里 CRC-32 都相同,只新增了 ManagedExtensionRecords.class 和它的 InvalidRecordException。src/main 里没有代码引用它们。
  • 真实 authority(core dist,两棵树各跑一次): 提交 session_metadata(对照)成功。提交 monitor_run 被拒("registered but not enabled"),journal 没有增长,base 与 PR 完全一致。唯一的差别在读取端(F1b):H0b 之前的 reader 会拒绝含这类记录的 Session,所以 H3 必须提高 minimumReader,和设计文档说的一样。

2. PR 自带测试在 macOS 上的结果,以及独立的 TS ↔ Java 差分(图 02)

  • PR 的平台表把 macOS 标为未运行。本次在 macOS 上:530 + 80 项测试、src/managed-runtime 全部 23 个文件(1,443 项)通过,tsc、ESLint 和 Prettier 无问题。mvn verify 通过 111 项测试,其中 7 项是新增的,Checkstyle 0 违规。这些数字与 PR 描述一致。
  • 差分: 3 个种子,每个 400,520 个用例:520 个 fixture,加上结构化随机生成的 grant、pin、run、monitor 和邻近修订对。1.0、-0、1e400 这类数字写法,以及恶意 id(孤立代理项、C0/DEL/C1 字符、512/513 字节、NFD),都以原始 JSON 文本交给各自的解析器。
    • 两个驱动都复现了全部 520 个 fixture 标注。
    • 两种语言都没有抛出契约拒绝之外的异常,Jackson 也读得了每一行。
    • 仅有的分歧就是 N1 的 371 例。
  • schema 对比模块(Ajv 2020,strict): 共 601,191 条单记录。schema 从未比模块更严;比模块宽松 5,998 次,全部落在文档列出的 JSON Schema 无法表达的规则里。
  • 语料中共有 139,982 条合法记录:其中每条 run、monitor、pin 都被接受为自身的后继,每条 grant 都被拒绝为自身的后继("完全相同的 grant 不替换任何东西")。只在写法上不同的修订(键顺序、1000.0、1e3、-0、重复键),两种语言结论一致,16/16。

3. 第 3 轮 fixture 缺口的实测(图 03,对应 F2)

4. 一个 Monitor 的完整生命周期,含 Runtime 丢失后的重建,两种语言都跑(图 04)

两种语言都接受全部 11 个合法修订:准入、intent、派发、attach、观测、Runtime 丢失、在 generation 2 下带新回执重建、max_events、结束后通知水位追平。两种语言也都拒绝全部 9 种捷径,包括重建时沿用 generation 1 的回执、watch 丢失期间还在观测,以及丢失的 watch 换了新 binding 后直接"结算"。

5. 文档事实、固定引用的参考设计与 CI 覆盖(图 05)

  • 在 doudouOUC/code_agent@6891216 上:
    • 存储设计 §3.1 的 33 个名称,与 MANAGED_SESSION_DOMAINS 名称一致、顺序一致,monitor_run 在 memory_job 之后。
    • 控制协议里 OperationGrant 的字段表,正好是 PR 的 9 个封闭键。
    • 扩展运行时 §10 为 monitor_run 列出的每一项都有对应字段。
  • 中英文设计文档的代码片段一致、数字一致,19/19 个标题对应。
  • 本 head 的 CI:
    • TS 测试只在 Test (ubuntu, Node 22) 中运行。
    • Java 契约测试只在两个 ubuntu 数据库任务中运行:"Runtime Broker and Managed Agent MariaDB / Java 21" 和 "Hosted no-tool processes / MySQL 8.4 / Java 21"。
    • 普通的 Java 11/17/21 任务以及 macOS、Windows 任务都不跑它,Node 的 macOS/Windows 任务被跳过。
    • 因此本次 macOS 运行是两种语言在该平台上的首次运行。

未覆盖

  • 没有在本地跑 Windows。
  • 没有端到端路径可驱动:在 H0c 之前没有代码导入这个模块。这里的"真实环境"指发布的 bundle、jar、真实的 Session authority 和真实的 Broker 账本。
  • 没有评判 PR 交给提案负责人的六个架构选择。

harness(生成器、两个驱动、变异体、探针)与原始结果见 assets-pr12837@f079755/pr12837,harness/README.md 说明了各部分如何重跑。

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head 18d83139

A cross-language contract module: managed-extension-record.ts (1060 lines) and its test (457), ManagedExtensionRecords.java (708) and its contract test (287), a 1818-line JSON schema, a 19,457-line shared fixture corpus, two design docs, and one token of live change — 'monitor_run' added to MANAGED_SESSION_DOMAINS, with the test that pins the list's length bumped 32 → 33. 10 files, +24348/-2.

Because the ratio of new code to shipped behaviour is so extreme, I spent the budget on the one token that executes and on whether the rest is genuinely inert.

No historical blocker

There are no reviews and no inline comments on this PR, so nothing has ever been filed against it and there is no thread to re-verify. Triage stage 2 at this head concludes "No Critical blockers". The author's own pre-submission audit reports one Critical in each of rounds 1 and 2 — a Monitor that failed after start having no recordable state, and Java throwing NumberFormatException on a number past double range so a successor check threw instead of returning false — both fixed before the PR was opened, with round 3 finding none. Neither stands at this head: watch_failed is in the stop-reason table, and non-finite numbers are refused with four fixtures carrying the raw 1e400.

The one shipped change is a reader-side widening, and it is the safe direction

MANAGED_SESSION_DOMAINS has exactly two references in the repository — its own module and its test — and inside the module it is consulted at one place, the case 'domain.committed': branch that validates a record's domain against the list. So adding monitor_run changes what a reader accepts; it does not add a writer, and no writer consults this list. Nothing in the tree produces a monitor_run record today, so the effect is that a journal carrying one parses instead of failing as an unknown domain, while every other unknown domain still fails closed. The author states the same split explicitly — a reader can now parse monitor_run in domain.committed while committing one is still refused — and amended the acceptance criterion to say so rather than claiming "existing record behaviour is unchanged". The writer-side obligation that this creates (a writer must gate on the reader's declared minimumReader version before emitting the new domain) belongs to H0c, is named in the design note, and cannot bite at this head because nothing emits it.

The rest is inert, verified rather than assumed

A repository-wide search for managed-extension-record returns zero importers: the module, its test, the schema and the fixtures are not reachable from any production path. That is what makes a 24k-line addition reviewable at all — nothing executes it, so a defect here is a defect in a specification and its reference implementations, discoverable by the contract test before any consumer exists.

Within the module I checked the defensive posture, since this code's whole job is to reject malformed untrusted structure: exported constants are Object.freezed (including the per-state transition arrays built by a loop that freezes each one), the object check rejects anything whose prototype is neither Object.prototype nor null with the comment noting an array fails there too, MAX_GENERATION is 2n ** 63n - 1n compared as a BigInt, limits are sourced from the existing MANAGED_SESSION_LIMITS rather than reinvented, depth_limit is a named rejection reason, phase names are pattern-constrained, and failures raise the module's typed ManagedSessionRecordError. That is the same closed-world, fail-closed shape the sibling managed-runtime modules use.

Cross-language agreement is pinned by a shared corpus, not by two sets of good intentions

The TS test and the Java contract test consume the same 19,457-line fixture file, which is the mechanism that keeps two implementations from drifting. Triage compared the constants one by one and reported them equal — 33 domains with monitor_run after memory_job, three state lines and every transition list, two delivery targets, five recovery and six quota reasons, seven monitor stop reasons, and a 2^53 − 2 count bound enforced by Number.isSafeInteger on one side and its Java equivalent on the other. The author's differential run at this head reports 1,201,560 cases with identical verdicts in both languages, and that spelling-level variants (key order, 1000.0, 1e3, -0, duplicate keys) agree 16/16.

The one divergence is documented, measured and fails closed. 371 of those cases disagree, all from a single Unicode-version gap: U+105D2 U+0307 composes to U+105C9 under Unicode 16, so JDK 21 (Unicode 15) accepts the pair as NFC while Node 24 — and Node 22.23, unlike 22.3 — refuses it. The consequence is rejection of an ID a peer would have accepted, not silent acceptance of one it would not, and the gap originates in the Managed Session stable-ID rule rather than in this contract. It is recorded in the design doc with the concrete pair and the version boundaries, which is the right handling for a divergence a contract module cannot fix on its own. I am naming it here so it is not mistaken for an unmeasured unknown.

CI

20 checks pass at this head and none has failed, including the Java lanes that run ManagedExtensionRecordContractTest (ubuntu-latest / Java 11, 17, 21, windows-latest, macos-latest, Runtime Broker and Managed Agent MariaDB, Hosted no-tool processes / MySQL 8.4, Real daemon E2E / Java 11) and the Node lanes that run the TypeScript module's test and typecheck it (Test (ubuntu-latest, Node 22.x), Lint & Static), plus Integration Tests (no-AK, No Sandbox), web-shell E2E Smoke, both Desktop Shell lanes, triage and the routing checks. review-pr is pending and is not a gate; a sandboxed verification had just been triggered and had not reported.

Scope

I read the live change and traced its only consumer, confirmed the module has no importers, and read the new module's exported surface, limits and rejection posture. I did not read the 1060-line TypeScript module or the 708-line Java class line by line, did not read the 1818-line schema, and sampled nothing of the 19,457-line corpus — for those I relied on the shared-corpus mechanism, the differential measurement and triage's constant-by-constant comparison, and I am naming that reliance rather than implying I covered it. The two doc sentences the author's verification flags as inaccurate (F1) and the nine fixture edges (F2) are documentation and coverage items, outside Critical-only scope.

Verdict: APPROVE — No Critical found and none was ever filed. Everything added except one array entry is unreachable from production code, that entry widens only what a reader parses while unknown domains still fail closed and no writer exists to emit one, both reference implementations are pinned to the same fixture corpus with a measured 1.2 M-case differential whose sole divergence is a documented Unicode-version gap that rejects rather than admits, and the two Criticals the author's own audit found were fixed before this head.

@wenshao
wenshao added this pull request to the merge queue Sep 27, 2026
Merged via the queue into main with commit e68822c Sep 27, 2026
117 of 118 checks passed
@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

The round-3 deferred items from the audit record above are tracked in #12846.

中文说明

上方审计记录中第 3 轮延后的事项由 #12846 跟进。

pull Bot pushed a commit to TKaxv-7S/qwen-code that referenced this pull request Sep 27, 2026
…nLM#12830)

* feat(managed-agent): Add the planned Stage H task contract (H0a)

Freeze the public task contract that the extension runtime design
requires before H0 is implemented. Every addition is marked planned, so
no route is mapped and the generated WebShell types do not change.

The public API gains a task list, task detail, task events behind an
opaque output cursor, and a cancel command that takes an
Idempotency-Key and answers 202 with a command operation. The WebShell
adapter mirrors them. The task view is SessionTaskView in the public
conventions; schema conditionals keep settled_at, started_at and the
advertised actions consistent with the task state, and closed objects
keep Runtime bindings, generations, PIDs and paths out of it. Cancel
reuses the command operation with a new task_cancel type, and a
completed cancel records acceptance, not a stop.

The MCP catalog, hook catalog, channel and automation resources are
named as planned routes without bodies, for H1, H2, H5 and H6 to fill
in. The contract test now also watches the channel and automation
prefixes, which a mapped planned route would otherwise pass unnoticed.

Part of QwenLM#12827.

* fix(managed-agent): Close review findings on the planned task contract

A cancel retry with the same Idempotency-Key now replays the original
operation before any state or capability check, so a lost 202 never
turns into a 409 once the task settles.

Task events carry their own cursor, so a consumer can checkpoint per
event instead of re-applying output after a crash, and a bad cursor is
invalid_event_cursor as on the Session event history. The task list is
newest first with an id tiebreaker, as the Session list is. The task
resource and its lists carry an object discriminator, the event list is
named PublicTaskEventList, and the event type is a shared enum.

recovery_blocked never offers send_input, completed needs started_at,
more tasks need a cursor and output chunks are never empty. A new test
checks these conditionals with valid and invalid instances, because the
API contract test validates only operations that are not planned.

The design notes record the departures from the Session event history,
the enum value that cannot be marked planned, the properties H0c must
also flip, and the Artifact attribution and Legacy state follow-ups.

* fix(managed-agent): Align task events and cancel replay with the API contract

Task events carry schema_version and projection_version, as the API
contract requires of public events, and their type is an open string, so
later types stay additive. A page's next_cursor is the position after
the last event the server examined, so events filtered out for a caller
cannot stall it, and the start of the retained events when after was
omitted. An expired cursor loses no output; how Artifacts and retained
events join without overlap is left to the H3 output segmentation.

Cancel replays the original operation for the same actor after the
caller's access is re-checked, as the contract's idempotency rules say,
and declares invalid_request and task_forbidden next to the 409s.

The route check reads every /v1/agent route, so later /v1/agent-*
resources are covered without another prefix. The planned-task test
runs every instance against the WebShell mirror as well, and checks the
WebShell cancel and event query requests.

* fix(managed-agent): Fix the task check order and cover the missed invariants

Cancel checks run in a fixed order: access (404, then 403), idempotent
replay, then task state. A missing Idempotency-Key is invalid_request
and a malformed one invalid_idempotency_key, as on the other idempotent
routes, and a Session that does not serve tasks answers
unsupported_feature. action_capabilities describes the task, not the
caller. The read routes drop 403, since a caller that cannot read gets
404, and a list cursor for more tasks is never empty.

The planned-task test now covers waiting, degraded and a failed task
that had started, and rejects state_changed without state and artifact
without artifact_id. The design note states the per-surface mutation
counts, the event type rule for later minor versions, and the reuse of
the command operation as the tradeoff behind the visible task_cancel.

* fix(managed-agent): Read retained task events before Artifacts on recovery

The recovery recipe after an expired cursor read the task Artifacts
first and the retained events second, which can lose an event that is
archived and expired between the two reads. Reading the retained events
first and the Artifacts second loses nothing, because every event that
expired before the event read began was archived before it. The
guarantee also needs artifact_refs to list every Artifact of the task.

The design note also stops attributing the ignore-unknown-types rule to
section 5 of the API contract, which only allows new event types, and
names the one public-only instance the WebShell mirror check skips.

* fix(managed-agent): Define the task event page cursor by what the page covers

next_cursor was the position after the last event the server examined.
A server that reads one row past the limit to compute has_more examines
an event it does not return, and a caller passing that cursor back would
skip it with no error. The cursor is now the position after the last
event the page covers: every earlier event was returned or filtered out
for this page, and an event held back by limit is never passed.

The replay guarantee now holds while the idempotency record is retained
rather than never failing, and the Artifact routing of high-volume
output cites design section 11 and API contract section 6, which is
where the rule comes from.

* fix(managed-agent): Expire task events only from the oldest end

The events route promised that a page never passes an event it did not
return or filter out, and that an expired cursor answers 409, but did
not say in which order events expire. If a later event could expire
before an earlier output event that is not yet in an Artifact, a caller
following the stream would skip it silently. Events now expire only
from the oldest end, so the retained events have no gaps and a cursor
older than the oldest retained event is the only way to miss one.

The spec's cancel description also carries the retention condition the
design note already states: a same-key retry replays the original
operation while its idempotency record is retained.

* fix(managed-agent): Fix read_output for a task's life and drop event filtering

A task whose read_output capability was withdrawn and restored would
have its output events filtered out while the cursor moved past them,
and a task without the capability kept output events that neither the
events route nor the Artifacts returned. read_output now does not change
during a task's life, and a task without it produces no output events,
so the events route never filters out an event that exists and the
no-loss guarantees hold for every task.

With nothing filtered, next_cursor is simply the cursor of the last
returned event. The capability text also says that only cancel needs an
authorization check, since reading output needs only read access, and
the design note points H0b at QwenLM#12837, which answers issue question 1.
pull Bot pushed a commit to bit-cook/qwen-code that referenced this pull request Sep 27, 2026
…wenLM#12863)

Close the items QwenLM#12837 deferred for the managed-extension-record/1
contract (H0b of QwenLM#12827), as QwenLM#12846 lists them. No production code
changes.

Correct two statements that H0c would build on. The Broker settles a
cancelled call that was never sent as cancelled, not not_started, and
no field of its record tells that apart from a call cancelled after it
was sent, so H0c must not map a call it cannot prove unsent to
not_started_proven. Readers now parse monitor_run in domain.committed,
and only domain-record submission checks that a domain is enabled, so
H0c must enforce that check on every path that commits domain records.

Add 34 fixture cases, and give one existing case a valid next record,
so that each targeted check has a case of its own: the six gaps the
issue lists and those four audit rounds found. Each targeted
clause-level mutant now fails in TypeScript and Java, and so does each
per-field mutant but three equivalent ones in TypeScript.

Record the decisions for H0c and H3: the v1 notification policy is
fixed; revocation is gate state outside the grant replacement rule;
only a purely observational Monitor is rebuilt, and only from
outcome_unknown. Add two open questions.

Closes QwenLM#12846

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Downgraded from Request changes to Comment: self-PR; CI failing: review-pr. Partially reviewed — gaps disclosed. Suggestions are inline. 27 Suggestion-level finding(s) could not be anchored to a changed line and were dropped; nothing further to act on here.

Not explored to full depth (tool budget reached): "agent reverse-audit (round 4)": the Java-side half of the mutation (compiling and replaying a widened case "running", "waiting" arm of ManagedExtensionRecords.requireRun ) was not executed …; "agent reverse-audit (round 2)": Rebuild 段落(:228)声称 rebuild 的 intent 与 dispatch "committed in that phase's physical ledger, keyed by Session, effect, phase and effect revision" —— 核对它需要打开 Broke…; "agent reverse-audit (round 2)": 第 256 行"A Python implementation written from this document … labeled every case, and the generator stops when it disagrees with a label" —— 该 oracle 明确"kept out…; "agent reverse-audit (round 3)": zh-CN prose parity beyond identifiers and numbers — a Chinese sentence in doc 171–279 that drops or adds a rule stated purely in prose (no backticks, no digits)…; "agent reverse-audit (round 3)": doc 189's "section 12 of the reference design" quota mapping and doc 205's "section 10 of the reference design" monitor-key coverage — the reference design is a…, and 3 more.

Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.

Test Plan (not a blocker): src/managed-runtime/managed-extension-record.test.ts — no such file or directory; src/managed-runtime/managed-session-records.test.ts — no such file or directory.

中文说明

⚠️ 已从请求修改降级为评论:self-PR; CI failing: review-pr。 仅完成部分审查,审查缺口已披露。 建议见行内评论。 27 条建议级发现无法锚定到改动行,已丢弃;此处无需进一步处理。

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 4)":the Java-side half of the mutation (compiling and replaying a widened case "running", "waiting" arm of ManagedExtensionRecords.requireRun ) was not executed …;"agent reverse-audit (round 2)":Rebuild 段落(:228)声称 rebuild 的 intent 与 dispatch "committed in that phase's physical ledger, keyed by Session, effect, phase and effect revision" —— 核对它需要打开 Broke…;"agent reverse-audit (round 2)":第 256 行"A Python implementation written from this document … labeled every case, and the generator stops when it disagrees with a label" —— 该 oracle 明确"kept out…;"agent reverse-audit (round 3)":zh-CN prose parity beyond identifiers and numbers — a Chinese sentence in doc 171–279 that drops or adds a rule stated purely in prose (no backticks, no digits)…;"agent reverse-audit (round 3)":doc 189's "section 12 of the reference design" quota mapping and doc 205's "section 10 of the reference design" monitor-key coverage — the reference design is a…,另有 3 条。

未审查:反向审计——在 5 轮的反审轮数上限内未收敛。

Test Plan(非阻断):src/managed-runtime/managed-extension-record.test.ts — no such file or directory; src/managed-runtime/managed-session-records.test.ts — no such file or directory。

— qwen3.8-max via Qwen Code /review (v0.24.7)

- In both languages, every malformed shape and every illegal transition in the fixtures is refused, and every valid case is accepted.
- The schema and the module disagree only on the listed rules that JSON Schema cannot state.
- `monitor_run` is in the v1 domain index and is still refused for submission.
- Every existing record, authority and store behavior is unchanged.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Critical] R1-1: [certifies-falsely] [new-surface] Both design-doc corrections the author committed to before merge are still false at the merged head, in both languages — and a human approval certified one of them as amended.

Two factual claims in this normative document are wrong at head 18d8313904. (a) This acceptance criterion — "Every existing record, authority and store behavior is unchanged." (zh-CN:264 「现有的记录、authority 与存储行为都不变。」) — is contradicted by the PR's own Risk section: domain.committed now parses monitor_run where the base refused it, so one reader path did widen. (b) Line 20 — "a cancelled PREPARED call settles as not_started without a dispatch" (zh-CN:20 same claim) — is contradicted by the Broker it describes: RuntimeBrokerService.java:1876-1878 states "A not_started status is accepted only as the Runtime's own terminal answer; the Broker never derives it", and the cancel path at :1641-1643 settles with executionStatus=cancelled. The triage thread confirmed both sentences at this head and recorded the author's commitment to fix them before merging; the approving review then asserted the criterion had been "amended … rather than claiming 'existing record behaviour is unchanged'" — git show HEAD: of both files shows it was not. Line 264 is the sentence an H0c implementer trusts to conclude no reader-compatibility work is needed (see also the minimumReader comment on Decision 1), and line 20 is the doc's stated reason the execution line is "not step for step" with the Broker ledger, so H0c would derive the Broker↔contract mapping from a false premise. The document is this contract-first slice's core deliverable, which is why this is graded as a blocker rather than a doc nit.

Witness:

git show HEAD:docs/design/...-contract.md (and the zh-CN twin) at 18d8313904:
  line 264: "- Every existing record, authority and store behavior is unchanged."  (present, both languages)
  line 20:  "a cancelled `PREPARED` call settles as `not_started` without a dispatch" (present, both languages)
BASE/PR probe on the domain.committed reader (base arm = the PR's single production line reverted):
  BASE: domain=monitor_run => REFUSED (must be one of: config_install, ...)
  PR:   domain=monitor_run => PARSED kind=domain.committed
  (control: session_metadata PARSED on both arms)

Suggested fix: in both languages, replace the line-264 criterion with the actual behaviour (domain.committed parses monitor_run; submission still refuses it), and correct line 20 to say a cancelled record settles as SETTLED with executionStatus=cancelled, that the Broker never derives not_started, and that EXECUTING is entered before the Runtime receives the call so it corresponds to dispatch_started.

中文说明

作者在合并前承诺修正的两处设计文档陈述,在已合并的 head 上仍然是错的,而且中英两版都错——其中一处还被人工 approve 认证为"已修正"。

(a) 第 264 行验收标准「Every existing record, authority and store behavior is unchanged.」/「现有的记录、authority 与存储行为都不变。」——PR 自己的 Risk 段就说明本 PR 改变了 reader 行为:domain.committed 现在能解析 monitor_run(base 上会拒绝)。(b) 第 20 行「a cancelled PREPARED call settles as not_started without a dispatch」/「被取消的 PREPARED 调用不经派发即以 not_started settled」——与它所描述的 Broker 相反:RuntimeBrokerService.java:1876-1878 写明 not_started 只作为 Runtime 自己的终态回答被接受、Broker 从不推导它;取消路径 :1641-1643 以 executionStatus=cancelled 结算。triage 线程在本 head 上确认了这两句,并记录了作者"合并前会修"的承诺;随后的 approve review 声称该标准"已改写"——但 git show HEAD: 显示两句在两种语言里都原封未动。

第 264 行是 H0c 实现者据以得出"无需 reader 兼容工作"结论的那句话;第 20 行是文档解释执行线为何与 Broker 台账"并非逐步对应"的依据,H0c 会从一个假前提出发推导 Broker↔契约映射。文档正是这个契约优先切片的核心交付物,因此按阻断级而非文档小疵定级。

建议修复:两种语言同步,把第 264 行改为实际行为(domain.committed 可解析 monitor_run,提交仍被拒绝),并把第 20 行改为:被取消的记录以 SETTLED+executionStatus=cancelled 结算,Broker 从不推导 not_started,EXECUTING 在 Runtime 收到调用之前进入、因此对应 dispatch_started。

— qwen3.8-max via Qwen Code /review (v0.24.7)


## Decisions

1. **`monitor_run` joins the closed v1 index** (question 1 of #12827). The storage design at the pinned commit lists it among the 33 v1 names and leaves only its validator to H0. No v1 reader can meet a `monitor_run` record yet, because the domain stays disabled, so the index keeps its version. A reader from before this change refuses the name, so when H3 enables the domain, a Session that holds such records must keep those readers out, for example by raising its `minimumReader`. The name goes after `memory_job`, in the design's order.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-2: Decision 1's only named mitigation — "for example by raising its minimumReader" — cannot be applied to the Sessions the sentence is about.

A Session's header is written once, at creation, from the module constant MANAGED_SESSION_MINIMUM_READER (managed-session-authority.ts:427-438, inside the if (header === undefined) branch); the create option's own docstring says "Supplied when creating a session; ignored once a header exists" (:271), and no header-rewrite or migration path exists anywhere under managed-runtime (grep: 0 hits). H3 enables the domain through a global switch (MANAGED_SESSION_ENABLED_DOMAINS), so every already-created Session becomes eligible to hold managed-monitor_run records while its header still says managed-session/1 — and a pre-H3 reader opens such a Session (the reader refuses only when the required version exceeds its own, managed-session-records.ts:1165-1169) and meets a domain name absent from its 32-name closed index: exactly the encounter this sentence says must be kept out. H3 must therefore either build a header-migration path — reader-compatibility work the acceptance criterion at line 264 says is not needed (R1-1) — or leave every pre-existing Session unable to use Monitor; an implementer who trusts this sentence plans for neither.

Witness:

production write sites of minimumReader in packages/core/src: exactly 1 (managed-session-authority.ts:438,
  inside the create branch at :427); header construction sites: exactly 1 (:471)
rewrite|upgrade|migrat under managed-runtime: 0 hits
reader refusal: only when requiredReaderVersion > currentReaderVersion (managed-session-records.ts:1165-1169)

Suggested fix: replace the clause with the mechanism that exists — e.g. "so when H3 enables the domain it must raise MANAGED_SESSION_MINIMUM_READER, which keeps pre-H3 readers out of every Session created afterwards; a Session created before the bump keeps the header it was written with, so H3 must also migrate that header or refuse to enable monitor_run for it" — and record the header migration as H3 work under Follow-up. The zh-CN twin at line 43 must move with it.

Fix constraint: export const MANAGED_SESSION_MINIMUM_READER = 'managed-session/1'; (managed-session-records.ts:22) is a module constant and the header is created only under if (header === undefined) (managed-session-authority.ts:427-438) — the corrected wording must not imply a per-Session raise exists.

中文说明

Decision 1 唯一给出的缓解手段——「例如提高其 minimumReader」——对这句话所要保护的那批 Session 根本不可用。

Session 的 header 只在创建时写入一次,取自模块常量 MANAGED_SESSION_MINIMUM_READER(managed-session-authority.ts:427-438 的 if (header === undefined) 分支);create 选项的文档字符串自己写着「Supplied when creating a session; ignored once a header exists」(:271),而 managed-runtime 下不存在任何 header 改写/迁移路径(grep 0 命中)。H3 通过全局开关(MANAGED_SESSION_ENABLED_DOMAINS)开放该 domain,于是每个已创建的 Session 都有资格持有 managed-monitor_run 记录,而它们的 header 仍是 managed-session/1——H3 之前的 reader 仍能打开这种 Session(读方只在所需版本超过自身时拒绝,managed-session-records.ts:1165-1169),并遇到一个不在其 32 名封闭索引里的 domain 名:这正是本句声称必须挡住的相遇。因此 H3 要么建一条 header 迁移路径(而第 264 行验收标准说无需 reader 兼容工作,见 R1-1),要么让所有既有 Session 无法使用 Monitor;相信这句话的实现者两者都不会规划。

建议修复:把该小句改成真实存在的机制——「H3 开放该 domain 时必须提高 MANAGED_SESSION_MINIMUM_READER,这只能把 H3 之前的 reader 挡在此后创建的 Session 之外;在提高之前创建的 Session 保留写入时的 header,因此 H3 还必须迁移这些 header,或拒绝为其开放 monitor_run」——并把 header 迁移记入 Follow-up 的 H3 工作。中文版第 43 行需在同一改动内同步。

修复约束:MANAGED_SESSION_MINIMUM_READER = 'managed-session/1'(managed-session-records.ts:22)是模块常量,header 只在 if (header === undefined) 下创建(managed-session-authority.ts:427-438)——修正后的措辞不得暗示存在逐 Session 的提高机制。

— qwen3.8-max via Qwen Code /review (v0.24.7)

- **Scope.** `recordRef` is the committed domain record that holds the operation's plan. Its kind must be `managed-<domain>` and its schema version 1, the pairing `domain.committed` already requires, so a grant cannot point at another domain's plan. `phases` lists the 1 to 16 distinct phases of that plan the grant admits. Which phases a domain has is registered with its body in H1–H6; H0b fixes their form.
- **No activation fields.** The record is closed, so an `epoch` or `activationId` is refused: a grant is not an activation.
- **Replacement.** A Runtime's per-operation gate may replace grant `a` with grant `b` only when both are valid, name the same Session, operation and domain, and either
- `b` renews `a`: the same revision and every field equal, phases in the same order, except a later `expiresAt`; or

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-3: The renewal rule makes phases array order load-bearing, while the scope rule that defines the same field (line 81) fixes only distinctness — so a semantically identical renewal can be refused as stale.

isOperationGrantSuccessor's same-revision branch implements line 84 literally: after.expiresAt > before.expiresAt && sameJson({...before, expiresAt: 0}, {...after, expiresAt: 0}) (managed-extension-record.ts:605-609), and sameJson is JSON.stringify text equality (:462-464) — a reordered resourceScope.phases fails it. The later-revision branch cannot rescue a renewal (it requires a strictly greater operationRevision, :612). Probed in both languages: a renewal presenting the same two phases in the other order with a later expiresAt is refused, while the same-order control is accepted. The corpus never exercises a pure reorder — the only renewal-scope case (renewal-changes-phases, fixtures:3848) drops a phase. The trigger is prospective (nothing commits grants in this slice), and the order requirement IS documented here — so this is spec coherence: line 81 and line 84 currently promise two different things about the same field, and the Java control plane holds phases as a Set<String> that Jackson serializes in hash order.

Witness:

TS:  parseOperationGrant(reordered phases) => ACCEPTED
     isOperationGrantSuccessor(before, reordered + later expiresAt) => false
     CONTROL (same order, later expiresAt) => true
Java (module's own compiled classes): same three verdicts

Suggested fix: choose one semantics and state it in both languages at lines 81/84 (and the zh-CN twins, which carry the identical rule 「phase 顺序也相同」): either make the order canonical and required ("in the plan's registered order") and enforce it in parseOperationGrant so a reordered list is not a well-formed grant at all; or keep the scope rule order-free and judge the renewal on the semantics it defines — compare phases as a set (equal length plus equal members) instead of by JSON.stringify of the whole grant.

Fix constraint: the parser's only phases property beyond PHASE_PATTERN and maxGrantPhases is the distinctness check (managed-extension-record.ts:546-548), matching doc line 81 — a canonical-order fix must land in the parser and the doc together, or the successor check will keep refusing grants the parser accepts.

Fix witness: a successor fixture pair with the same operationRevision, a later expiresAt, and ["query_receipt","send_segment"] against ["send_segment","query_receipt"], whose valid flag pins the chosen semantics in both replays — today no case presents a pure reorder, so adding or removing the order requirement moves no test in either language; with the new pair, removing the chosen rule reddens both.

中文说明

续租规则让 phases 数组的顺序成为判定要素,而定义同一字段的范围规则(第 81 行)只要求互不相同——于是一次语义完全相同的续租可能被当作陈旧而拒绝。

isOperationGrantSuccessor 的同修订分支逐字实现了第 84 行:after.expiresAt > before.expiresAt && sameJson({...before, expiresAt: 0}, {...after, expiresAt: 0})(managed-extension-record.ts:605-609),而 sameJson 是 JSON.stringify 文本相等(:462-464)——resourceScope.phases 换了顺序就过不了。更晚修订分支救不了续租(它要求 operationRevision 严格更大,:612)。两种语言实测:同样两个 phase 换个顺序、带更晚 expiresAt 的续租被拒,同序对照被接受。语料中没有任何纯换序用例——唯一的续租范围用例 renewal-changes-phases(fixtures:3848)是删掉一个 phase。触发是前瞻性的(本切片没有组件提交 grant),且顺序要求确实写在本行——所以这是规范自洽问题:第 81 行与第 84 行对同一字段承诺了两种语义,而 Java 控制面把 phases 存成 Set<String>,Jackson 按 hash 序序列化。

建议修复:二选一,并在中英两版的 81/84 行同时写明(中文版同一规则「phase 顺序也相同」):要么把顺序定为规范且必需(「按计划注册的顺序」),并在 parseOperationGrant 中强制,使换序列表根本不是合法 grant;要么保持范围规则与顺序无关,续租按其定义的语义判定——把 phases 当集合比较(长度相等加成员相等),而不是对整个 grant 做 JSON.stringify。

修复约束:解析器对 phases 除 PHASE_PATTERN 与 maxGrantPhases 外唯一的约束是去重检查(managed-extension-record.ts:546-548),与文档第 81 行一致——规范化顺序的修复必须同时落在解析器与文档,否则后继检查会继续拒绝解析器接受的 grant。

修复见证:新增一个后继 fixture 对——同 operationRevision、更晚 expiresAt、["query_receipt","send_segment"] 对 ["send_segment","query_receipt"]——其 valid 标签在两侧回放中钉住所选语义;今天没有纯换序用例,增删顺序要求都不会惊动任何测试,新用例落地后删除所选规则会使两侧变红。

— qwen3.8-max via Qwen Code /review (v0.24.7)

| ---------------------------------------------- | ------------------------------------------------------------- |
| `recovery_blocked` | a recovery reason, required |
| `running`, `waiting` | null, or a recovery reason: the run goes on in a degraded way |
| `failed` | null, or a quota reason |

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-4: No run that has ended can state that a recovery cause ended it — the five recovery reasons are usable only while the run is open.

The reasons table allows a recovery reason only on recovery_blocked, running and waiting, and forces null on reserved/admitted/settled/cancelled (line 167); assertReason implements exactly that (managed-extension-record.ts:671-673), and the Java validator mirrors it. But both transition tables allow recovery_blocked → failed on the run line and outcome_unknown → not_started_proven on the execution line — so H4's child_run whose Runtime was lost, and whose work is later proven never to have run, steps to failed and must drop its runtime_lost reason on that commit: a producer that keeps it produces a record both validators refuse (probed in both languages). The SessionTaskView that Decision 2 says H0c rebuilds from these blocks then reports a failed child run with no cause, at the one state the user actually reads, while the contract holds an exact word for what happened. The corpus has no case pairing a terminal state with a recovery reason in either direction.

Witness:

TS:  parseExtensionRun(recovery_blocked / runtime_lost / outcome_unknown) => ACCEPTED
     parseExtensionRun(failed + runtime_lost + not_started_proven) => REFUSED
       (run.reason does not fit the failed state.)
     parseExtensionRun(failed + reason null) => ACCEPTED
     successor(blocked -> failed+runtime_lost) => false;  (blocked -> failed+null) => true
Java: same four verdicts; all five recovery reasons REFUSED on failed

Suggested fix: widen the failed row (line 166 and zh-CN 167) to "null, a recovery reason or a quota reason" and mirror it in assertReason's failed arm (reason === null || isQuotaReason(reason) || isRecoveryReason(reason)), keeping the per-reason constraints (:675-694) untouched so no reason can attach to a state that contradicts it. If failed is deliberately reserved for quota causes, say so in the doc, name which record carries a recovery-caused ending, and state that a projection must read the cause from the previous revision.

Fix constraint: managed-extension-record.ts:678-680 (the execution_corrupt biconditional) and :778-784 (a terminal state requires an execution in PROVEN_EXECUTION_STATES = settled/not_started_proven) mean widening failed cannot admit failed+execution_corrupt or failed+outcome_unknown — only runtime_lost, dispatch_unknown and handler_unavailable become reachable there, which is the intended effect and must stay that way.

Fix witness: a run fixture with state: "failed", reason: "runtime_lost", a non-null runtime and execution: "not_started_proven", plus the successor pair recovery_blocked/runtime_lost → failed/runtime_lost — both are rejected in TS and Java today, so the case pins whichever verdict the author chooses and goes red if the failed arm is later changed without the doc.

中文说明

任何已结束的运行都无法陈述「是某个恢复原因使它结束」——五个恢复原因只能在运行仍打开时使用。

原因表只允许 recovery_blocked、running、waiting 携带恢复原因,并强制 reserved/admitted/settled/cancelled 为 null(第 167 行);assertReason 逐字实现(managed-extension-record.ts:671-673),Java 校验器镜像。但两张转移表都允许运行线 recovery_blocked → failed、执行线 outcome_unknown → not_started_proven——于是 H4 的 child_run:Runtime 丢失后被证明从未运行,走到 failed 时必须在那次提交里丢掉 runtime_lost;保留它的生产者会写出两种语言都拒绝的记录(两侧实测)。Decision 2 说 H0c 据这些 run block 重建 SessionTaskView,于是任务列表在用户唯一会读的那个状态上报告一个没有原因的失败运行,而契约里明明有描述此事的确切词汇。语料中没有任何用例在任一方向上把终态与恢复原因配对。

建议修复:把 failed 行(第 166 行与中文版 167 行)放宽为「null、一个恢复原因或一个配额原因」,并在 assertReason 的 failed 分支镜像(reason === null || isQuotaReason(reason) || isRecoveryReason(reason)),保持各原因自身的约束(:675-694)不动,使任何原因都不能出现在与其矛盾的状态上。若 failed 刻意只留给配额原因,请在文档中写明,指明哪条记录承载「因恢复原因而结束」,并说明投影必须从前一修订读取原因。

修复约束:managed-extension-record.ts:678-680(execution_corrupt 双条件)与 :778-784(终态要求执行处于 PROVEN_EXECUTION_STATES = settled/not_started_proven)决定了放宽 failed 不可能放进 failed+execution_corrupt 或 failed+outcome_unknown——只有 runtime_lost、dispatch_unknown、handler_unavailable 变得可达,这正是预期效果,且必须保持如此。

修复见证:新增一个 run fixture——state: "failed"、reason: "runtime_lost"、非空 runtime、execution: "not_started_proven"——外加后继对 recovery_blocked/runtime_lost → failed/runtime_lost;两者今天在 TS 与 Java 中都被拒绝,因此用例钉住作者选择的任一裁决,日后有人只改 failed 分支不改文档就会变红。

— qwen3.8-max via Qwen Code /review (v0.24.7)

| `failed` | null, or a quota reason |
| `reserved`, `admitted`, `settled`, `cancelled` | null |

- `outcome_unknown` requires an execution in `outcome_unknown`; `execution_corrupt` is the reason exactly when the execution is `corrupt`; `runtime_lost` requires the `runtime` it lost, and a run blocked on it is not `running_attached`; `dispatch_unknown` requires a `dispatchId`.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-5: runtime_lost's exclusion bars only running_attached — a run whose execution is already proven ended may stay recovery_blocked + runtime_lost forever, contradicting the reason's own definition.

The rule as implemented (TS :684-690, Java :350-354, schema :1231-1252, and this line in both languages) refuses only the attached shape. Executed against the reviewed commit on the corpus's own blocked-lost-runtime record with only execution changed: settled — parse ACCEPTED, successor true; not_started_proven — ACCEPTED, true; only running_attached is refused. recovery_blocked is not terminal, so the terminal-needs-proven-execution rule (:778-785) never fires on it. The chain is exactly the doc's Rebuild narrative minus the rebuild: Runtime lost → blocked, then the original system supplies its proof and the execution steps to settled/not_started_proven, and nothing requires the producer to move the run off recovery_blocked. The record then asserts both "the work is proven done (or proven never run)" and "blocked because the Runtime was lost and not re-attached" — contradicting the definition at line 181/zh-CN:184 ("A blocked run was not re-attached"). H0c's projection shows a task whose result already exists as permanently blocked; the not_started_proven variant is sharpest — proven never started, yet blocked on a lost Runtime with nothing to recover.

Witness:

base = corpus blocked-lost-runtime (valid), only `execution` changed:
  outcome_unknown     parse=ACCEPTED  successor=true
  settled             parse=ACCEPTED  successor=true   <-- contradiction ships
  not_started_proven  parse=ACCEPTED  successor=true   <-- contradiction ships
  running_attached    parse=REFUSED (run cannot be blocked on a lost Runtime while attached.)
fix arm (guard extended to PROVEN_EXECUTION_STATES): settled/not_started_proven REFUSED,
  outcome_unknown still ACCEPTED, 520-case replay digest unchanged (502a524cfb62f244), vitest 530 passed

Suggested fix: complete the rule in this line (and the zh-CN twin) — "runtime_lost requires the runtime it lost, and a run blocked on it has an execution that is neither running_attached nor proven to have ended" — implemented in both languages by extending the existing guard to run.execution === 'running_attached' || PROVEN_EXECUTION_STATES.includes(run.execution) under recovery_blocked, mirrored in the Java require, plus the schema guard (the rule is expressible: then.execution becomes not: {enum: [running_attached, settled, not_started_proven]}) and fixture cases in both directions.

Fix constraint: the guard must fire only under recovery_blocked — line 200/zh-CN:228 ("the run may keep runtime_lost as its reason to show that observations during the gap may be missing") requires that a rebuilt running + running_attached run may still carry runtime_lost; and the proven-ended set should reuse PROVEN_EXECUTION_STATES (managed-extension-record.ts:308-311), not a re-typed state list. The reused error message ("…while attached") is semantically wrong for the two new arms — give them their own.

Fix witness: new invalid runCases/runSuccessorCases entries (e.g. runtime-lost-after-proven-end) replayed by managed-extension-record.test.ts it.each and ManagedExtensionRecordContractTest.replay — removing the extended guard must flip them from refused to accepted and redden both suites; the pair case must flip isExtensionRunSuccessor true → false.

中文说明

runtime_lost 的排他规则只排除了 running_attached——执行已被证明结束(settled/not_started_proven)的运行仍可停留在 recovery_blocked + runtime_lost,与该原因自己的定义直接矛盾。

按实现的规则(TS :684-690、Java :350-354、schema :1231-1252 与本行两种语言)只拒绝 attached 形状。在被审提交上以语料自带的 blocked-lost-runtime 为基、只改 execution 实测:settled——parse 接受、后继 true;not_started_proven——同样;只有 running_attached 被拒。recovery_blocked 不是终态,所以「终态需要已证明结束的执行」规则(:778-785)对它不生效。这条链正是文档 Rebuild 叙事去掉重建的那一半:Runtime 丢失 → 阻塞,随后原系统给出证明、执行走到 settled/not_started_proven,而没有任何规则要求生产者把运行推离 recovery_blocked。记录于是同时断言「工作已被证明完成(或被证明从未运行)」和「因 Runtime 丢失且未重新 attach 而阻塞」——与第 181 行/中文版 184 行的定义(「阻塞的运行未能重新 attach」)矛盾。H0c 的投影会把一个结果已存在的任务永久显示为 blocked;not_started_proven 变体最荒谬——工作被证明从未开始,却仍以「Runtime 丢失」为由阻塞,无任何可恢复之物。

建议修复:把本行(及中文孪生)的规则补全——「runtime_lost 要求有它失去的 runtime;因它而阻塞的运行,其执行既不能是 running_attached,也不能是已被证明结束的执行」——两种语言把现有守卫扩为 recovery_blocked 下 run.execution === 'running_attached' || PROVEN_EXECUTION_STATES.includes(run.execution),Java 的 require 镜像,schema 同加守卫(该规则可表达:then.execution 改成 not: {enum: [running_attached, settled, not_started_proven]}),fixtures 双向补例。

修复约束:守卫必须只在 recovery_blocked 下生效——第 200 行/中文版 228 行(「运行可保留 runtime_lost 作为原因,以表明间隙期间的观测可能缺失」)要求重建后的 running + running_attached 仍可携带 runtime_lost;「已被证明结束」的集合应复用 PROVEN_EXECUTION_STATES(managed-extension-record.ts:308-311),不要重抄状态名。沿用的报错文案(「…while attached」)对两个新臂语义不对,应另给一条。

修复见证:新增 invalid 的 runCases/runSuccessorCases 用例(如 runtime-lost-after-proven-end),由 managed-extension-record.test.ts 的 it.each 与 ManagedExtensionRecordContractTest.replay 回放——去掉扩展守卫必须让它们从被拒翻成被接受、两侧套件变红;pair 用例还须让 isExtensionRunSuccessor 由 true 变 false。

— qwen3.8-max via Qwen Code /review (v0.24.7)

"kinds": {
"monitorRun": "managed-monitor_run",
"monitorOutput": "managed-tool-result-manifest"
},

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-6: The corpus uses managed-runtime-receipt and managed-monitor-observation as the kind vocabulary for startReceiptRef/lastObservationRef, but these two strings exist nowhere else in the repository and this kinds block does not declare them.

Repo-wide grep (excluding node_modules/dist/.git): managed-runtime-receipt 156 uses, managed-monitor-observation 146 — every hit inside this fixture; external hits 0. A kind→path mapping confirms each string appears only in startReceiptRef / lastObservationRef respectively. The design docs write only "durable ref" for these three refs (only outputRef gets a named kind); the schema's $defs/durableRef.properties.kind is a bare $ref to $defs/id with no const/enum; and parseMonitorRun validates kind/schemaVersion only for outputRef. The sibling contract's convention is that its kinds block IS the complete vocabulary (managed-tool-result-v1.fixtures.json declares 3 kinds and the corpus uses exactly those 3); this fixture uses 38 distinct kinds and declares 2. (commandRef's managed-tool-args is fine — it matches existing usages, e.g. managed-runtime-tool-v3-host.test-helper.ts:127.) When H3 lands, the TS Monitor tool and the Java Runtime that issues start receipts each pick their own kind string; both pass every validation and commit, and the disagreement surfaces only later, at the notification/resume path that parses the ref by kind.

Witness:

grep -rn (excl. node_modules/dist/.git): managed-runtime-receipt 156 | managed-monitor-observation 146
  -> every hit in this fixture; external hits 0
kind->path map: managed-runtime-receipt only in startReceiptRef; managed-monitor-observation only in lastObservationRef
sibling convention: managed-tool-result-v1.fixtures.json declares 3 kinds, corpus uses exactly 3
fixWitness mutation executed (add one line to this kinds block only):
  x validates the shared fixtures against the shared schema  x pins the domains, limits, kinds, ...
  (2 failed | 528 passed) — the block is pinned in three places

Suggested fix: add the two kinds to this kinds block (e.g. "startReceipt": "managed-runtime-receipt", "observation": "managed-monitor-observation"), mirrored in MANAGED_EXTENSION_RECORD_KINDS (managed-extension-record.ts:45-48) and the Java kinds constant; and either name them in the design docs' Monitor-run table (both languages) or state explicitly that these refs' kinds are opaque until H3, so readers do not mistake the fixture strings for normative vocabulary. Do NOT add a const to the schema's $defs/durableRef.kind — that $def is shared by commandRef, startReceiptRef, lastObservationRef, outputRef and recordRef, and one const would kill the other four.

Fix constraint: MANAGED_EXTENSION_RECORD_KINDS.monitorOutput is derived, not literal (monitorOutput: MANAGED_TOOL_RESULT_KINDS.manifest, managed-extension-record.ts:47; managed-tool-result.ts:28) — new entries must follow the same single-source style rather than rewriting the derived kind as a literal, or the two contracts drift independently.

Fix witness: managed-extension-record.test.ts "pins the domains, limits, kinds, state lines and reasons" (expect(fixtures.kinds).toStrictEqual({...MANAGED_EXTENSION_RECORD_KINDS})) and the Java whole-node pin — the executed mutation shows adding a line here without moving both constants reddens both tests.

中文说明

语料把 managed-runtime-receipt(156 次)和 managed-monitor-observation(146 次)当作 startReceiptRef/lastObservationRef 的 kind 词汇使用,但这两个字符串在整个仓库中只存在于这个 JSON 文件里,而文件自己的 kinds 常量块并没有声明它们。

设计文档对这三个 ref 只写「durable ref」(只有 outputRef 有具名 kind);schema 的 $defs/durableRef.properties.kind 是指向 $defs/id 的裸 $ref,无 const/enum;TS 侧 parseMonitorRun 也只对 outputRef 校验 kind/schemaVersion。兄弟契约的既有约定是「kinds 块 = 该契约完整词汇表」(managed-tool-result-v1.fixtures.json 声明 3 个、语料恰用这 3 个),本文件用了 38 个 distinct kind 却只声明 2 个。(commandRef 的 managed-tool-args 没问题——与既有代码用法一致,如 managed-runtime-tool-v3-host.test-helper.ts:127。)H3 落地时,TS Monitor 工具与签发 start receipt 的 Java Runtime 会各自挑一个 kind 字符串;两边都能通过全部校验并提交,不一致不会在提交时报错,只在通知/恢复路径按 kind 解析 ref 时失败。

建议修复:把这两个 kind 补进 kinds 块(如 "startReceipt": "managed-runtime-receipt"、"observation": "managed-monitor-observation"),并同步进 MANAGED_EXTENSION_RECORD_KINDS(managed-extension-record.ts:45-48)与 Java kinds 常量;同时在设计文档 §Monitor run 表格(中英两版)写明这两个 kind,或明确写下「这两个 ref 的 kind 在 H3 之前是不透明的」,以免读者把 fixture 里的字符串当成规范词汇。不要给 schema 的 $defs/durableRef.kind 加 const:该 $def 被 commandRef、startReceiptRef、lastObservationRef、outputRef、recordRef 共用,一个 const 会同时打死其余四个。

修复约束:MANAGED_EXTENSION_RECORD_KINDS.monitorOutput 是派生而非字面量(monitorOutput: MANAGED_TOOL_RESULT_KINDS.manifest,managed-extension-record.ts:47;managed-tool-result.ts:28),新增条目必须沿用这种单一来源写法,不能把已有的派生 kind 改写成字面量,否则两个契约会各自漂移。

修复见证:managed-extension-record.test.ts 的「pins the domains, limits, kinds, state lines and reasons」(expect(fixtures.kinds).toStrictEqual({...MANAGED_EXTENSION_RECORD_KINDS}))与 Java 的整节点 pin——已实跑的变异证明:只往这里的 kinds 块加一行而不动两侧常量,两个测试都会红。

— qwen3.8-max via Qwen Code /review (v0.24.7)

"valid": false,
"pin": {
"definitionId": "schedule-nightly",
"definitionRevision": 1e400,

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-7: Four boundary values are encoded as the JSON literal 1e400, which is outside IEEE-754 double range — the out-of-repo Python reference that regenerates this file cannot round-trip its own output, and the regenerated corpus is unloadable by both shipped parsers.

The four sites: grantCases/expiry-past-double-range.expiresAt (:1501), grantSuccessorCases/next-past-double-range.next.expiresAt (:4727), pinCases/revision-past-double-range.definitionRevision (:4870), monitorRunCases/max-events-past-double-range.maxEvents (:15185). Every double-based parser turns the literal into Infinity — a value that is not JSON. The doc names a Python implementation as the independent generator; json.load yields inf at all four places, json.dumps(..., allow_nan=False) raises ValueError: Out of range float values are not JSON compliant, and plain json.dumps writes the bare token Infinity — which both shipped parsers reject outright (measured: JSON.parse SyntaxError; the re-emitted file fails the real TS suite at fixture load). A second, independent round-trip blocker: the 11 deliberate lone-surrogate strings (e.g. "\ud800" at :2172) force ensure_ascii=True (or preserved escapes) in any regeneration path — replacing 1e400 alone does not make the corpus Python-round-trippable.

Witness:

python3: non-finite floats after json.load: 4 (the sites above)
  dumps(allow_nan=False): ValueError: Out of range float values are not JSON compliant
  plain dumps: emits bare Infinity x4 -> JSON.parse: SyntaxError: Unexpected token 'I'
re-emitted file vs the REAL TS suite: Test Files 1 failed, Tests no tests (fixture load)
fix arm (1e400 -> 1e308 x4, scratch tree): 530 passed; all four verdicts and messages byte-identical;
  next-past-double-range successor still false
lone surrogates: re-encode with ensure_ascii=False -> UnicodeEncodeError (surrogates not allowed)

Suggested fix: replace 1e400 with a finite literal far past every bound, e.g. 1e308, at all four sites (measured: no verdict or message changes in either language). If the genuinely non-finite behaviour must be pinned, pin it in each language's own unit test, where the value is constructed in memory instead of serialized into the shared file — note doc:63's "a number past the double range is not an integer" rule keeps its TS witness either way (Number.isSafeInteger rejects Infinity with the identical message), but the case ids would no longer claim a non-finite input the shared file cannot carry.

Fix constraint: the replacement must stay above the ceilings both languages enforce — maxTimeMs: 8_640_000_000_000_000 (managed-session-records.ts:36) and MAX_TIME = 8_640_000_000_000_000L (ManagedExtensionRecords.java:120) — and must not collide with the corpus's exact brackets (latest-expiry/expiry-past-maximum, largest-revision/revision-past-safe-count); 1e308 satisfies both.

Fix witness: a raw-text guard inside "validates the shared fixtures against the shared schema" — read the fixture file as text, match every JSON number literal, assert Number.isFinite(Number(match)) for each; restoring 1e400 at any of the four sites turns it red, while the existing it.each replays keep all four cases at valid: false.

中文说明

共享语料把四个边界值编码为 JSON 字面量 1e400,超出 IEEE-754 double 范围——文档指名的仓外 Python 参考实现无法往返自己的输出,再生成的语料两个 shipped 解析器都加载不了。

四处站点:grantCases/expiry-past-double-range.expiresAt(:1501)、grantSuccessorCases/next-past-double-range.next.expiresAt(:4727)、pinCases/revision-past-double-range.definitionRevision(:4870)、monitorRunCases/max-events-past-double-range.maxEvents(:15185)。任何基于 double 的解析器都会把该字面量变成 Infinity——一个不是 JSON 的值。实测:json.load 在四处得到 inf;json.dumps(..., allow_nan=False) 抛 ValueError;普通 json.dumps 写出裸 token Infinity——两个 shipped 解析器都直接拒绝该文件(JSON.parse SyntaxError;重 emit 的文件让真实 TS 套件在 fixture 加载阶段全灭)。另有第二个独立的往返障碍:11 处刻意的孤立代理项字符串(如 :2172 的 "\ud800")要求任何再生成路径使用 ensure_ascii=True(或保留原始转义)——只替换 1e400 并不能让语料可被 Python 往返。

建议修复:把四处 1e400 换成远离所有边界的有限字面量,例如 1e308(实测两种语言的判定与消息逐字不变)。若确实要钉住「非有限值」这一行为,请放进各语言自己的单元测试——值在内存中构造,而不是序列化进共享文件;注意 doc:63「超出 double 范围的数字不是整数」这条规则在 TS 侧的见证不受影响(Number.isSafeInteger 对 Infinity 报同样的消息),只是这些用例的 id 不再宣称共享文件根本无法承载的非有限输入。

修复约束:替换值必须仍高于两种语言强制的上限——maxTimeMs: 8_640_000_000_000_000(managed-session-records.ts:36)与 MAX_TIME = 8_640_000_000_000_000L(ManagedExtensionRecords.java:120)——且不得与语料既有的精确括号(latest-expiry/expiry-past-maximum、largest-revision/revision-past-safe-count)相撞;1e308 两者都满足。

修复见证:在「validates the shared fixtures against the shared schema」里加一个原文守卫——把 fixture 文件当文本读、匹配每个 JSON 数字字面量、断言 Number.isFinite(Number(match));在四个站点中任何一处恢复 1e400 都会让它变红,而现有 it.each 回放保持四例 valid: false。

— qwen3.8-max via Qwen Code /review (v0.24.7)

},
{
"id": "waiting-on-a-dispatch",
"valid": true,

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-8: The waiting row is the only row of the doc's Reasons table whose negative half the 520-case corpus never exercises — all six waiting × quota-reason cells are empty, so widening the waiting arm flips 0 of 520 in both languages.

All 440 run blocks in the corpus were enumerated: the waiting cells present are exactly dispatch_unknown (valid, this case), execution_corrupt (invalid, decided by the corrupt-execution rule, not the reason arm) and null. The sibling rows are complete in both directions — running has running-degraded (valid) and running-with-quota-reason (invalid); failed has six valid quota cases plus failed-with-recovery-reason (invalid). The mutation was executed: widening assertReason (managed-extension-record.ts:670-673) so waiting also permits a quota reason flips 0 of 520 verdicts (replay digest identical, suite 530 passed), while a direct probe shows waiting + rate_limit REFUSED by the real module and ACCEPTED by the mutant. The Java twin has the same shape (case "running", "waiting" -> reason == null || recovery, ManagedExtensionRecords.java:336); the schema does state the rule (:1077-1101), but the disagreement test compares module to schema only over corpus cases, so the third gate is blind too.

Witness:

state x quota-reason matrix over 440 run blocks: waiting row all-zero (the only such row)
BASE   waiting + rate_limit/count_limit/budget_exhausted -> REFUSED (run.reason does not fit the waiting state.)
MUTANT same three inputs -> ACCEPTED;  REPLAYED 520 CASES mismatches 0, verdict digest identical; vitest 530 passed
fix arm (add waiting-with-quota-reason, valid:false): MUTANT 1 failed | 530 passed; PRISTINE 531 passed;
  'agrees with the schema...' stays green (schema already refuses the shape -> no BEYOND_SCHEMA entry needed)

Suggested fix: add one invalid runCases entry, e.g. waiting-with-quota-reason, cloned from waiting-on-a-dispatch (:5241-5254, keeping dispatchId: "dispatch-1") with "reason": "count_limit" — probed: it is decided by the waiting arm alone, not swallowed by an earlier check (not an R1-20-style vacuous case).

Fix constraint: the doc states the per-list counts as fact in both languages ("520 cases: … run blocks (144) …", line 236 and the zh-CN twin) and nothing pins them (R1-19), so adding a case must update 144→145 and 520→521 in both files in the same change or the doc goes false silently.

Fix witness: it.each(fixtures.runCases) (managed-extension-record.test.ts:337) and replaysEveryRecordCase (ManagedExtensionRecordContractTest.java:152) — the new case reddens both under the executed widening mutation and is green on the pristine module (both arms measured).

中文说明

waiting 是文档 Reasons 表中唯一负面半边没有语料见证的行——520 例中没有任何 run block 把 state: "waiting" 与六个配额原因中的任何一个配对,因此把 waiting 臂放宽的变异在两种语言里都翻转 0/520。

枚举语料全部 440 个 run 块:waiting 出现的理由恰好是 dispatch_unknown(valid,即本用例)、execution_corrupt(invalid,由 corrupt-execution 规则判定而非理由臂)与 null。兄弟行双向完备——running 有 running-degraded(valid)与 running-with-quota-reason(invalid);failed 有六个 valid 配额用例加 failed-with-recovery-reason(invalid)。变异已实跑:放宽 assertReason(managed-extension-record.ts:670-673)使 waiting 也允许配额原因后,520 例判定 0 处翻转(回放摘要不变、套件 530 全绿),而直接探针显示 waiting + rate_limit 被真实模块拒绝、被变异体接受。Java 孪生同形(case "running", "waiting" -> reason == null || recovery,ManagedExtensionRecords.java:336);schema 倒是表达了该规则(:1077-1101),但分歧测试只在语料用例上比较模块与 schema,所以第三道闸同样是盲的。

建议修复:新增一条 invalid runCases,例如 waiting-with-quota-reason——从 waiting-on-a-dispatch(:5241-5254,保留 dispatchId: "dispatch-1")克隆并把 reason 改为 "count_limit"。已探针验证:它由 waiting 臂单独判定、不会被更早的检查吞掉(不是 R1-20 那类空洞用例)。

修复约束:文档在两种语言里把逐表计数写成事实(第 236 行「520 个用例:……运行块(144)……」及中文孪生),且没有测试钉住它们(R1-19),因此加例必须在同一改动中把两处计数同步改成 145/521,否则文档静默变假。

修复见证:it.each(fixtures.runCases)(managed-extension-record.test.ts:337)与 replaysEveryRecordCase(ManagedExtensionRecordContractTest.java:152)——在已实跑的放宽变异下新用例使两侧变红,在原始模块上保持绿(两臂均已实测)。

— qwen3.8-max via Qwen Code /review (v0.24.7)

"valid": false,
"run": {
"state": "settled",
"reason

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-9: The effectId half of the set-once identities' first-appearance allowance has no witness anywhere in the corpus — forbidding an effectId from ever appearing flips 0 of 520 cases in both replays.

Of the five identities the doc says "may be set once and never change after that" (managed-extension-record.ts:832-838), the corpus witnesses the first appearance of four — definition, executionCallId, dispatchId (identity-and-pin-set-later) and deliveryId (delivery-first-seen-planned, channel-delivery-planned-after-the-end). No valid runSuccessorCase takes effectId from null to a value, and it cannot be folded into identity-and-pin-set-later because parseExtensionRun refuses executionCallId and effectId together (:750-752). The mutation sweep was executed with sibling controls, so the harness's sensitivity is measured rather than assumed: every sibling limb is killed by exactly the cases named, while the effectId limb survives.

Witness:

baseline: FAILED=0 TOTAL=530
effectId setOnce -> sameJson:      FAILED=0   <-- survives
definition (control):              FAILED=1  'identity-and-pin-set-later'
executionCallId (control):         FAILED=2  'reserved-to-admitted', 'identity-and-pin-set-later'
dispatchId (control):              FAILED=1  'identity-and-pin-set-later'
deliveryId (control):              FAILED=2  'delivery-first-seen-planned', 'channel-delivery-planned-after-the-end'
startReceiptRef (control):         FAILED=1  'dispatched-to-attached'
corpus sweep: runSuccessorCases null->value for effectId: [] (all four siblings non-empty)

Suggested fix: add one valid runSuccessorCase, e.g. effect-id-set-later — previous = the admitted run with every identity, execution, runtime, definition and delivery null; next = the same with effectId: "effect-1" and execution: "intent". Both suites replay it with no code change.

Fix constraint: the case counts are normative prose in both languages (line 236: "520 cases: … run revisions (57) …") and no test pins them (R1-19), so adding a case must update 57→58 and 520→521 in both files or the doc goes false silently.

Fix witness: the TS it.each(fixtures.runSuccessorCases) and Java replayPairs("runSuccessorCases", …) on the new case id — both go red when setOnce is replaced by sameJson for effectId at :834, which is the mutation that currently flips nothing.

中文说明

五个「可设置一次、之后不得改变」的身份中,effectId 的首次出现许可在整份语料里没有任何见证——把「允许首次出现」收紧为「永远不得出现」的变异在两种语言的回放里都翻转 0/520。

语料见证了其中四个的首次出现——definition、executionCallId、dispatchId(identity-and-pin-set-later)与 deliveryId(delivery-first-seen-planned、channel-delivery-planned-after-the-end)。没有任何 valid 的 runSuccessorCase 把 effectId 从 null 变成有值,而且它无法折进 identity-and-pin-set-later,因为 parseExtensionRun 拒绝 executionCallId 与 effectId 同时存在(:750-752)。变异扫描连同兄弟对照一起实跑,harness 的敏感性是实测而非假设:每个兄弟分支都被点名的用例杀死,唯独 effectId 分支存活。

具体后果:某个 domain phase 先以 effectId: null 的 intent 提交运行、随后补上该 phase 的 effect 身份——在收紧后的规则下第二次修订会在 Runtime 门禁被拒,运行永远无法记录它实际执行所用的 effect,而两种语言的每个测试仍然通过。

建议修复:新增一条 valid 的 runSuccessorCase,例如 effect-id-set-later:previous = 每个身份、execution、runtime、definition、delivery 全为 null 的 admitted 运行;next = 同样内容加 effectId: "effect-1" 与 execution: "intent"。两套测试无需改代码即可回放。

修复约束:逐表计数是两种语言的规范性文字(第 236 行「520 个用例:……运行修订(57)……」),且没有测试钉住(R1-19),加例必须在同一改动中把两个文件的 57→58、520→521 一起改掉,否则文档静默变假。

修复见证:TS 的 it.each(fixtures.runSuccessorCases) 与 Java 的 replayPairs("runSuccessorCases", …) 回放新用例 id——把 :834 的 effectId 从 setOnce 换成 sameJson 时两者都必须变红,而这正是今天翻转不了任何判定的那个变异。

— qwen3.8-max via Qwen Code /review (v0.24.7)

"valid": false,
"run": {
"state": "settled",
"reason

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-10: This case is rejected during parse before the successor rule it names is ever reached, so it pins nothing — mutation-verified in both directions.

The next block carries both "reason": "runtime_lost" and "runtime": null, so assertReason fails first (managed-extension-record.ts:682: "run.reason runtime_lost needs the Runtime binding it lost.") and isExtensionRunSuccessor returns false at its parse guard without ever evaluating the re-attach rule the id names ("a re-attach under a later generation must not clear the Runtime binding", :866). This violates the doc's own contract at line 236: "Each invalid case is aimed at one rule." A sweep of all 123 pair cases confirms this is the ONLY one whose id names a pair rule yet stops at parse (the other nine parse-stopped cases are named invalid-*/next-past-double-range — their parse failure IS the rule).

Witness:

1) intact:   parseExtensionRun(next) threw 'run.reason runtime_lost needs the Runtime binding it lost.'
             reason:null variant parses, and isExtensionRunSuccessor still returns false (via the real rule)
2) mutant (delete 'after.runtime !== null &&' at :866), fixture as-is: 530 passed — the guard is unpinned
3) mutant + reason:null fix: 1 failed — TypeError: Cannot read properties of null (reading 'generation')
4) fix only, mutant reverted: 530 passed — the fix itself is safe

Suggested fix: change this case's next.reason from "runtime_lost" to null (a running state allows a null reason). Measured: next then parses, isExtensionRunSuccessor(previous, next) still returns false, valid: false stands, but the rejection now comes from the successor rule itself. If the author prefers to keep the "a rebuilt run may still carry runtime_lost" semantics, add a separate next.reason: null case rather than leaving this one stopped at parse.

Fix constraint: managed-extension-record.ts:682 fails runtime_lost when runtime is null, so the fixed next cannot keep runtime_lost with a null runtime; keeping the reason requires supplying a binding, which stops testing the clears-binding rule.

Fix witness: this case's replay in managed-extension-record.test.ts it.each(fixtures.runSuccessorCases) and ManagedExtensionRecordContractTest.replayPairs — the executed mutation proof: deleting after.runtime !== null && at :865-870 keeps 520 cases green today; after the fix the same mutation must redden this case with the TypeError above.

中文说明

该用例的 next 块在 parseExtensionRun 阶段就被拒绝,isExtensionRunSuccessor 从未走到它名字所指的那条规则,因此它没有钉住任何东西。

next 同时写了 "reason": "runtime_lost" 和 "runtime": null,assertReason 先命中 managed-extension-record.ts:682(「run.reason runtime_lost needs the Runtime binding it lost.」),于是「更晚 generation 下重新 attach 不得清空 Runtime binding」这条后继规则(:866)在本用例上完全未被执行——违反文档第 236 行自己的约定「每个无效用例都针对一条规则」。对全部 123 个 pair 用例的扫描确认:这是唯一一个 id 指向 pair 规则、却停在 parse 的用例(其余九个停在 parse 的用例本来就以 invalid-*/next-past-double-range 命名——parse 失败正是它们要钉的规则)。

变异已双向实跑:现状下删掉 :866 的 after.runtime !== null &&,520 例 0 处翻转(守卫无人钉住);按建议修复后再做同一变异,该用例以 TypeError: Cannot read properties of null (reading 'generation') 变红;只改 fixture、还原变异,530 全绿(修复本身安全)。

建议修复:把该用例 next 块里的 "reason": "runtime_lost" 改为 "reason": null(running 状态允许 reason 为 null)。实测改后:next 可正常 parse,isExtensionRunSuccessor(previous, next) 仍为 false,"valid": false 不变,但拒绝理由变成了后继规则本身。若作者更想保留「重建后仍可带 runtime_lost」的语义,应另加一条 next.reason: null 的用例,而不是让这条停在 parse 阶段。

修复约束:managed-extension-record.ts:682 在 runtime 为 null 时拒绝 runtime_lost,所以修复后的 next 不能在 runtime 为 null 时继续保留 runtime_lost;若要保留该 reason,就必须同时给出 binding,那样便不再检验「清空 binding」这条规则。

修复见证:该用例在 managed-extension-record.test.ts 的 it.each(fixtures.runSuccessorCases) 与 ManagedExtensionRecordContractTest.replayPairs 中的回放——变异证明:删除 :865-870 的 after.runtime !== null &&,今天 520 例全绿;修复后同一变异必须让本用例以上述 TypeError 变红。

— qwen3.8-max via Qwen Code /review (v0.24.7)

wenshao added a commit that referenced this pull request Oct 4, 2026
…eader

Round-5 real-stack verification's cross-version matrix disproved the
v2 stamp's premise: monitor_run has been a parsed, known domain since
#12837 (v0.24.7), and a v1 reader opens a Session holding monitor_run
records without complaint. Stamping minimumReader: managed-session/2 on
every new Session — while monitor_run remains disabled — would make a
rollback or a mixed-version rollout lose access to every Session
created in between, and buys protection no deployed binary needs.

Revert the stamp to managed-session/1 and drop the per-domain header
refusal; the enablement list remains the single domain gate. The stamp
rises only when a change genuinely breaks an older reader mid-scan, and
only one release after a tolerant reader ships. The golden fixture's
genesis header regenerates to v1; the records and authority suites gain
v1-header admittance witnesses in place of the refusal cases; the
design's reader section now records the reversal with its evidence.
yiliang114 pushed a commit to yiliang114/qwen-code that referenced this pull request Oct 6, 2026
…13265)

* docs(managed-agent): design H3 background Shell and Monitor runtime

* feat(core): add managed child_run (kind shell) record body for H3

* feat(managed-agent): mirror child_run shell record body in Java store

* test(core): commit and rebuild child_run shell chains through the session authority

* feat(managed-agent): close child_run reference closure and add stop-request draining

* feat(core): add named cgroup unit attach and managed child-run supervisor

* feat(cli): admit v3 background shell under the supervised worker registry

* feat(core): add open-ended shell stream capture with manifest revisions

* fix(core): type-narrow stream capture finalize and refused publish double

* fix(cli): close background shell type gaps after the main merge

* fix(managed-agent): answer the R1 review round on the child_run contract

* fix(managed-agent): repair the CI-only cgroup root and env-guard failures

How: wrap the delegated-root probes so a missing or unreadable root answers
the documented isolation error on Linux too (macOS refused at the platform
check first, which is why only CI saw the raw ENOENT), and document the
background Shell environment allowlist in the process.env guard.

Why: the Test lane on 4fb9f0a9c8 went red on exactly these two items; both
are this branch's own changes, not flake.

Test: hook-command-cgroup and process-env-guard suites green; tsc clean.

* feat(managed-agent): inject the child-run supervisor at worker boot

How: registerManagedContextRoutes builds a ManagedChildRunSupervisor from
the delegated cgroup root the Hook commands already use
(QWEN_MANAGED_HOOK_CGROUP_ROOT) and passes it as the executor's fifth
argument; a boot without the delegation keeps the executor's committed
refusal instead of failing to boot.

Why: executeV3Background landed in the previous increment with the
supervisor parameter unwired, so every real-stack background start answered
the committed no-supervisor refusal.

Test: new wiring case asserts the supervisor is injected exactly when the
root is set; managed-context-worker, process-env-guard and
managed-background-shell suites green (775); tsc clean.

* docs(managed-agent): pin the two-row ledger shape for background processes

How: the runtime-ownership section now says the ledger holds two rows —
the foreground-length start invocation settling with the handle, and a
second execution admitted at start that carries no model result and stays
non-terminal until physical exit evidence.

Why: the Java ledger writes a result exactly when an execution settles, so
a single row that is both handle-delivered and active cannot exist; the
two-row shape keeps the settled-if-and-only-if-result invariant while
preserving every cited behavior (Runtime holds, evidence-only settle).

* feat(managed-agent): commit child_run shell lines from the hosted side

How: HostedChildRunSession funnels every record write through one
serialized, replay-safe commitExtensionRecord path — admit first with the
start call's args as commandRef, dispatch and the set-once
managed-runtime-receipt after the physical start, outputRef advance-only,
and settlement only on proven exit evidence, a proven failure, or an
honored stop request. Live revisions alternate the run line (no
self-loops); the authority freeze refuses anything after a terminal
revision.

Why: the dual path puts product records on the hosted authority, but
nothing commits child_run from the hosted side yet — the worker-owned
registry intentionally does not touch records.

Test: four authority-level cases — full chain with projection, stop
request to cancelled, pre-start failure frozen on not_started_proven, and
serialization with deep-equal skips.

* feat(managed-agent): add the detached capture family to Tool v3 results

How: a background Shell start now settles success with a capture object
whose status is 'detached' — no reason, no manifest — because the live
output streams through the child_run record's growing manifest rather
than the result; the shared schema pins both invariants, the TS parser
mirrors them, the shared fixtures gain one valid and two invalid envelope
cases replayed on both sides, and the Java projector accepts the missing
manifest for unavailable or detached captures. Admission refusals settle
not_started with a null capture, the shape the durable unstarted family
already owns.

Why: every Session tool.receipt event lands in the Java delivery
projection, which requires a capture object and a committed-if-manifest
pairing for anything that started — a success result with a null capture
fails that projection and corrupts replay.

Test: core contract suite 519 (three new envelope cases), TS serve suites
33 + turn/harness 323 + background 6, Java publication contract and
projector suites 8; tsc clean on both packages.

* fix(managed-agent): answer the R2 review round on worker, capture and record lines

How — the four Criticals first: the stream capture publishes an ended
stream's revision only after its seal decision, so a sealed-or-incomplete
descriptor never changes again (R2-3); finalize's settle path degrades to
an unavailable capture instead of throwing past a writability failure, and
the last published revision stands like the worker-cut cap (R2-5); the
child_run body now refuses a settled or cancelled run that is not settled
execution, a failed run off its two ending lines, start_failed outside
not_started_proven, and a process-level failure without a settled
execution, on both languages with new witnesses replayed from the shared
fixtures (R2-6, R1-30); and the supervision suite skips win32 instead of
spawning a shell that cannot exist (R2-7).

How — the Suggestions land as: supervisor start now proves membership by
cgroup.procs with an fd-3 status channel, failing closed as isolation
(R2-37); the bounded EOF wait caps daemon-inherited pipes instead of
hanging, and a spawn-time error drains the entry (R2-1/R2-14); registry
hasHolds requires its session and setProcessResult carries the full result
shape (R2-13/R2-15); create() validates its caller-named unit like attach()
(R2-26); the executor checks the journal before any effect, mirrors the
closing/is_active re-check after prepare, gates the ninth live Shell as a
committed quota refusal, and keeps cgroup membership out of an empty
QWEN_MANAGED_HOOK_CGROUP_ROOT (R2-23/R2-19/R2-24/R2-18); the spawn uses
the configured shell and the session-context environment (R2-21/R2-20);
invalid fixture cases each pin their refusing clause on both languages
(R2-36); projection fixtures pin the draining precedence rows and the Java
replay reads stopRequested (R2-28/R2-38); the draining Javadoc states the
rule the code implements (R2-39); the duplicated ref helper is gone (R2-40);
the stream-capture suite is typed and gains the seal-before-publish,
page-cursor, late-failure and settle-degrade witnesses (R2-29/30/31/32);
and the design doc names the cgroup switch decision, the cursor's single
ordinal space, the unconditional host-scope evidence gate and the zh
stopped-object (R2-8/9/10/11/12).

Test: core suites 788, cli serve suites 776 + background 11, Java record/
projection/store suites 24; tsc clean.

* feat(managed-agent): thread the child-run orchestrator and the detached receipt family

How: HostedSession gains its session-scoped childRuns orchestrator next to
hooks/mcp, the tool turn receives it stored-ahead of the admission branch,
the reopen verifier and the recovery replay both accept the third durable
receipt family — a blocked delivery with a detached capture, beside the
complete and not-started ones — with the publication store correctly left
out of its proof, since its durable truth is the child_run record.

Why: the H3 background start settles with a handle whose output lives on
the record's growing manifest; without the third family, any Session
holding a background receipt could never reopen or be replayed.

Test: recovery-session suite 52 with two new detached-family cases;
harness-session and tool-turn suites green; tsc and lint clean.

* test(cli): type the background-shell capture double

How: drop the as-unknown cast on the FakeSink return and implement the
failCapture member the cast had been hiding (R2-29's cli location).

Why: an unchecked double can drift from the interface it claims to match —
and it already had.

Test: background-shell suite 10/10; tsc clean.

* feat(managed-agent): admit background Shell starts from the hosted tool turn

How: the turn replaces its two deliberate is_background refusals — exactly
when the Session owns its child_run orchestrator and the domain is
enabled — with orchestration that mirrors the shared line: the record
intent precedes every physical effect, dispatch_started lands the moment
the checkpoint commits, a settled detached handle attaches the physical
start to the record from the same facts and lands as the third durable
receipt family (blocked delivery with a null manifest, replay-validated
exactly like the not-started family), and a proven-unstarted refusal
settles start_failed on not_started_proven while riding the existing
unstarted family unchanged. Malformed is_background values keep their old
refusal, and the family stays out while the domain is disabled — the same
probe governs both paths.

Why: the background start is the first record-bearing, user-visible H3
effect, and without a hosted history family for the handle a Session that
ran one could never reopen or be replayed.

Test: three new authority-level turn cases — admitted detached flow with
admit/dispatch/attach order and the blocked receipt family, a
proven-unstarted refuse closing NOT_STARTED, and the disabled-domain
refusal preserving its exact text; turn suite 146, plus harness-session,
recovery, background and child-run suites 246; tsc clean.

* feat(managed-agent): admit background Tool v3 dispatches and acknowledge their detached settle

How: the v3 start admission flips from "foreground Shell only" to
"Shell only" — background dispatches now begin the same way; the poll
loop accepts the settled detached family through the status answer (the
only durable settle such a result can ever produce, since nothing is
published for a background handle), and the acknowledgement path
canonizes the settled handle envelope: a matching blocked receipt with a
null manifest and no history revision forwards exactly itself, without
consulting the publication receipt.

Why: without the detached branch, a background start settled a queued
handle only to mark the execution UNKNOWN after the 30-minute
foreground-shaped deadline — every background dispatch would wedge the
Runtime on its first call, and its acknowledgement would die on the
missing publication it was never going to have.

Test: new witness drives the detached start end-to-end — reservation,
v3 dispatch, settled handle via status, and ack with a null-manifest
blocked receipt, with the publication receipt path provably untouched;
service suite 139/139, module suite 600/600.

* feat(managed-agent): answer shell-status and shell-terminate from the worker registry

How: a new private maintenance route sibling to ManagedHookProtocol
(/internal/managed-runtime/v3/shells) backed by the process registry —
status answers running or unknown (never claimed without evidence),
terminate drains with the supervisor's own rules and answers exited with
its proof, and an end that cannot be proven is unknown with the hold
still registered. The unit name derives from the runtime invocation
identity through the shared shellUnitNameOf, so Java, the hosted turn
and the worker derive it identically; the executor takes the registry as
a constructor injection so the route and the journal share one.

Why: every later maintenance verb — recovery reconciliation, the ordered
close drain, the reconcile-settle — has to ask the only physical owner
these questions; answering them from anything but the registry would
invent truth.

Test: five worker-side cases — closed-key refusals, unknown for
unowned, running scoped to the registered Session, terminate-with-evidence
answering exited, and an unproven terminate answering unknown with the
hold intact; shell, context-worker, v3-route and child-run suites 784;
tsc clean.

* feat(managed-agent): admit the background process row and answer its maintenance protocol

How: the start dispatching of a background Shell admits a second ledger
execution beside the invocation row — the new process row carries no
model result and stays PREPARED (non-terminal) until physical proof, so
hasActiveBy* holds the Runtime exactly while the process lives. The
maintenance protocol mirrors ManagedHookProtocol on its own route:
shell-status and shell-terminate with a targetOperationId (both recovery
kinds), wire-validated on both transports and accepted through the same
workspace-generation and ownership gates. observeBackgroundProcess asks
the physical owner and settles the process row with proven evidence —
settle through the new repository verb settlePrepared, which mirrors the
PREPARED-requestCancel shape on both repositories; anything unproven
leaves the row active with its hold, never a claimed end.

Why: the ledger writes a result exactly when an execution settles, so a
single background execution carrying its handle at start cannot exist;
two rows keep the settled-if-and-only-if-result invariant while the
Runtime holds precisely, and the maintenance route gives every later
verb the worker's physical answer instead of an invented one.

Test: two new witnesses — a background start that admits the process row
(the invocation SETTLED, the process PREPARED, the Runtime busy until
status arrives, then exit evidence settling it and release succeeding),
and an unproven status keeping its hold; the runtime also gained the row
identity invariant fix and the settle-through-settlePrepared shape;
service suite 141/141, module 603/603.

* fix(managed-agent): answer maintenance views under the target identity

How: the shell-status/shell-terminate view now carries the target
operation's identity, exactly as the Hook lookup family does — the
requester asks about the process, never about the asking call, and the
Java wire validator's expectedId rule requires it.

Why: this repair was encoded in the route design but landed uncommitted
after ②e-a; the self-audit caught it before it could ride a later batch.

* feat(managed-agent): run the background exit leg through the record

How: the hosted publisher gains the background family — its registration
marker admits the open-ended capture through the session-level admission
(proven start on the record, never the model call), the open manifest
publishes at prepare, every awaited write or finish forwards the newest
manifest revision to the record's outputRef, and the exit finalizes with
the same evidence object settling the record. The client acknowledged
receipt keeps the detached-blocked shape; drain skips background stores
until the ordered close drain wires them; the register guard only binds
the foreground model-call family.

Why: the start handle already committed, so a background Shell's whole
physical truth is output and exit evidence — anything that reached the
record another way would let a second tool.receipt contradict its own
line, and an open-ended manifest published anywhere but the Session's
own store is an unreadable closure for the store-side commit.

Test: a full end-to-end authority case — admit-dispatch-attach, bytes
through the real rpc, outputRef advancing live, the final sealed manifest
reading the same exit (code 3, both streams sealed, manifest complete),
the record settled with exactly that evidence, and zero tool.receipt
events for the outcome; publisher, turn and child-run suites 174; tsc
clean.

* feat(managed-agent): close monitor_run revisions over their cited Session resources

The H0c closure rule — every ref a record cites must exist in the
Session's resource store at commit time — now covers monitor_run on both
stores: the authority reads commandRef, startReceiptRef, outputRef and
lastObservationRef through verifyExtensionResources, and the Java
applyRevision gains the same four-field branch beside child_run. This is
the precondition for enabling the monitor domains: without it a monitor
revision could anchor its chain to resources nobody committed, and the
replay would project a watch whose evidence was never verified.

Both fixture rigs stop citing fictional resource identities. The
authority test publishes its chain refs as real bodies once per harness
and mirrors each fixture-cited ref as a real same-kind placeholder of
the stated length, memoized so chain identity stays fixed across
revisions. The journal test harness does the same at its generic
request entry point, rewriting each cited ref to the placeholder's real
digest and appending the mirrored resources, but only for strictly
well-formed JSON monitor bodies — a body with trailing content keeps the
raw path so the store's grammar refusal still fires first over the
closure's.

Each side gains a witness that pins the closure itself: a monitor run
citing a resource outside its commit is refused with
managed_session_resource_missing, and each witness was observed red with
its branch disabled before landing.

* feat(managed-agent): gate monitor_run behind the managed-session/2 reader

The H3 design left the reader-gating mechanism to the implementation;
the journal settles it. Headers are immutable — the scanner refuses a
repeated managed_session_header_v1 line and any unknown subtype after
the header — so no transaction can legally raise a Session's
requirement in-journal, and a superseding header line would change the
journal format in both languages for a narrow mixed-version window.
New Sessions therefore stamp minimumReader managed-session/2 at
creation, and every path that commits a monitor_run record —
commitDomainRecord, commitExtensionRecord and the generic append guard
— additionally refuses when the Session's header names less. Readers
still accept requirements up to their own version, so pre-H3 Sessions
stay openable everywhere and simply can never receive monitor_run.

child_run is deliberately not header-gated: a managed-session/1 reader
already parses its name, and the asymmetric hazard lives on the server
side, which the established deploy-server-first sequencing covers.

The shared Stage H golden is regenerated for the v2 genesis bytes and
the Java integration replay passes against it unchanged. New witnesses:
new Sessions stamp v2, a v1 Session refuses monitor_run with the
precise error (observed red with the gate removed), and both header
sides of the version comparison are pinned in the records tests. The
design doc records the decision and the now-paid closure precondition
in both languages.

* feat(managed-agent): funnel monitor_run revisions through a hosted orchestrator

The mirror of HostedChildRunSession for the Monitor domain. The
managed-runtime worker will own the watch loop, but the dual path puts
every product record on the hosted authority, so HostedMonitorSession
commits the record line as watch facts arrive: revision 1 with the
intent before any side effect, dispatch onto a Runtime binding, the
set-once start receipt after the watch starts, observation revisions
only while attached with a watermark that never goes back, output
advancement only forward, and the terminal shapes — settled by
exited/max_events/idle_timeout, failed by a start failure on
not_started_proven or a settled watch failure, cancelled by a stop
request whose notification watermark may still advance after.

Writes serialize per Session and replay-safe by command id with
deep-equal skip, exactly like the Shell funnel. The suite drives the
whole line through a real authority: projection at admission and after
settle, max_events refused before its quota and accepted at it, a
second start receipt refused, an observation before attach refused,
and an idempotent output advance committing nothing. Quota
(enforcement), the debounce-floor observation loop with notification
composition, and rebuild-after-loss stay with the runner and route
increments that follow.

* feat(managed-agent): drive monitor observations through a debounced loop

The loop half of the Monitor runner, on top of the hosted funnel: an
admitted Monitor dispatches, starts an injected watch executor and runs
its own time. Stdout lines aggregate into one observation revision per
debounce window with a one-second floor, so the document's reopen bound
holds for any debounce a watch asks for; the run settles itself on the
contract's own terminal conditions — the observation quota, silence
past idle_timeout, a natural exit with its buffered lines flushed
first — or is stopped, which terminates the watch and lands cancelled.

The executor contract is typed for the worker's cgroup owner (onLine,
onExit after the start resolves, a terminate handle); the suite drives
the whole line through a real authority with a fake executor and a
manual clock: windowed aggregation and its floor, quota settlement and
watch termination exactly at max_events, idle settlement re-armed by
each accepted observation, exit flush order, start and mid-run failure
shapes, and a stop that ignores late lines. Every cross-boundary
settle — timer or executor callback — is counted by the loop's done
handle, which the close drain will await; window flushes count too, so
done never reports idle while an observation commit is in flight.
Notification composition and its wake stay with the next increment;
quota enforcement and rebuild-after-loss stay with the route slice.

* feat(managed-agent): notify from monitor observations in the same transaction

The H0c machinery now runs for Monitors: every accepted observation
commits its revision together with a notification input and the wake
the authority generates for it, and the run's notifiedThrough watermark
advances to exactly that observation's sequence. The loop composes the
input — monitor-1:notify:N ids, the window's joined lines as the
content a woken turn will read, an empty admission, no deadline — and
the funnel threads it through commitExtensionRecord, so a recovery
replay re-runs nothing the watermark already covers.

Dedupe semantics are pinned at both levels: the funnel advances the
watermark only when a notification rides the revision (an observation
without one keeps it where it was), and the loop suite shows the three
events — domain.committed, input.accepted, wake.requested carrying the
input's accepted id as its source — landing in one transaction with the
joined lines as its content. Consumption of that wake by the hosted
scheduler, and settling the input when no turn can take it, come next
(open question 6 of H0c).

* fix(managed-agent): keep a finished background Shell's receipt answerable

H7 of the real-stack rounds: the Broker settles its :process row only
when shell-status answers exited, but the worker's registry deleted an
entry the moment the Shell ended — a natural exit could never be
answered again, and the ledger row would hold the Runtime against
release and close forever. The registry now retains each completed
unit's receipt until the worker ends: the hold drops exactly as before,
but shell-status and shell-terminate answer a proven end from the
retained evidence, idempotently, for the owning Session scope only. An
end without evidence keeps nothing and answers unknown as before; the
existing live flows are untouched. Suite: natural exit keeps answering
exited with its evidence for both kinds after the hold dropped, the
wrong Session still hears unknown, and an evidence-less end stays
unknown. The close-drain/reconcile callers that consume this, and the
Java-side settle-before-busy on workspace release, land with the next
batch.

* fix(managed-agent): settle provable background exits before release's busy check

The second half of H7. The worker now keeps a finished Shell's evidence
answerable, so the Broker has someone to ask — but observeBackgroundProcess
still had no production caller, so a naturally exited Shell left its
:process row non-terminal and every release wedged on
runtime_session_busy. releaseSession now sweeps the background process
rows this Broker admitted for the Session before it computes busy: each
unproven row asks its physical owner once through the same
shell-status → settlePrepared leg observeBackgroundProcess already owns,
a proven exit settles with its evidence, and any lookup that fails or
cannot prove an end simply keeps the row so busy stays the accurate
answer. The sweep tracks per-context admissions because the
reconciliation scan deliberately excludes PREPARED rows, and mutates the
admission set under the context lock.

Witnesses: releaseSettlesAnExitedBackgroundProcessBeforeBusy proves an
exited row settles with its evidence and the release succeeds (observed
red with the sweep bypassed); the pre-existing
unprovenProcessStatusKeepsItsHold still proves an answer without proof
keeps the row and the busy refusal. Full module 604/604. The workspace
close drain still terminates shells in order rather than sweeping — its
ordered stop belongs to the close-drain slice, and these two verbs are
exactly what it will consume.

* fix(managed-agent): check the journaled receipt before attaching a background Shell's start

H6 of the real-stack rounds: acceptBackgroundShell called childRuns.attach
before its tool.receipt replay check. A retried accept — the leg recovering
pending results re-drives — would mint a second start receipt on the same
identity before reaching the replay path, and the record's set-once rule
would refuse it as a rerun shape, so a recoverable retry crashed as a
conflict. The lookup now runs first: a journaled receipt replays its
validated history with no attach at all, and a fresh accept attaches only
when this identity still lacks a start receipt — the record id IS the
execution identity, so an existing one is always our own crash-window
attach between record and journal, which the retry now tolerates instead
of minting a rerun. Witnesses: a retried accept answers from the journal
with exactly one attach and one tool.receipt line, and an
already-attached-but-unjournaled record completes its accept without
attaching again. The full hosted turn suite stays 148/148.

* fix(managed-agent): backpressure the background Shell output pipes

G5 of the real-stack rounds: the executor fed the bounded capture with a
fire-and-forget sink.write — every chunk queued behind asynchronous
publication, so a fast producer flooded memory (a 256 MiB run peaked near
315 MiB RSS at a 20 MiB/s store). The feed now counts bytes in flight and
pauses stdout and stderr when the drain falls sixteen MiB behind, resuming
at four MiB — enough to cover a store several seconds slower than the
process produces without stalling a steady producer against the floor.
Completion or rejection of a write both release their bytes, so a broken
capture can never leave the pipes paused. The witness drives a controlled
sink: seventeen MiB of chunks pauses a stream exactly once, draining past
the watermark but above it resumes nothing, and draining the rest resumes
exactly once. The outputRef-live-advance half of G5 belongs to the
turn-publisher's promotion to activation scope, sized with the close-drain
slice, and stays queued there.

* feat(managed-agent): promote the Shell publisher to Session scope

The ②d-out debt, and with it the live-advance half of G5. The Shell
publisher was a per-turn object: its server died with the turn, while
background Shell traffic flows across turns — between turns (or after a
close) the worker's retained descriptor pointed at a dead endpoint, so
no live output advance and no exit leg could ever complete past one
turn. The hosted Session now owns one long-lived publisher across
turns (the turn lazily structures it on the shell options and never
closes it; the Session's ordered close in this PR's fifth slice drains
it). The two small wire-level fixes that the promotion exposes:

- The worker's background execute stamps the background marker into its
  prepare request: the closed six-field wire capture has no such field,
  so a remote prepare used to arrive unmarked and die at admission; the
  executor derives the marker from the background input itself.
- register() takes the registering turn's prompt for the foreground
  admission check, so one instance admits the foreground of any turn of
  its Session while a foreground pretending another turn's identity is
  still refused; the background arm keeps its record-based admission.

Witnesses: a single instance admits a background watch plus the
foreground of two turns and refuses a misattributed one; the worker rig
pins the stamped marker on its prepare; the four touched suites stay
184/184 and both package typechecks pass (after a core dist rebuild for
the capture type's optional background marker). The publisher's own
descriptor installs idempotently per turn (identical descriptors are
accepted), so consecutive turns never fight for the worker slot.

* feat(managed-agent): drain a Session's background Shells ahead of release

The close-drain's missing legs. Until now a Session with a live
background Shell could never close: the worker's activation gate refused
release with managed_activation_conflict forever, and the CLI-side close
route skipped the background stores. The ordered sequence the design
asks for now exists end to end:

- The registry gains stopSession: terminate one Session's entries with
  the supervisor's evidence rules (TERM, escalate, empty-check), await
  each completion so its finalization — last manifest, sealed envelope,
  record settle through the session publisher — lands before the caller
  proceeds; an unproven stop keeps its hold for the caller to see.
- The worker's activation release gate drains before it refuses: active
  background work no longer wedges a close that can prove its stops,
  and anything still unproven keeps the exact same conflict.
- The hosted close route now closes the Session-scoped publisher right
  after the broker release returns — which by then has drained the
  Shells and settled their records through that publisher — and the
  drain closes every capture store, background families included.

Witnesses: a registry drain settles one Session's shells with evidence,
both kinds, while leaving another's holds untouched, and an unprovable
terminate under a session drain still preserves its hold; the 766-test
context-worker suite stays green through the release gate's new async
drain branch. What remains deliberately unproven locally: the
full-chain release-with-live-shell drain needs a real cgroup supervisor,
pinned as the Linux head acceptance instead of a mocked supervisor.

* docs(managed-agent): sync the H3 design's status line with landed slices

The status line still claimed publication and orchestration remained
design while most of them landed across this PR's batches; refresh both
language versions with the landed list and keep the remaining-design
list accurate: the Monitor physical side (route, rebuild, cgroup watch)
and wake consumer, both submission enablements, task events, and the
Linux physical acceptance.

* feat(managed-agent): spawn Monitor watches through a cgroup watcher

The physical half of the Monitor runner's executor interface, for the
worker: ManagedMonitorWatcher spawns the watch command under a unit
named for the execution identity, through the same supervisor (its
membership attestation unchanged), splits stdout into the observation
lines the loop buffers — boundaries respected, the Legacy partial-line
cap honored, the tail flushed as the final line at exit — and reports
the physical end exactly once, natural exit versus mid-run failure, so
the loop's settled and watch_failed shapes both hold. stderr rides only
the future output leg, never observations; terminate defers to the
supervisor's TERM-escalate-empty rules. The suite drives a supervisor
double: unit/executable/cwd derivation, chunk-boundary splits with the
final flush, the cap drop, spawn-error and membership-failure shapes,
and terminate delegation.

* feat(managed-agent): answer monitor status and stop from a worker registry

The monitor half of the maintenance route, mirroring the Shell's: a
worker-side registry of supervised watches that holds each Session's
Runtime while a watch runs, answers an end only with the supervisor's
own evidence, and retains a proven exit's receipt until the worker ends
so maintainers can learn the truth idempotently. The route sibling at
/internal/managed-runtime/v3/monitors serves monitor-status and
monitor-stop with the same closed wire shape, target-identity rule, and
unknown-but-never-unproven discipline as the shells' route; it joins
the worker's route list so the v2 envelope admits it. Registration
arrives with the Monitor start admission later — status/stop already
have consumers in the release drain and the rebuild slice that follows.
This also repairs the watcher test arity that broke the previous push's
build lanes (the mock's 2-arg onExit omission, seen as
error TS2554 (155,34) across every CI lane at that head).

* feat(managed-agent): admit monitor watches into the open-ended capture session

The γ3-α precondition for every later Monitor leg:
LocalShellStreamResultSession's prove-a-live-start admission reads the
capture family's record domain instead of being pinned to child_run —
same Session/binding/background-marker checks, same live-run gate,
same identity captureId, per-domain labels so a refusal names its own
family. The producer side names its domain explicitly (record presence
is never consulted where both domains could in principle compete), so
a monitor watch carries its monitor_run record into exactly the same
bounded stream the background Shell already rides. Witnesses: shell
admission unchanged; monitor admission succeeds under its domain;
wrong-domain proven start refuses; missing record and foreign binding
refuse with their own errors.

* feat(managed-agent): admit Monitor watches through the v3 executor

The workhorse of the Monitor leg. The worker's v3 executor gains its
is_monitor branch beside is_background: journaled admission first, then
refuses (never as a transport error), then the watch spawns through the
ManagedMonitorWatcher under a unit named for the execution identity,
streams stdout into the SAME bounded family (16/4 MiB backpressure on
the pipe, with an explicit refusal if the watcher reports no supervised
unit), and settles with its detached handle while a registry keeps the
Runtime hold. Three watch-level facts are done physically now:

- The monitor registry is a wrapper around the background Shell
  registry, never a copy: supervisors, EOF drains and publisher finishes
  stay in one place, while monitor holds and the four-watch quota bucket
  remain entirely their own. Its maintenance route answers run/exited/
  unknown with identical semantics to shells'.
- The execute path gates on the generated four-live-watch cap with
  `Session already runs 4 Monitor watches.` and additionally on a
  missing cgroup root and a missing capture service.
- Each prepare carries background+monitoring markers, so the hosted
  side of γ3-α knows the monitor_run domain to check.

Suite: 28 tests across the three touched files — the fork settles a
detached start under the right unit and marker, lines land as monitor
output, the fifth watch is refused, and supervisor-less boot refuses
with its dedicated text. The turn admission arm (funnel-driven
admit→dispatch→attach through HostedMonitorSession), the watcher→loop
observation feed, and rebuild still land in the following increments.

* feat(managed-agent): admit Monitor watches through the hosted tool turn

The turn-side arm of the Monitor leg. A call named monitor is now
admitted through its own three-state gate (requested/unavailable/ill-
formed, with the Shell-profile Monitor declaration now advertised),
implements the same record-first discipline as background Shells do —
the funnel admit commits monitor_run revision 1 before any side effect,
dispatchStarted follows the durable checkpoint, and the watch attaches
when its v3 execution settles. The granted publication, journal intent,
dispatch and accept forks are widened to carry is_monitor alongside
is_background, and acceptMonitor mirrors acceptBackgroundShell end to
end (not_started → start_failed with the unstarted family; detached →
same blocked receipt family with exactly one tool.receipt line, the H6
journal-verdict-before-attach order, and the blocked ack completing the
row). HostedSession gains a session-scoped HostedMonitorSession beside
childRuns; turn construction forwards it at both sites.

Two existing contract pins move deliberately: the Shell profile's
advertised list now names monitor (the harness declaration test names
it explicitly), and the three publisher-close witnesses now assert the
Session-scoped lifecycle settled by the earlier promotion — no close
at any turn ending (completed/error), the server still answering, close
exactly once at the Session's own delete. Suite: 3 new admission-arm
tests through a full hosted turn (funnel order, clamped defaults with
the Legacy monitor caps, start_failed shape, unavailable text), the
whole turn suite stays 151/151 and the harness session suite 180/180.
The observation fan-out (worker lines → hosted loop), quota witness at
the turn, and rebuild-after-runtime_lost land in the next increments.

* feat(managed-agent): fan a monitor watch's observations into the hosted loop

The observation channel of the Monitor leg. The session publisher now
marks monitor captures with recordDomain monitor_run through γ3-α, fans
each published stdout chunk through the Legacy line-split semantics —
boundaries kept, remainder held across chunks with the 4096-byte
partial cap — into a per-capture observer, and closes it with
onExit(failed) at finalization with the tail flushed first. The turn
gains resumeMonitorWatch: a fresh accept creates one HostedMonitorLoop
per Session report and resumes it from the already-attached record, fed
through a HostedMonitorRemoteExecutor whose start binds those observer
callbacks; a replay starts nothing twice, exactly like it never
re-attaches. Globally, loop timers unref so observation windows never
anchor the process.

Monitor capture finalize settles through the monitor funnel itself
(watch_failed when no evidence, exited via settleQuiet), beside the
shell's unchanged path. The publisher carries monitors next to
childRuns, and the turn's own construction forwards it. Suite: the
admission-arm tests now also pin one Session-level loop starting on a
fresh accept with the record-first flow intact; wide suites of 189
tests pass with cli tsc, prettier and eslint clean. Monitor output
before the observer registers stays durable in the open stream, with
observation starting at attach — noted in the design not as a flaw but
as the attach seam, a line crossing it already lands in the output
Artifact readers see.

* feat(managed-agent): rebuild a read-only Monitor after its Runtime is lost

The record side of the rebuild promise. HostMonitorSession gains the two
verbs the H3 lifecycle needs around runtime loss: blockedRuntimeLost
lands the run in recovery_blocked with the runtime_lost reason and the
lost binding still named, so degrading is what the projection shows
until something decides otherwise; rebuildFromRuntimeLost accepts only
the outcome_unknown/runtime_lost line the H0b rule allows, starts a
fresh watch under a newer Runtime generation with its own receipt, and
keeps the reason for honest projection while observation positions
never moves — the support case never restarts what the watermark holds.
monitorRebuildAllowed routes every candidate command through the Legacy
AST read-only check with its walk, so a side-effecting command stays
blocked accurately. Five witnesses cover the full rebuild chain
(recovery-block, rebuild with receipt remint and generation increase, a
rebuild refused from a live run and from an ended one, and the gate's
read-only/side-effect/empty answers). Binding-replacement detection and
re-drive of the new physical watch come with the recovery slice after
the wake consumer.

* feat(managed-agent): serve the task events routes with their SQL journal

H3 flips listSessionTaskEvents and queryWebShellTaskEvents from planned
to partial on both surfaces.

The Java control plane gains a bounded per-task event journal (V39): one
row per committed task event, written in the Session commit transaction,
with the per-task sequence allocated under the journal head lock in
commit order. Record revisions journal their task view changes at the
H0c announcement point. The durable retention floor lives in a per-task
cursor row, so it survives an empty retained set, restarts and
projection rebuilds; the Artifact visibility barrier pins it behind
unarchived output, and the backlog bound refuses past capacity instead
of silently discarding. Output events carry their per-stream segment
ordinals with overlap/gap refusals, and artifact references stop
fail-loud at the 100 bound. output_cursor/outputCursor lift with the
routes and the task views serve them.

The section 6.1 demonstrations run as store-level contract traffic in
the new journal contract test, the route probes and record gates cover
the flips, and the #12847 C15/C16 instances close alongside. WebShell
types regenerated from the bumped 1.30.0 contract.

* docs(managed-agent): add the task events lane's build report

Flyway numbering findings (V39 taken; V35-V38 claimed by in-flight PRs
on main, V15/V29 Java-only burns untouchable), implementation decisions
against the H3 design text, the CLI/daemon flow-surface finding (none
exists, regeneration is the only surface change), the full test ledger
and the owed items.

* feat(managed-agent): build the monitor wake envelope and the pending-input intake

Two support pieces for the H3 wake consumer (β2-b1a):

- monitorNotificationText wraps one due observation window in the exact
  Legacy task-notification envelope — task-id, optional tool-use-id,
  kind/status/event-count, summary and result — applying the same
  truncateNotificationLabel/stripDisplayControlChars/escapeXml pipeline
  the Legacy Monitor applies, so a turn learns nothing new.
- pendingSessionInputs derives the Session's still-pending inputs from
  the journal alone: an accepted input is consumed exactly when a turn
  settles under its turnId, order-insensitively, so a replayed admission
  after a restart is not re-driven.

* feat(managed-agent): deliver the Legacy notification envelope on monitor wakes

The loop's notification input now publishes the exact task-notification
envelope a Legacy Monitor's wake carried — task-id, tool-use-id, kind,
status, event-count, summary and the window's lines — so the turn the
wake raises reads nothing it has not already read. The description is
the tool call's, falling back to the watched command. A witness pins
the exact wire text of the first observation's input content.

* feat(managed-agent): run the monitor wake through an embedded scheduler

H3 closes the wake loop: an accepted observation already commits its
notification input and wake.requested in one transaction; this slice
makes that wake effective.

The pump re-derives the pending monitor inputs from the journal alone —
nothing is held in memory, so a restart re-derives the set. On an idle
Session the oldest notification runs as an ordinary text turn with the
Legacy envelope as its prompt, and the turn's settle consumes the input.
A busy Session queues in the journal exactly like channel and Goal
inputs; the retry loop redelivers. A Session that is parked, blocked or
mid-recovery keeps its reminders pending and reports them. On the close
path every pending notification settles cancelled without a model turn,
under the turn-result record's own idempotency key, so no wedged
notification ever holds a Session as hosted_turn_recovery_required at
its next open: the parked-turn scans now skip monitor inputs, which the
pump owns exclusively.

A wake turn that dies without settling re-drives on reopen — the input
is consumed exactly when a turn settles under its turnId, and a crashed
turn never settles — matching the queue semantics channel and Goal
inputs already honor.

* fix(managed-agent): keep every new Session on the managed-session/1 reader

Round-5 real-stack verification's cross-version matrix disproved the
v2 stamp's premise: monitor_run has been a parsed, known domain since
#12837 (v0.24.7), and a v1 reader opens a Session holding monitor_run
records without complaint. Stamping minimumReader: managed-session/2 on
every new Session — while monitor_run remains disabled — would make a
rollback or a mixed-version rollout lose access to every Session
created in between, and buys protection no deployed binary needs.

Revert the stamp to managed-session/1 and drop the per-domain header
refusal; the enablement list remains the single domain gate. The stamp
rises only when a change genuinely breaks an older reader mid-scan, and
only one release after a tolerant reader ships. The golden fixture's
genesis header regenerates to v1; the records and authority suites gain
v1-header admittance witnesses in place of the refusal cases; the
design's reader section now records the reversal with its evidence.

* fix(managed-agent): keep monitor output byte-true and advance only forward

Round-5 real-stack verification found three facts about the monitor
output path:

- A capture rebuilt from decoded lines cannot be byte-true: per-chunk
  toString split multi-byte runes into U+FFFD, and line-rounding dropped
  blank lines and the final unterminated line. The watcher now forwards
  raw stdout chunks to the capture on a new onChunk channel, so the
  durable Artifact reproduces the command's stdout exactly, while the
  observation lines keep the Legacy emit-path semantics (a StringDecoder
  holds a split rune until it completes; a blank line consumes no
  observation).
- The watcher registered its exit listener only after the supervisor's
  start resolved, so a fast watch could end unheard and the debounced
  loop never settled. A latched end with a check-phase sweep covers a
  watch that exited before its listeners attached, after the host is
  known to be listening.
- The funnels' advanceOutput promised "only ever to a newer revision"
  in prose and accepted an older manifest in code. Both hosted funnels
  now read the revision of both manifests and refuse a regression
  before anything commits; witnesses pin the refusal in place.

* feat(managed-agent): drain a busy background Shell at release, and answer the asked operation

Round-5 close-path findings (H8): with a running background Shell the
Broker refused busy before transport.release — yet that HTTP call is the
only route to the worker's drain — so a `tail -f` held session release
and workspace close forever, with nothing able to stop it.

The release path now distinguishes busy kinds: anything active that is
not one of the Session's own background process rows keeps the exact
busy refusal with zero worker hops, while proven-running rows trigger a
dedicated release call first whose 409 busy answer is swallowed (the
worker's route drains before refusing, and its maintenance routes are
not activation-gated, so the follow-up sweep still proves what the
drain ended). Rows settle from the drain's evidence and the complete
release converges on what is genuinely still running — one call when
the drain finishes in time, one retry when it does not. Two witnesses
pin the contract: a running Shell reaches the worker (the round-5 probe
found the refusal never did) and a non-background busy refuses without
a worker hop.

Also from the round: the shell maintenance validator computed the
expected operationId from targetOperationId but never compared it with
the answer's own operationId, and the sweep witness answered with the
process-row id where the real route echoes the call id. The validator
now refuses a view answering another operation, a new protocol suite
pins it, and the witnesses answer the shape the route really sends.

* fix(managed-agent): re-assert the output pause behind every flushed chunk

Round-5 G5 finding: Node's flushStdio resumes both paused pipes when
the background launcher exits, and the executor tracked the pause in
its own flag rather than the stream's, so a writer that outlives its
launcher pushed 230 MiB into the queue and climbed RSS by +200 MiB.

The onOutput accounting now re-asserts the pause behind every delivered
chunk, and applyPause is idempotent through the stream's isPaused state
instead of the local flag — the stop-and-go the backpressure intended
holds through the launcher exit, on both the Shell and the Monitor
paths. The existing 17-MiB pipe pacing witness keeps its single pause
call.

* test(integration): expect the monitor tool on the hosted shell profile

The hosted shell profile now advertises `monitor` beside
`run_shell_command`, so the fake-model driver's tool-name assertion —
which the Hosted workspace tool-turn IT checks on every model call —
must expect it. The file profile's expectation is unchanged: the
declaration is shell-profile only.

Fixes the single red IT in the MySQL fault-gates lane
(HostedWorkspaceToolTurnIT.packagedHarnessUsesSavedWorkspacesThroughReal
BrokerWorkerAndSqlStore), whose only difference was the extra
'monitor' entry.

* fix(managed-agent): read the whole journal when deriving pending monitor wakes

`readEvents()` is a bounded read: with no arguments it returns the first
`defaultReadEvents` (100) events, capped at 256. Both consumers of the
pending-input derivation called it bare, so on any Session whose log runs
past that page — every real Session with tool calls — the derivation saw
only the log's head:

- the wake scheduler's `next()` never found a notification input, so a
  monitor wake was never delivered as a turn once the Session passed the
  page. The feature was dead in production while every rig stayed green,
  because the rigs commit a handful of events;
- `settlePendingMonitorInputs` never found the owed inputs at close, so
  they stayed unsettled and the Session reopened as
  `hosted_turn_recovery_required` — exactly the wedge this close-path
  settle exists to prevent.

Both now read the full committed prefix with
`eventsInSequenceRange(1, committedSequence)`, the discipline every other
journal scan in the hosted session already uses. The witness commits 110
observations so the notification lands beyond the default page, asserts
the log really exceeds it, and was verified red under the inverse edit
(the settle returns 0 and the input stays owed).

* fix(managed-agent): stop a full task journal from wedging record commits

The task event journal is a derived, bounded feed: its backlog bound
refuses the next event past capacity rather than silently discarding it,
and an unarchived output event pins the retention floor so the automatic
pass cannot expire anything. The record commit called `appendStateChange`
inline, so once a task's journal was pinned and full, that refusal
propagated out of the extension-record commit — every later revision of
that task failed, which would take the Shell and Monitor record lines
with it.

The record row is the authoritative state and is already written when the
journal append runs, so a refusing journal now degrades the feed (warned
with tenant, session, task, revision and the refusal code) and never the
commit. The store still throws for callers that can apply backpressure —
which is where the output producer's retry belongs.

The witness drives the real wedge: it pins the floor with an unarchived
output event, fills the journal to the bound, asserts the next append
refuses with `managed_task_event_backlog_full`, then commits a record
revision whose view changes and asserts it lands. Verified red under the
inverse edit, where the commit dies with "The task's event journal holds
512 retained events".

* refactor(managed-agent): drop the shell view's never-set error field

The maintenance view declared an optional `error` object, the Java
response validator carried a branch for it, and no producer on either
side ever set one: the worker answers `running`, `exited` or `unknown`,
and a route-level failure travels as an HTTP status with its own code
body. A field that is declared and validated but never written is a dead
switch, so it comes out of the TypeScript view and the Java field set —
which now also refuses an answer that carries it, instead of silently
accepting a shape nothing produces.

* fix(managed-agent): let a replayed output advance pass instead of refusing

The forward-only check refused any advance whose revision was not
strictly newer, which included a redelivery of the very reference the
record already holds. The publisher guards against that in memory, but
the guard does not survive a worker restart, so a replayed advance could
throw inside the background exit leg and wedge a Shell or Monitor that
had already ended.

Both funnels now treat an identical reference as the no-op the deep-equal
skip already owns, and still refuse any other reference that is not
strictly newer. For a Shell the replay commits one liveness step, since
the live run line alternates by design; for a Monitor it commits nothing.
Witnesses pin both, beside the existing refusal witness.

* chore(managed-agent): keep the task-events lane report out of the PR diff

A one-off construction report does not belong at the repository root:
the project's directory rules put working artifacts under the ignored
.qwen tree and keep the tracked docs/ for design and plans. The report
is preserved at .qwen/investigations/2026-10-04-task-events-lane-report.md,
and its Flyway numbering note is superseded by the V40 renumber the
merge commit records.

* fix(managed-agent): quiet the wake pump's consume guard at a blocked owner

runTurn's recovery-blocked path marks the Session blocked and returns
'settled' without the turn ever existing, so the input stays owed — the
accurate answer the whole design gives for a parked Session. The pump's
after-settle check read that as a consumption bug and threw, which both
marked the already-blocked Session again and logged a spurious error on
every such observation. It now stops at an accurate blocked state and
still throws when the input simply did not move: that is the programming
error the guard exists for.

* fix(managed-agent): meld a monitor watch through one terminal step

Round-6 found the shape the Monitor's final leg really had: prepare's
open-then-advance could never survive the start order (output before a
start receipt is a parse-level refusal), the manifest still advanced
through the Shell funnel whatever the capture's recordDomain said (a
Monitor never builds its own Shell record), and the finalize settled the
record ahead of the loop's own exit chain, so a watch's last window died
in denial between two async hops — with its swallow queued to silence it.

Three changes, each with its own witness that goes red under the inverse
edit:

- The manifest advance reads the record's start receipt first and only
  advances past it; a missing record still fails loudly instead of the
  old "no record to revise" at whatever funnel happened to be there.
- It routes by recordDomain: monitor captures land on the Monitor
  funnel, Shell captures on the Shell's.
- The loop's onExit now returns a chain carrying the exit's own
  flush → settle in order, and the publisher awaits it rather than
  settling itself — a final window can no longer be dropped ahead of
  the settle, and a commit failure of the chain surfaces inside the
  finalize instead of returning an empty observation state silently.

The end-to-end witness drives the production shape on a real authority,
orchestrator and Publisher over HTTP: prepare ahead of attach, output
advancing only after the receipt, then one terminal step holding
observationSequence 1 with the settled mark behind it.

* fix(managed-agent): settle the process row when the Runtime proves no start

Round-6 found the permanent wedge: the `:process` row admits PREPARED
ahead of dispatch, but an admitted-refusal answer
(`{executionStatus: "not_started", capture: null}` — cgroup root missing,
quota exceeded, and friends) settled only its invocation. The process
was never registered by the Runtime, so every later shell-status can
only be unknown, the exit-only settle path never applies, and every
release carries the hold forever.

A proven never-started is proof: when the invocation settles with
executionStatus not_started, the sibling row now settles with the same
not_started mark, its admitted refusal as the evidence. Busy computed
after that never counts a process that could never have existed; an
unproven-lost connection still wedges exactly as the declaration says.
The witness drives the admitted refusal shape end-to-end: the process
row lands SETTLED with not_started, and release passes without the busy
answer. Verified red under the inverse edit, which reproduces the
report's `expected: <SETTLED> but was: <PREPARED>` verbatim.

* fix(managed-agent): attribute wake receipts to their own turn and stop re-driving it

Two wake-path defects from the round-6 audit family, plus the monitor
registry the maintenance route actually shares:

- `recoverShellReceipts` exempted monitor inputs so thoroughly that a
  wake turn's receipts could never be attributed: any `tool.receipt`
  journaled under a wake turn would attach to the previous prompt or be
  skipped. Receipt attribution now advances on every accepted input,
  while the pending set itself still exempts monitor inputs as before —
  nothing reframes them as parked Turns.
- A wake turn that started, journaled records and died was re-executed
  text-only on reopen, minting a second user record and leaving an
  unanswered first call in the model's history. `wakeHasPriorAttempt`
  names the condition simply; the pump now parks such turns accurately
  blocked for the recovery fleet instead of redriving.
- The worker constructed its executor with no monitor registry, so the
  executor built a private one while the monitor maintenance routes got
  another empty one — every status and stop answered unknown. One
  registry is shared by executor and routes now, mirroring what the
  Shell path already does.

* fix(managed-agent): read the task event floor after the events, not before

Two stores-and-reads of an event page ran as separate autocommit queries:
the retention floor was checked before `events.read`, so an expiry mid-
read produced a 200 that silently skipped events under an unchanged
nextCursor, the exact gap the contract's `409 cursor_expired` exists to
answer. The replayable-events family already documents the sound order:
events first, floor after — a monotonic floor that is not above the
cursor after the read was not above it during the read either.

The event route now decodes the cursor, reads the page, and only then
answers `cursor_expired` when the post-read floor has crossed it;
past-the-floor omitted-after pages still start at whatever floor stands
then, and an empty page's cursor derives from that post-read floor. Two
witnesses cover both routes: real-journal paging and an explicit cursor
that races the floor's progression with the floor rising at the read
itself — verified red under the inverse order ("Expecting code to raise
a throwable"), exactly the silent-gap shape the race once returned.

* fix(managed-agent): move the task artifact references by compare-and-set

The ref list was read, then rebuilt in Java and upserted — two writers
ona task both satisfying the same bound could each write a single
addition of the other's list and silently age one reference out, the
exact eviction the hundred-ref bound exists to refuse. The append now
moves only by compare-and-set against the list it just read: the first
writer's update lands, the second sees zero rows changed and retries
against what actually stands, and the missing-row path retries a raced
first insert against the comparable shape instead of overwriting it.
Exhausted retries refuse, they do not lose. Witnesses keep every
reference in order under back-to-back appends and prove a stale snapshot
can no longer rewrite the list (the exact clobber the old write shape
carried). Duplicates and the hundred-ref bound still refuse with the
same error, never a silent eviction.

* test(managed-agent): straighten three small voice and fixture drifts

- The record-level integer refusal message now pairs with the Java
  store's identical boundary, so a cross-language failure reads alike
  wherever it was caught ("must be an integer from X to Y").
- The design docs (both languages) name `V40` and the real trigger of a
  `state_changed` event: a task view change, never every record revision.
- The monitor tool declaration says what `max_events` actually counts —
  one-second-or-longer debounce windows, not lines — so a quota set
  from line expectations lands right; `idle_timeout_ms` names the
  quietness it stops on.
- The two task-list contract checks make themselves valid before the
  unknown field lands, so the refusal measures the schema rule and
  never borrows the cursor rule's answer.

* fix(managed-agent): admit named cgroup units and keep fast-exit evidence

- The unit-name guard compared against the empty string, which every
  name contains: every named `create` refused and every `attach`
  answered undefined, so no supervised command could ever start under
  its stable identity. The intended NUL check is now an escape, and the
  test stops embedding a NUL byte itself so the file diffs as text.
- The launcher announces `joined` on fd 3 once the kernel accepts its
  membership write, and the prober takes that marker beside the
  `cgroup.procs` listing: a command that joined, ran and exited inside
  the first listing window was previously journaled as a missing cgroup
  delegation — a refusal with invented evidence.
- Exit evidence survives the membership window: the getters consult the
  child's own `exitCode`/`signalCode` when the exit ran before the
  listener, and `terminate` answers through the same read, so a fast
  process can no longer leak its Runtime hold into a never-resolving
  `once('exit')`. A spawn-time `error` handler covers the window before
  the process object exists.

Each witness ran red under its inverse: the named-unit refusal reaches
`resolveRoot` only after the fix, the marker acceptance with an
unreadable process list refused after two seconds without it, and the
in-window exit lost its evidence without the fallback.

* fix(managed-agent): account capture pages from the batch they publish

`flushPage` counted and cleared the live pending array after its
publish await, so a segment landing mid-window from the other stream's
queue was priced into a page that never carried it and then discarded
with the array — a silently self-inconsistent manifest whose
`segmentCount` exceeds the segments the page holds. The flush now
splices the batch up front and derives every field from that batch.

The both-stream flush ahead of every manifest revision stays: the
manifest validator requires each descriptor's pages to add up to its
stored byte length. A segment arriving between the splice and the
revision now wedges the capture loudly (`storage_failed`) instead of
producing a manifest that claims bytes its pages do not hold; the
witness pins the fail-stop, and the old loss shape ran red under the
inverse (`brokenReason` stayed null while the fold silently completed).

* fix(managed-agent): answer the detached capture family in the validator

The broker's hand-rolled v3 capture validator predated the `detached`
capture status this branch adds to the shared envelope: three clauses
rejected the background start's settle, and this branch's own consumer
(`isDetachedCapture`) was unreachable over the real transport even
though the shared schema, the projector and the TS rules all moved.

The validator now mirrors the TS rules — reasonless like `complete`,
manifest exempt and required to be null — plus the two refusal
directions (a manifest or a reason present while detached). Witnesses:
a detached settle round-trips with its blocked delivery, and both
refusals throw. The accept case ran red under the inverse ("Managed
Runtime status capture is invalid.").

* fix(managed-agent): harden the admission, capture and hold edges of the worker

Four slice-C findings, one family each:

- An admission that fails after its dispatch entry is journaled
  (prepare rejection, or the close recheck landing mid-flight) parked
  the entry `prepared` forever — holding the Session's Runtime, wedging
  close-admission, answering re-issues with a phantom wait. Both
  background paths now settle the refusal `not_started`, like every
  other pre-effect refusal already does.
- The Monitor hold family was half-wired: holds consulted Shells only,
 …
yiliang114 pushed a commit to everyoneexe/qwen-code that referenced this pull request Oct 6, 2026
…LM#13300 (QwenLM#13355)

* docs(managed-agent): design H3 background Shell and Monitor runtime

* feat(core): add managed child_run (kind shell) record body for H3

* feat(managed-agent): mirror child_run shell record body in Java store

* test(core): commit and rebuild child_run shell chains through the session authority

* feat(managed-agent): close child_run reference closure and add stop-request draining

* feat(core): add named cgroup unit attach and managed child-run supervisor

* feat(cli): admit v3 background shell under the supervised worker registry

* feat(core): add open-ended shell stream capture with manifest revisions

* fix(core): type-narrow stream capture finalize and refused publish double

* fix(cli): close background shell type gaps after the main merge

* fix(managed-agent): answer the R1 review round on the child_run contract

* fix(managed-agent): close H0c critical follow-ups R3-1/R3-2/R3-3

PR 1 of #13300, fixing the three Critical findings from the round-3
review of #12855.

R3-1: the Broker-record execution mapping now reads the record's
dispatchGeneration: SETTLED/cancelled with generation 0 (the record's
own proof it was never claimed, per ToolExecutionRecord) maps to
not_started_proven instead of settled, removing the illegal
intent -> settled shape that left a Monitor cancelled before dispatch
with no committable settling revision. A shared brokerExecutionCases
row (settled-cancelled-unclaimed) replays it in both languages, the
TS divergence pin gains the second wire-reading divergence, and
Decision 10 of the authority design is reconciled with the H0b
record-contract doc in both language pairs.

R3-2: the store now runs the event-level half of the envelope checks
for every event line at commit time: closed envelope with an optional
subject, version, declared sequence, well-formed event id outside the
reserved <domain>:<n> namespace, this Session's closed key, valid time,
and kind from a mirrored EVENT_KINDS vocabulary; every domain.committed
payload is validated whether or not a body is registered (reusing the
pinned DOMAINS); unknown-subtype lines follow the scanner's "after the
Managed header" condition; a transaction carries at most one Stage H
record; and per-kind byte caps (maxEventBytes, maxCommitMarkerBytes)
are pinned in the shared limits contract and enforced per line kind.
All refusals answer the existing 409 so the commit rolls back.
Replay commits return before any of this runs, as before.

R3-3: task.updated rides its own outbox (managed_agent_task_event,
V35) written in the same commit transaction and keyed by a unique
source key, drained later by the task-events slice; managed_agent_event
stays turn and lifecycle events, so an announcement between two text
deltas can no longer split a message Part under the frozen projection
version. The route description is scoped to what the server does
(a Session being deleted announces nothing while its tasks stay
readable), and the design's replay-idempotency sentence is corrected.

Test helpers that wrote journals a real authority could not read back
(ActionJournal's reserved event ids, bare event/marker lines in the
publication and session-store integration suites) now write
well-formed lines. Mutations of every new check were verified red
against the suites.

* fix(managed-agent): repair the CI-only cgroup root and env-guard failures

How: wrap the delegated-root probes so a missing or unreadable root answers
the documented isolation error on Linux too (macOS refused at the platform
check first, which is why only CI saw the raw ENOENT), and document the
background Shell environment allowlist in the process.env guard.

Why: the Test lane on 4fb9f0a9c8 went red on exactly these two items; both
are this branch's own changes, not flake.

Test: hook-command-cgroup and process-env-guard suites green; tsc clean.

* feat(managed-agent): inject the child-run supervisor at worker boot

How: registerManagedContextRoutes builds a ManagedChildRunSupervisor from
the delegated cgroup root the Hook commands already use
(QWEN_MANAGED_HOOK_CGROUP_ROOT) and passes it as the executor's fifth
argument; a boot without the delegation keeps the executor's committed
refusal instead of failing to boot.

Why: executeV3Background landed in the previous increment with the
supervisor parameter unwired, so every real-stack background start answered
the committed no-supervisor refusal.

Test: new wiring case asserts the supervisor is injected exactly when the
root is set; managed-context-worker, process-env-guard and
managed-background-shell suites green (775); tsc clean.

* docs(managed-agent): pin the two-row ledger shape for background processes

How: the runtime-ownership section now says the ledger holds two rows —
the foreground-length start invocation settling with the handle, and a
second execution admitted at start that carries no model result and stays
non-terminal until physical exit evidence.

Why: the Java ledger writes a result exactly when an execution settles, so
a single row that is both handle-delivered and active cannot exist; the
two-row shape keeps the settled-if-and-only-if-result invariant while
preserving every cited behavior (Runtime holds, evidence-only settle).

* feat(managed-agent): commit child_run shell lines from the hosted side

How: HostedChildRunSession funnels every record write through one
serialized, replay-safe commitExtensionRecord path — admit first with the
start call's args as commandRef, dispatch and the set-once
managed-runtime-receipt after the physical start, outputRef advance-only,
and settlement only on proven exit evidence, a proven failure, or an
honored stop request. Live revisions alternate the run line (no
self-loops); the authority freeze refuses anything after a terminal
revision.

Why: the dual path puts product records on the hosted authority, but
nothing commits child_run from the hosted side yet — the worker-owned
registry intentionally does not touch records.

Test: four authority-level cases — full chain with projection, stop
request to cancelled, pre-start failure frozen on not_started_proven, and
serialization with deep-equal skips.

* feat(managed-agent): add the detached capture family to Tool v3 results

How: a background Shell start now settles success with a capture object
whose status is 'detached' — no reason, no manifest — because the live
output streams through the child_run record's growing manifest rather
than the result; the shared schema pins both invariants, the TS parser
mirrors them, the shared fixtures gain one valid and two invalid envelope
cases replayed on both sides, and the Java projector accepts the missing
manifest for unavailable or detached captures. Admission refusals settle
not_started with a null capture, the shape the durable unstarted family
already owns.

Why: every Session tool.receipt event lands in the Java delivery
projection, which requires a capture object and a committed-if-manifest
pairing for anything that started — a success result with a null capture
fails that projection and corrupts replay.

Test: core contract suite 519 (three new envelope cases), TS serve suites
33 + turn/harness 323 + background 6, Java publication contract and
projector suites 8; tsc clean on both packages.

* fix(managed-agent): answer the R2 review round on worker, capture and record lines

How — the four Criticals first: the stream capture publishes an ended
stream's revision only after its seal decision, so a sealed-or-incomplete
descriptor never changes again (R2-3); finalize's settle path degrades to
an unavailable capture instead of throwing past a writability failure, and
the last published revision stands like the worker-cut cap (R2-5); the
child_run body now refuses a settled or cancelled run that is not settled
execution, a failed run off its two ending lines, start_failed outside
not_started_proven, and a process-level failure without a settled
execution, on both languages with new witnesses replayed from the shared
fixtures (R2-6, R1-30); and the supervision suite skips win32 instead of
spawning a shell that cannot exist (R2-7).

How — the Suggestions land as: supervisor start now proves membership by
cgroup.procs with an fd-3 status channel, failing closed as isolation
(R2-37); the bounded EOF wait caps daemon-inherited pipes instead of
hanging, and a spawn-time error drains the entry (R2-1/R2-14); registry
hasHolds requires its session and setProcessResult carries the full result
shape (R2-13/R2-15); create() validates its caller-named unit like attach()
(R2-26); the executor checks the journal before any effect, mirrors the
closing/is_active re-check after prepare, gates the ninth live Shell as a
committed quota refusal, and keeps cgroup membership out of an empty
QWEN_MANAGED_HOOK_CGROUP_ROOT (R2-23/R2-19/R2-24/R2-18); the spawn uses
the configured shell and the session-context environment (R2-21/R2-20);
invalid fixture cases each pin their refusing clause on both languages
(R2-36); projection fixtures pin the draining precedence rows and the Java
replay reads stopRequested (R2-28/R2-38); the draining Javadoc states the
rule the code implements (R2-39); the duplicated ref helper is gone (R2-40);
the stream-capture suite is typed and gains the seal-before-publish,
page-cursor, late-failure and settle-degrade witnesses (R2-29/30/31/32);
and the design doc names the cgroup switch decision, the cursor's single
ordinal space, the unconditional host-scope evidence gate and the zh
stopped-object (R2-8/9/10/11/12).

Test: core suites 788, cli serve suites 776 + background 11, Java record/
projection/store suites 24; tsc clean.

* feat(managed-agent): thread the child-run orchestrator and the detached receipt family

How: HostedSession gains its session-scoped childRuns orchestrator next to
hooks/mcp, the tool turn receives it stored-ahead of the admission branch,
the reopen verifier and the recovery replay both accept the third durable
receipt family — a blocked delivery with a detached capture, beside the
complete and not-started ones — with the publication store correctly left
out of its proof, since its durable truth is the child_run record.

Why: the H3 background start settles with a handle whose output lives on
the record's growing manifest; without the third family, any Session
holding a background receipt could never reopen or be replayed.

Test: recovery-session suite 52 with two new detached-family cases;
harness-session and tool-turn suites green; tsc and lint clean.

* test(cli): type the background-shell capture double

How: drop the as-unknown cast on the FakeSink return and implement the
failCapture member the cast had been hiding (R2-29's cli location).

Why: an unchecked double can drift from the interface it claims to match —
and it already had.

Test: background-shell suite 10/10; tsc clean.

* feat(managed-agent): admit background Shell starts from the hosted tool turn

How: the turn replaces its two deliberate is_background refusals — exactly
when the Session owns its child_run orchestrator and the domain is
enabled — with orchestration that mirrors the shared line: the record
intent precedes every physical effect, dispatch_started lands the moment
the checkpoint commits, a settled detached handle attaches the physical
start to the record from the same facts and lands as the third durable
receipt family (blocked delivery with a null manifest, replay-validated
exactly like the not-started family), and a proven-unstarted refusal
settles start_failed on not_started_proven while riding the existing
unstarted family unchanged. Malformed is_background values keep their old
refusal, and the family stays out while the domain is disabled — the same
probe governs both paths.

Why: the background start is the first record-bearing, user-visible H3
effect, and without a hosted history family for the handle a Session that
ran one could never reopen or be replayed.

Test: three new authority-level turn cases — admitted detached flow with
admit/dispatch/attach order and the blocked receipt family, a
proven-unstarted refuse closing NOT_STARTED, and the disabled-domain
refusal preserving its exact text; turn suite 146, plus harness-session,
recovery, background and child-run suites 246; tsc clean.

* feat(managed-agent): admit background Tool v3 dispatches and acknowledge their detached settle

How: the v3 start admission flips from "foreground Shell only" to
"Shell only" — background dispatches now begin the same way; the poll
loop accepts the settled detached family through the status answer (the
only durable settle such a result can ever produce, since nothing is
published for a background handle), and the acknowledgement path
canonizes the settled handle envelope: a matching blocked receipt with a
null manifest and no history revision forwards exactly itself, without
consulting the publication receipt.

Why: without the detached branch, a background start settled a queued
handle only to mark the execution UNKNOWN after the 30-minute
foreground-shaped deadline — every background dispatch would wedge the
Runtime on its first call, and its acknowledgement would die on the
missing publication it was never going to have.

Test: new witness drives the detached start end-to-end — reservation,
v3 dispatch, settled handle via status, and ack with a null-manifest
blocked receipt, with the publication receipt path provably untouched;
service suite 139/139, module suite 600/600.

* feat(managed-agent): answer shell-status and shell-terminate from the worker registry

How: a new private maintenance route sibling to ManagedHookProtocol
(/internal/managed-runtime/v3/shells) backed by the process registry —
status answers running or unknown (never claimed without evidence),
terminate drains with the supervisor's own rules and answers exited with
its proof, and an end that cannot be proven is unknown with the hold
still registered. The unit name derives from the runtime invocation
identity through the shared shellUnitNameOf, so Java, the hosted turn
and the worker derive it identically; the executor takes the registry as
a constructor injection so the route and the journal share one.

Why: every later maintenance verb — recovery reconciliation, the ordered
close drain, the reconcile-settle — has to ask the only physical owner
these questions; answering them from anything but the registry would
invent truth.

Test: five worker-side cases — closed-key refusals, unknown for
unowned, running scoped to the registered Session, terminate-with-evidence
answering exited, and an unproven terminate answering unknown with the
hold intact; shell, context-worker, v3-route and child-run suites 784;
tsc clean.

* feat(managed-agent): admit the background process row and answer its maintenance protocol

How: the start dispatching of a background Shell admits a second ledger
execution beside the invocation row — the new process row carries no
model result and stays PREPARED (non-terminal) until physical proof, so
hasActiveBy* holds the Runtime exactly while the process lives. The
maintenance protocol mirrors ManagedHookProtocol on its own route:
shell-status and shell-terminate with a targetOperationId (both recovery
kinds), wire-validated on both transports and accepted through the same
workspace-generation and ownership gates. observeBackgroundProcess asks
the physical owner and settles the process row with proven evidence —
settle through the new repository verb settlePrepared, which mirrors the
PREPARED-requestCancel shape on both repositories; anything unproven
leaves the row active with its hold, never a claimed end.

Why: the ledger writes a result exactly when an execution settles, so a
single background execution carrying its handle at start cannot exist;
two rows keep the settled-if-and-only-if-result invariant while the
Runtime holds precisely, and the maintenance route gives every later
verb the worker's physical answer instead of an invented one.

Test: two new witnesses — a background start that admits the process row
(the invocation SETTLED, the process PREPARED, the Runtime busy until
status arrives, then exit evidence settling it and release succeeding),
and an unproven status keeping its hold; the runtime also gained the row
identity invariant fix and the settle-through-settlePrepared shape;
service suite 141/141, module 603/603.

* fix(managed-agent): answer maintenance views under the target identity

How: the shell-status/shell-terminate view now carries the target
operation's identity, exactly as the Hook lookup family does — the
requester asks about the process, never about the asking call, and the
Java wire validator's expectedId rule requires it.

Why: this repair was encoded in the route design but landed uncommitted
after ②e-a; the self-audit caught it before it could ride a later batch.

* feat(managed-agent): run the background exit leg through the record

How: the hosted publisher gains the background family — its registration
marker admits the open-ended capture through the session-level admission
(proven start on the record, never the model call), the open manifest
publishes at prepare, every awaited write or finish forwards the newest
manifest revision to the record's outputRef, and the exit finalizes with
the same evidence object settling the record. The client acknowledged
receipt keeps the detached-blocked shape; drain skips background stores
until the ordered close drain wires them; the register guard only binds
the foreground model-call family.

Why: the start handle already committed, so a background Shell's whole
physical truth is output and exit evidence — anything that reached the
record another way would let a second tool.receipt contradict its own
line, and an open-ended manifest published anywhere but the Session's
own store is an unreadable closure for the store-side commit.

Test: a full end-to-end authority case — admit-dispatch-attach, bytes
through the real rpc, outputRef advancing live, the final sealed manifest
reading the same exit (code 3, both streams sealed, manifest complete),
the record settled with exactly that evidence, and zero tool.receipt
events for the outcome; publisher, turn and child-run suites 174; tsc
clean.

* feat(managed-agent): close monitor_run revisions over their cited Session resources

The H0c closure rule — every ref a record cites must exist in the
Session's resource store at commit time — now covers monitor_run on both
stores: the authority reads commandRef, startReceiptRef, outputRef and
lastObservationRef through verifyExtensionResources, and the Java
applyRevision gains the same four-field branch beside child_run. This is
the precondition for enabling the monitor domains: without it a monitor
revision could anchor its chain to resources nobody committed, and the
replay would project a watch whose evidence was never verified.

Both fixture rigs stop citing fictional resource identities. The
authority test publishes its chain refs as real bodies once per harness
and mirrors each fixture-cited ref as a real same-kind placeholder of
the stated length, memoized so chain identity stays fixed across
revisions. The journal test harness does the same at its generic
request entry point, rewriting each cited ref to the placeholder's real
digest and appending the mirrored resources, but only for strictly
well-formed JSON monitor bodies — a body with trailing content keeps the
raw path so the store's grammar refusal still fires first over the
closure's.

Each side gains a witness that pins the closure itself: a monitor run
citing a resource outside its commit is refused with
managed_session_resource_missing, and each witness was observed red with
its branch disabled before landing.

* feat(managed-agent): gate monitor_run behind the managed-session/2 reader

The H3 design left the reader-gating mechanism to the implementation;
the journal settles it. Headers are immutable — the scanner refuses a
repeated managed_session_header_v1 line and any unknown subtype after
the header — so no transaction can legally raise a Session's
requirement in-journal, and a superseding header line would change the
journal format in both languages for a narrow mixed-version window.
New Sessions therefore stamp minimumReader managed-session/2 at
creation, and every path that commits a monitor_run record —
commitDomainRecord, commitExtensionRecord and the generic append guard
— additionally refuses when the Session's header names less. Readers
still accept requirements up to their own version, so pre-H3 Sessions
stay openable everywhere and simply can never receive monitor_run.

child_run is deliberately not header-gated: a managed-session/1 reader
already parses its name, and the asymmetric hazard lives on the server
side, which the established deploy-server-first sequencing covers.

The shared Stage H golden is regenerated for the v2 genesis bytes and
the Java integration replay passes against it unchanged. New witnesses:
new Sessions stamp v2, a v1 Session refuses monitor_run with the
precise error (observed red with the gate removed), and both header
sides of the version comparison are pinned in the records tests. The
design doc records the decision and the now-paid closure precondition
in both languages.

* feat(managed-agent): funnel monitor_run revisions through a hosted orchestrator

The mirror of HostedChildRunSession for the Monitor domain. The
managed-runtime worker will own the watch loop, but the dual path puts
every product record on the hosted authority, so HostedMonitorSession
commits the record line as watch facts arrive: revision 1 with the
intent before any side effect, dispatch onto a Runtime binding, the
set-once start receipt after the watch starts, observation revisions
only while attached with a watermark that never goes back, output
advancement only forward, and the terminal shapes — settled by
exited/max_events/idle_timeout, failed by a start failure on
not_started_proven or a settled watch failure, cancelled by a stop
request whose notification watermark may still advance after.

Writes serialize per Session and replay-safe by command id with
deep-equal skip, exactly like the Shell funnel. The suite drives the
whole line through a real authority: projection at admission and after
settle, max_events refused before its quota and accepted at it, a
second start receipt refused, an observation before attach refused,
and an idempotent output advance committing nothing. Quota
(enforcement), the debounce-floor observation loop with notification
composition, and rebuild-after-loss stay with the runner and route
increments that follow.

* feat(managed-agent): drive monitor observations through a debounced loop

The loop half of the Monitor runner, on top of the hosted funnel: an
admitted Monitor dispatches, starts an injected watch executor and runs
its own time. Stdout lines aggregate into one observation revision per
debounce window with a one-second floor, so the document's reopen bound
holds for any debounce a watch asks for; the run settles itself on the
contract's own terminal conditions — the observation quota, silence
past idle_timeout, a natural exit with its buffered lines flushed
first — or is stopped, which terminates the watch and lands cancelled.

The executor contract is typed for the worker's cgroup owner (onLine,
onExit after the start resolves, a terminate handle); the suite drives
the whole line through a real authority with a fake executor and a
manual clock: windowed aggregation and its floor, quota settlement and
watch termination exactly at max_events, idle settlement re-armed by
each accepted observation, exit flush order, start and mid-run failure
shapes, and a stop that ignores late lines. Every cross-boundary
settle — timer or executor callback — is counted by the loop's done
handle, which the close drain will await; window flushes count too, so
done never reports idle while an observation commit is in flight.
Notification composition and its wake stay with the next increment;
quota enforcement and rebuild-after-loss stay with the route slice.

* feat(managed-agent): notify from monitor observations in the same transaction

The H0c machinery now runs for Monitors: every accepted observation
commits its revision together with a notification input and the wake
the authority generates for it, and the run's notifiedThrough watermark
advances to exactly that observation's sequence. The loop composes the
input — monitor-1:notify:N ids, the window's joined lines as the
content a woken turn will read, an empty admission, no deadline — and
the funnel threads it through commitExtensionRecord, so a recovery
replay re-runs nothing the watermark already covers.

Dedupe semantics are pinned at both levels: the funnel advances the
watermark only when a notification rides the revision (an observation
without one keeps it where it was), and the loop suite shows the three
events — domain.committed, input.accepted, wake.requested carrying the
input's accepted id as its source — landing in one transaction with the
joined lines as its content. Consumption of that wake by the hosted
scheduler, and settling the input when no turn can take it, come next
(open question 6 of H0c).

* fix(managed-agent): keep a finished background Shell's receipt answerable

H7 of the real-stack rounds: the Broker settles its :process row only
when shell-status answers exited, but the worker's registry deleted an
entry the moment the Shell ended — a natural exit could never be
answered again, and the ledger row would hold the Runtime against
release and close forever. The registry now retains each completed
unit's receipt until the worker ends: the hold drops exactly as before,
but shell-status and shell-terminate answer a proven end from the
retained evidence, idempotently, for the owning Session scope only. An
end without evidence keeps nothing and answers unknown as before; the
existing live flows are untouched. Suite: natural exit keeps answering
exited with its evidence for both kinds after the hold dropped, the
wrong Session still hears unknown, and an evidence-less end stays
unknown. The close-drain/reconcile callers that consume this, and the
Java-side settle-before-busy on workspace release, land with the next
batch.

* fix(managed-agent): renumber the task-event outbox to V36 and narrow the commit-time guarantee

#13301 landed V35__managed_session_tool_profile.sql on main, so this
branch's V35__managed_task_event.sql now shares its version: the merged
tree fails at startup with "Found more than one migration with version
35". Move the outbox migration to V36 and update the authority design
doc in both languages.

The commit-time validator mirrors the checks the authority applies to
every line whatever its kind (envelope, vocabularies, domain checks,
reserved ids, byte caps), not the per-kind payload schemas, subject
rules or commit-marker digests, which stay the authority's contract.
Reword the apply() Javadoc, the test comment and the design doc so they
no longer promise that no commit can brick the Session, and say what a
stored reserved id breaks: it collides with the id of that domain's next
Stage H record rather than failing the next open.

Pin the event id and operation id shape checks with two negative cases;
disabling either check previously left every test green.

* fix(managed-agent): settle provable background exits before release's busy check

The second half of H7. The worker now keeps a finished Shell's evidence
answerable, so the Broker has someone to ask — but observeBackgroundProcess
still had no production caller, so a naturally exited Shell left its
:process row non-terminal and every release wedged on
runtime_session_busy. releaseSession now sweeps the background process
rows this Broker admitted for the Session before it computes busy: each
unproven row asks its physical owner once through the same
shell-status → settlePrepared leg observeBackgroundProcess already owns,
a proven exit settles with its evidence, and any lookup that fails or
cannot prove an end simply keeps the row so busy stays the accurate
answer. The sweep tracks per-context admissions because the
reconciliation scan deliberately excludes PREPARED rows, and mutates the
admission set under the context lock.

Witnesses: releaseSettlesAnExitedBackgroundProcessBeforeBusy proves an
exited row settles with its evidence and the release succeeds (observed
red with the sweep bypassed); the pre-existing
unprovenProcessStatusKeepsItsHold still proves an answer without proof
keeps the row and the busy refusal. Full module 604/604. The workspace
close drain still terminates shells in order rather than sweeping — its
ordered stop belongs to the close-drain slice, and these two verbs are
exactly what it will consume.

* fix(managed-agent): check the journaled receipt before attaching a background Shell's start

H6 of the real-stack rounds: acceptBackgroundShell called childRuns.attach
before its tool.receipt replay check. A retried accept — the leg recovering
pending results re-drives — would mint a second start receipt on the same
identity before reaching the replay path, and the record's set-once rule
would refuse it as a rerun shape, so a recoverable retry crashed as a
conflict. The lookup now runs first: a journaled receipt replays its
validated history with no attach at all, and a fresh accept attaches only
when this identity still lacks a start receipt — the record id IS the
execution identity, so an existing one is always our own crash-window
attach between record and journal, which the retry now tolerates instead
of minting a rerun. Witnesses: a retried accept answers from the journal
with exactly one attach and one tool.receipt line, and an
already-attached-but-unjournaled record completes its accept without
attaching again. The full hosted turn suite stays 148/148.

* fix(managed-agent): backpressure the background Shell output pipes

G5 of the real-stack rounds: the executor fed the bounded capture with a
fire-and-forget sink.write — every chunk queued behind asynchronous
publication, so a fast producer flooded memory (a 256 MiB run peaked near
315 MiB RSS at a 20 MiB/s store). The feed now counts bytes in flight and
pauses stdout and stderr when the drain falls sixteen MiB behind, resuming
at four MiB — enough to cover a store several seconds slower than the
process produces without stalling a steady producer against the floor.
Completion or rejection of a write both release their bytes, so a broken
capture can never leave the pipes paused. The witness drives a controlled
sink: seventeen MiB of chunks pauses a stream exactly once, draining past
the watermark but above it resumes nothing, and draining the rest resumes
exactly once. The outputRef-live-advance half of G5 belongs to the
turn-publisher's promotion to activation scope, sized with the close-drain
slice, and stays queued there.

* feat(managed-agent): promote the Shell publisher to Session scope

The ②d-out debt, and with it the live-advance half of G5. The Shell
publisher was a per-turn object: its server died with the turn, while
background Shell traffic flows across turns — between turns (or after a
close) the worker's retained descriptor pointed at a dead endpoint, so
no live output advance and no exit leg could ever complete past one
turn. The hosted Session now owns one long-lived publisher across
turns (the turn lazily structures it on the shell options and never
closes it; the Session's ordered close in this PR's fifth slice drains
it). The two small wire-level fixes that the promotion exposes:

- The worker's background execute stamps the background marker into its
  prepare request: the closed six-field wire capture has no such field,
  so a remote prepare used to arrive unmarked and die at admission; the
  executor derives the marker from the background input itself.
- register() takes the registering turn's prompt for the foreground
  admission check, so one instance admits the foreground of any turn of
  its Session while a foreground pretending another turn's identity is
  still refused; the background arm keeps its record-based admission.

Witnesses: a single instance admits a background watch plus the
foreground of two turns and refuses a misattributed one; the worker rig
pins the stamped marker on its prepare; the four touched suites stay
184/184 and both package typechecks pass (after a core dist rebuild for
the capture type's optional background marker). The publisher's own
descriptor installs idempotently per turn (identical descriptors are
accepted), so consecutive turns never fight for the worker slot.

* feat(managed-agent): drain a Session's background Shells ahead of release

The close-drain's missing legs. Until now a Session with a live
background Shell could never close: the worker's activation gate refused
release with managed_activation_conflict forever, and the CLI-side close
route skipped the background stores. The ordered sequence the design
asks for now exists end to end:

- The registry gains stopSession: terminate one Session's entries with
  the supervisor's evidence rules (TERM, escalate, empty-check), await
  each completion so its finalization — last manifest, sealed envelope,
  record settle through the session publisher — lands before the caller
  proceeds; an unproven stop keeps its hold for the caller to see.
- The worker's activation release gate drains before it refuses: active
  background work no longer wedges a close that can prove its stops,
  and anything still unproven keeps the exact same conflict.
- The hosted close route now closes the Session-scoped publisher right
  after the broker release returns — which by then has drained the
  Shells and settled their records through that publisher — and the
  drain closes every capture store, background families included.

Witnesses: a registry drain settles one Session's shells with evidence,
both kinds, while leaving another's holds untouched, and an unprovable
terminate under a session drain still preserves its hold; the 766-test
context-worker suite stays green through the release gate's new async
drain branch. What remains deliberately unproven locally: the
full-chain release-with-live-shell drain needs a real cgroup supervisor,
pinned as the Linux head acceptance instead of a mocked supervisor.

* docs(managed-agent): sync the H3 design's status line with landed slices

The status line still claimed publication and orchestration remained
design while most of them landed across this PR's batches; refresh both
language versions with the landed list and keep the remaining-design
list accurate: the Monitor physical side (route, rebuild, cgroup watch)
and wake consumer, both submission enablements, task events, and the
Linux physical acceptance.

* feat(managed-agent): spawn Monitor watches through a cgroup watcher

The physical half of the Monitor runner's executor interface, for the
worker: ManagedMonitorWatcher spawns the watch command under a unit
named for the execution identity, through the same supervisor (its
membership attestation unchanged), splits stdout into the observation
lines the loop buffers — boundaries respected, the Legacy partial-line
cap honored, the tail flushed as the final line at exit — and reports
the physical end exactly once, natural exit versus mid-run failure, so
the loop's settled and watch_failed shapes both hold. stderr rides only
the future output leg, never observations; terminate defers to the
supervisor's TERM-escalate-empty rules. The suite drives a supervisor
double: unit/executable/cwd derivation, chunk-boundary splits with the
final flush, the cap drop, spawn-error and membership-failure shapes,
and terminate delegation.

* fix(managed-agent): accept message.retracted at commit time

#13351 added the message.retracted event kind on main. The hosted text
delta stream writes it when a restarted model attempt replaces a prefix
it already published. The commit-time vocabulary this branch mirrors from
the authority still listed sixteen kinds, so after the merge with main
the store refused every retraction with 409 ("event.kind must be one
of [...]") and left the orphaned deltas in place.

Add the kind to ManagedExtensionRecords.EVENT_KINDS and to the shared
store fixture, so the parity pins in both languages agree again.

* feat(managed-agent): answer monitor status and stop from a worker registry

The monitor half of the maintenance route, mirroring the Shell's: a
worker-side registry of supervised watches that holds each Session's
Runtime while a watch runs, answers an end only with the supervisor's
own evidence, and retains a proven exit's receipt until the worker ends
so maintainers can learn the truth idempotently. The route sibling at
/internal/managed-runtime/v3/monitors serves monitor-status and
monitor-stop with the same closed wire shape, target-identity rule, and
unknown-but-never-unproven discipline as the shells' route; it joins
the worker's route list so the v2 envelope admits it. Registration
arrives with the Monitor start admission later — status/stop already
have consumers in the release drain and the rebuild slice that follows.
This also repairs the watcher test arity that broke the previous push's
build lanes (the mock's 2-arg onExit omission, seen as
error TS2554 (155,34) across every CI lane at that head).

* feat(managed-agent): admit monitor watches into the open-ended capture session

The γ3-α precondition for every later Monitor leg:
LocalShellStreamResultSession's prove-a-live-start admission reads the
capture family's record domain instead of being pinned to child_run —
same Session/binding/background-marker checks, same live-run gate,
same identity captureId, per-domain labels so a refusal names its own
family. The producer side names its domain explicitly (record presence
is never consulted where both domains could in principle compete), so
a monitor watch carries its monitor_run record into exactly the same
bounded stream the background Shell already rides. Witnesses: shell
admission unchanged; monitor admission succeeds under its domain;
wrong-domain proven start refuses; missing record and foreign binding
refuse with their own errors.

* feat(managed-agent): admit Monitor watches through the v3 executor

The workhorse of the Monitor leg. The worker's v3 executor gains its
is_monitor branch beside is_background: journaled admission first, then
refuses (never as a transport error), then the watch spawns through the
ManagedMonitorWatcher under a unit named for the execution identity,
streams stdout into the SAME bounded family (16/4 MiB backpressure on
the pipe, with an explicit refusal if the watcher reports no supervised
unit), and settles with its detached handle while a registry keeps the
Runtime hold. Three watch-level facts are done physically now:

- The monitor registry is a wrapper around the background Shell
  registry, never a copy: supervisors, EOF drains and publisher finishes
  stay in one place, while monitor holds and the four-watch quota bucket
  remain entirely their own. Its maintenance route answers run/exited/
  unknown with identical semantics to shells'.
- The execute path gates on the generated four-live-watch cap with
  `Session already runs 4 Monitor watches.` and additionally on a
  missing cgroup root and a missing capture service.
- Each prepare carries background+monitoring markers, so the hosted
  side of γ3-α knows the monitor_run domain to check.

Suite: 28 tests across the three touched files — the fork settles a
detached start under the right unit and marker, lines land as monitor
output, the fifth watch is refused, and supervisor-less boot refuses
with its dedicated text. The turn admission arm (funnel-driven
admit→dispatch→attach through HostedMonitorSession), the watcher→loop
observation feed, and rebuild still land in the following increments.

* fix(managed-agent): refuse a domain.committed without a textual domain

A domain.committed line whose payload.domain was missing or not a string
read as "no domain" and skipped requireDomainCommitted, so it committed
200 and the authority refused the Session at the next open. Decide that
a line is domain.committed from its kind alone, and take the domain from
the vocabulary check before building the expected record kind.

Also address the rest of the first /review round:
- The commit marker is the only line apply() measures. validateUtf8JsonLines
  already caps every line at MAX_EVENT_BYTES, which it now names.
- The outbox's MAX(sequence_id) stays a plain read. A locking read
  deadlocks concurrent first announcements of different Sessions on MySQL
  8.4 (85 of 480 in a probe). The comments now name the guarantee that
  actually holds: the Session store's journal-head lock is taken before
  the transaction's first plain read.
- Stop claiming that the Session event stream carries only turn and
  lifecycle events. V36 is not applied anywhere yet, so its comment can
  still change.
- Restore the shipped v1.19 OpenAPI sentence and record the move to the
  outbox as v1.30 (info.version 1.30.0).
- Reconcile both design docs: Decision 10's state table and the scope of
  its wire-only absolute, the goals, the public-contract bullet, the
  validation plan, the files list, and the restored caveat that the record
  contract does not catch an unprovable not_started_proven.
- Tests: refuse a non-textual and a missing domain, commit a
  message.retracted line, pin the final "events, then its commit marker"
  guard again, assert the empty outbox in the MySQL IT, and give the
  oversized-resource case bytes at its own sequence.

* feat(managed-agent): admit Monitor watches through the hosted tool turn

The turn-side arm of the Monitor leg. A call named monitor is now
admitted through its own three-state gate (requested/unavailable/ill-
formed, with the Shell-profile Monitor declaration now advertised),
implements the same record-first discipline as background Shells do —
the funnel admit commits monitor_run revision 1 before any side effect,
dispatchStarted follows the durable checkpoint, and the watch attaches
when its v3 execution settles. The granted publication, journal intent,
dispatch and accept forks are widened to carry is_monitor alongside
is_background, and acceptMonitor mirrors acceptBackgroundShell end to
end (not_started → start_failed with the unstarted family; detached →
same blocked receipt family with exactly one tool.receipt line, the H6
journal-verdict-before-attach order, and the blocked ack completing the
row). HostedSession gains a session-scoped HostedMonitorSession beside
childRuns; turn construction forwards it at both sites.

Two existing contract pins move deliberately: the Shell profile's
advertised list now names monitor (the harness declaration test names
it explicitly), and the three publisher-close witnesses now assert the
Session-scoped lifecycle settled by the earlier promotion — no close
at any turn ending (completed/error), the server still answering, close
exactly once at the Session's own delete. Suite: 3 new admission-arm
tests through a full hosted turn (funnel order, clamped defaults with
the Legacy monitor caps, start_failed shape, unavailable text), the
whole turn suite stays 151/151 and the harness session suite 180/180.
The observation fan-out (worker lines → hosted loop), quota witness at
the turn, and rebuild-after-runtime_lost land in the next increments.

* feat(managed-agent): fan a monitor watch's observations into the hosted loop

The observation channel of the Monitor leg. The session publisher now
marks monitor captures with recordDomain monitor_run through γ3-α, fans
each published stdout chunk through the Legacy line-split semantics —
boundaries kept, remainder held across chunks with the 4096-byte
partial cap — into a per-capture observer, and closes it with
onExit(failed) at finalization with the tail flushed first. The turn
gains resumeMonitorWatch: a fresh accept creates one HostedMonitorLoop
per Session report and resumes it from the already-attached record, fed
through a HostedMonitorRemoteExecutor whose start binds those observer
callbacks; a replay starts nothing twice, exactly like it never
re-attaches. Globally, loop timers unref so observation windows never
anchor the process.

Monitor capture finalize settles through the monitor funnel itself
(watch_failed when no evidence, exited via settleQuiet), beside the
shell's unchanged path. The publisher carries monitors next to
childRuns, and the turn's own construction forwards it. Suite: the
admission-arm tests now also pin one Session-level loop starting on a
fresh accept with the record-first flow intact; wide suites of 189
tests pass with cli tsc, prettier and eslint clean. Monitor output
before the observer registers stays durable in the open stream, with
observation starting at attach — noted in the design not as a flaw but
as the attach seam, a line crossing it already lands in the output
Artifact readers see.

* feat(managed-agent): rebuild a read-only Monitor after its Runtime is lost

The record side of the rebuild promise. HostMonitorSession gains the two
verbs the H3 lifecycle needs around runtime loss: blockedRuntimeLost
lands the run in recovery_blocked with the runtime_lost reason and the
lost binding still named, so degrading is what the projection shows
until something decides otherwise; rebuildFromRuntimeLost accepts only
the outcome_unknown/runtime_lost line the H0b rule allows, starts a
fresh watch under a newer Runtime generation with its own receipt, and
keeps the reason for honest projection while observation positions
never moves — the support case never restarts what the watermark holds.
monitorRebuildAllowed routes every candidate command through the Legacy
AST read-only check with its walk, so a side-effecting command stays
blocked accurately. Five witnesses cover the full rebuild chain
(recovery-block, rebuild with receipt remint and generation increase, a
rebuild refused from a live run and from an ended one, and the gate's
read-only/side-effect/empty answers). Binding-replacement detection and
re-drive of the new physical watch come with the recovery slice after
the wake consumer.

* feat(managed-agent): serve the task events routes with their SQL journal

H3 flips listSessionTaskEvents and queryWebShellTaskEvents from planned
to partial on both surfaces.

The Java control plane gains a bounded per-task event journal (V39): one
row per committed task event, written in the Session commit transaction,
with the per-task sequence allocated under the journal head lock in
commit order. Record revisions journal their task view changes at the
H0c announcement point. The durable retention floor lives in a per-task
cursor row, so it survives an empty retained set, restarts and
projection rebuilds; the Artifact visibility barrier pins it behind
unarchived output, and the backlog bound refuses past capacity instead
of silently discarding. Output events carry their per-stream segment
ordinals with overlap/gap refusals, and artifact references stop
fail-loud at the 100 bound. output_cursor/outputCursor lift with the
routes and the task views serve them.

The section 6.1 demonstrations run as store-level contract traffic in
the new journal contract test, the route probes and record gates cover
the flips, and the #12847 C15/C16 instances close alongside. WebShell
types regenerated from the bumped 1.30.0 contract.

* docs(managed-agent): add the task events lane's build report

Flyway numbering findings (V39 taken; V35-V38 claimed by in-flight PRs
on main, V15/V29 Java-only burns untouchable), implementation decisions
against the H3 design text, the CLI/daemon flow-surface finding (none
exists, regeneration is the only surface change), the full test ledger
and the owed items.

* feat(managed-agent): build the monitor wake envelope and the pending-input intake

Two support pieces for the H3 wake consumer (β2-b1a):

- monitorNotificationText wraps one due observation window in the exact
  Legacy task-notification envelope — task-id, optional tool-use-id,
  kind/status/event-count, summary and result — applying the same
  truncateNotificationLabel/stripDisplayControlChars/escapeXml pipeline
  the Legacy Monitor applies, so a turn learns nothing new.
- pendingSessionInputs derives the Session's still-pending inputs from
  the journal alone: an accepted input is consumed exactly when a turn
  settles under its turnId, order-insensitively, so a replayed admission
  after a restart is not re-driven.

* feat(managed-agent): deliver the Legacy notification envelope on monitor wakes

The loop's notification input now publishes the exact task-notification
envelope a Legacy Monitor's wake carried — task-id, tool-use-id, kind,
status, event-count, summary and the window's lines — so the turn the
wake raises reads nothing it has not already read. The description is
the tool call's, falling back to the watched command. A witness pins
the exact wire text of the first observation's input content.

* feat(managed-agent): run the monitor wake through an embedded scheduler

H3 closes the wake loop: an accepted observation already commits its
notification input and wake.requested in one transaction; this slice
makes that wake effective.

The pump re-derives the pending monitor inputs from the journal alone —
nothing is held in memory, so a restart re-derives the set. On an idle
Session the oldest notification runs as an ordinary text turn with the
Legacy envelope as its prompt, and the turn's settle consumes the input.
A busy Session queues in the journal exactly like channel and Goal
inputs; the retry loop redelivers. A Session that is parked, blocked or
mid-recovery keeps its reminders pending and reports them. On the close
path every pending notification settles cancelled without a model turn,
under the turn-result record's own idempotency key, so no wedged
notification ever holds a Session as hosted_turn_recovery_required at
its next open: the parked-turn scans now skip monitor inputs, which the
pump owns exclusively.

A wake turn that dies without settling re-drives on reopen — the input
is consumed exactly when a turn settles under its turnId, and a crashed
turn never settles — matching the queue semantics channel and Goal
inputs already honor.

* fix(managed-agent): keep every new Session on the managed-session/1 reader

Round-5 real-stack verification's cross-version matrix disproved the
v2 stamp's premise: monitor_run has been a parsed, known domain since
#12837 (v0.24.7), and a v1 reader opens a Session holding monitor_run
records without complaint. Stamping minimumReader: managed-session/2 on
every new Session — while monitor_run remains disabled — would make a
rollback or a mixed-version rollout lose access to every Session
created in between, and buys protection no deployed binary needs.

Revert the stamp to managed-session/1 and drop the per-domain header
refusal; the enablement list remains the single domain gate. The stamp
rises only when a change genuinely breaks an older reader mid-scan, and
only one release after a tolerant reader ships. The golden fixture's
genesis header regenerates to v1; the records and authority suites gain
v1-header admittance witnesses in place of the refusal cases; the
design's reader section now records the reversal with its evidence.

* fix(managed-agent): keep monitor output byte-true and advance only forward

Round-5 real-stack verification found three facts about the monitor
output path:

- A capture rebuilt from decoded lines cannot be byte-true: per-chunk
  toString split multi-byte runes into U+FFFD, and line-rounding dropped
  blank lines and the final unterminated line. The watcher now forwards
  raw stdout chunks to the capture on a new onChunk channel, so the
  durable Artifact reproduces the command's stdout exactly, while the
  observation lines keep the Legacy emit-path semantics (a StringDecoder
  holds a split rune until it completes; a blank line consumes no
  observation).
- The watcher registered its exit listener only after the supervisor's
  start resolved, so a fast watch could end unheard and the debounced
  loop never settled. A latched end with a check-phase sweep covers a
  watch that exited before its listeners attached, after the host is
  known to be listening.
- The funnels' advanceOutput promised "only ever to a newer revision"
  in prose and accepted an older manifest in code. Both hosted funnels
  now read the revision of both manifests and refuse a regression
  before anything commits; witnesses pin the refusal in place.

* feat(managed-agent): drain a busy background Shell at release, and answer the asked operation

Round-5 close-path findings (H8): with a running background Shell the
Broker refused busy before transport.release — yet that HTTP call is the
only route to the worker's drain — so a `tail -f` held session release
and workspace close forever, with nothing able to stop it.

The release path now distinguishes busy kinds: anything active that is
not one of the Session's own background process rows keeps the exact
busy refusal with zero worker hops, while proven-running rows trigger a
dedicated release call first whose 409 busy answer is swallowed (the
worker's route drains before refusing, and its maintenance routes are
not activation-gated, so the follow-up sweep still proves what the
drain ended). Rows settle from the drain's evidence and the complete
release converges on what is genuinely still running — one call when
the drain finishes in time, one retry when it does not. Two witnesses
pin the contract: a running Shell reaches the worker (the round-5 probe
found the refusal never did) and a non-background busy refuses without
a worker hop.

Also from the round: the shell maintenance validator computed the
expected operationId from targetOperationId but never compared it with
the answer's own operationId, and the sweep witness answered with the
process-row id where the real route echoes the call id. The validator
now refuses a view answering another operation, a new protocol suite
pins it, and the witnesses answer the shape the route really sends.

* fix(managed-agent): re-assert the output pause behind every flushed chunk

Round-5 G5 finding: Node's flushStdio resumes both paused pipes when
the background launcher exits, and the executor tracked the pause in
its own flag rather than the stream's, so a writer that outlives its
launcher pushed 230 MiB into the queue and climbed RSS by +200 MiB.

The onOutput accounting now re-asserts the pause behind every delivered
chunk, and applyPause is idempotent through the stream's isPaused state
instead of the local flag — the stop-and-go the backpressure intended
holds through the launcher exit, on both the Shell and the Monitor
paths. The existing 17-MiB pipe pacing witness keeps its single pause
call.

* fix(managed-agent): mirror message.retracted from the H3 event vocabulary in the commit-side kind mirror

* fix(managed-agent): validate every domain.committed payload, with or without a textual domain

* test(integration): expect the monitor tool on the hosted shell profile

The hosted shell profile now advertises `monitor` beside
`run_shell_command`, so the fake-model driver's tool-name assertion —
which the Hosted workspace tool-turn IT checks on every model call —
must expect it. The file profile's expectation is unchanged: the
declaration is shell-profile only.

Fixes the single red IT in the MySQL fault-gates lane
(HostedWorkspaceToolTurnIT.packagedHarnessUsesSavedWorkspacesThroughReal
BrokerWorkerAndSqlStore), whose only difference was the extra
'monitor' entry.

* fix(managed-agent): read the whole journal when deriving pending monitor wakes

`readEvents()` is a bounded read: with no arguments it returns the first
`defaultReadEvents` (100) events, capped at 256. Both consumers of the
pending-input derivation called it bare, so on any Session whose log runs
past that page — every real Session with tool calls — the derivation saw
only the log's head:

- the wake scheduler's `next()` never found a notification input, so a
  monitor wake was never delivered as a turn once the Session passed the
  page. The feature was dead in production while every rig stayed green,
  because the rigs commit a handful of events;
- `settlePendingMonitorInputs` never found the owed inputs at close, so
  they stayed unsettled and the Session reopened as
  `hosted_turn_recovery_required` — exactly the wedge this close-path
  settle exists to prevent.

Both now read the full committed prefix with
`eventsInSequenceRange(1, committedSequence)`, the discipline every other
journal scan in the hosted session already uses. The witness commits 110
observations so the notification lands beyond the default page, asserts
the log really exceeds it, and was verified red under the inverse edit
(the settle returns 0 and the input stays owed).

* fix(managed-agent): stop a full task journal from wedging record commits

The task event journal is a derived, bounded feed: its backlog bound
refuses the next event past capacity rather than silently discarding it,
and an unarchived output event pins the retention floor so the automatic
pass cannot expire anything. The record commit called `appendStateChange`
inline, so once a task's journal was pinned and full, that refusal
propagated out of the extension-record commit — every later revision of
that task failed, which would take the Shell and Monitor record lines
with it.

The record row is the authoritative state and is already written when the
journal append runs, so a refusing journal now degrades the feed (warned
with tenant, session, task, revision and the refusal code) and never the
commit. The store still throws for callers that can apply backpressure —
which is where the output producer's retry belongs.

The witness drives the real wedge: it pins the floor with an unarchived
output event, fills the journal to the bound, asserts the next append
refuses with `managed_task_event_backlog_full`, then commits a record
revision whose view changes and asserts it lands. Verified red under the
inverse edit, where the commit dies with "The task's event journal holds
512 retained events".

* refactor(managed-agent): drop the shell view's never-set error field

The maintenance view declared an optional `error` object, the Java
response validator carried a branch for it, and no producer on either
side ever set one: the worker answers `running`, `exited` or `unknown`,
and a route-level failure travels as an HTTP status with its own code
body. A field that is declared and validated but never written is a dead
switch, so it comes out of the TypeScript view and the Java field set —
which now also refuses an answer that carries it, instead of silently
accepting a shape nothing produces.

* fix(managed-agent): let a replayed output advance pass instead of refusing

The forward-only check refused any advance whose revision was not
strictly newer, which included a redelivery of the very reference the
record already holds. The publisher guards against that in memory, but
the guard does not survive a worker restart, so a replayed advance could
throw inside the background exit leg and wedge a Shell or Monitor that
had already ended.

Both funnels now treat an identical reference as the no-op the deep-equal
skip already owns, and still refuse any other reference that is not
strictly newer. For a Shell the replay commits one liveness step, since
the live run line alternates by design; for a Monitor it commits nothing.
Witnesses pin both, beside the existing refusal witness.

* chore(managed-agent): keep the task-events lane report out of the PR diff

A one-off construction report does not belong at the repository root:
the project's directory rules put working artifacts under the ignored
.qwen tree and keep the tracked docs/ for design and plans. The report
is preserved at .qwen/investigations/2026-10-04-task-events-lane-report.md,
and its Flyway numbering note is superseded by the V40 renumber the
merge commit records.

* merge: absorb review-round state and retarget onto the H3 stack

* fix(managed-agent): quiet the wake pump's consume guard at a blocked owner

runTurn's recovery-blocked path marks the Session blocked and returns
'settled' without the turn ever existing, so the input stays owed — the
accurate answer the whole design gives for a parked Session. The pump's
after-settle check read that as a consumption bug and threw, which both
marked the already-blocked Session again and logged a spurious error on
every such observation. It now stops at an accurate blocked state and
still throws when the input simply did not move: that is the programming
error the guard exists for.

* fix(managed-agent): meld a monitor watch through one terminal step

Round-6 found the shape the Monitor's final leg really had: prepare's
open-then-advance could never survive the start order (output before a
start receipt is a parse-level refusal), the manifest still advanced
through the Shell funnel whatever the capture's recordDomain said (a
Monitor never builds its own Shell record), and the finalize settled the
record ahead of the loop's own exit chain, so a watch's last window died
in denial between two async hops — with its swallow queued to silence it.

Three changes, each with its own witness that goes red under the inverse
edit:

- The manifest advance reads the record's start receipt first and only
  advances past it; a missing record still fails loudly instead of the
  old "no record to revise" at whatever funnel happened to be there.
- It routes by recordDomain: monitor captures land on the Monitor
  funnel, Shell captures on the Shell's.
- The loop's onExit now returns a chain carrying the exit's own
  flush → settle in order, and the publisher awaits it rather than
  settling itself — a final window can no longer be dropped ahead of
  the settle, and a commit failure of the chain surfaces inside the
  finalize instead of returning an empty observation state silently.

The end-to-end witness drives the production shape on a real authority,
orchestrator and Publisher over HTTP: prepare ahead of attach, output
advancing only after the receipt, then one terminal step holding
observationSequence 1 with the settled mark behind it.

* fix(managed-agent): settle the process row when the Runtime proves no start

Round-6 found the permanent wedge: the `:process` row admits PREPARED
ahead of dispatch, but an admitted-refusal answer
(`{executionStatus: "not_started", capture: null}` — cgroup root missing,
quota exceeded, and friends) settled only its invocation. The process
was never registered by the Runtime, so every later shell-status can
only be unknown, the exit-only settle path never applies, and every
release carries the hold forever.

A proven never-started is proof: when the invocation settles with
executionStatus not_started, the sibling row now settles with the same
not_started mark, its admitted refusal as the evidence. Busy computed
after that never counts a process that could never have existed; an
unproven-lost connection still wedges exactly as the declaration says.
The witness drives the admitted refusal shape end-to-end: the process
row lands SETTLED with not_started, and release passes without the busy
answer. Verified red under the inverse edit, which reproduces the
report's `expected: <SETTLED> but was: <PREPARED>` verbatim.

* fix(managed-agent): attribute wake receipts to their own turn and stop re-driving it

Two wake-path defects from the round-6 audit family, plus the monitor
registry the maintenance route actually shares:

- `recoverShellReceipts` exempted monitor inputs so thoroughly that a
  wake turn's receipts could never be attributed: any `tool.receipt`
  journaled under a wake turn would attach to the previous prompt or be
  skipped. Receipt attribution now advances on every accepted input,
  while the pending set itself still exempts monitor inputs as before —
  nothing reframes them as parked Turns.
- A wake turn that started, journaled records and died was re-executed
  text-only on reopen, minting a second user record and leaving an
  unanswered first call in the model's history. `wakeHasPriorAttempt`
  names the condition simply; the pump now parks such turns accurately
  blocked for the recovery fleet instead of redriving.
- The worker constructed its executor with no monitor registry, so the
  executor built a private one while the monitor maintenance routes got
  another empty one — every status and stop answered unknown. One
  registry is shared by executor and routes now, mirroring what the
  Shell path already does.

* fix(managed-agent): read the task event floor after the events, not before

Two stores-and-reads of an event page ran as separate autocommit queries:
the retention floor was checked before `events.read`, so an expiry mid-
read produced a 200 that silently skipped events under an unchanged
nextCursor, the exact gap the contract's `409 cursor_expired` exists to
answer. The replayable-events family already documents the sound order:
events first, floor after — a monotoni…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants