Repository navigation
feat(managed-agent): Define the managed-extension-record/1 contract (H0b) - #12837
Conversation
…H0b) Add the records that the Stage H capabilities share, before anything commits them: the OperationGrant and when a Runtime gate may replace one; the logical run, physical execution and delivery lines with their single steps and the rules that tie them together, as one run block that every Stage H record embeds; recovery and quota reasons; stable identities and the Runtime binding; the definitionRevision pin; and the monitor_run record body with its revision rule. One shared schema and fixture file pins the contract. A TypeScript module in core and a Java validator in managed-agent-server replay every case, and both check every pair of states on every line. monitor_run joins the closed v1 domain index but stays disabled for submission. Refs #12827
Audit rounds before this PREach round froze the diff and gave it to two fresh auditors with no other context: one read the diff undirected, the other started from the design and tried to break the TypeScript module, the Java class and the schema with real probes, fuzzing and mutations. Rounds 1 and 2 fixed every valid finding; from round 3 only a Critical finding would have changed code. Round 1 — 1 Critical, fixed
Round 2 — 1 Critical, fixed
Round 3 — no Critical; deferred
Across the three rounds, the auditors' differential fuzzing over several million generated records and revision pairs found no case where TypeScript and Java disagree. The schema differs from them only on the rules the design lists as unstatable. 中文说明提交本 PR 之前的审计轮次每一轮都冻结 diff,交给两个没有其他上下文的新审计者:一个无方向地通读 diff,另一个从设计出发,用真实探针、模糊测试和变异尝试攻破 TypeScript 模块、Java 类和 schema。第 1、2 轮修复了所有有效发现;从第 3 轮起只有 Critical 才会改动代码。 第 1 轮:1 个 Critical,已修复
第 2 轮:1 个 Critical,已修复
第 3 轮:无 Critical;延后项
三轮中,审计者对数百万个生成的记录和修订对做了差分模糊测试,没有发现 TypeScript 与 Java 结论不同的情况。schema 与两者的差异只出现在设计列为无法表达的规则上。 |
…filtering A task whose read_output capability was withdrawn and restored would have its output events filtered out while the cursor moved past them, and a task without the capability kept output events that neither the events route nor the Artifacts returned. read_output now does not change during a task's life, and a task without it produces no output events, so the events route never filters out an event that exists and the no-loss guarantees hold for every task. With nothing filtered, next_cursor is simply the cursor of the last returned event. The capability text also says that only cancel needs an authorization check, since reading output needs only read access, and the design note points H0b at #12837, which answers issue question 1.
Local verification on macOS: built artifacts, head
|
| # | Kind | Finding |
|---|---|---|
| F1 | Doc, fix before merge | Two sentences are wrong in both languages. Round 3 and the triage bot flagged both statically; I ran both. (a) The Broker's real InMemoryToolExecutionRepository settles a cancelled PREPARED call as SETTLED with executionStatus=cancelled, not not_started. JdbcToolExecutionRepository:241 writes the same value. (b) The real LocalManagedSessionAuthority on base refuses a domain.committed event for monitor_run ("payload.domain must be one of …"), and this PR parses it. So the acceptance criterion "every existing record, authority and store behavior is unchanged" needs the exception "…except that readers now parse monitor_run in domain.committed; submitting it is still refused." |
| F2 | Test coverage, non-blocking | I wrote 9 targeted mutants for the gaps round 3 deferred: T1–T5 in TypeScript, J1–J4 in Java. All 9 pass the PR's own suites (530/530 and 7/7). All 9 are killed by my differential run or by a directed pair, so both languages behave correctly today; the shared fixtures just don't pin these edges. t3.jsonl in the harness is a ready-made case for T3, the one that random mutation missed. |
| Q1 | Question for H0c | The contract has no rule for a record's first revision. The parse functions judge one record and the successor functions judge a pair, so revision 1 can be any valid state, such as a run born with execution running_attached (the canonical monitorRun fixture is one) or a Monitor born settled with max_events. "An execution first appears as intent" therefore constrains revisions, not creation. Should H0c add an initial-revision predicate (for example: state reserved/admitted, execution null/intent, delivery null/planned), or is creating a record mid-life intended, for instance to adopt a Legacy monitor? Either answer works; one sentence in the doc would settle it before H0c commits records. |
| N1 | Informational | The documented NFC gap is concretely U+105D2 U+0307. Unicode 16 made that pair compose to U+105C9. JDK 21 (Unicode 15) accepts it as NFC; Node 24.18.1 (Unicode 17) refuses it. The TS verdict also changes within Node 22: 22.3.0 (Unicode 15.1) accepts, like JDK 21, and 22.23.2 (Unicode 17) refuses. So a TS writer and a TS reader on different Node 22 minors can disagree too. The gap comes from the Managed Session stable-ID rule rather than from this PR, and "identifiers must stay within Unicode 15" is not enforced by either validator. The doc could name the concrete pair and the Node-minor dependence. |
| N2 | Nit | The TS lease bounds are literals, but core already exports the same bounds as HTTP_MANAGED_SESSION_STORE_CONTRACT.minimumLeaseDurationMs/maximumLeaseDurationMs (1000/300000) in http-managed-session-store.ts, the TS counterpart of Java's ManagedSessionStoreModels. The triage bot's "there is no shared TS constant to reuse" is therefore not accurate. Optional. |
1. What ships: one token
- CLI bundle. I compared both
dist/trees (1,319 files each) with esbuild content-hash names normalized. They differ in two files: one chunk gains"monitor_run",and one carries the git-commit stamp.managed-extension-recordis not in the bundle. - Jar. All 361 base entries have identical CRC-32 in the PR jar; the only additions are
ManagedExtensionRecords.classand itsInvalidRecordException. Nothing insrc/mainreferences them. - Real authority, from core
dist, on both trees. Committingsession_metadata(the control) works. Committingmonitor_runis refused with "registered but not enabled" and the journal does not grow, identically on base and PR. Only the reader side differs (F1b): a pre-H0b reader refuses such a Session, so H3 must raiseminimumReader, as the design doc says.
2. The PR's suites on macOS, and an independent TS ↔ Java differential
- The PR's platform table marks macOS as not run. Here: 530 + 80 tests, all 23
src/managed-runtimefiles (1,443 tests),tsc, ESLint and Prettier are clean.mvn verifypasses 111 tests, 7 of them new, with 0 Checkstyle violations. These numbers match the PR description. - Differential: 3 seeds × 400,520 cases. Each seed is all 520 fixtures plus structured random grants, pins, runs, monitors and near-neighbour revision pairs. Number spellings such as
1.0,-0and1e400, and hostile ids (lone surrogates, C0/DEL/C1 characters, 512/513 bytes, NFD), reach each parser as raw text.- Both drivers reproduce all 520 fixture labels.
- Neither language threw anything other than a contract refusal, and Jackson read every line.
- The only disagreements are the 371 N1 cases.
- Schema vs module (Ajv 2020, strict): 601,191 single records. The schema was never stricter than the module. It was looser 5,998 times, all within the doc's list of rules JSON Schema cannot state.
- Of the 139,982 valid records in the corpus, every run, monitor and pin is accepted as its own successor, and every grant is refused as its own successor ("an identical grant replaces nothing"). Revisions that differ only in spelling (key order,
1000.0,1e3,-0, duplicate keys) get the same verdict in both languages, 16/16.
3. The round-3 fixture gaps, measured (F2)
4. A Monitor's life, including a Runtime-loss rebuild, in both languages
Both languages accept all 11 legal revisions: admission, intent, dispatch, attach, observation, Runtime loss, rebuild under generation 2 with a new receipt, max_events, and the watermark catching up after the end. Both refuse all 9 shortcuts, including reusing the generation-1 receipt on rebuild, observing while the watch is lost, and a lost watch "settling" under a new binding.
5. Doc facts, the pinned references, and CI coverage
- At
doudouOUC/code_agent@6891216:- Storage design §3.1 lists the same 33 names in the same order as
MANAGED_SESSION_DOMAINS, withmonitor_runaftermemory_job. - The control protocol's
OperationGrantfield list is exactly the PR's 9 closed keys. - Every item extension-runtime §10 lists for
monitor_runhas a field.
- Storage design §3.1 lists the same 33 names in the same order as
- The EN and ZH design docs have identical code spans, identical numbers and 19/19 headings.
- CI on this head:
- The TS suites ran only in Test (ubuntu, Node 22).
- The Java contract test ran only in the two ubuntu database jobs, "Runtime Broker and Managed Agent MariaDB / Java 21" and "Hosted no-tool processes / MySQL 8.4 / Java 21".
- The plain Java 11/17/21, macOS and Windows jobs don't run it, and the Node macOS/Windows jobs were skipped.
- So this macOS run is the first for both languages on that platform.
Not covered
- Windows was not run locally.
- There is no end-to-end path to drive: nothing imports the module until H0c. "Real environment" here means the shipped bundle, the jar, the real Session authority and the real Broker ledger.
- I did not weigh in on the six architectural choices the PR routes to the proposal owner.
Harness (generator, both drivers, mutants, probes) and raw results: assets-pr12837@f079755/pr12837. harness/README.md explains how to rerun each piece.
中文版
macOS 本地验证:用构建产物验证,head 18d8313904
结论:没有发现阻止合并的问题。 实际发布的内容只多了一个 token,即 domain 索引里的 "monitor_run",。在 1,201,560 个用例上,TypeScript 模块与 Java 校验器每次结论都相同,唯一的例外是设计文档已经写明的 Unicode 版本差异,共 371 例。schema 从未比模块更严格。建议合并前改掉 F1 的两句文档;F2 的 9 个 fixture 边界可以在本 PR 补,也可以在 H0c 之前补;Q1 是留给 H0c 的问题。
环境:macOS 26 arm64、Node 24.18.1、pnpm 11.24.0、JDK 21.0.12、Maven 3.9.16。base 88d881491b 和 PR 18d8313904 两棵树都用 npm run build && npm run bundle 和 mvn package 构建。A/B 对比、各个探针和差分都跑在这些构建产物上(packages/core/dist、jar 或 target/classes),不是在 vitest 下跑源码。
发现
| # | 类型 | 内容 |
|---|---|---|
| F1 | 文档,建议合并前修 | 有两句话在中英文里都不准确。第 3 轮审计和 triage bot 都做过静态指认,我实际跑了这两处。(a) Broker 真实的 InMemoryToolExecutionRepository 把被取消的 PREPARED 调用结算为 SETTLED、executionStatus=cancelled,而不是 not_started;JdbcToolExecutionRepository:241 写入的值相同。(b) base 上真实的 LocalManagedSessionAuthority 会拒绝 monitor_run 的 domain.committed 事件("payload.domain must be one of …"),本 PR 能解析它。因此验收标准"现有记录、authority 和 store 行为不变"要加一个例外:"……但 reader 现在能在 domain.committed 中解析 monitor_run,提交仍被拒绝。" |
| F2 | 测试覆盖,不阻塞 | 针对第 3 轮延后的缺口,我写了 9 个定向变异体:TypeScript 5 个(T1–T5),Java 4 个(J1–J4)。9 个全部通过 PR 自带的测试(530/530 和 7/7)。9 个都能被我的差分或一个定向用例对杀死,说明两种语言现在的行为是对的,只是共享 fixtures 没有固定这些边界。随机变异没有触达 T3,harness 里的 t3.jsonl 就是现成的 T3 用例。 |
| Q1 | 留给 H0c 的问题 | 契约没有规定一条记录的首个修订。解析函数判断单条记录,后继函数判断一对记录,所以修订 1 可以是任何合法状态,例如 execution 一出生就是 running_attached 的 run(规范 monitorRun fixture 就是这样),或者一出生就以 max_events settled 的 Monitor。"execution 首次出现必须是 intent"因此只约束修订,不约束创建。H0c 是否要加一个首修订判定(例如 state 为 reserved/admitted,execution 为 null/intent,delivery 为 null/planned),还是有意允许中途创建记录,比如接管 Legacy monitor?两种答案都可以,但最好在 H0c 开始提交记录之前,在文档里用一句话写明。 |
| N1 | 信息 | 文档里提到的 NFC 差异,具体就是 U+105D2 U+0307:Unicode 16 让这对字符组合成 U+105C9。JDK 21(Unicode 15)认为它是 NFC 并接受,Node 24.18.1(Unicode 17)拒绝。在 Node 22 内部,TS 的结论也会变:22.3.0(Unicode 15.1)与 JDK 21 一样接受,22.23.2(Unicode 17)拒绝。所以同为 TS,写入方和读取方跑在不同的 Node 22 小版本上也可能结论不同。这个差异来自 Managed Session 稳定 ID 规则,不是本 PR 引入的;"标识符须限于 Unicode 15 已分配字符"这条约束,两个校验器都没有强制。文档可以写明这个具体字符对,以及结论随 Node 小版本变化这一点。 |
| N2 | 小建议 | TS 的租约上下界写成了字面量,但 core 已经在 http-managed-session-store.ts 导出了同样的上下界 HTTP_MANAGED_SESSION_STORE_CONTRACT.minimumLeaseDurationMs/maximumLeaseDurationMs(1000/300000),它对应 Java 的 ManagedSessionStoreModels。所以 triage bot"没有可复用的共享 TS 常量"的说法不准确。可选。 |
1. 实际发布的内容:只多一个 token(图 01)
- CLI bundle: 我对比了两边的
dist/(各 1,319 个文件),先把 esbuild 的内容哈希文件名归一化。两边只有两个文件不同:一个 chunk 多了"monitor_run",,另一个是 git commit 戳。managed-extension-record不在 bundle 里。 - jar: base 的全部 361 个条目在 PR 的 jar 里 CRC-32 都相同,只新增了
ManagedExtensionRecords.class和它的InvalidRecordException。src/main里没有代码引用它们。 - 真实 authority(core
dist,两棵树各跑一次): 提交session_metadata(对照)成功。提交monitor_run被拒("registered but not enabled"),journal 没有增长,base 与 PR 完全一致。唯一的差别在读取端(F1b):H0b 之前的 reader 会拒绝含这类记录的 Session,所以 H3 必须提高minimumReader,和设计文档说的一样。
2. PR 自带测试在 macOS 上的结果,以及独立的 TS ↔ Java 差分(图 02)
- PR 的平台表把 macOS 标为未运行。本次在 macOS 上:530 + 80 项测试、
src/managed-runtime全部 23 个文件(1,443 项)通过,tsc、ESLint 和 Prettier 无问题。mvn verify通过 111 项测试,其中 7 项是新增的,Checkstyle 0 违规。这些数字与 PR 描述一致。 - 差分: 3 个种子,每个 400,520 个用例:520 个 fixture,加上结构化随机生成的 grant、pin、run、monitor 和邻近修订对。
1.0、-0、1e400这类数字写法,以及恶意 id(孤立代理项、C0/DEL/C1 字符、512/513 字节、NFD),都以原始 JSON 文本交给各自的解析器。- 两个驱动都复现了全部 520 个 fixture 标注。
- 两种语言都没有抛出契约拒绝之外的异常,Jackson 也读得了每一行。
- 仅有的分歧就是 N1 的 371 例。
- schema 对比模块(Ajv 2020,strict): 共 601,191 条单记录。schema 从未比模块更严;比模块宽松 5,998 次,全部落在文档列出的 JSON Schema 无法表达的规则里。
- 语料中共有 139,982 条合法记录:其中每条 run、monitor、pin 都被接受为自身的后继,每条 grant 都被拒绝为自身的后继("完全相同的 grant 不替换任何东西")。只在写法上不同的修订(键顺序、
1000.0、1e3、-0、重复键),两种语言结论一致,16/16。
3. 第 3 轮 fixture 缺口的实测(图 03,对应 F2)
4. 一个 Monitor 的完整生命周期,含 Runtime 丢失后的重建,两种语言都跑(图 04)
两种语言都接受全部 11 个合法修订:准入、intent、派发、attach、观测、Runtime 丢失、在 generation 2 下带新回执重建、max_events、结束后通知水位追平。两种语言也都拒绝全部 9 种捷径,包括重建时沿用 generation 1 的回执、watch 丢失期间还在观测,以及丢失的 watch 换了新 binding 后直接"结算"。
5. 文档事实、固定引用的参考设计与 CI 覆盖(图 05)
- 在
doudouOUC/code_agent@6891216上:- 存储设计 §3.1 的 33 个名称,与
MANAGED_SESSION_DOMAINS名称一致、顺序一致,monitor_run在memory_job之后。 - 控制协议里
OperationGrant的字段表,正好是 PR 的 9 个封闭键。 - 扩展运行时 §10 为
monitor_run列出的每一项都有对应字段。
- 存储设计 §3.1 的 33 个名称,与
- 中英文设计文档的代码片段一致、数字一致,19/19 个标题对应。
- 本 head 的 CI:
- TS 测试只在 Test (ubuntu, Node 22) 中运行。
- Java 契约测试只在两个 ubuntu 数据库任务中运行:"Runtime Broker and Managed Agent MariaDB / Java 21" 和 "Hosted no-tool processes / MySQL 8.4 / Java 21"。
- 普通的 Java 11/17/21 任务以及 macOS、Windows 任务都不跑它,Node 的 macOS/Windows 任务被跳过。
- 因此本次 macOS 运行是两种语言在该平台上的首次运行。
未覆盖
- 没有在本地跑 Windows。
- 没有端到端路径可驱动:在 H0c 之前没有代码导入这个模块。这里的"真实环境"指发布的 bundle、jar、真实的 Session authority 和真实的 Broker 账本。
- 没有评判 PR 交给提案负责人的六个架构选择。
harness(生成器、两个驱动、变异体、探针)与原始结果见 assets-pr12837@f079755/pr12837,harness/README.md 说明了各部分如何重跑。
|
@qwen-code /triage |
qqqys
left a comment
There was a problem hiding this comment.
Independent Critical-only review — head 18d83139
A cross-language contract module: managed-extension-record.ts (1060 lines) and its test (457), ManagedExtensionRecords.java (708) and its contract test (287), a 1818-line JSON schema, a 19,457-line shared fixture corpus, two design docs, and one token of live change — 'monitor_run' added to MANAGED_SESSION_DOMAINS, with the test that pins the list's length bumped 32 → 33. 10 files, +24348/-2.
Because the ratio of new code to shipped behaviour is so extreme, I spent the budget on the one token that executes and on whether the rest is genuinely inert.
No historical blocker
There are no reviews and no inline comments on this PR, so nothing has ever been filed against it and there is no thread to re-verify. Triage stage 2 at this head concludes "No Critical blockers". The author's own pre-submission audit reports one Critical in each of rounds 1 and 2 — a Monitor that failed after start having no recordable state, and Java throwing NumberFormatException on a number past double range so a successor check threw instead of returning false — both fixed before the PR was opened, with round 3 finding none. Neither stands at this head: watch_failed is in the stop-reason table, and non-finite numbers are refused with four fixtures carrying the raw 1e400.
The one shipped change is a reader-side widening, and it is the safe direction
MANAGED_SESSION_DOMAINS has exactly two references in the repository — its own module and its test — and inside the module it is consulted at one place, the case 'domain.committed': branch that validates a record's domain against the list. So adding monitor_run changes what a reader accepts; it does not add a writer, and no writer consults this list. Nothing in the tree produces a monitor_run record today, so the effect is that a journal carrying one parses instead of failing as an unknown domain, while every other unknown domain still fails closed. The author states the same split explicitly — a reader can now parse monitor_run in domain.committed while committing one is still refused — and amended the acceptance criterion to say so rather than claiming "existing record behaviour is unchanged". The writer-side obligation that this creates (a writer must gate on the reader's declared minimumReader version before emitting the new domain) belongs to H0c, is named in the design note, and cannot bite at this head because nothing emits it.
The rest is inert, verified rather than assumed
A repository-wide search for managed-extension-record returns zero importers: the module, its test, the schema and the fixtures are not reachable from any production path. That is what makes a 24k-line addition reviewable at all — nothing executes it, so a defect here is a defect in a specification and its reference implementations, discoverable by the contract test before any consumer exists.
Within the module I checked the defensive posture, since this code's whole job is to reject malformed untrusted structure: exported constants are Object.freezed (including the per-state transition arrays built by a loop that freezes each one), the object check rejects anything whose prototype is neither Object.prototype nor null with the comment noting an array fails there too, MAX_GENERATION is 2n ** 63n - 1n compared as a BigInt, limits are sourced from the existing MANAGED_SESSION_LIMITS rather than reinvented, depth_limit is a named rejection reason, phase names are pattern-constrained, and failures raise the module's typed ManagedSessionRecordError. That is the same closed-world, fail-closed shape the sibling managed-runtime modules use.
Cross-language agreement is pinned by a shared corpus, not by two sets of good intentions
The TS test and the Java contract test consume the same 19,457-line fixture file, which is the mechanism that keeps two implementations from drifting. Triage compared the constants one by one and reported them equal — 33 domains with monitor_run after memory_job, three state lines and every transition list, two delivery targets, five recovery and six quota reasons, seven monitor stop reasons, and a 2^53 − 2 count bound enforced by Number.isSafeInteger on one side and its Java equivalent on the other. The author's differential run at this head reports 1,201,560 cases with identical verdicts in both languages, and that spelling-level variants (key order, 1000.0, 1e3, -0, duplicate keys) agree 16/16.
The one divergence is documented, measured and fails closed. 371 of those cases disagree, all from a single Unicode-version gap: U+105D2 U+0307 composes to U+105C9 under Unicode 16, so JDK 21 (Unicode 15) accepts the pair as NFC while Node 24 — and Node 22.23, unlike 22.3 — refuses it. The consequence is rejection of an ID a peer would have accepted, not silent acceptance of one it would not, and the gap originates in the Managed Session stable-ID rule rather than in this contract. It is recorded in the design doc with the concrete pair and the version boundaries, which is the right handling for a divergence a contract module cannot fix on its own. I am naming it here so it is not mistaken for an unmeasured unknown.
CI
20 checks pass at this head and none has failed, including the Java lanes that run ManagedExtensionRecordContractTest (ubuntu-latest / Java 11, 17, 21, windows-latest, macos-latest, Runtime Broker and Managed Agent MariaDB, Hosted no-tool processes / MySQL 8.4, Real daemon E2E / Java 11) and the Node lanes that run the TypeScript module's test and typecheck it (Test (ubuntu-latest, Node 22.x), Lint & Static), plus Integration Tests (no-AK, No Sandbox), web-shell E2E Smoke, both Desktop Shell lanes, triage and the routing checks. review-pr is pending and is not a gate; a sandboxed verification had just been triggered and had not reported.
Scope
I read the live change and traced its only consumer, confirmed the module has no importers, and read the new module's exported surface, limits and rejection posture. I did not read the 1060-line TypeScript module or the 708-line Java class line by line, did not read the 1818-line schema, and sampled nothing of the 19,457-line corpus — for those I relied on the shared-corpus mechanism, the differential measurement and triage's constant-by-constant comparison, and I am naming that reliance rather than implying I covered it. The two doc sentences the author's verification flags as inaccurate (F1) and the nine fixture edges (F2) are documentation and coverage items, outside Critical-only scope.
Verdict: APPROVE — No Critical found and none was ever filed. Everything added except one array entry is unreachable from production code, that entry widens only what a reader parses while unknown domains still fail closed and no writer exists to emit one, both reference implementations are pinned to the same fixture corpus with a measured 1.2 M-case differential whose sole divergence is a documented Unicode-version gap that rejects rather than admits, and the two Criticals the author's own audit found were fixed before this head.
…nLM#12830) * feat(managed-agent): Add the planned Stage H task contract (H0a) Freeze the public task contract that the extension runtime design requires before H0 is implemented. Every addition is marked planned, so no route is mapped and the generated WebShell types do not change. The public API gains a task list, task detail, task events behind an opaque output cursor, and a cancel command that takes an Idempotency-Key and answers 202 with a command operation. The WebShell adapter mirrors them. The task view is SessionTaskView in the public conventions; schema conditionals keep settled_at, started_at and the advertised actions consistent with the task state, and closed objects keep Runtime bindings, generations, PIDs and paths out of it. Cancel reuses the command operation with a new task_cancel type, and a completed cancel records acceptance, not a stop. The MCP catalog, hook catalog, channel and automation resources are named as planned routes without bodies, for H1, H2, H5 and H6 to fill in. The contract test now also watches the channel and automation prefixes, which a mapped planned route would otherwise pass unnoticed. Part of QwenLM#12827. * fix(managed-agent): Close review findings on the planned task contract A cancel retry with the same Idempotency-Key now replays the original operation before any state or capability check, so a lost 202 never turns into a 409 once the task settles. Task events carry their own cursor, so a consumer can checkpoint per event instead of re-applying output after a crash, and a bad cursor is invalid_event_cursor as on the Session event history. The task list is newest first with an id tiebreaker, as the Session list is. The task resource and its lists carry an object discriminator, the event list is named PublicTaskEventList, and the event type is a shared enum. recovery_blocked never offers send_input, completed needs started_at, more tasks need a cursor and output chunks are never empty. A new test checks these conditionals with valid and invalid instances, because the API contract test validates only operations that are not planned. The design notes record the departures from the Session event history, the enum value that cannot be marked planned, the properties H0c must also flip, and the Artifact attribution and Legacy state follow-ups. * fix(managed-agent): Align task events and cancel replay with the API contract Task events carry schema_version and projection_version, as the API contract requires of public events, and their type is an open string, so later types stay additive. A page's next_cursor is the position after the last event the server examined, so events filtered out for a caller cannot stall it, and the start of the retained events when after was omitted. An expired cursor loses no output; how Artifacts and retained events join without overlap is left to the H3 output segmentation. Cancel replays the original operation for the same actor after the caller's access is re-checked, as the contract's idempotency rules say, and declares invalid_request and task_forbidden next to the 409s. The route check reads every /v1/agent route, so later /v1/agent-* resources are covered without another prefix. The planned-task test runs every instance against the WebShell mirror as well, and checks the WebShell cancel and event query requests. * fix(managed-agent): Fix the task check order and cover the missed invariants Cancel checks run in a fixed order: access (404, then 403), idempotent replay, then task state. A missing Idempotency-Key is invalid_request and a malformed one invalid_idempotency_key, as on the other idempotent routes, and a Session that does not serve tasks answers unsupported_feature. action_capabilities describes the task, not the caller. The read routes drop 403, since a caller that cannot read gets 404, and a list cursor for more tasks is never empty. The planned-task test now covers waiting, degraded and a failed task that had started, and rejects state_changed without state and artifact without artifact_id. The design note states the per-surface mutation counts, the event type rule for later minor versions, and the reuse of the command operation as the tradeoff behind the visible task_cancel. * fix(managed-agent): Read retained task events before Artifacts on recovery The recovery recipe after an expired cursor read the task Artifacts first and the retained events second, which can lose an event that is archived and expired between the two reads. Reading the retained events first and the Artifacts second loses nothing, because every event that expired before the event read began was archived before it. The guarantee also needs artifact_refs to list every Artifact of the task. The design note also stops attributing the ignore-unknown-types rule to section 5 of the API contract, which only allows new event types, and names the one public-only instance the WebShell mirror check skips. * fix(managed-agent): Define the task event page cursor by what the page covers next_cursor was the position after the last event the server examined. A server that reads one row past the limit to compute has_more examines an event it does not return, and a caller passing that cursor back would skip it with no error. The cursor is now the position after the last event the page covers: every earlier event was returned or filtered out for this page, and an event held back by limit is never passed. The replay guarantee now holds while the idempotency record is retained rather than never failing, and the Artifact routing of high-volume output cites design section 11 and API contract section 6, which is where the rule comes from. * fix(managed-agent): Expire task events only from the oldest end The events route promised that a page never passes an event it did not return or filter out, and that an expired cursor answers 409, but did not say in which order events expire. If a later event could expire before an earlier output event that is not yet in an Artifact, a caller following the stream would skip it silently. Events now expire only from the oldest end, so the retained events have no gaps and a cursor older than the oldest retained event is the only way to miss one. The spec's cancel description also carries the retention condition the design note already states: a same-key retry replays the original operation while its idempotency record is retained. * fix(managed-agent): Fix read_output for a task's life and drop event filtering A task whose read_output capability was withdrawn and restored would have its output events filtered out while the cursor moved past them, and a task without the capability kept output events that neither the events route nor the Artifacts returned. read_output now does not change during a task's life, and a task without it produces no output events, so the events route never filters out an event that exists and the no-loss guarantees hold for every task. With nothing filtered, next_cursor is simply the cursor of the last returned event. The capability text also says that only cancel needs an authorization check, since reading output needs only read access, and the design note points H0b at QwenLM#12837, which answers issue question 1.
…wenLM#12863) Close the items QwenLM#12837 deferred for the managed-extension-record/1 contract (H0b of QwenLM#12827), as QwenLM#12846 lists them. No production code changes. Correct two statements that H0c would build on. The Broker settles a cancelled call that was never sent as cancelled, not not_started, and no field of its record tells that apart from a call cancelled after it was sent, so H0c must not map a call it cannot prove unsent to not_started_proven. Readers now parse monitor_run in domain.committed, and only domain-record submission checks that a domain is enabled, so H0c must enforce that check on every path that commits domain records. Add 34 fixture cases, and give one existing case a valid next record, so that each targeted check has a case of its own: the six gaps the issue lists and those four audit rounds found. Each targeted clause-level mutant now fails in TypeScript and Java, and so does each per-field mutant but three equivalent ones in TypeScript. Record the decisions for H0c and H3: the v1 notification policy is fixed; revocation is gate state outside the grant replacement rule; only a purely observational Monitor is rebuilt, and only from outcome_unknown. Add two open questions. Closes QwenLM#12846
wenshao
left a comment
There was a problem hiding this comment.
Not explored to full depth (tool budget reached): "agent reverse-audit (round 4)": the Java-side half of the mutation (compiling and replaying a widened case "running", "waiting" arm of ManagedExtensionRecords.requireRun ) was not executed …; "agent reverse-audit (round 2)": Rebuild 段落(:228)声称 rebuild 的 intent 与 dispatch "committed in that phase's physical ledger, keyed by Session, effect, phase and effect revision" —— 核对它需要打开 Broke…; "agent reverse-audit (round 2)": 第 256 行"A Python implementation written from this document … labeled every case, and the generator stops when it disagrees with a label" —— 该 oracle 明确"kept out…; "agent reverse-audit (round 3)": zh-CN prose parity beyond identifiers and numbers — a Chinese sentence in doc 171–279 that drops or adds a rule stated purely in prose (no backticks, no digits)…; "agent reverse-audit (round 3)": doc 189's "section 12 of the reference design" quota mapping and doc 205's "section 10 of the reference design" monitor-key coverage — the reference design is a…, and 3 more.
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.
Test Plan (not a blocker): src/managed-runtime/managed-extension-record.test.ts — no such file or directory; src/managed-runtime/managed-session-records.test.ts — no such file or directory.
中文说明
未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 4)":the Java-side half of the mutation (compiling and replaying a widened case "running", "waiting" arm of ManagedExtensionRecords.requireRun ) was not executed …;"agent reverse-audit (round 2)":Rebuild 段落(:228)声称 rebuild 的 intent 与 dispatch "committed in that phase's physical ledger, keyed by Session, effect, phase and effect revision" —— 核对它需要打开 Broke…;"agent reverse-audit (round 2)":第 256 行"A Python implementation written from this document … labeled every case, and the generator stops when it disagrees with a label" —— 该 oracle 明确"kept out…;"agent reverse-audit (round 3)":zh-CN prose parity beyond identifiers and numbers — a Chinese sentence in doc 171–279 that drops or adds a rule stated purely in prose (no backticks, no digits)…;"agent reverse-audit (round 3)":doc 189's "section 12 of the reference design" quota mapping and doc 205's "section 10 of the reference design" monitor-key coverage — the reference design is a…,另有 3 条。
未审查:反向审计——在 5 轮的反审轮数上限内未收敛。
Test Plan(非阻断):src/managed-runtime/managed-extension-record.test.ts — no such file or directory; src/managed-runtime/managed-session-records.test.ts — no such file or directory。
— qwen3.8-max via Qwen Code /review (v0.24.7)
| - In both languages, every malformed shape and every illegal transition in the fixtures is refused, and every valid case is accepted. | ||
| - The schema and the module disagree only on the listed rules that JSON Schema cannot state. | ||
| - `monitor_run` is in the v1 domain index and is still refused for submission. | ||
| - Every existing record, authority and store behavior is unchanged. |
There was a problem hiding this comment.
[Critical] R1-1: [certifies-falsely] [new-surface] Both design-doc corrections the author committed to before merge are still false at the merged head, in both languages — and a human approval certified one of them as amended.
Two factual claims in this normative document are wrong at head 18d8313904. (a) This acceptance criterion — "Every existing record, authority and store behavior is unchanged." (zh-CN:264 「现有的记录、authority 与存储行为都不变。」) — is contradicted by the PR's own Risk section: domain.committed now parses monitor_run where the base refused it, so one reader path did widen. (b) Line 20 — "a cancelled PREPARED call settles as not_started without a dispatch" (zh-CN:20 same claim) — is contradicted by the Broker it describes: RuntimeBrokerService.java:1876-1878 states "A not_started status is accepted only as the Runtime's own terminal answer; the Broker never derives it", and the cancel path at :1641-1643 settles with executionStatus=cancelled. The triage thread confirmed both sentences at this head and recorded the author's commitment to fix them before merging; the approving review then asserted the criterion had been "amended … rather than claiming 'existing record behaviour is unchanged'" — git show HEAD: of both files shows it was not. Line 264 is the sentence an H0c implementer trusts to conclude no reader-compatibility work is needed (see also the minimumReader comment on Decision 1), and line 20 is the doc's stated reason the execution line is "not step for step" with the Broker ledger, so H0c would derive the Broker↔contract mapping from a false premise. The document is this contract-first slice's core deliverable, which is why this is graded as a blocker rather than a doc nit.
Witness:
git show HEAD:docs/design/...-contract.md (and the zh-CN twin) at 18d8313904:
line 264: "- Every existing record, authority and store behavior is unchanged." (present, both languages)
line 20: "a cancelled `PREPARED` call settles as `not_started` without a dispatch" (present, both languages)
BASE/PR probe on the domain.committed reader (base arm = the PR's single production line reverted):
BASE: domain=monitor_run => REFUSED (must be one of: config_install, ...)
PR: domain=monitor_run => PARSED kind=domain.committed
(control: session_metadata PARSED on both arms)
Suggested fix: in both languages, replace the line-264 criterion with the actual behaviour (domain.committed parses monitor_run; submission still refuses it), and correct line 20 to say a cancelled record settles as SETTLED with executionStatus=cancelled, that the Broker never derives not_started, and that EXECUTING is entered before the Runtime receives the call so it corresponds to dispatch_started.
中文说明
作者在合并前承诺修正的两处设计文档陈述,在已合并的 head 上仍然是错的,而且中英两版都错——其中一处还被人工 approve 认证为"已修正"。
(a) 第 264 行验收标准「Every existing record, authority and store behavior is unchanged.」/「现有的记录、authority 与存储行为都不变。」——PR 自己的 Risk 段就说明本 PR 改变了 reader 行为:domain.committed 现在能解析 monitor_run(base 上会拒绝)。(b) 第 20 行「a cancelled PREPARED call settles as not_started without a dispatch」/「被取消的 PREPARED 调用不经派发即以 not_started settled」——与它所描述的 Broker 相反:RuntimeBrokerService.java:1876-1878 写明 not_started 只作为 Runtime 自己的终态回答被接受、Broker 从不推导它;取消路径 :1641-1643 以 executionStatus=cancelled 结算。triage 线程在本 head 上确认了这两句,并记录了作者"合并前会修"的承诺;随后的 approve review 声称该标准"已改写"——但 git show HEAD: 显示两句在两种语言里都原封未动。
第 264 行是 H0c 实现者据以得出"无需 reader 兼容工作"结论的那句话;第 20 行是文档解释执行线为何与 Broker 台账"并非逐步对应"的依据,H0c 会从一个假前提出发推导 Broker↔契约映射。文档正是这个契约优先切片的核心交付物,因此按阻断级而非文档小疵定级。
建议修复:两种语言同步,把第 264 行改为实际行为(domain.committed 可解析 monitor_run,提交仍被拒绝),并把第 20 行改为:被取消的记录以 SETTLED+executionStatus=cancelled 结算,Broker 从不推导 not_started,EXECUTING 在 Runtime 收到调用之前进入、因此对应 dispatch_started。
— qwen3.8-max via Qwen Code /review (v0.24.7)
|
|
||
| ## Decisions | ||
|
|
||
| 1. **`monitor_run` joins the closed v1 index** (question 1 of #12827). The storage design at the pinned commit lists it among the 33 v1 names and leaves only its validator to H0. No v1 reader can meet a `monitor_run` record yet, because the domain stays disabled, so the index keeps its version. A reader from before this change refuses the name, so when H3 enables the domain, a Session that holds such records must keep those readers out, for example by raising its `minimumReader`. The name goes after `memory_job`, in the design's order. |
There was a problem hiding this comment.
[Suggestion] R1-2: Decision 1's only named mitigation — "for example by raising its minimumReader" — cannot be applied to the Sessions the sentence is about.
A Session's header is written once, at creation, from the module constant MANAGED_SESSION_MINIMUM_READER (managed-session-authority.ts:427-438, inside the if (header === undefined) branch); the create option's own docstring says "Supplied when creating a session; ignored once a header exists" (:271), and no header-rewrite or migration path exists anywhere under managed-runtime (grep: 0 hits). H3 enables the domain through a global switch (MANAGED_SESSION_ENABLED_DOMAINS), so every already-created Session becomes eligible to hold managed-monitor_run records while its header still says managed-session/1 — and a pre-H3 reader opens such a Session (the reader refuses only when the required version exceeds its own, managed-session-records.ts:1165-1169) and meets a domain name absent from its 32-name closed index: exactly the encounter this sentence says must be kept out. H3 must therefore either build a header-migration path — reader-compatibility work the acceptance criterion at line 264 says is not needed (R1-1) — or leave every pre-existing Session unable to use Monitor; an implementer who trusts this sentence plans for neither.
Witness:
production write sites of minimumReader in packages/core/src: exactly 1 (managed-session-authority.ts:438,
inside the create branch at :427); header construction sites: exactly 1 (:471)
rewrite|upgrade|migrat under managed-runtime: 0 hits
reader refusal: only when requiredReaderVersion > currentReaderVersion (managed-session-records.ts:1165-1169)
Suggested fix: replace the clause with the mechanism that exists — e.g. "so when H3 enables the domain it must raise MANAGED_SESSION_MINIMUM_READER, which keeps pre-H3 readers out of every Session created afterwards; a Session created before the bump keeps the header it was written with, so H3 must also migrate that header or refuse to enable monitor_run for it" — and record the header migration as H3 work under Follow-up. The zh-CN twin at line 43 must move with it.
Fix constraint: export const MANAGED_SESSION_MINIMUM_READER = 'managed-session/1'; (managed-session-records.ts:22) is a module constant and the header is created only under if (header === undefined) (managed-session-authority.ts:427-438) — the corrected wording must not imply a per-Session raise exists.
中文说明
Decision 1 唯一给出的缓解手段——「例如提高其 minimumReader」——对这句话所要保护的那批 Session 根本不可用。
Session 的 header 只在创建时写入一次,取自模块常量 MANAGED_SESSION_MINIMUM_READER(managed-session-authority.ts:427-438 的 if (header === undefined) 分支);create 选项的文档字符串自己写着「Supplied when creating a session; ignored once a header exists」(:271),而 managed-runtime 下不存在任何 header 改写/迁移路径(grep 0 命中)。H3 通过全局开关(MANAGED_SESSION_ENABLED_DOMAINS)开放该 domain,于是每个已创建的 Session 都有资格持有 managed-monitor_run 记录,而它们的 header 仍是 managed-session/1——H3 之前的 reader 仍能打开这种 Session(读方只在所需版本超过自身时拒绝,managed-session-records.ts:1165-1169),并遇到一个不在其 32 名封闭索引里的 domain 名:这正是本句声称必须挡住的相遇。因此 H3 要么建一条 header 迁移路径(而第 264 行验收标准说无需 reader 兼容工作,见 R1-1),要么让所有既有 Session 无法使用 Monitor;相信这句话的实现者两者都不会规划。
建议修复:把该小句改成真实存在的机制——「H3 开放该 domain 时必须提高 MANAGED_SESSION_MINIMUM_READER,这只能把 H3 之前的 reader 挡在此后创建的 Session 之外;在提高之前创建的 Session 保留写入时的 header,因此 H3 还必须迁移这些 header,或拒绝为其开放 monitor_run」——并把 header 迁移记入 Follow-up 的 H3 工作。中文版第 43 行需在同一改动内同步。
修复约束:MANAGED_SESSION_MINIMUM_READER = 'managed-session/1'(managed-session-records.ts:22)是模块常量,header 只在 if (header === undefined) 下创建(managed-session-authority.ts:427-438)——修正后的措辞不得暗示存在逐 Session 的提高机制。
— qwen3.8-max via Qwen Code /review (v0.24.7)
| - **Scope.** `recordRef` is the committed domain record that holds the operation's plan. Its kind must be `managed-<domain>` and its schema version 1, the pairing `domain.committed` already requires, so a grant cannot point at another domain's plan. `phases` lists the 1 to 16 distinct phases of that plan the grant admits. Which phases a domain has is registered with its body in H1–H6; H0b fixes their form. | ||
| - **No activation fields.** The record is closed, so an `epoch` or `activationId` is refused: a grant is not an activation. | ||
| - **Replacement.** A Runtime's per-operation gate may replace grant `a` with grant `b` only when both are valid, name the same Session, operation and domain, and either | ||
| - `b` renews `a`: the same revision and every field equal, phases in the same order, except a later `expiresAt`; or |
There was a problem hiding this comment.
[Suggestion] R1-3: The renewal rule makes phases array order load-bearing, while the scope rule that defines the same field (line 81) fixes only distinctness — so a semantically identical renewal can be refused as stale.
isOperationGrantSuccessor's same-revision branch implements line 84 literally: after.expiresAt > before.expiresAt && sameJson({...before, expiresAt: 0}, {...after, expiresAt: 0}) (managed-extension-record.ts:605-609), and sameJson is JSON.stringify text equality (:462-464) — a reordered resourceScope.phases fails it. The later-revision branch cannot rescue a renewal (it requires a strictly greater operationRevision, :612). Probed in both languages: a renewal presenting the same two phases in the other order with a later expiresAt is refused, while the same-order control is accepted. The corpus never exercises a pure reorder — the only renewal-scope case (renewal-changes-phases, fixtures:3848) drops a phase. The trigger is prospective (nothing commits grants in this slice), and the order requirement IS documented here — so this is spec coherence: line 81 and line 84 currently promise two different things about the same field, and the Java control plane holds phases as a Set<String> that Jackson serializes in hash order.
Witness:
TS: parseOperationGrant(reordered phases) => ACCEPTED
isOperationGrantSuccessor(before, reordered + later expiresAt) => false
CONTROL (same order, later expiresAt) => true
Java (module's own compiled classes): same three verdicts
Suggested fix: choose one semantics and state it in both languages at lines 81/84 (and the zh-CN twins, which carry the identical rule 「phase 顺序也相同」): either make the order canonical and required ("in the plan's registered order") and enforce it in parseOperationGrant so a reordered list is not a well-formed grant at all; or keep the scope rule order-free and judge the renewal on the semantics it defines — compare phases as a set (equal length plus equal members) instead of by JSON.stringify of the whole grant.
Fix constraint: the parser's only phases property beyond PHASE_PATTERN and maxGrantPhases is the distinctness check (managed-extension-record.ts:546-548), matching doc line 81 — a canonical-order fix must land in the parser and the doc together, or the successor check will keep refusing grants the parser accepts.
Fix witness: a successor fixture pair with the same operationRevision, a later expiresAt, and ["query_receipt","send_segment"] against ["send_segment","query_receipt"], whose valid flag pins the chosen semantics in both replays — today no case presents a pure reorder, so adding or removing the order requirement moves no test in either language; with the new pair, removing the chosen rule reddens both.
中文说明
续租规则让 phases 数组的顺序成为判定要素,而定义同一字段的范围规则(第 81 行)只要求互不相同——于是一次语义完全相同的续租可能被当作陈旧而拒绝。
isOperationGrantSuccessor 的同修订分支逐字实现了第 84 行:after.expiresAt > before.expiresAt && sameJson({...before, expiresAt: 0}, {...after, expiresAt: 0})(managed-extension-record.ts:605-609),而 sameJson 是 JSON.stringify 文本相等(:462-464)——resourceScope.phases 换了顺序就过不了。更晚修订分支救不了续租(它要求 operationRevision 严格更大,:612)。两种语言实测:同样两个 phase 换个顺序、带更晚 expiresAt 的续租被拒,同序对照被接受。语料中没有任何纯换序用例——唯一的续租范围用例 renewal-changes-phases(fixtures:3848)是删掉一个 phase。触发是前瞻性的(本切片没有组件提交 grant),且顺序要求确实写在本行——所以这是规范自洽问题:第 81 行与第 84 行对同一字段承诺了两种语义,而 Java 控制面把 phases 存成 Set<String>,Jackson 按 hash 序序列化。
建议修复:二选一,并在中英两版的 81/84 行同时写明(中文版同一规则「phase 顺序也相同」):要么把顺序定为规范且必需(「按计划注册的顺序」),并在 parseOperationGrant 中强制,使换序列表根本不是合法 grant;要么保持范围规则与顺序无关,续租按其定义的语义判定——把 phases 当集合比较(长度相等加成员相等),而不是对整个 grant 做 JSON.stringify。
修复约束:解析器对 phases 除 PHASE_PATTERN 与 maxGrantPhases 外唯一的约束是去重检查(managed-extension-record.ts:546-548),与文档第 81 行一致——规范化顺序的修复必须同时落在解析器与文档,否则后继检查会继续拒绝解析器接受的 grant。
修复见证:新增一个后继 fixture 对——同 operationRevision、更晚 expiresAt、["query_receipt","send_segment"] 对 ["send_segment","query_receipt"]——其 valid 标签在两侧回放中钉住所选语义;今天没有纯换序用例,增删顺序要求都不会惊动任何测试,新用例落地后删除所选规则会使两侧变红。
— qwen3.8-max via Qwen Code /review (v0.24.7)
| | ---------------------------------------------- | ------------------------------------------------------------- | | ||
| | `recovery_blocked` | a recovery reason, required | | ||
| | `running`, `waiting` | null, or a recovery reason: the run goes on in a degraded way | | ||
| | `failed` | null, or a quota reason | |
There was a problem hiding this comment.
[Suggestion] R1-4: No run that has ended can state that a recovery cause ended it — the five recovery reasons are usable only while the run is open.
The reasons table allows a recovery reason only on recovery_blocked, running and waiting, and forces null on reserved/admitted/settled/cancelled (line 167); assertReason implements exactly that (managed-extension-record.ts:671-673), and the Java validator mirrors it. But both transition tables allow recovery_blocked → failed on the run line and outcome_unknown → not_started_proven on the execution line — so H4's child_run whose Runtime was lost, and whose work is later proven never to have run, steps to failed and must drop its runtime_lost reason on that commit: a producer that keeps it produces a record both validators refuse (probed in both languages). The SessionTaskView that Decision 2 says H0c rebuilds from these blocks then reports a failed child run with no cause, at the one state the user actually reads, while the contract holds an exact word for what happened. The corpus has no case pairing a terminal state with a recovery reason in either direction.
Witness:
TS: parseExtensionRun(recovery_blocked / runtime_lost / outcome_unknown) => ACCEPTED
parseExtensionRun(failed + runtime_lost + not_started_proven) => REFUSED
(run.reason does not fit the failed state.)
parseExtensionRun(failed + reason null) => ACCEPTED
successor(blocked -> failed+runtime_lost) => false; (blocked -> failed+null) => true
Java: same four verdicts; all five recovery reasons REFUSED on failed
Suggested fix: widen the failed row (line 166 and zh-CN 167) to "null, a recovery reason or a quota reason" and mirror it in assertReason's failed arm (reason === null || isQuotaReason(reason) || isRecoveryReason(reason)), keeping the per-reason constraints (:675-694) untouched so no reason can attach to a state that contradicts it. If failed is deliberately reserved for quota causes, say so in the doc, name which record carries a recovery-caused ending, and state that a projection must read the cause from the previous revision.
Fix constraint: managed-extension-record.ts:678-680 (the execution_corrupt biconditional) and :778-784 (a terminal state requires an execution in PROVEN_EXECUTION_STATES = settled/not_started_proven) mean widening failed cannot admit failed+execution_corrupt or failed+outcome_unknown — only runtime_lost, dispatch_unknown and handler_unavailable become reachable there, which is the intended effect and must stay that way.
Fix witness: a run fixture with state: "failed", reason: "runtime_lost", a non-null runtime and execution: "not_started_proven", plus the successor pair recovery_blocked/runtime_lost → failed/runtime_lost — both are rejected in TS and Java today, so the case pins whichever verdict the author chooses and goes red if the failed arm is later changed without the doc.
中文说明
任何已结束的运行都无法陈述「是某个恢复原因使它结束」——五个恢复原因只能在运行仍打开时使用。
原因表只允许 recovery_blocked、running、waiting 携带恢复原因,并强制 reserved/admitted/settled/cancelled 为 null(第 167 行);assertReason 逐字实现(managed-extension-record.ts:671-673),Java 校验器镜像。但两张转移表都允许运行线 recovery_blocked → failed、执行线 outcome_unknown → not_started_proven——于是 H4 的 child_run:Runtime 丢失后被证明从未运行,走到 failed 时必须在那次提交里丢掉 runtime_lost;保留它的生产者会写出两种语言都拒绝的记录(两侧实测)。Decision 2 说 H0c 据这些 run block 重建 SessionTaskView,于是任务列表在用户唯一会读的那个状态上报告一个没有原因的失败运行,而契约里明明有描述此事的确切词汇。语料中没有任何用例在任一方向上把终态与恢复原因配对。
建议修复:把 failed 行(第 166 行与中文版 167 行)放宽为「null、一个恢复原因或一个配额原因」,并在 assertReason 的 failed 分支镜像(reason === null || isQuotaReason(reason) || isRecoveryReason(reason)),保持各原因自身的约束(:675-694)不动,使任何原因都不能出现在与其矛盾的状态上。若 failed 刻意只留给配额原因,请在文档中写明,指明哪条记录承载「因恢复原因而结束」,并说明投影必须从前一修订读取原因。
修复约束:managed-extension-record.ts:678-680(execution_corrupt 双条件)与 :778-784(终态要求执行处于 PROVEN_EXECUTION_STATES = settled/not_started_proven)决定了放宽 failed 不可能放进 failed+execution_corrupt 或 failed+outcome_unknown——只有 runtime_lost、dispatch_unknown、handler_unavailable 变得可达,这正是预期效果,且必须保持如此。
修复见证:新增一个 run fixture——state: "failed"、reason: "runtime_lost"、非空 runtime、execution: "not_started_proven"——外加后继对 recovery_blocked/runtime_lost → failed/runtime_lost;两者今天在 TS 与 Java 中都被拒绝,因此用例钉住作者选择的任一裁决,日后有人只改 failed 分支不改文档就会变红。
— qwen3.8-max via Qwen Code /review (v0.24.7)
| | `failed` | null, or a quota reason | | ||
| | `reserved`, `admitted`, `settled`, `cancelled` | null | | ||
|
|
||
| - `outcome_unknown` requires an execution in `outcome_unknown`; `execution_corrupt` is the reason exactly when the execution is `corrupt`; `runtime_lost` requires the `runtime` it lost, and a run blocked on it is not `running_attached`; `dispatch_unknown` requires a `dispatchId`. |
There was a problem hiding this comment.
[Suggestion] R1-5: runtime_lost's exclusion bars only running_attached — a run whose execution is already proven ended may stay recovery_blocked + runtime_lost forever, contradicting the reason's own definition.
The rule as implemented (TS :684-690, Java :350-354, schema :1231-1252, and this line in both languages) refuses only the attached shape. Executed against the reviewed commit on the corpus's own blocked-lost-runtime record with only execution changed: settled — parse ACCEPTED, successor true; not_started_proven — ACCEPTED, true; only running_attached is refused. recovery_blocked is not terminal, so the terminal-needs-proven-execution rule (:778-785) never fires on it. The chain is exactly the doc's Rebuild narrative minus the rebuild: Runtime lost → blocked, then the original system supplies its proof and the execution steps to settled/not_started_proven, and nothing requires the producer to move the run off recovery_blocked. The record then asserts both "the work is proven done (or proven never run)" and "blocked because the Runtime was lost and not re-attached" — contradicting the definition at line 181/zh-CN:184 ("A blocked run was not re-attached"). H0c's projection shows a task whose result already exists as permanently blocked; the not_started_proven variant is sharpest — proven never started, yet blocked on a lost Runtime with nothing to recover.
Witness:
base = corpus blocked-lost-runtime (valid), only `execution` changed:
outcome_unknown parse=ACCEPTED successor=true
settled parse=ACCEPTED successor=true <-- contradiction ships
not_started_proven parse=ACCEPTED successor=true <-- contradiction ships
running_attached parse=REFUSED (run cannot be blocked on a lost Runtime while attached.)
fix arm (guard extended to PROVEN_EXECUTION_STATES): settled/not_started_proven REFUSED,
outcome_unknown still ACCEPTED, 520-case replay digest unchanged (502a524cfb62f244), vitest 530 passed
Suggested fix: complete the rule in this line (and the zh-CN twin) — "runtime_lost requires the runtime it lost, and a run blocked on it has an execution that is neither running_attached nor proven to have ended" — implemented in both languages by extending the existing guard to run.execution === 'running_attached' || PROVEN_EXECUTION_STATES.includes(run.execution) under recovery_blocked, mirrored in the Java require, plus the schema guard (the rule is expressible: then.execution becomes not: {enum: [running_attached, settled, not_started_proven]}) and fixture cases in both directions.
Fix constraint: the guard must fire only under recovery_blocked — line 200/zh-CN:228 ("the run may keep runtime_lost as its reason to show that observations during the gap may be missing") requires that a rebuilt running + running_attached run may still carry runtime_lost; and the proven-ended set should reuse PROVEN_EXECUTION_STATES (managed-extension-record.ts:308-311), not a re-typed state list. The reused error message ("…while attached") is semantically wrong for the two new arms — give them their own.
Fix witness: new invalid runCases/runSuccessorCases entries (e.g. runtime-lost-after-proven-end) replayed by managed-extension-record.test.ts it.each and ManagedExtensionRecordContractTest.replay — removing the extended guard must flip them from refused to accepted and redden both suites; the pair case must flip isExtensionRunSuccessor true → false.
中文说明
runtime_lost 的排他规则只排除了 running_attached——执行已被证明结束(settled/not_started_proven)的运行仍可停留在 recovery_blocked + runtime_lost,与该原因自己的定义直接矛盾。
按实现的规则(TS :684-690、Java :350-354、schema :1231-1252 与本行两种语言)只拒绝 attached 形状。在被审提交上以语料自带的 blocked-lost-runtime 为基、只改 execution 实测:settled——parse 接受、后继 true;not_started_proven——同样;只有 running_attached 被拒。recovery_blocked 不是终态,所以「终态需要已证明结束的执行」规则(:778-785)对它不生效。这条链正是文档 Rebuild 叙事去掉重建的那一半:Runtime 丢失 → 阻塞,随后原系统给出证明、执行走到 settled/not_started_proven,而没有任何规则要求生产者把运行推离 recovery_blocked。记录于是同时断言「工作已被证明完成(或被证明从未运行)」和「因 Runtime 丢失且未重新 attach 而阻塞」——与第 181 行/中文版 184 行的定义(「阻塞的运行未能重新 attach」)矛盾。H0c 的投影会把一个结果已存在的任务永久显示为 blocked;not_started_proven 变体最荒谬——工作被证明从未开始,却仍以「Runtime 丢失」为由阻塞,无任何可恢复之物。
建议修复:把本行(及中文孪生)的规则补全——「runtime_lost 要求有它失去的 runtime;因它而阻塞的运行,其执行既不能是 running_attached,也不能是已被证明结束的执行」——两种语言把现有守卫扩为 recovery_blocked 下 run.execution === 'running_attached' || PROVEN_EXECUTION_STATES.includes(run.execution),Java 的 require 镜像,schema 同加守卫(该规则可表达:then.execution 改成 not: {enum: [running_attached, settled, not_started_proven]}),fixtures 双向补例。
修复约束:守卫必须只在 recovery_blocked 下生效——第 200 行/中文版 228 行(「运行可保留 runtime_lost 作为原因,以表明间隙期间的观测可能缺失」)要求重建后的 running + running_attached 仍可携带 runtime_lost;「已被证明结束」的集合应复用 PROVEN_EXECUTION_STATES(managed-extension-record.ts:308-311),不要重抄状态名。沿用的报错文案(「…while attached」)对两个新臂语义不对,应另给一条。
修复见证:新增 invalid 的 runCases/runSuccessorCases 用例(如 runtime-lost-after-proven-end),由 managed-extension-record.test.ts 的 it.each 与 ManagedExtensionRecordContractTest.replay 回放——去掉扩展守卫必须让它们从被拒翻成被接受、两侧套件变红;pair 用例还须让 isExtensionRunSuccessor 由 true 变 false。
— qwen3.8-max via Qwen Code /review (v0.24.7)
| "kinds": { | ||
| "monitorRun": "managed-monitor_run", | ||
| "monitorOutput": "managed-tool-result-manifest" | ||
| }, |
There was a problem hiding this comment.
[Suggestion] R1-6: The corpus uses managed-runtime-receipt and managed-monitor-observation as the kind vocabulary for startReceiptRef/lastObservationRef, but these two strings exist nowhere else in the repository and this kinds block does not declare them.
Repo-wide grep (excluding node_modules/dist/.git): managed-runtime-receipt 156 uses, managed-monitor-observation 146 — every hit inside this fixture; external hits 0. A kind→path mapping confirms each string appears only in startReceiptRef / lastObservationRef respectively. The design docs write only "durable ref" for these three refs (only outputRef gets a named kind); the schema's $defs/durableRef.properties.kind is a bare $ref to $defs/id with no const/enum; and parseMonitorRun validates kind/schemaVersion only for outputRef. The sibling contract's convention is that its kinds block IS the complete vocabulary (managed-tool-result-v1.fixtures.json declares 3 kinds and the corpus uses exactly those 3); this fixture uses 38 distinct kinds and declares 2. (commandRef's managed-tool-args is fine — it matches existing usages, e.g. managed-runtime-tool-v3-host.test-helper.ts:127.) When H3 lands, the TS Monitor tool and the Java Runtime that issues start receipts each pick their own kind string; both pass every validation and commit, and the disagreement surfaces only later, at the notification/resume path that parses the ref by kind.
Witness:
grep -rn (excl. node_modules/dist/.git): managed-runtime-receipt 156 | managed-monitor-observation 146
-> every hit in this fixture; external hits 0
kind->path map: managed-runtime-receipt only in startReceiptRef; managed-monitor-observation only in lastObservationRef
sibling convention: managed-tool-result-v1.fixtures.json declares 3 kinds, corpus uses exactly 3
fixWitness mutation executed (add one line to this kinds block only):
x validates the shared fixtures against the shared schema x pins the domains, limits, kinds, ...
(2 failed | 528 passed) — the block is pinned in three places
Suggested fix: add the two kinds to this kinds block (e.g. "startReceipt": "managed-runtime-receipt", "observation": "managed-monitor-observation"), mirrored in MANAGED_EXTENSION_RECORD_KINDS (managed-extension-record.ts:45-48) and the Java kinds constant; and either name them in the design docs' Monitor-run table (both languages) or state explicitly that these refs' kinds are opaque until H3, so readers do not mistake the fixture strings for normative vocabulary. Do NOT add a const to the schema's $defs/durableRef.kind — that $def is shared by commandRef, startReceiptRef, lastObservationRef, outputRef and recordRef, and one const would kill the other four.
Fix constraint: MANAGED_EXTENSION_RECORD_KINDS.monitorOutput is derived, not literal (monitorOutput: MANAGED_TOOL_RESULT_KINDS.manifest, managed-extension-record.ts:47; managed-tool-result.ts:28) — new entries must follow the same single-source style rather than rewriting the derived kind as a literal, or the two contracts drift independently.
Fix witness: managed-extension-record.test.ts "pins the domains, limits, kinds, state lines and reasons" (expect(fixtures.kinds).toStrictEqual({...MANAGED_EXTENSION_RECORD_KINDS})) and the Java whole-node pin — the executed mutation shows adding a line here without moving both constants reddens both tests.
中文说明
语料把 managed-runtime-receipt(156 次)和 managed-monitor-observation(146 次)当作 startReceiptRef/lastObservationRef 的 kind 词汇使用,但这两个字符串在整个仓库中只存在于这个 JSON 文件里,而文件自己的 kinds 常量块并没有声明它们。
设计文档对这三个 ref 只写「durable ref」(只有 outputRef 有具名 kind);schema 的 $defs/durableRef.properties.kind 是指向 $defs/id 的裸 $ref,无 const/enum;TS 侧 parseMonitorRun 也只对 outputRef 校验 kind/schemaVersion。兄弟契约的既有约定是「kinds 块 = 该契约完整词汇表」(managed-tool-result-v1.fixtures.json 声明 3 个、语料恰用这 3 个),本文件用了 38 个 distinct kind 却只声明 2 个。(commandRef 的 managed-tool-args 没问题——与既有代码用法一致,如 managed-runtime-tool-v3-host.test-helper.ts:127。)H3 落地时,TS Monitor 工具与签发 start receipt 的 Java Runtime 会各自挑一个 kind 字符串;两边都能通过全部校验并提交,不一致不会在提交时报错,只在通知/恢复路径按 kind 解析 ref 时失败。
建议修复:把这两个 kind 补进 kinds 块(如 "startReceipt": "managed-runtime-receipt"、"observation": "managed-monitor-observation"),并同步进 MANAGED_EXTENSION_RECORD_KINDS(managed-extension-record.ts:45-48)与 Java kinds 常量;同时在设计文档 §Monitor run 表格(中英两版)写明这两个 kind,或明确写下「这两个 ref 的 kind 在 H3 之前是不透明的」,以免读者把 fixture 里的字符串当成规范词汇。不要给 schema 的 $defs/durableRef.kind 加 const:该 $def 被 commandRef、startReceiptRef、lastObservationRef、outputRef、recordRef 共用,一个 const 会同时打死其余四个。
修复约束:MANAGED_EXTENSION_RECORD_KINDS.monitorOutput 是派生而非字面量(monitorOutput: MANAGED_TOOL_RESULT_KINDS.manifest,managed-extension-record.ts:47;managed-tool-result.ts:28),新增条目必须沿用这种单一来源写法,不能把已有的派生 kind 改写成字面量,否则两个契约会各自漂移。
修复见证:managed-extension-record.test.ts 的「pins the domains, limits, kinds, state lines and reasons」(expect(fixtures.kinds).toStrictEqual({...MANAGED_EXTENSION_RECORD_KINDS}))与 Java 的整节点 pin——已实跑的变异证明:只往这里的 kinds 块加一行而不动两侧常量,两个测试都会红。
— qwen3.8-max via Qwen Code /review (v0.24.7)
| "valid": false, | ||
| "pin": { | ||
| "definitionId": "schedule-nightly", | ||
| "definitionRevision": 1e400, |
There was a problem hiding this comment.
[Suggestion] R1-7: Four boundary values are encoded as the JSON literal 1e400, which is outside IEEE-754 double range — the out-of-repo Python reference that regenerates this file cannot round-trip its own output, and the regenerated corpus is unloadable by both shipped parsers.
The four sites: grantCases/expiry-past-double-range.expiresAt (:1501), grantSuccessorCases/next-past-double-range.next.expiresAt (:4727), pinCases/revision-past-double-range.definitionRevision (:4870), monitorRunCases/max-events-past-double-range.maxEvents (:15185). Every double-based parser turns the literal into Infinity — a value that is not JSON. The doc names a Python implementation as the independent generator; json.load yields inf at all four places, json.dumps(..., allow_nan=False) raises ValueError: Out of range float values are not JSON compliant, and plain json.dumps writes the bare token Infinity — which both shipped parsers reject outright (measured: JSON.parse SyntaxError; the re-emitted file fails the real TS suite at fixture load). A second, independent round-trip blocker: the 11 deliberate lone-surrogate strings (e.g. "\ud800" at :2172) force ensure_ascii=True (or preserved escapes) in any regeneration path — replacing 1e400 alone does not make the corpus Python-round-trippable.
Witness:
python3: non-finite floats after json.load: 4 (the sites above)
dumps(allow_nan=False): ValueError: Out of range float values are not JSON compliant
plain dumps: emits bare Infinity x4 -> JSON.parse: SyntaxError: Unexpected token 'I'
re-emitted file vs the REAL TS suite: Test Files 1 failed, Tests no tests (fixture load)
fix arm (1e400 -> 1e308 x4, scratch tree): 530 passed; all four verdicts and messages byte-identical;
next-past-double-range successor still false
lone surrogates: re-encode with ensure_ascii=False -> UnicodeEncodeError (surrogates not allowed)
Suggested fix: replace 1e400 with a finite literal far past every bound, e.g. 1e308, at all four sites (measured: no verdict or message changes in either language). If the genuinely non-finite behaviour must be pinned, pin it in each language's own unit test, where the value is constructed in memory instead of serialized into the shared file — note doc:63's "a number past the double range is not an integer" rule keeps its TS witness either way (Number.isSafeInteger rejects Infinity with the identical message), but the case ids would no longer claim a non-finite input the shared file cannot carry.
Fix constraint: the replacement must stay above the ceilings both languages enforce — maxTimeMs: 8_640_000_000_000_000 (managed-session-records.ts:36) and MAX_TIME = 8_640_000_000_000_000L (ManagedExtensionRecords.java:120) — and must not collide with the corpus's exact brackets (latest-expiry/expiry-past-maximum, largest-revision/revision-past-safe-count); 1e308 satisfies both.
Fix witness: a raw-text guard inside "validates the shared fixtures against the shared schema" — read the fixture file as text, match every JSON number literal, assert Number.isFinite(Number(match)) for each; restoring 1e400 at any of the four sites turns it red, while the existing it.each replays keep all four cases at valid: false.
中文说明
共享语料把四个边界值编码为 JSON 字面量 1e400,超出 IEEE-754 double 范围——文档指名的仓外 Python 参考实现无法往返自己的输出,再生成的语料两个 shipped 解析器都加载不了。
四处站点:grantCases/expiry-past-double-range.expiresAt(:1501)、grantSuccessorCases/next-past-double-range.next.expiresAt(:4727)、pinCases/revision-past-double-range.definitionRevision(:4870)、monitorRunCases/max-events-past-double-range.maxEvents(:15185)。任何基于 double 的解析器都会把该字面量变成 Infinity——一个不是 JSON 的值。实测:json.load 在四处得到 inf;json.dumps(..., allow_nan=False) 抛 ValueError;普通 json.dumps 写出裸 token Infinity——两个 shipped 解析器都直接拒绝该文件(JSON.parse SyntaxError;重 emit 的文件让真实 TS 套件在 fixture 加载阶段全灭)。另有第二个独立的往返障碍:11 处刻意的孤立代理项字符串(如 :2172 的 "\ud800")要求任何再生成路径使用 ensure_ascii=True(或保留原始转义)——只替换 1e400 并不能让语料可被 Python 往返。
建议修复:把四处 1e400 换成远离所有边界的有限字面量,例如 1e308(实测两种语言的判定与消息逐字不变)。若确实要钉住「非有限值」这一行为,请放进各语言自己的单元测试——值在内存中构造,而不是序列化进共享文件;注意 doc:63「超出 double 范围的数字不是整数」这条规则在 TS 侧的见证不受影响(Number.isSafeInteger 对 Infinity 报同样的消息),只是这些用例的 id 不再宣称共享文件根本无法承载的非有限输入。
修复约束:替换值必须仍高于两种语言强制的上限——maxTimeMs: 8_640_000_000_000_000(managed-session-records.ts:36)与 MAX_TIME = 8_640_000_000_000_000L(ManagedExtensionRecords.java:120)——且不得与语料既有的精确括号(latest-expiry/expiry-past-maximum、largest-revision/revision-past-safe-count)相撞;1e308 两者都满足。
修复见证:在「validates the shared fixtures against the shared schema」里加一个原文守卫——把 fixture 文件当文本读、匹配每个 JSON 数字字面量、断言 Number.isFinite(Number(match));在四个站点中任何一处恢复 1e400 都会让它变红,而现有 it.each 回放保持四例 valid: false。
— qwen3.8-max via Qwen Code /review (v0.24.7)
| }, | ||
| { | ||
| "id": "waiting-on-a-dispatch", | ||
| "valid": true, |
There was a problem hiding this comment.
[Suggestion] R1-8: The waiting row is the only row of the doc's Reasons table whose negative half the 520-case corpus never exercises — all six waiting × quota-reason cells are empty, so widening the waiting arm flips 0 of 520 in both languages.
All 440 run blocks in the corpus were enumerated: the waiting cells present are exactly dispatch_unknown (valid, this case), execution_corrupt (invalid, decided by the corrupt-execution rule, not the reason arm) and null. The sibling rows are complete in both directions — running has running-degraded (valid) and running-with-quota-reason (invalid); failed has six valid quota cases plus failed-with-recovery-reason (invalid). The mutation was executed: widening assertReason (managed-extension-record.ts:670-673) so waiting also permits a quota reason flips 0 of 520 verdicts (replay digest identical, suite 530 passed), while a direct probe shows waiting + rate_limit REFUSED by the real module and ACCEPTED by the mutant. The Java twin has the same shape (case "running", "waiting" -> reason == null || recovery, ManagedExtensionRecords.java:336); the schema does state the rule (:1077-1101), but the disagreement test compares module to schema only over corpus cases, so the third gate is blind too.
Witness:
state x quota-reason matrix over 440 run blocks: waiting row all-zero (the only such row)
BASE waiting + rate_limit/count_limit/budget_exhausted -> REFUSED (run.reason does not fit the waiting state.)
MUTANT same three inputs -> ACCEPTED; REPLAYED 520 CASES mismatches 0, verdict digest identical; vitest 530 passed
fix arm (add waiting-with-quota-reason, valid:false): MUTANT 1 failed | 530 passed; PRISTINE 531 passed;
'agrees with the schema...' stays green (schema already refuses the shape -> no BEYOND_SCHEMA entry needed)
Suggested fix: add one invalid runCases entry, e.g. waiting-with-quota-reason, cloned from waiting-on-a-dispatch (:5241-5254, keeping dispatchId: "dispatch-1") with "reason": "count_limit" — probed: it is decided by the waiting arm alone, not swallowed by an earlier check (not an R1-20-style vacuous case).
Fix constraint: the doc states the per-list counts as fact in both languages ("520 cases: … run blocks (144) …", line 236 and the zh-CN twin) and nothing pins them (R1-19), so adding a case must update 144→145 and 520→521 in both files in the same change or the doc goes false silently.
Fix witness: it.each(fixtures.runCases) (managed-extension-record.test.ts:337) and replaysEveryRecordCase (ManagedExtensionRecordContractTest.java:152) — the new case reddens both under the executed widening mutation and is green on the pristine module (both arms measured).
中文说明
waiting 是文档 Reasons 表中唯一负面半边没有语料见证的行——520 例中没有任何 run block 把 state: "waiting" 与六个配额原因中的任何一个配对,因此把 waiting 臂放宽的变异在两种语言里都翻转 0/520。
枚举语料全部 440 个 run 块:waiting 出现的理由恰好是 dispatch_unknown(valid,即本用例)、execution_corrupt(invalid,由 corrupt-execution 规则判定而非理由臂)与 null。兄弟行双向完备——running 有 running-degraded(valid)与 running-with-quota-reason(invalid);failed 有六个 valid 配额用例加 failed-with-recovery-reason(invalid)。变异已实跑:放宽 assertReason(managed-extension-record.ts:670-673)使 waiting 也允许配额原因后,520 例判定 0 处翻转(回放摘要不变、套件 530 全绿),而直接探针显示 waiting + rate_limit 被真实模块拒绝、被变异体接受。Java 孪生同形(case "running", "waiting" -> reason == null || recovery,ManagedExtensionRecords.java:336);schema 倒是表达了该规则(:1077-1101),但分歧测试只在语料用例上比较模块与 schema,所以第三道闸同样是盲的。
建议修复:新增一条 invalid runCases,例如 waiting-with-quota-reason——从 waiting-on-a-dispatch(:5241-5254,保留 dispatchId: "dispatch-1")克隆并把 reason 改为 "count_limit"。已探针验证:它由 waiting 臂单独判定、不会被更早的检查吞掉(不是 R1-20 那类空洞用例)。
修复约束:文档在两种语言里把逐表计数写成事实(第 236 行「520 个用例:……运行块(144)……」及中文孪生),且没有测试钉住它们(R1-19),因此加例必须在同一改动中把两处计数同步改成 145/521,否则文档静默变假。
修复见证:it.each(fixtures.runCases)(managed-extension-record.test.ts:337)与 replaysEveryRecordCase(ManagedExtensionRecordContractTest.java:152)——在已实跑的放宽变异下新用例使两侧变红,在原始模块上保持绿(两臂均已实测)。
— qwen3.8-max via Qwen Code /review (v0.24.7)
| "valid": false, | ||
| "run": { | ||
| "state": "settled", | ||
| "reason |
There was a problem hiding this comment.
[Suggestion] R1-9: The effectId half of the set-once identities' first-appearance allowance has no witness anywhere in the corpus — forbidding an effectId from ever appearing flips 0 of 520 cases in both replays.
Of the five identities the doc says "may be set once and never change after that" (managed-extension-record.ts:832-838), the corpus witnesses the first appearance of four — definition, executionCallId, dispatchId (identity-and-pin-set-later) and deliveryId (delivery-first-seen-planned, channel-delivery-planned-after-the-end). No valid runSuccessorCase takes effectId from null to a value, and it cannot be folded into identity-and-pin-set-later because parseExtensionRun refuses executionCallId and effectId together (:750-752). The mutation sweep was executed with sibling controls, so the harness's sensitivity is measured rather than assumed: every sibling limb is killed by exactly the cases named, while the effectId limb survives.
Witness:
baseline: FAILED=0 TOTAL=530
effectId setOnce -> sameJson: FAILED=0 <-- survives
definition (control): FAILED=1 'identity-and-pin-set-later'
executionCallId (control): FAILED=2 'reserved-to-admitted', 'identity-and-pin-set-later'
dispatchId (control): FAILED=1 'identity-and-pin-set-later'
deliveryId (control): FAILED=2 'delivery-first-seen-planned', 'channel-delivery-planned-after-the-end'
startReceiptRef (control): FAILED=1 'dispatched-to-attached'
corpus sweep: runSuccessorCases null->value for effectId: [] (all four siblings non-empty)
Suggested fix: add one valid runSuccessorCase, e.g. effect-id-set-later — previous = the admitted run with every identity, execution, runtime, definition and delivery null; next = the same with effectId: "effect-1" and execution: "intent". Both suites replay it with no code change.
Fix constraint: the case counts are normative prose in both languages (line 236: "520 cases: … run revisions (57) …") and no test pins them (R1-19), so adding a case must update 57→58 and 520→521 in both files or the doc goes false silently.
Fix witness: the TS it.each(fixtures.runSuccessorCases) and Java replayPairs("runSuccessorCases", …) on the new case id — both go red when setOnce is replaced by sameJson for effectId at :834, which is the mutation that currently flips nothing.
中文说明
五个「可设置一次、之后不得改变」的身份中,effectId 的首次出现许可在整份语料里没有任何见证——把「允许首次出现」收紧为「永远不得出现」的变异在两种语言的回放里都翻转 0/520。
语料见证了其中四个的首次出现——definition、executionCallId、dispatchId(identity-and-pin-set-later)与 deliveryId(delivery-first-seen-planned、channel-delivery-planned-after-the-end)。没有任何 valid 的 runSuccessorCase 把 effectId 从 null 变成有值,而且它无法折进 identity-and-pin-set-later,因为 parseExtensionRun 拒绝 executionCallId 与 effectId 同时存在(:750-752)。变异扫描连同兄弟对照一起实跑,harness 的敏感性是实测而非假设:每个兄弟分支都被点名的用例杀死,唯独 effectId 分支存活。
具体后果:某个 domain phase 先以 effectId: null 的 intent 提交运行、随后补上该 phase 的 effect 身份——在收紧后的规则下第二次修订会在 Runtime 门禁被拒,运行永远无法记录它实际执行所用的 effect,而两种语言的每个测试仍然通过。
建议修复:新增一条 valid 的 runSuccessorCase,例如 effect-id-set-later:previous = 每个身份、execution、runtime、definition、delivery 全为 null 的 admitted 运行;next = 同样内容加 effectId: "effect-1" 与 execution: "intent"。两套测试无需改代码即可回放。
修复约束:逐表计数是两种语言的规范性文字(第 236 行「520 个用例:……运行修订(57)……」),且没有测试钉住(R1-19),加例必须在同一改动中把两个文件的 57→58、520→521 一起改掉,否则文档静默变假。
修复见证:TS 的 it.each(fixtures.runSuccessorCases) 与 Java 的 replayPairs("runSuccessorCases", …) 回放新用例 id——把 :834 的 effectId 从 setOnce 换成 sameJson 时两者都必须变红,而这正是今天翻转不了任何判定的那个变异。
— qwen3.8-max via Qwen Code /review (v0.24.7)
| "valid": false, | ||
| "run": { | ||
| "state": "settled", | ||
| "reason |
There was a problem hiding this comment.
[Suggestion] R1-10: This case is rejected during parse before the successor rule it names is ever reached, so it pins nothing — mutation-verified in both directions.
The next block carries both "reason": "runtime_lost" and "runtime": null, so assertReason fails first (managed-extension-record.ts:682: "run.reason runtime_lost needs the Runtime binding it lost.") and isExtensionRunSuccessor returns false at its parse guard without ever evaluating the re-attach rule the id names ("a re-attach under a later generation must not clear the Runtime binding", :866). This violates the doc's own contract at line 236: "Each invalid case is aimed at one rule." A sweep of all 123 pair cases confirms this is the ONLY one whose id names a pair rule yet stops at parse (the other nine parse-stopped cases are named invalid-*/next-past-double-range — their parse failure IS the rule).
Witness:
1) intact: parseExtensionRun(next) threw 'run.reason runtime_lost needs the Runtime binding it lost.'
reason:null variant parses, and isExtensionRunSuccessor still returns false (via the real rule)
2) mutant (delete 'after.runtime !== null &&' at :866), fixture as-is: 530 passed — the guard is unpinned
3) mutant + reason:null fix: 1 failed — TypeError: Cannot read properties of null (reading 'generation')
4) fix only, mutant reverted: 530 passed — the fix itself is safe
Suggested fix: change this case's next.reason from "runtime_lost" to null (a running state allows a null reason). Measured: next then parses, isExtensionRunSuccessor(previous, next) still returns false, valid: false stands, but the rejection now comes from the successor rule itself. If the author prefers to keep the "a rebuilt run may still carry runtime_lost" semantics, add a separate next.reason: null case rather than leaving this one stopped at parse.
Fix constraint: managed-extension-record.ts:682 fails runtime_lost when runtime is null, so the fixed next cannot keep runtime_lost with a null runtime; keeping the reason requires supplying a binding, which stops testing the clears-binding rule.
Fix witness: this case's replay in managed-extension-record.test.ts it.each(fixtures.runSuccessorCases) and ManagedExtensionRecordContractTest.replayPairs — the executed mutation proof: deleting after.runtime !== null && at :865-870 keeps 520 cases green today; after the fix the same mutation must redden this case with the TypeError above.
中文说明
该用例的 next 块在 parseExtensionRun 阶段就被拒绝,isExtensionRunSuccessor 从未走到它名字所指的那条规则,因此它没有钉住任何东西。
next 同时写了 "reason": "runtime_lost" 和 "runtime": null,assertReason 先命中 managed-extension-record.ts:682(「run.reason runtime_lost needs the Runtime binding it lost.」),于是「更晚 generation 下重新 attach 不得清空 Runtime binding」这条后继规则(:866)在本用例上完全未被执行——违反文档第 236 行自己的约定「每个无效用例都针对一条规则」。对全部 123 个 pair 用例的扫描确认:这是唯一一个 id 指向 pair 规则、却停在 parse 的用例(其余九个停在 parse 的用例本来就以 invalid-*/next-past-double-range 命名——parse 失败正是它们要钉的规则)。
变异已双向实跑:现状下删掉 :866 的 after.runtime !== null &&,520 例 0 处翻转(守卫无人钉住);按建议修复后再做同一变异,该用例以 TypeError: Cannot read properties of null (reading 'generation') 变红;只改 fixture、还原变异,530 全绿(修复本身安全)。
建议修复:把该用例 next 块里的 "reason": "runtime_lost" 改为 "reason": null(running 状态允许 reason 为 null)。实测改后:next 可正常 parse,isExtensionRunSuccessor(previous, next) 仍为 false,"valid": false 不变,但拒绝理由变成了后继规则本身。若作者更想保留「重建后仍可带 runtime_lost」的语义,应另加一条 next.reason: null 的用例,而不是让这条停在 parse 阶段。
修复约束:managed-extension-record.ts:682 在 runtime 为 null 时拒绝 runtime_lost,所以修复后的 next 不能在 runtime 为 null 时继续保留 runtime_lost;若要保留该 reason,就必须同时给出 binding,那样便不再检验「清空 binding」这条规则。
修复见证:该用例在 managed-extension-record.test.ts 的 it.each(fixtures.runSuccessorCases) 与 ManagedExtensionRecordContractTest.replayPairs 中的回放——变异证明:删除 :865-870 的 after.runtime !== null &&,今天 520 例全绿;修复后同一变异必须让本用例以上述 TypeError 变红。
— qwen3.8-max via Qwen Code /review (v0.24.7)
…eader Round-5 real-stack verification's cross-version matrix disproved the v2 stamp's premise: monitor_run has been a parsed, known domain since #12837 (v0.24.7), and a v1 reader opens a Session holding monitor_run records without complaint. Stamping minimumReader: managed-session/2 on every new Session — while monitor_run remains disabled — would make a rollback or a mixed-version rollout lose access to every Session created in between, and buys protection no deployed binary needs. Revert the stamp to managed-session/1 and drop the per-domain header refusal; the enablement list remains the single domain gate. The stamp rises only when a change genuinely breaks an older reader mid-scan, and only one release after a tolerant reader ships. The golden fixture's genesis header regenerates to v1; the records and authority suites gain v1-header admittance witnesses in place of the refusal cases; the design's reader section now records the reversal with its evidence.
…13265) * docs(managed-agent): design H3 background Shell and Monitor runtime * feat(core): add managed child_run (kind shell) record body for H3 * feat(managed-agent): mirror child_run shell record body in Java store * test(core): commit and rebuild child_run shell chains through the session authority * feat(managed-agent): close child_run reference closure and add stop-request draining * feat(core): add named cgroup unit attach and managed child-run supervisor * feat(cli): admit v3 background shell under the supervised worker registry * feat(core): add open-ended shell stream capture with manifest revisions * fix(core): type-narrow stream capture finalize and refused publish double * fix(cli): close background shell type gaps after the main merge * fix(managed-agent): answer the R1 review round on the child_run contract * fix(managed-agent): repair the CI-only cgroup root and env-guard failures How: wrap the delegated-root probes so a missing or unreadable root answers the documented isolation error on Linux too (macOS refused at the platform check first, which is why only CI saw the raw ENOENT), and document the background Shell environment allowlist in the process.env guard. Why: the Test lane on 4fb9f0a9c8 went red on exactly these two items; both are this branch's own changes, not flake. Test: hook-command-cgroup and process-env-guard suites green; tsc clean. * feat(managed-agent): inject the child-run supervisor at worker boot How: registerManagedContextRoutes builds a ManagedChildRunSupervisor from the delegated cgroup root the Hook commands already use (QWEN_MANAGED_HOOK_CGROUP_ROOT) and passes it as the executor's fifth argument; a boot without the delegation keeps the executor's committed refusal instead of failing to boot. Why: executeV3Background landed in the previous increment with the supervisor parameter unwired, so every real-stack background start answered the committed no-supervisor refusal. Test: new wiring case asserts the supervisor is injected exactly when the root is set; managed-context-worker, process-env-guard and managed-background-shell suites green (775); tsc clean. * docs(managed-agent): pin the two-row ledger shape for background processes How: the runtime-ownership section now says the ledger holds two rows — the foreground-length start invocation settling with the handle, and a second execution admitted at start that carries no model result and stays non-terminal until physical exit evidence. Why: the Java ledger writes a result exactly when an execution settles, so a single row that is both handle-delivered and active cannot exist; the two-row shape keeps the settled-if-and-only-if-result invariant while preserving every cited behavior (Runtime holds, evidence-only settle). * feat(managed-agent): commit child_run shell lines from the hosted side How: HostedChildRunSession funnels every record write through one serialized, replay-safe commitExtensionRecord path — admit first with the start call's args as commandRef, dispatch and the set-once managed-runtime-receipt after the physical start, outputRef advance-only, and settlement only on proven exit evidence, a proven failure, or an honored stop request. Live revisions alternate the run line (no self-loops); the authority freeze refuses anything after a terminal revision. Why: the dual path puts product records on the hosted authority, but nothing commits child_run from the hosted side yet — the worker-owned registry intentionally does not touch records. Test: four authority-level cases — full chain with projection, stop request to cancelled, pre-start failure frozen on not_started_proven, and serialization with deep-equal skips. * feat(managed-agent): add the detached capture family to Tool v3 results How: a background Shell start now settles success with a capture object whose status is 'detached' — no reason, no manifest — because the live output streams through the child_run record's growing manifest rather than the result; the shared schema pins both invariants, the TS parser mirrors them, the shared fixtures gain one valid and two invalid envelope cases replayed on both sides, and the Java projector accepts the missing manifest for unavailable or detached captures. Admission refusals settle not_started with a null capture, the shape the durable unstarted family already owns. Why: every Session tool.receipt event lands in the Java delivery projection, which requires a capture object and a committed-if-manifest pairing for anything that started — a success result with a null capture fails that projection and corrupts replay. Test: core contract suite 519 (three new envelope cases), TS serve suites 33 + turn/harness 323 + background 6, Java publication contract and projector suites 8; tsc clean on both packages. * fix(managed-agent): answer the R2 review round on worker, capture and record lines How — the four Criticals first: the stream capture publishes an ended stream's revision only after its seal decision, so a sealed-or-incomplete descriptor never changes again (R2-3); finalize's settle path degrades to an unavailable capture instead of throwing past a writability failure, and the last published revision stands like the worker-cut cap (R2-5); the child_run body now refuses a settled or cancelled run that is not settled execution, a failed run off its two ending lines, start_failed outside not_started_proven, and a process-level failure without a settled execution, on both languages with new witnesses replayed from the shared fixtures (R2-6, R1-30); and the supervision suite skips win32 instead of spawning a shell that cannot exist (R2-7). How — the Suggestions land as: supervisor start now proves membership by cgroup.procs with an fd-3 status channel, failing closed as isolation (R2-37); the bounded EOF wait caps daemon-inherited pipes instead of hanging, and a spawn-time error drains the entry (R2-1/R2-14); registry hasHolds requires its session and setProcessResult carries the full result shape (R2-13/R2-15); create() validates its caller-named unit like attach() (R2-26); the executor checks the journal before any effect, mirrors the closing/is_active re-check after prepare, gates the ninth live Shell as a committed quota refusal, and keeps cgroup membership out of an empty QWEN_MANAGED_HOOK_CGROUP_ROOT (R2-23/R2-19/R2-24/R2-18); the spawn uses the configured shell and the session-context environment (R2-21/R2-20); invalid fixture cases each pin their refusing clause on both languages (R2-36); projection fixtures pin the draining precedence rows and the Java replay reads stopRequested (R2-28/R2-38); the draining Javadoc states the rule the code implements (R2-39); the duplicated ref helper is gone (R2-40); the stream-capture suite is typed and gains the seal-before-publish, page-cursor, late-failure and settle-degrade witnesses (R2-29/30/31/32); and the design doc names the cgroup switch decision, the cursor's single ordinal space, the unconditional host-scope evidence gate and the zh stopped-object (R2-8/9/10/11/12). Test: core suites 788, cli serve suites 776 + background 11, Java record/ projection/store suites 24; tsc clean. * feat(managed-agent): thread the child-run orchestrator and the detached receipt family How: HostedSession gains its session-scoped childRuns orchestrator next to hooks/mcp, the tool turn receives it stored-ahead of the admission branch, the reopen verifier and the recovery replay both accept the third durable receipt family — a blocked delivery with a detached capture, beside the complete and not-started ones — with the publication store correctly left out of its proof, since its durable truth is the child_run record. Why: the H3 background start settles with a handle whose output lives on the record's growing manifest; without the third family, any Session holding a background receipt could never reopen or be replayed. Test: recovery-session suite 52 with two new detached-family cases; harness-session and tool-turn suites green; tsc and lint clean. * test(cli): type the background-shell capture double How: drop the as-unknown cast on the FakeSink return and implement the failCapture member the cast had been hiding (R2-29's cli location). Why: an unchecked double can drift from the interface it claims to match — and it already had. Test: background-shell suite 10/10; tsc clean. * feat(managed-agent): admit background Shell starts from the hosted tool turn How: the turn replaces its two deliberate is_background refusals — exactly when the Session owns its child_run orchestrator and the domain is enabled — with orchestration that mirrors the shared line: the record intent precedes every physical effect, dispatch_started lands the moment the checkpoint commits, a settled detached handle attaches the physical start to the record from the same facts and lands as the third durable receipt family (blocked delivery with a null manifest, replay-validated exactly like the not-started family), and a proven-unstarted refusal settles start_failed on not_started_proven while riding the existing unstarted family unchanged. Malformed is_background values keep their old refusal, and the family stays out while the domain is disabled — the same probe governs both paths. Why: the background start is the first record-bearing, user-visible H3 effect, and without a hosted history family for the handle a Session that ran one could never reopen or be replayed. Test: three new authority-level turn cases — admitted detached flow with admit/dispatch/attach order and the blocked receipt family, a proven-unstarted refuse closing NOT_STARTED, and the disabled-domain refusal preserving its exact text; turn suite 146, plus harness-session, recovery, background and child-run suites 246; tsc clean. * feat(managed-agent): admit background Tool v3 dispatches and acknowledge their detached settle How: the v3 start admission flips from "foreground Shell only" to "Shell only" — background dispatches now begin the same way; the poll loop accepts the settled detached family through the status answer (the only durable settle such a result can ever produce, since nothing is published for a background handle), and the acknowledgement path canonizes the settled handle envelope: a matching blocked receipt with a null manifest and no history revision forwards exactly itself, without consulting the publication receipt. Why: without the detached branch, a background start settled a queued handle only to mark the execution UNKNOWN after the 30-minute foreground-shaped deadline — every background dispatch would wedge the Runtime on its first call, and its acknowledgement would die on the missing publication it was never going to have. Test: new witness drives the detached start end-to-end — reservation, v3 dispatch, settled handle via status, and ack with a null-manifest blocked receipt, with the publication receipt path provably untouched; service suite 139/139, module suite 600/600. * feat(managed-agent): answer shell-status and shell-terminate from the worker registry How: a new private maintenance route sibling to ManagedHookProtocol (/internal/managed-runtime/v3/shells) backed by the process registry — status answers running or unknown (never claimed without evidence), terminate drains with the supervisor's own rules and answers exited with its proof, and an end that cannot be proven is unknown with the hold still registered. The unit name derives from the runtime invocation identity through the shared shellUnitNameOf, so Java, the hosted turn and the worker derive it identically; the executor takes the registry as a constructor injection so the route and the journal share one. Why: every later maintenance verb — recovery reconciliation, the ordered close drain, the reconcile-settle — has to ask the only physical owner these questions; answering them from anything but the registry would invent truth. Test: five worker-side cases — closed-key refusals, unknown for unowned, running scoped to the registered Session, terminate-with-evidence answering exited, and an unproven terminate answering unknown with the hold intact; shell, context-worker, v3-route and child-run suites 784; tsc clean. * feat(managed-agent): admit the background process row and answer its maintenance protocol How: the start dispatching of a background Shell admits a second ledger execution beside the invocation row — the new process row carries no model result and stays PREPARED (non-terminal) until physical proof, so hasActiveBy* holds the Runtime exactly while the process lives. The maintenance protocol mirrors ManagedHookProtocol on its own route: shell-status and shell-terminate with a targetOperationId (both recovery kinds), wire-validated on both transports and accepted through the same workspace-generation and ownership gates. observeBackgroundProcess asks the physical owner and settles the process row with proven evidence — settle through the new repository verb settlePrepared, which mirrors the PREPARED-requestCancel shape on both repositories; anything unproven leaves the row active with its hold, never a claimed end. Why: the ledger writes a result exactly when an execution settles, so a single background execution carrying its handle at start cannot exist; two rows keep the settled-if-and-only-if-result invariant while the Runtime holds precisely, and the maintenance route gives every later verb the worker's physical answer instead of an invented one. Test: two new witnesses — a background start that admits the process row (the invocation SETTLED, the process PREPARED, the Runtime busy until status arrives, then exit evidence settling it and release succeeding), and an unproven status keeping its hold; the runtime also gained the row identity invariant fix and the settle-through-settlePrepared shape; service suite 141/141, module 603/603. * fix(managed-agent): answer maintenance views under the target identity How: the shell-status/shell-terminate view now carries the target operation's identity, exactly as the Hook lookup family does — the requester asks about the process, never about the asking call, and the Java wire validator's expectedId rule requires it. Why: this repair was encoded in the route design but landed uncommitted after ②e-a; the self-audit caught it before it could ride a later batch. * feat(managed-agent): run the background exit leg through the record How: the hosted publisher gains the background family — its registration marker admits the open-ended capture through the session-level admission (proven start on the record, never the model call), the open manifest publishes at prepare, every awaited write or finish forwards the newest manifest revision to the record's outputRef, and the exit finalizes with the same evidence object settling the record. The client acknowledged receipt keeps the detached-blocked shape; drain skips background stores until the ordered close drain wires them; the register guard only binds the foreground model-call family. Why: the start handle already committed, so a background Shell's whole physical truth is output and exit evidence — anything that reached the record another way would let a second tool.receipt contradict its own line, and an open-ended manifest published anywhere but the Session's own store is an unreadable closure for the store-side commit. Test: a full end-to-end authority case — admit-dispatch-attach, bytes through the real rpc, outputRef advancing live, the final sealed manifest reading the same exit (code 3, both streams sealed, manifest complete), the record settled with exactly that evidence, and zero tool.receipt events for the outcome; publisher, turn and child-run suites 174; tsc clean. * feat(managed-agent): close monitor_run revisions over their cited Session resources The H0c closure rule — every ref a record cites must exist in the Session's resource store at commit time — now covers monitor_run on both stores: the authority reads commandRef, startReceiptRef, outputRef and lastObservationRef through verifyExtensionResources, and the Java applyRevision gains the same four-field branch beside child_run. This is the precondition for enabling the monitor domains: without it a monitor revision could anchor its chain to resources nobody committed, and the replay would project a watch whose evidence was never verified. Both fixture rigs stop citing fictional resource identities. The authority test publishes its chain refs as real bodies once per harness and mirrors each fixture-cited ref as a real same-kind placeholder of the stated length, memoized so chain identity stays fixed across revisions. The journal test harness does the same at its generic request entry point, rewriting each cited ref to the placeholder's real digest and appending the mirrored resources, but only for strictly well-formed JSON monitor bodies — a body with trailing content keeps the raw path so the store's grammar refusal still fires first over the closure's. Each side gains a witness that pins the closure itself: a monitor run citing a resource outside its commit is refused with managed_session_resource_missing, and each witness was observed red with its branch disabled before landing. * feat(managed-agent): gate monitor_run behind the managed-session/2 reader The H3 design left the reader-gating mechanism to the implementation; the journal settles it. Headers are immutable — the scanner refuses a repeated managed_session_header_v1 line and any unknown subtype after the header — so no transaction can legally raise a Session's requirement in-journal, and a superseding header line would change the journal format in both languages for a narrow mixed-version window. New Sessions therefore stamp minimumReader managed-session/2 at creation, and every path that commits a monitor_run record — commitDomainRecord, commitExtensionRecord and the generic append guard — additionally refuses when the Session's header names less. Readers still accept requirements up to their own version, so pre-H3 Sessions stay openable everywhere and simply can never receive monitor_run. child_run is deliberately not header-gated: a managed-session/1 reader already parses its name, and the asymmetric hazard lives on the server side, which the established deploy-server-first sequencing covers. The shared Stage H golden is regenerated for the v2 genesis bytes and the Java integration replay passes against it unchanged. New witnesses: new Sessions stamp v2, a v1 Session refuses monitor_run with the precise error (observed red with the gate removed), and both header sides of the version comparison are pinned in the records tests. The design doc records the decision and the now-paid closure precondition in both languages. * feat(managed-agent): funnel monitor_run revisions through a hosted orchestrator The mirror of HostedChildRunSession for the Monitor domain. The managed-runtime worker will own the watch loop, but the dual path puts every product record on the hosted authority, so HostedMonitorSession commits the record line as watch facts arrive: revision 1 with the intent before any side effect, dispatch onto a Runtime binding, the set-once start receipt after the watch starts, observation revisions only while attached with a watermark that never goes back, output advancement only forward, and the terminal shapes — settled by exited/max_events/idle_timeout, failed by a start failure on not_started_proven or a settled watch failure, cancelled by a stop request whose notification watermark may still advance after. Writes serialize per Session and replay-safe by command id with deep-equal skip, exactly like the Shell funnel. The suite drives the whole line through a real authority: projection at admission and after settle, max_events refused before its quota and accepted at it, a second start receipt refused, an observation before attach refused, and an idempotent output advance committing nothing. Quota (enforcement), the debounce-floor observation loop with notification composition, and rebuild-after-loss stay with the runner and route increments that follow. * feat(managed-agent): drive monitor observations through a debounced loop The loop half of the Monitor runner, on top of the hosted funnel: an admitted Monitor dispatches, starts an injected watch executor and runs its own time. Stdout lines aggregate into one observation revision per debounce window with a one-second floor, so the document's reopen bound holds for any debounce a watch asks for; the run settles itself on the contract's own terminal conditions — the observation quota, silence past idle_timeout, a natural exit with its buffered lines flushed first — or is stopped, which terminates the watch and lands cancelled. The executor contract is typed for the worker's cgroup owner (onLine, onExit after the start resolves, a terminate handle); the suite drives the whole line through a real authority with a fake executor and a manual clock: windowed aggregation and its floor, quota settlement and watch termination exactly at max_events, idle settlement re-armed by each accepted observation, exit flush order, start and mid-run failure shapes, and a stop that ignores late lines. Every cross-boundary settle — timer or executor callback — is counted by the loop's done handle, which the close drain will await; window flushes count too, so done never reports idle while an observation commit is in flight. Notification composition and its wake stay with the next increment; quota enforcement and rebuild-after-loss stay with the route slice. * feat(managed-agent): notify from monitor observations in the same transaction The H0c machinery now runs for Monitors: every accepted observation commits its revision together with a notification input and the wake the authority generates for it, and the run's notifiedThrough watermark advances to exactly that observation's sequence. The loop composes the input — monitor-1:notify:N ids, the window's joined lines as the content a woken turn will read, an empty admission, no deadline — and the funnel threads it through commitExtensionRecord, so a recovery replay re-runs nothing the watermark already covers. Dedupe semantics are pinned at both levels: the funnel advances the watermark only when a notification rides the revision (an observation without one keeps it where it was), and the loop suite shows the three events — domain.committed, input.accepted, wake.requested carrying the input's accepted id as its source — landing in one transaction with the joined lines as its content. Consumption of that wake by the hosted scheduler, and settling the input when no turn can take it, come next (open question 6 of H0c). * fix(managed-agent): keep a finished background Shell's receipt answerable H7 of the real-stack rounds: the Broker settles its :process row only when shell-status answers exited, but the worker's registry deleted an entry the moment the Shell ended — a natural exit could never be answered again, and the ledger row would hold the Runtime against release and close forever. The registry now retains each completed unit's receipt until the worker ends: the hold drops exactly as before, but shell-status and shell-terminate answer a proven end from the retained evidence, idempotently, for the owning Session scope only. An end without evidence keeps nothing and answers unknown as before; the existing live flows are untouched. Suite: natural exit keeps answering exited with its evidence for both kinds after the hold dropped, the wrong Session still hears unknown, and an evidence-less end stays unknown. The close-drain/reconcile callers that consume this, and the Java-side settle-before-busy on workspace release, land with the next batch. * fix(managed-agent): settle provable background exits before release's busy check The second half of H7. The worker now keeps a finished Shell's evidence answerable, so the Broker has someone to ask — but observeBackgroundProcess still had no production caller, so a naturally exited Shell left its :process row non-terminal and every release wedged on runtime_session_busy. releaseSession now sweeps the background process rows this Broker admitted for the Session before it computes busy: each unproven row asks its physical owner once through the same shell-status → settlePrepared leg observeBackgroundProcess already owns, a proven exit settles with its evidence, and any lookup that fails or cannot prove an end simply keeps the row so busy stays the accurate answer. The sweep tracks per-context admissions because the reconciliation scan deliberately excludes PREPARED rows, and mutates the admission set under the context lock. Witnesses: releaseSettlesAnExitedBackgroundProcessBeforeBusy proves an exited row settles with its evidence and the release succeeds (observed red with the sweep bypassed); the pre-existing unprovenProcessStatusKeepsItsHold still proves an answer without proof keeps the row and the busy refusal. Full module 604/604. The workspace close drain still terminates shells in order rather than sweeping — its ordered stop belongs to the close-drain slice, and these two verbs are exactly what it will consume. * fix(managed-agent): check the journaled receipt before attaching a background Shell's start H6 of the real-stack rounds: acceptBackgroundShell called childRuns.attach before its tool.receipt replay check. A retried accept — the leg recovering pending results re-drives — would mint a second start receipt on the same identity before reaching the replay path, and the record's set-once rule would refuse it as a rerun shape, so a recoverable retry crashed as a conflict. The lookup now runs first: a journaled receipt replays its validated history with no attach at all, and a fresh accept attaches only when this identity still lacks a start receipt — the record id IS the execution identity, so an existing one is always our own crash-window attach between record and journal, which the retry now tolerates instead of minting a rerun. Witnesses: a retried accept answers from the journal with exactly one attach and one tool.receipt line, and an already-attached-but-unjournaled record completes its accept without attaching again. The full hosted turn suite stays 148/148. * fix(managed-agent): backpressure the background Shell output pipes G5 of the real-stack rounds: the executor fed the bounded capture with a fire-and-forget sink.write — every chunk queued behind asynchronous publication, so a fast producer flooded memory (a 256 MiB run peaked near 315 MiB RSS at a 20 MiB/s store). The feed now counts bytes in flight and pauses stdout and stderr when the drain falls sixteen MiB behind, resuming at four MiB — enough to cover a store several seconds slower than the process produces without stalling a steady producer against the floor. Completion or rejection of a write both release their bytes, so a broken capture can never leave the pipes paused. The witness drives a controlled sink: seventeen MiB of chunks pauses a stream exactly once, draining past the watermark but above it resumes nothing, and draining the rest resumes exactly once. The outputRef-live-advance half of G5 belongs to the turn-publisher's promotion to activation scope, sized with the close-drain slice, and stays queued there. * feat(managed-agent): promote the Shell publisher to Session scope The ②d-out debt, and with it the live-advance half of G5. The Shell publisher was a per-turn object: its server died with the turn, while background Shell traffic flows across turns — between turns (or after a close) the worker's retained descriptor pointed at a dead endpoint, so no live output advance and no exit leg could ever complete past one turn. The hosted Session now owns one long-lived publisher across turns (the turn lazily structures it on the shell options and never closes it; the Session's ordered close in this PR's fifth slice drains it). The two small wire-level fixes that the promotion exposes: - The worker's background execute stamps the background marker into its prepare request: the closed six-field wire capture has no such field, so a remote prepare used to arrive unmarked and die at admission; the executor derives the marker from the background input itself. - register() takes the registering turn's prompt for the foreground admission check, so one instance admits the foreground of any turn of its Session while a foreground pretending another turn's identity is still refused; the background arm keeps its record-based admission. Witnesses: a single instance admits a background watch plus the foreground of two turns and refuses a misattributed one; the worker rig pins the stamped marker on its prepare; the four touched suites stay 184/184 and both package typechecks pass (after a core dist rebuild for the capture type's optional background marker). The publisher's own descriptor installs idempotently per turn (identical descriptors are accepted), so consecutive turns never fight for the worker slot. * feat(managed-agent): drain a Session's background Shells ahead of release The close-drain's missing legs. Until now a Session with a live background Shell could never close: the worker's activation gate refused release with managed_activation_conflict forever, and the CLI-side close route skipped the background stores. The ordered sequence the design asks for now exists end to end: - The registry gains stopSession: terminate one Session's entries with the supervisor's evidence rules (TERM, escalate, empty-check), await each completion so its finalization — last manifest, sealed envelope, record settle through the session publisher — lands before the caller proceeds; an unproven stop keeps its hold for the caller to see. - The worker's activation release gate drains before it refuses: active background work no longer wedges a close that can prove its stops, and anything still unproven keeps the exact same conflict. - The hosted close route now closes the Session-scoped publisher right after the broker release returns — which by then has drained the Shells and settled their records through that publisher — and the drain closes every capture store, background families included. Witnesses: a registry drain settles one Session's shells with evidence, both kinds, while leaving another's holds untouched, and an unprovable terminate under a session drain still preserves its hold; the 766-test context-worker suite stays green through the release gate's new async drain branch. What remains deliberately unproven locally: the full-chain release-with-live-shell drain needs a real cgroup supervisor, pinned as the Linux head acceptance instead of a mocked supervisor. * docs(managed-agent): sync the H3 design's status line with landed slices The status line still claimed publication and orchestration remained design while most of them landed across this PR's batches; refresh both language versions with the landed list and keep the remaining-design list accurate: the Monitor physical side (route, rebuild, cgroup watch) and wake consumer, both submission enablements, task events, and the Linux physical acceptance. * feat(managed-agent): spawn Monitor watches through a cgroup watcher The physical half of the Monitor runner's executor interface, for the worker: ManagedMonitorWatcher spawns the watch command under a unit named for the execution identity, through the same supervisor (its membership attestation unchanged), splits stdout into the observation lines the loop buffers — boundaries respected, the Legacy partial-line cap honored, the tail flushed as the final line at exit — and reports the physical end exactly once, natural exit versus mid-run failure, so the loop's settled and watch_failed shapes both hold. stderr rides only the future output leg, never observations; terminate defers to the supervisor's TERM-escalate-empty rules. The suite drives a supervisor double: unit/executable/cwd derivation, chunk-boundary splits with the final flush, the cap drop, spawn-error and membership-failure shapes, and terminate delegation. * feat(managed-agent): answer monitor status and stop from a worker registry The monitor half of the maintenance route, mirroring the Shell's: a worker-side registry of supervised watches that holds each Session's Runtime while a watch runs, answers an end only with the supervisor's own evidence, and retains a proven exit's receipt until the worker ends so maintainers can learn the truth idempotently. The route sibling at /internal/managed-runtime/v3/monitors serves monitor-status and monitor-stop with the same closed wire shape, target-identity rule, and unknown-but-never-unproven discipline as the shells' route; it joins the worker's route list so the v2 envelope admits it. Registration arrives with the Monitor start admission later — status/stop already have consumers in the release drain and the rebuild slice that follows. This also repairs the watcher test arity that broke the previous push's build lanes (the mock's 2-arg onExit omission, seen as error TS2554 (155,34) across every CI lane at that head). * feat(managed-agent): admit monitor watches into the open-ended capture session The γ3-α precondition for every later Monitor leg: LocalShellStreamResultSession's prove-a-live-start admission reads the capture family's record domain instead of being pinned to child_run — same Session/binding/background-marker checks, same live-run gate, same identity captureId, per-domain labels so a refusal names its own family. The producer side names its domain explicitly (record presence is never consulted where both domains could in principle compete), so a monitor watch carries its monitor_run record into exactly the same bounded stream the background Shell already rides. Witnesses: shell admission unchanged; monitor admission succeeds under its domain; wrong-domain proven start refuses; missing record and foreign binding refuse with their own errors. * feat(managed-agent): admit Monitor watches through the v3 executor The workhorse of the Monitor leg. The worker's v3 executor gains its is_monitor branch beside is_background: journaled admission first, then refuses (never as a transport error), then the watch spawns through the ManagedMonitorWatcher under a unit named for the execution identity, streams stdout into the SAME bounded family (16/4 MiB backpressure on the pipe, with an explicit refusal if the watcher reports no supervised unit), and settles with its detached handle while a registry keeps the Runtime hold. Three watch-level facts are done physically now: - The monitor registry is a wrapper around the background Shell registry, never a copy: supervisors, EOF drains and publisher finishes stay in one place, while monitor holds and the four-watch quota bucket remain entirely their own. Its maintenance route answers run/exited/ unknown with identical semantics to shells'. - The execute path gates on the generated four-live-watch cap with `Session already runs 4 Monitor watches.` and additionally on a missing cgroup root and a missing capture service. - Each prepare carries background+monitoring markers, so the hosted side of γ3-α knows the monitor_run domain to check. Suite: 28 tests across the three touched files — the fork settles a detached start under the right unit and marker, lines land as monitor output, the fifth watch is refused, and supervisor-less boot refuses with its dedicated text. The turn admission arm (funnel-driven admit→dispatch→attach through HostedMonitorSession), the watcher→loop observation feed, and rebuild still land in the following increments. * feat(managed-agent): admit Monitor watches through the hosted tool turn The turn-side arm of the Monitor leg. A call named monitor is now admitted through its own three-state gate (requested/unavailable/ill- formed, with the Shell-profile Monitor declaration now advertised), implements the same record-first discipline as background Shells do — the funnel admit commits monitor_run revision 1 before any side effect, dispatchStarted follows the durable checkpoint, and the watch attaches when its v3 execution settles. The granted publication, journal intent, dispatch and accept forks are widened to carry is_monitor alongside is_background, and acceptMonitor mirrors acceptBackgroundShell end to end (not_started → start_failed with the unstarted family; detached → same blocked receipt family with exactly one tool.receipt line, the H6 journal-verdict-before-attach order, and the blocked ack completing the row). HostedSession gains a session-scoped HostedMonitorSession beside childRuns; turn construction forwards it at both sites. Two existing contract pins move deliberately: the Shell profile's advertised list now names monitor (the harness declaration test names it explicitly), and the three publisher-close witnesses now assert the Session-scoped lifecycle settled by the earlier promotion — no close at any turn ending (completed/error), the server still answering, close exactly once at the Session's own delete. Suite: 3 new admission-arm tests through a full hosted turn (funnel order, clamped defaults with the Legacy monitor caps, start_failed shape, unavailable text), the whole turn suite stays 151/151 and the harness session suite 180/180. The observation fan-out (worker lines → hosted loop), quota witness at the turn, and rebuild-after-runtime_lost land in the next increments. * feat(managed-agent): fan a monitor watch's observations into the hosted loop The observation channel of the Monitor leg. The session publisher now marks monitor captures with recordDomain monitor_run through γ3-α, fans each published stdout chunk through the Legacy line-split semantics — boundaries kept, remainder held across chunks with the 4096-byte partial cap — into a per-capture observer, and closes it with onExit(failed) at finalization with the tail flushed first. The turn gains resumeMonitorWatch: a fresh accept creates one HostedMonitorLoop per Session report and resumes it from the already-attached record, fed through a HostedMonitorRemoteExecutor whose start binds those observer callbacks; a replay starts nothing twice, exactly like it never re-attaches. Globally, loop timers unref so observation windows never anchor the process. Monitor capture finalize settles through the monitor funnel itself (watch_failed when no evidence, exited via settleQuiet), beside the shell's unchanged path. The publisher carries monitors next to childRuns, and the turn's own construction forwards it. Suite: the admission-arm tests now also pin one Session-level loop starting on a fresh accept with the record-first flow intact; wide suites of 189 tests pass with cli tsc, prettier and eslint clean. Monitor output before the observer registers stays durable in the open stream, with observation starting at attach — noted in the design not as a flaw but as the attach seam, a line crossing it already lands in the output Artifact readers see. * feat(managed-agent): rebuild a read-only Monitor after its Runtime is lost The record side of the rebuild promise. HostMonitorSession gains the two verbs the H3 lifecycle needs around runtime loss: blockedRuntimeLost lands the run in recovery_blocked with the runtime_lost reason and the lost binding still named, so degrading is what the projection shows until something decides otherwise; rebuildFromRuntimeLost accepts only the outcome_unknown/runtime_lost line the H0b rule allows, starts a fresh watch under a newer Runtime generation with its own receipt, and keeps the reason for honest projection while observation positions never moves — the support case never restarts what the watermark holds. monitorRebuildAllowed routes every candidate command through the Legacy AST read-only check with its walk, so a side-effecting command stays blocked accurately. Five witnesses cover the full rebuild chain (recovery-block, rebuild with receipt remint and generation increase, a rebuild refused from a live run and from an ended one, and the gate's read-only/side-effect/empty answers). Binding-replacement detection and re-drive of the new physical watch come with the recovery slice after the wake consumer. * feat(managed-agent): serve the task events routes with their SQL journal H3 flips listSessionTaskEvents and queryWebShellTaskEvents from planned to partial on both surfaces. The Java control plane gains a bounded per-task event journal (V39): one row per committed task event, written in the Session commit transaction, with the per-task sequence allocated under the journal head lock in commit order. Record revisions journal their task view changes at the H0c announcement point. The durable retention floor lives in a per-task cursor row, so it survives an empty retained set, restarts and projection rebuilds; the Artifact visibility barrier pins it behind unarchived output, and the backlog bound refuses past capacity instead of silently discarding. Output events carry their per-stream segment ordinals with overlap/gap refusals, and artifact references stop fail-loud at the 100 bound. output_cursor/outputCursor lift with the routes and the task views serve them. The section 6.1 demonstrations run as store-level contract traffic in the new journal contract test, the route probes and record gates cover the flips, and the #12847 C15/C16 instances close alongside. WebShell types regenerated from the bumped 1.30.0 contract. * docs(managed-agent): add the task events lane's build report Flyway numbering findings (V39 taken; V35-V38 claimed by in-flight PRs on main, V15/V29 Java-only burns untouchable), implementation decisions against the H3 design text, the CLI/daemon flow-surface finding (none exists, regeneration is the only surface change), the full test ledger and the owed items. * feat(managed-agent): build the monitor wake envelope and the pending-input intake Two support pieces for the H3 wake consumer (β2-b1a): - monitorNotificationText wraps one due observation window in the exact Legacy task-notification envelope — task-id, optional tool-use-id, kind/status/event-count, summary and result — applying the same truncateNotificationLabel/stripDisplayControlChars/escapeXml pipeline the Legacy Monitor applies, so a turn learns nothing new. - pendingSessionInputs derives the Session's still-pending inputs from the journal alone: an accepted input is consumed exactly when a turn settles under its turnId, order-insensitively, so a replayed admission after a restart is not re-driven. * feat(managed-agent): deliver the Legacy notification envelope on monitor wakes The loop's notification input now publishes the exact task-notification envelope a Legacy Monitor's wake carried — task-id, tool-use-id, kind, status, event-count, summary and the window's lines — so the turn the wake raises reads nothing it has not already read. The description is the tool call's, falling back to the watched command. A witness pins the exact wire text of the first observation's input content. * feat(managed-agent): run the monitor wake through an embedded scheduler H3 closes the wake loop: an accepted observation already commits its notification input and wake.requested in one transaction; this slice makes that wake effective. The pump re-derives the pending monitor inputs from the journal alone — nothing is held in memory, so a restart re-derives the set. On an idle Session the oldest notification runs as an ordinary text turn with the Legacy envelope as its prompt, and the turn's settle consumes the input. A busy Session queues in the journal exactly like channel and Goal inputs; the retry loop redelivers. A Session that is parked, blocked or mid-recovery keeps its reminders pending and reports them. On the close path every pending notification settles cancelled without a model turn, under the turn-result record's own idempotency key, so no wedged notification ever holds a Session as hosted_turn_recovery_required at its next open: the parked-turn scans now skip monitor inputs, which the pump owns exclusively. A wake turn that dies without settling re-drives on reopen — the input is consumed exactly when a turn settles under its turnId, and a crashed turn never settles — matching the queue semantics channel and Goal inputs already honor. * fix(managed-agent): keep every new Session on the managed-session/1 reader Round-5 real-stack verification's cross-version matrix disproved the v2 stamp's premise: monitor_run has been a parsed, known domain since #12837 (v0.24.7), and a v1 reader opens a Session holding monitor_run records without complaint. Stamping minimumReader: managed-session/2 on every new Session — while monitor_run remains disabled — would make a rollback or a mixed-version rollout lose access to every Session created in between, and buys protection no deployed binary needs. Revert the stamp to managed-session/1 and drop the per-domain header refusal; the enablement list remains the single domain gate. The stamp rises only when a change genuinely breaks an older reader mid-scan, and only one release after a tolerant reader ships. The golden fixture's genesis header regenerates to v1; the records and authority suites gain v1-header admittance witnesses in place of the refusal cases; the design's reader section now records the reversal with its evidence. * fix(managed-agent): keep monitor output byte-true and advance only forward Round-5 real-stack verification found three facts about the monitor output path: - A capture rebuilt from decoded lines cannot be byte-true: per-chunk toString split multi-byte runes into U+FFFD, and line-rounding dropped blank lines and the final unterminated line. The watcher now forwards raw stdout chunks to the capture on a new onChunk channel, so the durable Artifact reproduces the command's stdout exactly, while the observation lines keep the Legacy emit-path semantics (a StringDecoder holds a split rune until it completes; a blank line consumes no observation). - The watcher registered its exit listener only after the supervisor's start resolved, so a fast watch could end unheard and the debounced loop never settled. A latched end with a check-phase sweep covers a watch that exited before its listeners attached, after the host is known to be listening. - The funnels' advanceOutput promised "only ever to a newer revision" in prose and accepted an older manifest in code. Both hosted funnels now read the revision of both manifests and refuse a regression before anything commits; witnesses pin the refusal in place. * feat(managed-agent): drain a busy background Shell at release, and answer the asked operation Round-5 close-path findings (H8): with a running background Shell the Broker refused busy before transport.release — yet that HTTP call is the only route to the worker's drain — so a `tail -f` held session release and workspace close forever, with nothing able to stop it. The release path now distinguishes busy kinds: anything active that is not one of the Session's own background process rows keeps the exact busy refusal with zero worker hops, while proven-running rows trigger a dedicated release call first whose 409 busy answer is swallowed (the worker's route drains before refusing, and its maintenance routes are not activation-gated, so the follow-up sweep still proves what the drain ended). Rows settle from the drain's evidence and the complete release converges on what is genuinely still running — one call when the drain finishes in time, one retry when it does not. Two witnesses pin the contract: a running Shell reaches the worker (the round-5 probe found the refusal never did) and a non-background busy refuses without a worker hop. Also from the round: the shell maintenance validator computed the expected operationId from targetOperationId but never compared it with the answer's own operationId, and the sweep witness answered with the process-row id where the real route echoes the call id. The validator now refuses a view answering another operation, a new protocol suite pins it, and the witnesses answer the shape the route really sends. * fix(managed-agent): re-assert the output pause behind every flushed chunk Round-5 G5 finding: Node's flushStdio resumes both paused pipes when the background launcher exits, and the executor tracked the pause in its own flag rather than the stream's, so a writer that outlives its launcher pushed 230 MiB into the queue and climbed RSS by +200 MiB. The onOutput accounting now re-asserts the pause behind every delivered chunk, and applyPause is idempotent through the stream's isPaused state instead of the local flag — the stop-and-go the backpressure intended holds through the launcher exit, on both the Shell and the Monitor paths. The existing 17-MiB pipe pacing witness keeps its single pause call. * test(integration): expect the monitor tool on the hosted shell profile The hosted shell profile now advertises `monitor` beside `run_shell_command`, so the fake-model driver's tool-name assertion — which the Hosted workspace tool-turn IT checks on every model call — must expect it. The file profile's expectation is unchanged: the declaration is shell-profile only. Fixes the single red IT in the MySQL fault-gates lane (HostedWorkspaceToolTurnIT.packagedHarnessUsesSavedWorkspacesThroughReal BrokerWorkerAndSqlStore), whose only difference was the extra 'monitor' entry. * fix(managed-agent): read the whole journal when deriving pending monitor wakes `readEvents()` is a bounded read: with no arguments it returns the first `defaultReadEvents` (100) events, capped at 256. Both consumers of the pending-input derivation called it bare, so on any Session whose log runs past that page — every real Session with tool calls — the derivation saw only the log's head: - the wake scheduler's `next()` never found a notification input, so a monitor wake was never delivered as a turn once the Session passed the page. The feature was dead in production while every rig stayed green, because the rigs commit a handful of events; - `settlePendingMonitorInputs` never found the owed inputs at close, so they stayed unsettled and the Session reopened as `hosted_turn_recovery_required` — exactly the wedge this close-path settle exists to prevent. Both now read the full committed prefix with `eventsInSequenceRange(1, committedSequence)`, the discipline every other journal scan in the hosted session already uses. The witness commits 110 observations so the notification lands beyond the default page, asserts the log really exceeds it, and was verified red under the inverse edit (the settle returns 0 and the input stays owed). * fix(managed-agent): stop a full task journal from wedging record commits The task event journal is a derived, bounded feed: its backlog bound refuses the next event past capacity rather than silently discarding it, and an unarchived output event pins the retention floor so the automatic pass cannot expire anything. The record commit called `appendStateChange` inline, so once a task's journal was pinned and full, that refusal propagated out of the extension-record commit — every later revision of that task failed, which would take the Shell and Monitor record lines with it. The record row is the authoritative state and is already written when the journal append runs, so a refusing journal now degrades the feed (warned with tenant, session, task, revision and the refusal code) and never the commit. The store still throws for callers that can apply backpressure — which is where the output producer's retry belongs. The witness drives the real wedge: it pins the floor with an unarchived output event, fills the journal to the bound, asserts the next append refuses with `managed_task_event_backlog_full`, then commits a record revision whose view changes and asserts it lands. Verified red under the inverse edit, where the commit dies with "The task's event journal holds 512 retained events". * refactor(managed-agent): drop the shell view's never-set error field The maintenance view declared an optional `error` object, the Java response validator carried a branch for it, and no producer on either side ever set one: the worker answers `running`, `exited` or `unknown`, and a route-level failure travels as an HTTP status with its own code body. A field that is declared and validated but never written is a dead switch, so it comes out of the TypeScript view and the Java field set — which now also refuses an answer that carries it, instead of silently accepting a shape nothing produces. * fix(managed-agent): let a replayed output advance pass instead of refusing The forward-only check refused any advance whose revision was not strictly newer, which included a redelivery of the very reference the record already holds. The publisher guards against that in memory, but the guard does not survive a worker restart, so a replayed advance could throw inside the background exit leg and wedge a Shell or Monitor that had already ended. Both funnels now treat an identical reference as the no-op the deep-equal skip already owns, and still refuse any other reference that is not strictly newer. For a Shell the replay commits one liveness step, since the live run line alternates by design; for a Monitor it commits nothing. Witnesses pin both, beside the existing refusal witness. * chore(managed-agent): keep the task-events lane report out of the PR diff A one-off construction report does not belong at the repository root: the project's directory rules put working artifacts under the ignored .qwen tree and keep the tracked docs/ for design and plans. The report is preserved at .qwen/investigations/2026-10-04-task-events-lane-report.md, and its Flyway numbering note is superseded by the V40 renumber the merge commit records. * fix(managed-agent): quiet the wake pump's consume guard at a blocked owner runTurn's recovery-blocked path marks the Session blocked and returns 'settled' without the turn ever existing, so the input stays owed — the accurate answer the whole design gives for a parked Session. The pump's after-settle check read that as a consumption bug and threw, which both marked the already-blocked Session again and logged a spurious error on every such observation. It now stops at an accurate blocked state and still throws when the input simply did not move: that is the programming error the guard exists for. * fix(managed-agent): meld a monitor watch through one terminal step Round-6 found the shape the Monitor's final leg really had: prepare's open-then-advance could never survive the start order (output before a start receipt is a parse-level refusal), the manifest still advanced through the Shell funnel whatever the capture's recordDomain said (a Monitor never builds its own Shell record), and the finalize settled the record ahead of the loop's own exit chain, so a watch's last window died in denial between two async hops — with its swallow queued to silence it. Three changes, each with its own witness that goes red under the inverse edit: - The manifest advance reads the record's start receipt first and only advances past it; a missing record still fails loudly instead of the old "no record to revise" at whatever funnel happened to be there. - It routes by recordDomain: monitor captures land on the Monitor funnel, Shell captures on the Shell's. - The loop's onExit now returns a chain carrying the exit's own flush → settle in order, and the publisher awaits it rather than settling itself — a final window can no longer be dropped ahead of the settle, and a commit failure of the chain surfaces inside the finalize instead of returning an empty observation state silently. The end-to-end witness drives the production shape on a real authority, orchestrator and Publisher over HTTP: prepare ahead of attach, output advancing only after the receipt, then one terminal step holding observationSequence 1 with the settled mark behind it. * fix(managed-agent): settle the process row when the Runtime proves no start Round-6 found the permanent wedge: the `:process` row admits PREPARED ahead of dispatch, but an admitted-refusal answer (`{executionStatus: "not_started", capture: null}` — cgroup root missing, quota exceeded, and friends) settled only its invocation. The process was never registered by the Runtime, so every later shell-status can only be unknown, the exit-only settle path never applies, and every release carries the hold forever. A proven never-started is proof: when the invocation settles with executionStatus not_started, the sibling row now settles with the same not_started mark, its admitted refusal as the evidence. Busy computed after that never counts a process that could never have existed; an unproven-lost connection still wedges exactly as the declaration says. The witness drives the admitted refusal shape end-to-end: the process row lands SETTLED with not_started, and release passes without the busy answer. Verified red under the inverse edit, which reproduces the report's `expected: <SETTLED> but was: <PREPARED>` verbatim. * fix(managed-agent): attribute wake receipts to their own turn and stop re-driving it Two wake-path defects from the round-6 audit family, plus the monitor registry the maintenance route actually shares: - `recoverShellReceipts` exempted monitor inputs so thoroughly that a wake turn's receipts could never be attributed: any `tool.receipt` journaled under a wake turn would attach to the previous prompt or be skipped. Receipt attribution now advances on every accepted input, while the pending set itself still exempts monitor inputs as before — nothing reframes them as parked Turns. - A wake turn that started, journaled records and died was re-executed text-only on reopen, minting a second user record and leaving an unanswered first call in the model's history. `wakeHasPriorAttempt` names the condition simply; the pump now parks such turns accurately blocked for the recovery fleet instead of redriving. - The worker constructed its executor with no monitor registry, so the executor built a private one while the monitor maintenance routes got another empty one — every status and stop answered unknown. One registry is shared by executor and routes now, mirroring what the Shell path already does. * fix(managed-agent): read the task event floor after the events, not before Two stores-and-reads of an event page ran as separate autocommit queries: the retention floor was checked before `events.read`, so an expiry mid- read produced a 200 that silently skipped events under an unchanged nextCursor, the exact gap the contract's `409 cursor_expired` exists to answer. The replayable-events family already documents the sound order: events first, floor after — a monotonic floor that is not above the cursor after the read was not above it during the read either. The event route now decodes the cursor, reads the page, and only then answers `cursor_expired` when the post-read floor has crossed it; past-the-floor omitted-after pages still start at whatever floor stands then, and an empty page's cursor derives from that post-read floor. Two witnesses cover both routes: real-journal paging and an explicit cursor that races the floor's progression with the floor rising at the read itself — verified red under the inverse order ("Expecting code to raise a throwable"), exactly the silent-gap shape the race once returned. * fix(managed-agent): move the task artifact references by compare-and-set The ref list was read, then rebuilt in Java and upserted — two writers ona task both satisfying the same bound could each write a single addition of the other's list and silently age one reference out, the exact eviction the hundred-ref bound exists to refuse. The append now moves only by compare-and-set against the list it just read: the first writer's update lands, the second sees zero rows changed and retries against what actually stands, and the missing-row path retries a raced first insert against the comparable shape instead of overwriting it. Exhausted retries refuse, they do not lose. Witnesses keep every reference in order under back-to-back appends and prove a stale snapshot can no longer rewrite the list (the exact clobber the old write shape carried). Duplicates and the hundred-ref bound still refuse with the same error, never a silent eviction. * test(managed-agent): straighten three small voice and fixture drifts - The record-level integer refusal message now pairs with the Java store's identical boundary, so a cross-language failure reads alike wherever it was caught ("must be an integer from X to Y"). - The design docs (both languages) name `V40` and the real trigger of a `state_changed` event: a task view change, never every record revision. - The monitor tool declaration says what `max_events` actually counts — one-second-or-longer debounce windows, not lines — so a quota set from line expectations lands right; `idle_timeout_ms` names the quietness it stops on. - The two task-list contract checks make themselves valid before the unknown field lands, so the refusal measures the schema rule and never borrows the cursor rule's answer. * fix(managed-agent): admit named cgroup units and keep fast-exit evidence - The unit-name guard compared against the empty string, which every name contains: every named `create` refused and every `attach` answered undefined, so no supervised command could ever start under its stable identity. The intended NUL check is now an escape, and the test stops embedding a NUL byte itself so the file diffs as text. - The launcher announces `joined` on fd 3 once the kernel accepts its membership write, and the prober takes that marker beside the `cgroup.procs` listing: a command that joined, ran and exited inside the first listing window was previously journaled as a missing cgroup delegation — a refusal with invented evidence. - Exit evidence survives the membership window: the getters consult the child's own `exitCode`/`signalCode` when the exit ran before the listener, and `terminate` answers through the same read, so a fast process can no longer leak its Runtime hold into a never-resolving `once('exit')`. A spawn-time `error` handler covers the window before the process object exists. Each witness ran red under its inverse: the named-unit refusal reaches `resolveRoot` only after the fix, the marker acceptance with an unreadable process list refused after two seconds without it, and the in-window exit lost its evidence without the fallback. * fix(managed-agent): account capture pages from the batch they publish `flushPage` counted and cleared the live pending array after its publish await, so a segment landing mid-window from the other stream's queue was priced into a page that never carried it and then discarded with the array — a silently self-inconsistent manifest whose `segmentCount` exceeds the segments the page holds. The flush now splices the batch up front and derives every field from that batch. The both-stream flush ahead of every manifest revision stays: the manifest validator requires each descriptor's pages to add up to its stored byte length. A segment arriving between the splice and the revision now wedges the capture loudly (`storage_failed`) instead of producing a manifest that claims bytes its pages do not hold; the witness pins the fail-stop, and the old loss shape ran red under the inverse (`brokenReason` stayed null while the fold silently completed). * fix(managed-agent): answer the detached capture family in the validator The broker's hand-rolled v3 capture validator predated the `detached` capture status this branch adds to the shared envelope: three clauses rejected the background start's settle, and this branch's own consumer (`isDetachedCapture`) was unreachable over the real transport even though the shared schema, the projector and the TS rules all moved. The validator now mirrors the TS rules — reasonless like `complete`, manifest exempt and required to be null — plus the two refusal directions (a manifest or a reason present while detached). Witnesses: a detached settle round-trips with its blocked delivery, and both refusals throw. The accept case ran red under the inverse ("Managed Runtime status capture is invalid."). * fix(managed-agent): harden the admission, capture and hold edges of the worker Four slice-C findings, one family each: - An admission that fails after its dispatch entry is journaled (prepare rejection, or the close recheck landing mid-flight) parked the entry `prepared` forever — holding the Session's Runtime, wedging close-admission, answering re-issues with a phantom wait. Both background paths now settle the refusal `not_started`, like every other pre-effect refusal already does. - The Monitor hold family was half-wired: holds consulted Shells only, …
…LM#13300 (QwenLM#13355) * docs(managed-agent): design H3 background Shell and Monitor runtime * feat(core): add managed child_run (kind shell) record body for H3 * feat(managed-agent): mirror child_run shell record body in Java store * test(core): commit and rebuild child_run shell chains through the session authority * feat(managed-agent): close child_run reference closure and add stop-request draining * feat(core): add named cgroup unit attach and managed child-run supervisor * feat(cli): admit v3 background shell under the supervised worker registry * feat(core): add open-ended shell stream capture with manifest revisions * fix(core): type-narrow stream capture finalize and refused publish double * fix(cli): close background shell type gaps after the main merge * fix(managed-agent): answer the R1 review round on the child_run contract * fix(managed-agent): close H0c critical follow-ups R3-1/R3-2/R3-3 PR 1 of #13300, fixing the three Critical findings from the round-3 review of #12855. R3-1: the Broker-record execution mapping now reads the record's dispatchGeneration: SETTLED/cancelled with generation 0 (the record's own proof it was never claimed, per ToolExecutionRecord) maps to not_started_proven instead of settled, removing the illegal intent -> settled shape that left a Monitor cancelled before dispatch with no committable settling revision. A shared brokerExecutionCases row (settled-cancelled-unclaimed) replays it in both languages, the TS divergence pin gains the second wire-reading divergence, and Decision 10 of the authority design is reconciled with the H0b record-contract doc in both language pairs. R3-2: the store now runs the event-level half of the envelope checks for every event line at commit time: closed envelope with an optional subject, version, declared sequence, well-formed event id outside the reserved <domain>:<n> namespace, this Session's closed key, valid time, and kind from a mirrored EVENT_KINDS vocabulary; every domain.committed payload is validated whether or not a body is registered (reusing the pinned DOMAINS); unknown-subtype lines follow the scanner's "after the Managed header" condition; a transaction carries at most one Stage H record; and per-kind byte caps (maxEventBytes, maxCommitMarkerBytes) are pinned in the shared limits contract and enforced per line kind. All refusals answer the existing 409 so the commit rolls back. Replay commits return before any of this runs, as before. R3-3: task.updated rides its own outbox (managed_agent_task_event, V35) written in the same commit transaction and keyed by a unique source key, drained later by the task-events slice; managed_agent_event stays turn and lifecycle events, so an announcement between two text deltas can no longer split a message Part under the frozen projection version. The route description is scoped to what the server does (a Session being deleted announces nothing while its tasks stay readable), and the design's replay-idempotency sentence is corrected. Test helpers that wrote journals a real authority could not read back (ActionJournal's reserved event ids, bare event/marker lines in the publication and session-store integration suites) now write well-formed lines. Mutations of every new check were verified red against the suites. * fix(managed-agent): repair the CI-only cgroup root and env-guard failures How: wrap the delegated-root probes so a missing or unreadable root answers the documented isolation error on Linux too (macOS refused at the platform check first, which is why only CI saw the raw ENOENT), and document the background Shell environment allowlist in the process.env guard. Why: the Test lane on 4fb9f0a9c8 went red on exactly these two items; both are this branch's own changes, not flake. Test: hook-command-cgroup and process-env-guard suites green; tsc clean. * feat(managed-agent): inject the child-run supervisor at worker boot How: registerManagedContextRoutes builds a ManagedChildRunSupervisor from the delegated cgroup root the Hook commands already use (QWEN_MANAGED_HOOK_CGROUP_ROOT) and passes it as the executor's fifth argument; a boot without the delegation keeps the executor's committed refusal instead of failing to boot. Why: executeV3Background landed in the previous increment with the supervisor parameter unwired, so every real-stack background start answered the committed no-supervisor refusal. Test: new wiring case asserts the supervisor is injected exactly when the root is set; managed-context-worker, process-env-guard and managed-background-shell suites green (775); tsc clean. * docs(managed-agent): pin the two-row ledger shape for background processes How: the runtime-ownership section now says the ledger holds two rows — the foreground-length start invocation settling with the handle, and a second execution admitted at start that carries no model result and stays non-terminal until physical exit evidence. Why: the Java ledger writes a result exactly when an execution settles, so a single row that is both handle-delivered and active cannot exist; the two-row shape keeps the settled-if-and-only-if-result invariant while preserving every cited behavior (Runtime holds, evidence-only settle). * feat(managed-agent): commit child_run shell lines from the hosted side How: HostedChildRunSession funnels every record write through one serialized, replay-safe commitExtensionRecord path — admit first with the start call's args as commandRef, dispatch and the set-once managed-runtime-receipt after the physical start, outputRef advance-only, and settlement only on proven exit evidence, a proven failure, or an honored stop request. Live revisions alternate the run line (no self-loops); the authority freeze refuses anything after a terminal revision. Why: the dual path puts product records on the hosted authority, but nothing commits child_run from the hosted side yet — the worker-owned registry intentionally does not touch records. Test: four authority-level cases — full chain with projection, stop request to cancelled, pre-start failure frozen on not_started_proven, and serialization with deep-equal skips. * feat(managed-agent): add the detached capture family to Tool v3 results How: a background Shell start now settles success with a capture object whose status is 'detached' — no reason, no manifest — because the live output streams through the child_run record's growing manifest rather than the result; the shared schema pins both invariants, the TS parser mirrors them, the shared fixtures gain one valid and two invalid envelope cases replayed on both sides, and the Java projector accepts the missing manifest for unavailable or detached captures. Admission refusals settle not_started with a null capture, the shape the durable unstarted family already owns. Why: every Session tool.receipt event lands in the Java delivery projection, which requires a capture object and a committed-if-manifest pairing for anything that started — a success result with a null capture fails that projection and corrupts replay. Test: core contract suite 519 (three new envelope cases), TS serve suites 33 + turn/harness 323 + background 6, Java publication contract and projector suites 8; tsc clean on both packages. * fix(managed-agent): answer the R2 review round on worker, capture and record lines How — the four Criticals first: the stream capture publishes an ended stream's revision only after its seal decision, so a sealed-or-incomplete descriptor never changes again (R2-3); finalize's settle path degrades to an unavailable capture instead of throwing past a writability failure, and the last published revision stands like the worker-cut cap (R2-5); the child_run body now refuses a settled or cancelled run that is not settled execution, a failed run off its two ending lines, start_failed outside not_started_proven, and a process-level failure without a settled execution, on both languages with new witnesses replayed from the shared fixtures (R2-6, R1-30); and the supervision suite skips win32 instead of spawning a shell that cannot exist (R2-7). How — the Suggestions land as: supervisor start now proves membership by cgroup.procs with an fd-3 status channel, failing closed as isolation (R2-37); the bounded EOF wait caps daemon-inherited pipes instead of hanging, and a spawn-time error drains the entry (R2-1/R2-14); registry hasHolds requires its session and setProcessResult carries the full result shape (R2-13/R2-15); create() validates its caller-named unit like attach() (R2-26); the executor checks the journal before any effect, mirrors the closing/is_active re-check after prepare, gates the ninth live Shell as a committed quota refusal, and keeps cgroup membership out of an empty QWEN_MANAGED_HOOK_CGROUP_ROOT (R2-23/R2-19/R2-24/R2-18); the spawn uses the configured shell and the session-context environment (R2-21/R2-20); invalid fixture cases each pin their refusing clause on both languages (R2-36); projection fixtures pin the draining precedence rows and the Java replay reads stopRequested (R2-28/R2-38); the draining Javadoc states the rule the code implements (R2-39); the duplicated ref helper is gone (R2-40); the stream-capture suite is typed and gains the seal-before-publish, page-cursor, late-failure and settle-degrade witnesses (R2-29/30/31/32); and the design doc names the cgroup switch decision, the cursor's single ordinal space, the unconditional host-scope evidence gate and the zh stopped-object (R2-8/9/10/11/12). Test: core suites 788, cli serve suites 776 + background 11, Java record/ projection/store suites 24; tsc clean. * feat(managed-agent): thread the child-run orchestrator and the detached receipt family How: HostedSession gains its session-scoped childRuns orchestrator next to hooks/mcp, the tool turn receives it stored-ahead of the admission branch, the reopen verifier and the recovery replay both accept the third durable receipt family — a blocked delivery with a detached capture, beside the complete and not-started ones — with the publication store correctly left out of its proof, since its durable truth is the child_run record. Why: the H3 background start settles with a handle whose output lives on the record's growing manifest; without the third family, any Session holding a background receipt could never reopen or be replayed. Test: recovery-session suite 52 with two new detached-family cases; harness-session and tool-turn suites green; tsc and lint clean. * test(cli): type the background-shell capture double How: drop the as-unknown cast on the FakeSink return and implement the failCapture member the cast had been hiding (R2-29's cli location). Why: an unchecked double can drift from the interface it claims to match — and it already had. Test: background-shell suite 10/10; tsc clean. * feat(managed-agent): admit background Shell starts from the hosted tool turn How: the turn replaces its two deliberate is_background refusals — exactly when the Session owns its child_run orchestrator and the domain is enabled — with orchestration that mirrors the shared line: the record intent precedes every physical effect, dispatch_started lands the moment the checkpoint commits, a settled detached handle attaches the physical start to the record from the same facts and lands as the third durable receipt family (blocked delivery with a null manifest, replay-validated exactly like the not-started family), and a proven-unstarted refusal settles start_failed on not_started_proven while riding the existing unstarted family unchanged. Malformed is_background values keep their old refusal, and the family stays out while the domain is disabled — the same probe governs both paths. Why: the background start is the first record-bearing, user-visible H3 effect, and without a hosted history family for the handle a Session that ran one could never reopen or be replayed. Test: three new authority-level turn cases — admitted detached flow with admit/dispatch/attach order and the blocked receipt family, a proven-unstarted refuse closing NOT_STARTED, and the disabled-domain refusal preserving its exact text; turn suite 146, plus harness-session, recovery, background and child-run suites 246; tsc clean. * feat(managed-agent): admit background Tool v3 dispatches and acknowledge their detached settle How: the v3 start admission flips from "foreground Shell only" to "Shell only" — background dispatches now begin the same way; the poll loop accepts the settled detached family through the status answer (the only durable settle such a result can ever produce, since nothing is published for a background handle), and the acknowledgement path canonizes the settled handle envelope: a matching blocked receipt with a null manifest and no history revision forwards exactly itself, without consulting the publication receipt. Why: without the detached branch, a background start settled a queued handle only to mark the execution UNKNOWN after the 30-minute foreground-shaped deadline — every background dispatch would wedge the Runtime on its first call, and its acknowledgement would die on the missing publication it was never going to have. Test: new witness drives the detached start end-to-end — reservation, v3 dispatch, settled handle via status, and ack with a null-manifest blocked receipt, with the publication receipt path provably untouched; service suite 139/139, module suite 600/600. * feat(managed-agent): answer shell-status and shell-terminate from the worker registry How: a new private maintenance route sibling to ManagedHookProtocol (/internal/managed-runtime/v3/shells) backed by the process registry — status answers running or unknown (never claimed without evidence), terminate drains with the supervisor's own rules and answers exited with its proof, and an end that cannot be proven is unknown with the hold still registered. The unit name derives from the runtime invocation identity through the shared shellUnitNameOf, so Java, the hosted turn and the worker derive it identically; the executor takes the registry as a constructor injection so the route and the journal share one. Why: every later maintenance verb — recovery reconciliation, the ordered close drain, the reconcile-settle — has to ask the only physical owner these questions; answering them from anything but the registry would invent truth. Test: five worker-side cases — closed-key refusals, unknown for unowned, running scoped to the registered Session, terminate-with-evidence answering exited, and an unproven terminate answering unknown with the hold intact; shell, context-worker, v3-route and child-run suites 784; tsc clean. * feat(managed-agent): admit the background process row and answer its maintenance protocol How: the start dispatching of a background Shell admits a second ledger execution beside the invocation row — the new process row carries no model result and stays PREPARED (non-terminal) until physical proof, so hasActiveBy* holds the Runtime exactly while the process lives. The maintenance protocol mirrors ManagedHookProtocol on its own route: shell-status and shell-terminate with a targetOperationId (both recovery kinds), wire-validated on both transports and accepted through the same workspace-generation and ownership gates. observeBackgroundProcess asks the physical owner and settles the process row with proven evidence — settle through the new repository verb settlePrepared, which mirrors the PREPARED-requestCancel shape on both repositories; anything unproven leaves the row active with its hold, never a claimed end. Why: the ledger writes a result exactly when an execution settles, so a single background execution carrying its handle at start cannot exist; two rows keep the settled-if-and-only-if-result invariant while the Runtime holds precisely, and the maintenance route gives every later verb the worker's physical answer instead of an invented one. Test: two new witnesses — a background start that admits the process row (the invocation SETTLED, the process PREPARED, the Runtime busy until status arrives, then exit evidence settling it and release succeeding), and an unproven status keeping its hold; the runtime also gained the row identity invariant fix and the settle-through-settlePrepared shape; service suite 141/141, module 603/603. * fix(managed-agent): answer maintenance views under the target identity How: the shell-status/shell-terminate view now carries the target operation's identity, exactly as the Hook lookup family does — the requester asks about the process, never about the asking call, and the Java wire validator's expectedId rule requires it. Why: this repair was encoded in the route design but landed uncommitted after ②e-a; the self-audit caught it before it could ride a later batch. * feat(managed-agent): run the background exit leg through the record How: the hosted publisher gains the background family — its registration marker admits the open-ended capture through the session-level admission (proven start on the record, never the model call), the open manifest publishes at prepare, every awaited write or finish forwards the newest manifest revision to the record's outputRef, and the exit finalizes with the same evidence object settling the record. The client acknowledged receipt keeps the detached-blocked shape; drain skips background stores until the ordered close drain wires them; the register guard only binds the foreground model-call family. Why: the start handle already committed, so a background Shell's whole physical truth is output and exit evidence — anything that reached the record another way would let a second tool.receipt contradict its own line, and an open-ended manifest published anywhere but the Session's own store is an unreadable closure for the store-side commit. Test: a full end-to-end authority case — admit-dispatch-attach, bytes through the real rpc, outputRef advancing live, the final sealed manifest reading the same exit (code 3, both streams sealed, manifest complete), the record settled with exactly that evidence, and zero tool.receipt events for the outcome; publisher, turn and child-run suites 174; tsc clean. * feat(managed-agent): close monitor_run revisions over their cited Session resources The H0c closure rule — every ref a record cites must exist in the Session's resource store at commit time — now covers monitor_run on both stores: the authority reads commandRef, startReceiptRef, outputRef and lastObservationRef through verifyExtensionResources, and the Java applyRevision gains the same four-field branch beside child_run. This is the precondition for enabling the monitor domains: without it a monitor revision could anchor its chain to resources nobody committed, and the replay would project a watch whose evidence was never verified. Both fixture rigs stop citing fictional resource identities. The authority test publishes its chain refs as real bodies once per harness and mirrors each fixture-cited ref as a real same-kind placeholder of the stated length, memoized so chain identity stays fixed across revisions. The journal test harness does the same at its generic request entry point, rewriting each cited ref to the placeholder's real digest and appending the mirrored resources, but only for strictly well-formed JSON monitor bodies — a body with trailing content keeps the raw path so the store's grammar refusal still fires first over the closure's. Each side gains a witness that pins the closure itself: a monitor run citing a resource outside its commit is refused with managed_session_resource_missing, and each witness was observed red with its branch disabled before landing. * feat(managed-agent): gate monitor_run behind the managed-session/2 reader The H3 design left the reader-gating mechanism to the implementation; the journal settles it. Headers are immutable — the scanner refuses a repeated managed_session_header_v1 line and any unknown subtype after the header — so no transaction can legally raise a Session's requirement in-journal, and a superseding header line would change the journal format in both languages for a narrow mixed-version window. New Sessions therefore stamp minimumReader managed-session/2 at creation, and every path that commits a monitor_run record — commitDomainRecord, commitExtensionRecord and the generic append guard — additionally refuses when the Session's header names less. Readers still accept requirements up to their own version, so pre-H3 Sessions stay openable everywhere and simply can never receive monitor_run. child_run is deliberately not header-gated: a managed-session/1 reader already parses its name, and the asymmetric hazard lives on the server side, which the established deploy-server-first sequencing covers. The shared Stage H golden is regenerated for the v2 genesis bytes and the Java integration replay passes against it unchanged. New witnesses: new Sessions stamp v2, a v1 Session refuses monitor_run with the precise error (observed red with the gate removed), and both header sides of the version comparison are pinned in the records tests. The design doc records the decision and the now-paid closure precondition in both languages. * feat(managed-agent): funnel monitor_run revisions through a hosted orchestrator The mirror of HostedChildRunSession for the Monitor domain. The managed-runtime worker will own the watch loop, but the dual path puts every product record on the hosted authority, so HostedMonitorSession commits the record line as watch facts arrive: revision 1 with the intent before any side effect, dispatch onto a Runtime binding, the set-once start receipt after the watch starts, observation revisions only while attached with a watermark that never goes back, output advancement only forward, and the terminal shapes — settled by exited/max_events/idle_timeout, failed by a start failure on not_started_proven or a settled watch failure, cancelled by a stop request whose notification watermark may still advance after. Writes serialize per Session and replay-safe by command id with deep-equal skip, exactly like the Shell funnel. The suite drives the whole line through a real authority: projection at admission and after settle, max_events refused before its quota and accepted at it, a second start receipt refused, an observation before attach refused, and an idempotent output advance committing nothing. Quota (enforcement), the debounce-floor observation loop with notification composition, and rebuild-after-loss stay with the runner and route increments that follow. * feat(managed-agent): drive monitor observations through a debounced loop The loop half of the Monitor runner, on top of the hosted funnel: an admitted Monitor dispatches, starts an injected watch executor and runs its own time. Stdout lines aggregate into one observation revision per debounce window with a one-second floor, so the document's reopen bound holds for any debounce a watch asks for; the run settles itself on the contract's own terminal conditions — the observation quota, silence past idle_timeout, a natural exit with its buffered lines flushed first — or is stopped, which terminates the watch and lands cancelled. The executor contract is typed for the worker's cgroup owner (onLine, onExit after the start resolves, a terminate handle); the suite drives the whole line through a real authority with a fake executor and a manual clock: windowed aggregation and its floor, quota settlement and watch termination exactly at max_events, idle settlement re-armed by each accepted observation, exit flush order, start and mid-run failure shapes, and a stop that ignores late lines. Every cross-boundary settle — timer or executor callback — is counted by the loop's done handle, which the close drain will await; window flushes count too, so done never reports idle while an observation commit is in flight. Notification composition and its wake stay with the next increment; quota enforcement and rebuild-after-loss stay with the route slice. * feat(managed-agent): notify from monitor observations in the same transaction The H0c machinery now runs for Monitors: every accepted observation commits its revision together with a notification input and the wake the authority generates for it, and the run's notifiedThrough watermark advances to exactly that observation's sequence. The loop composes the input — monitor-1:notify:N ids, the window's joined lines as the content a woken turn will read, an empty admission, no deadline — and the funnel threads it through commitExtensionRecord, so a recovery replay re-runs nothing the watermark already covers. Dedupe semantics are pinned at both levels: the funnel advances the watermark only when a notification rides the revision (an observation without one keeps it where it was), and the loop suite shows the three events — domain.committed, input.accepted, wake.requested carrying the input's accepted id as its source — landing in one transaction with the joined lines as its content. Consumption of that wake by the hosted scheduler, and settling the input when no turn can take it, come next (open question 6 of H0c). * fix(managed-agent): keep a finished background Shell's receipt answerable H7 of the real-stack rounds: the Broker settles its :process row only when shell-status answers exited, but the worker's registry deleted an entry the moment the Shell ended — a natural exit could never be answered again, and the ledger row would hold the Runtime against release and close forever. The registry now retains each completed unit's receipt until the worker ends: the hold drops exactly as before, but shell-status and shell-terminate answer a proven end from the retained evidence, idempotently, for the owning Session scope only. An end without evidence keeps nothing and answers unknown as before; the existing live flows are untouched. Suite: natural exit keeps answering exited with its evidence for both kinds after the hold dropped, the wrong Session still hears unknown, and an evidence-less end stays unknown. The close-drain/reconcile callers that consume this, and the Java-side settle-before-busy on workspace release, land with the next batch. * fix(managed-agent): renumber the task-event outbox to V36 and narrow the commit-time guarantee #13301 landed V35__managed_session_tool_profile.sql on main, so this branch's V35__managed_task_event.sql now shares its version: the merged tree fails at startup with "Found more than one migration with version 35". Move the outbox migration to V36 and update the authority design doc in both languages. The commit-time validator mirrors the checks the authority applies to every line whatever its kind (envelope, vocabularies, domain checks, reserved ids, byte caps), not the per-kind payload schemas, subject rules or commit-marker digests, which stay the authority's contract. Reword the apply() Javadoc, the test comment and the design doc so they no longer promise that no commit can brick the Session, and say what a stored reserved id breaks: it collides with the id of that domain's next Stage H record rather than failing the next open. Pin the event id and operation id shape checks with two negative cases; disabling either check previously left every test green. * fix(managed-agent): settle provable background exits before release's busy check The second half of H7. The worker now keeps a finished Shell's evidence answerable, so the Broker has someone to ask — but observeBackgroundProcess still had no production caller, so a naturally exited Shell left its :process row non-terminal and every release wedged on runtime_session_busy. releaseSession now sweeps the background process rows this Broker admitted for the Session before it computes busy: each unproven row asks its physical owner once through the same shell-status → settlePrepared leg observeBackgroundProcess already owns, a proven exit settles with its evidence, and any lookup that fails or cannot prove an end simply keeps the row so busy stays the accurate answer. The sweep tracks per-context admissions because the reconciliation scan deliberately excludes PREPARED rows, and mutates the admission set under the context lock. Witnesses: releaseSettlesAnExitedBackgroundProcessBeforeBusy proves an exited row settles with its evidence and the release succeeds (observed red with the sweep bypassed); the pre-existing unprovenProcessStatusKeepsItsHold still proves an answer without proof keeps the row and the busy refusal. Full module 604/604. The workspace close drain still terminates shells in order rather than sweeping — its ordered stop belongs to the close-drain slice, and these two verbs are exactly what it will consume. * fix(managed-agent): check the journaled receipt before attaching a background Shell's start H6 of the real-stack rounds: acceptBackgroundShell called childRuns.attach before its tool.receipt replay check. A retried accept — the leg recovering pending results re-drives — would mint a second start receipt on the same identity before reaching the replay path, and the record's set-once rule would refuse it as a rerun shape, so a recoverable retry crashed as a conflict. The lookup now runs first: a journaled receipt replays its validated history with no attach at all, and a fresh accept attaches only when this identity still lacks a start receipt — the record id IS the execution identity, so an existing one is always our own crash-window attach between record and journal, which the retry now tolerates instead of minting a rerun. Witnesses: a retried accept answers from the journal with exactly one attach and one tool.receipt line, and an already-attached-but-unjournaled record completes its accept without attaching again. The full hosted turn suite stays 148/148. * fix(managed-agent): backpressure the background Shell output pipes G5 of the real-stack rounds: the executor fed the bounded capture with a fire-and-forget sink.write — every chunk queued behind asynchronous publication, so a fast producer flooded memory (a 256 MiB run peaked near 315 MiB RSS at a 20 MiB/s store). The feed now counts bytes in flight and pauses stdout and stderr when the drain falls sixteen MiB behind, resuming at four MiB — enough to cover a store several seconds slower than the process produces without stalling a steady producer against the floor. Completion or rejection of a write both release their bytes, so a broken capture can never leave the pipes paused. The witness drives a controlled sink: seventeen MiB of chunks pauses a stream exactly once, draining past the watermark but above it resumes nothing, and draining the rest resumes exactly once. The outputRef-live-advance half of G5 belongs to the turn-publisher's promotion to activation scope, sized with the close-drain slice, and stays queued there. * feat(managed-agent): promote the Shell publisher to Session scope The ②d-out debt, and with it the live-advance half of G5. The Shell publisher was a per-turn object: its server died with the turn, while background Shell traffic flows across turns — between turns (or after a close) the worker's retained descriptor pointed at a dead endpoint, so no live output advance and no exit leg could ever complete past one turn. The hosted Session now owns one long-lived publisher across turns (the turn lazily structures it on the shell options and never closes it; the Session's ordered close in this PR's fifth slice drains it). The two small wire-level fixes that the promotion exposes: - The worker's background execute stamps the background marker into its prepare request: the closed six-field wire capture has no such field, so a remote prepare used to arrive unmarked and die at admission; the executor derives the marker from the background input itself. - register() takes the registering turn's prompt for the foreground admission check, so one instance admits the foreground of any turn of its Session while a foreground pretending another turn's identity is still refused; the background arm keeps its record-based admission. Witnesses: a single instance admits a background watch plus the foreground of two turns and refuses a misattributed one; the worker rig pins the stamped marker on its prepare; the four touched suites stay 184/184 and both package typechecks pass (after a core dist rebuild for the capture type's optional background marker). The publisher's own descriptor installs idempotently per turn (identical descriptors are accepted), so consecutive turns never fight for the worker slot. * feat(managed-agent): drain a Session's background Shells ahead of release The close-drain's missing legs. Until now a Session with a live background Shell could never close: the worker's activation gate refused release with managed_activation_conflict forever, and the CLI-side close route skipped the background stores. The ordered sequence the design asks for now exists end to end: - The registry gains stopSession: terminate one Session's entries with the supervisor's evidence rules (TERM, escalate, empty-check), await each completion so its finalization — last manifest, sealed envelope, record settle through the session publisher — lands before the caller proceeds; an unproven stop keeps its hold for the caller to see. - The worker's activation release gate drains before it refuses: active background work no longer wedges a close that can prove its stops, and anything still unproven keeps the exact same conflict. - The hosted close route now closes the Session-scoped publisher right after the broker release returns — which by then has drained the Shells and settled their records through that publisher — and the drain closes every capture store, background families included. Witnesses: a registry drain settles one Session's shells with evidence, both kinds, while leaving another's holds untouched, and an unprovable terminate under a session drain still preserves its hold; the 766-test context-worker suite stays green through the release gate's new async drain branch. What remains deliberately unproven locally: the full-chain release-with-live-shell drain needs a real cgroup supervisor, pinned as the Linux head acceptance instead of a mocked supervisor. * docs(managed-agent): sync the H3 design's status line with landed slices The status line still claimed publication and orchestration remained design while most of them landed across this PR's batches; refresh both language versions with the landed list and keep the remaining-design list accurate: the Monitor physical side (route, rebuild, cgroup watch) and wake consumer, both submission enablements, task events, and the Linux physical acceptance. * feat(managed-agent): spawn Monitor watches through a cgroup watcher The physical half of the Monitor runner's executor interface, for the worker: ManagedMonitorWatcher spawns the watch command under a unit named for the execution identity, through the same supervisor (its membership attestation unchanged), splits stdout into the observation lines the loop buffers — boundaries respected, the Legacy partial-line cap honored, the tail flushed as the final line at exit — and reports the physical end exactly once, natural exit versus mid-run failure, so the loop's settled and watch_failed shapes both hold. stderr rides only the future output leg, never observations; terminate defers to the supervisor's TERM-escalate-empty rules. The suite drives a supervisor double: unit/executable/cwd derivation, chunk-boundary splits with the final flush, the cap drop, spawn-error and membership-failure shapes, and terminate delegation. * fix(managed-agent): accept message.retracted at commit time #13351 added the message.retracted event kind on main. The hosted text delta stream writes it when a restarted model attempt replaces a prefix it already published. The commit-time vocabulary this branch mirrors from the authority still listed sixteen kinds, so after the merge with main the store refused every retraction with 409 ("event.kind must be one of [...]") and left the orphaned deltas in place. Add the kind to ManagedExtensionRecords.EVENT_KINDS and to the shared store fixture, so the parity pins in both languages agree again. * feat(managed-agent): answer monitor status and stop from a worker registry The monitor half of the maintenance route, mirroring the Shell's: a worker-side registry of supervised watches that holds each Session's Runtime while a watch runs, answers an end only with the supervisor's own evidence, and retains a proven exit's receipt until the worker ends so maintainers can learn the truth idempotently. The route sibling at /internal/managed-runtime/v3/monitors serves monitor-status and monitor-stop with the same closed wire shape, target-identity rule, and unknown-but-never-unproven discipline as the shells' route; it joins the worker's route list so the v2 envelope admits it. Registration arrives with the Monitor start admission later — status/stop already have consumers in the release drain and the rebuild slice that follows. This also repairs the watcher test arity that broke the previous push's build lanes (the mock's 2-arg onExit omission, seen as error TS2554 (155,34) across every CI lane at that head). * feat(managed-agent): admit monitor watches into the open-ended capture session The γ3-α precondition for every later Monitor leg: LocalShellStreamResultSession's prove-a-live-start admission reads the capture family's record domain instead of being pinned to child_run — same Session/binding/background-marker checks, same live-run gate, same identity captureId, per-domain labels so a refusal names its own family. The producer side names its domain explicitly (record presence is never consulted where both domains could in principle compete), so a monitor watch carries its monitor_run record into exactly the same bounded stream the background Shell already rides. Witnesses: shell admission unchanged; monitor admission succeeds under its domain; wrong-domain proven start refuses; missing record and foreign binding refuse with their own errors. * feat(managed-agent): admit Monitor watches through the v3 executor The workhorse of the Monitor leg. The worker's v3 executor gains its is_monitor branch beside is_background: journaled admission first, then refuses (never as a transport error), then the watch spawns through the ManagedMonitorWatcher under a unit named for the execution identity, streams stdout into the SAME bounded family (16/4 MiB backpressure on the pipe, with an explicit refusal if the watcher reports no supervised unit), and settles with its detached handle while a registry keeps the Runtime hold. Three watch-level facts are done physically now: - The monitor registry is a wrapper around the background Shell registry, never a copy: supervisors, EOF drains and publisher finishes stay in one place, while monitor holds and the four-watch quota bucket remain entirely their own. Its maintenance route answers run/exited/ unknown with identical semantics to shells'. - The execute path gates on the generated four-live-watch cap with `Session already runs 4 Monitor watches.` and additionally on a missing cgroup root and a missing capture service. - Each prepare carries background+monitoring markers, so the hosted side of γ3-α knows the monitor_run domain to check. Suite: 28 tests across the three touched files — the fork settles a detached start under the right unit and marker, lines land as monitor output, the fifth watch is refused, and supervisor-less boot refuses with its dedicated text. The turn admission arm (funnel-driven admit→dispatch→attach through HostedMonitorSession), the watcher→loop observation feed, and rebuild still land in the following increments. * fix(managed-agent): refuse a domain.committed without a textual domain A domain.committed line whose payload.domain was missing or not a string read as "no domain" and skipped requireDomainCommitted, so it committed 200 and the authority refused the Session at the next open. Decide that a line is domain.committed from its kind alone, and take the domain from the vocabulary check before building the expected record kind. Also address the rest of the first /review round: - The commit marker is the only line apply() measures. validateUtf8JsonLines already caps every line at MAX_EVENT_BYTES, which it now names. - The outbox's MAX(sequence_id) stays a plain read. A locking read deadlocks concurrent first announcements of different Sessions on MySQL 8.4 (85 of 480 in a probe). The comments now name the guarantee that actually holds: the Session store's journal-head lock is taken before the transaction's first plain read. - Stop claiming that the Session event stream carries only turn and lifecycle events. V36 is not applied anywhere yet, so its comment can still change. - Restore the shipped v1.19 OpenAPI sentence and record the move to the outbox as v1.30 (info.version 1.30.0). - Reconcile both design docs: Decision 10's state table and the scope of its wire-only absolute, the goals, the public-contract bullet, the validation plan, the files list, and the restored caveat that the record contract does not catch an unprovable not_started_proven. - Tests: refuse a non-textual and a missing domain, commit a message.retracted line, pin the final "events, then its commit marker" guard again, assert the empty outbox in the MySQL IT, and give the oversized-resource case bytes at its own sequence. * feat(managed-agent): admit Monitor watches through the hosted tool turn The turn-side arm of the Monitor leg. A call named monitor is now admitted through its own three-state gate (requested/unavailable/ill- formed, with the Shell-profile Monitor declaration now advertised), implements the same record-first discipline as background Shells do — the funnel admit commits monitor_run revision 1 before any side effect, dispatchStarted follows the durable checkpoint, and the watch attaches when its v3 execution settles. The granted publication, journal intent, dispatch and accept forks are widened to carry is_monitor alongside is_background, and acceptMonitor mirrors acceptBackgroundShell end to end (not_started → start_failed with the unstarted family; detached → same blocked receipt family with exactly one tool.receipt line, the H6 journal-verdict-before-attach order, and the blocked ack completing the row). HostedSession gains a session-scoped HostedMonitorSession beside childRuns; turn construction forwards it at both sites. Two existing contract pins move deliberately: the Shell profile's advertised list now names monitor (the harness declaration test names it explicitly), and the three publisher-close witnesses now assert the Session-scoped lifecycle settled by the earlier promotion — no close at any turn ending (completed/error), the server still answering, close exactly once at the Session's own delete. Suite: 3 new admission-arm tests through a full hosted turn (funnel order, clamped defaults with the Legacy monitor caps, start_failed shape, unavailable text), the whole turn suite stays 151/151 and the harness session suite 180/180. The observation fan-out (worker lines → hosted loop), quota witness at the turn, and rebuild-after-runtime_lost land in the next increments. * feat(managed-agent): fan a monitor watch's observations into the hosted loop The observation channel of the Monitor leg. The session publisher now marks monitor captures with recordDomain monitor_run through γ3-α, fans each published stdout chunk through the Legacy line-split semantics — boundaries kept, remainder held across chunks with the 4096-byte partial cap — into a per-capture observer, and closes it with onExit(failed) at finalization with the tail flushed first. The turn gains resumeMonitorWatch: a fresh accept creates one HostedMonitorLoop per Session report and resumes it from the already-attached record, fed through a HostedMonitorRemoteExecutor whose start binds those observer callbacks; a replay starts nothing twice, exactly like it never re-attaches. Globally, loop timers unref so observation windows never anchor the process. Monitor capture finalize settles through the monitor funnel itself (watch_failed when no evidence, exited via settleQuiet), beside the shell's unchanged path. The publisher carries monitors next to childRuns, and the turn's own construction forwards it. Suite: the admission-arm tests now also pin one Session-level loop starting on a fresh accept with the record-first flow intact; wide suites of 189 tests pass with cli tsc, prettier and eslint clean. Monitor output before the observer registers stays durable in the open stream, with observation starting at attach — noted in the design not as a flaw but as the attach seam, a line crossing it already lands in the output Artifact readers see. * feat(managed-agent): rebuild a read-only Monitor after its Runtime is lost The record side of the rebuild promise. HostMonitorSession gains the two verbs the H3 lifecycle needs around runtime loss: blockedRuntimeLost lands the run in recovery_blocked with the runtime_lost reason and the lost binding still named, so degrading is what the projection shows until something decides otherwise; rebuildFromRuntimeLost accepts only the outcome_unknown/runtime_lost line the H0b rule allows, starts a fresh watch under a newer Runtime generation with its own receipt, and keeps the reason for honest projection while observation positions never moves — the support case never restarts what the watermark holds. monitorRebuildAllowed routes every candidate command through the Legacy AST read-only check with its walk, so a side-effecting command stays blocked accurately. Five witnesses cover the full rebuild chain (recovery-block, rebuild with receipt remint and generation increase, a rebuild refused from a live run and from an ended one, and the gate's read-only/side-effect/empty answers). Binding-replacement detection and re-drive of the new physical watch come with the recovery slice after the wake consumer. * feat(managed-agent): serve the task events routes with their SQL journal H3 flips listSessionTaskEvents and queryWebShellTaskEvents from planned to partial on both surfaces. The Java control plane gains a bounded per-task event journal (V39): one row per committed task event, written in the Session commit transaction, with the per-task sequence allocated under the journal head lock in commit order. Record revisions journal their task view changes at the H0c announcement point. The durable retention floor lives in a per-task cursor row, so it survives an empty retained set, restarts and projection rebuilds; the Artifact visibility barrier pins it behind unarchived output, and the backlog bound refuses past capacity instead of silently discarding. Output events carry their per-stream segment ordinals with overlap/gap refusals, and artifact references stop fail-loud at the 100 bound. output_cursor/outputCursor lift with the routes and the task views serve them. The section 6.1 demonstrations run as store-level contract traffic in the new journal contract test, the route probes and record gates cover the flips, and the #12847 C15/C16 instances close alongside. WebShell types regenerated from the bumped 1.30.0 contract. * docs(managed-agent): add the task events lane's build report Flyway numbering findings (V39 taken; V35-V38 claimed by in-flight PRs on main, V15/V29 Java-only burns untouchable), implementation decisions against the H3 design text, the CLI/daemon flow-surface finding (none exists, regeneration is the only surface change), the full test ledger and the owed items. * feat(managed-agent): build the monitor wake envelope and the pending-input intake Two support pieces for the H3 wake consumer (β2-b1a): - monitorNotificationText wraps one due observation window in the exact Legacy task-notification envelope — task-id, optional tool-use-id, kind/status/event-count, summary and result — applying the same truncateNotificationLabel/stripDisplayControlChars/escapeXml pipeline the Legacy Monitor applies, so a turn learns nothing new. - pendingSessionInputs derives the Session's still-pending inputs from the journal alone: an accepted input is consumed exactly when a turn settles under its turnId, order-insensitively, so a replayed admission after a restart is not re-driven. * feat(managed-agent): deliver the Legacy notification envelope on monitor wakes The loop's notification input now publishes the exact task-notification envelope a Legacy Monitor's wake carried — task-id, tool-use-id, kind, status, event-count, summary and the window's lines — so the turn the wake raises reads nothing it has not already read. The description is the tool call's, falling back to the watched command. A witness pins the exact wire text of the first observation's input content. * feat(managed-agent): run the monitor wake through an embedded scheduler H3 closes the wake loop: an accepted observation already commits its notification input and wake.requested in one transaction; this slice makes that wake effective. The pump re-derives the pending monitor inputs from the journal alone — nothing is held in memory, so a restart re-derives the set. On an idle Session the oldest notification runs as an ordinary text turn with the Legacy envelope as its prompt, and the turn's settle consumes the input. A busy Session queues in the journal exactly like channel and Goal inputs; the retry loop redelivers. A Session that is parked, blocked or mid-recovery keeps its reminders pending and reports them. On the close path every pending notification settles cancelled without a model turn, under the turn-result record's own idempotency key, so no wedged notification ever holds a Session as hosted_turn_recovery_required at its next open: the parked-turn scans now skip monitor inputs, which the pump owns exclusively. A wake turn that dies without settling re-drives on reopen — the input is consumed exactly when a turn settles under its turnId, and a crashed turn never settles — matching the queue semantics channel and Goal inputs already honor. * fix(managed-agent): keep every new Session on the managed-session/1 reader Round-5 real-stack verification's cross-version matrix disproved the v2 stamp's premise: monitor_run has been a parsed, known domain since #12837 (v0.24.7), and a v1 reader opens a Session holding monitor_run records without complaint. Stamping minimumReader: managed-session/2 on every new Session — while monitor_run remains disabled — would make a rollback or a mixed-version rollout lose access to every Session created in between, and buys protection no deployed binary needs. Revert the stamp to managed-session/1 and drop the per-domain header refusal; the enablement list remains the single domain gate. The stamp rises only when a change genuinely breaks an older reader mid-scan, and only one release after a tolerant reader ships. The golden fixture's genesis header regenerates to v1; the records and authority suites gain v1-header admittance witnesses in place of the refusal cases; the design's reader section now records the reversal with its evidence. * fix(managed-agent): keep monitor output byte-true and advance only forward Round-5 real-stack verification found three facts about the monitor output path: - A capture rebuilt from decoded lines cannot be byte-true: per-chunk toString split multi-byte runes into U+FFFD, and line-rounding dropped blank lines and the final unterminated line. The watcher now forwards raw stdout chunks to the capture on a new onChunk channel, so the durable Artifact reproduces the command's stdout exactly, while the observation lines keep the Legacy emit-path semantics (a StringDecoder holds a split rune until it completes; a blank line consumes no observation). - The watcher registered its exit listener only after the supervisor's start resolved, so a fast watch could end unheard and the debounced loop never settled. A latched end with a check-phase sweep covers a watch that exited before its listeners attached, after the host is known to be listening. - The funnels' advanceOutput promised "only ever to a newer revision" in prose and accepted an older manifest in code. Both hosted funnels now read the revision of both manifests and refuse a regression before anything commits; witnesses pin the refusal in place. * feat(managed-agent): drain a busy background Shell at release, and answer the asked operation Round-5 close-path findings (H8): with a running background Shell the Broker refused busy before transport.release — yet that HTTP call is the only route to the worker's drain — so a `tail -f` held session release and workspace close forever, with nothing able to stop it. The release path now distinguishes busy kinds: anything active that is not one of the Session's own background process rows keeps the exact busy refusal with zero worker hops, while proven-running rows trigger a dedicated release call first whose 409 busy answer is swallowed (the worker's route drains before refusing, and its maintenance routes are not activation-gated, so the follow-up sweep still proves what the drain ended). Rows settle from the drain's evidence and the complete release converges on what is genuinely still running — one call when the drain finishes in time, one retry when it does not. Two witnesses pin the contract: a running Shell reaches the worker (the round-5 probe found the refusal never did) and a non-background busy refuses without a worker hop. Also from the round: the shell maintenance validator computed the expected operationId from targetOperationId but never compared it with the answer's own operationId, and the sweep witness answered with the process-row id where the real route echoes the call id. The validator now refuses a view answering another operation, a new protocol suite pins it, and the witnesses answer the shape the route really sends. * fix(managed-agent): re-assert the output pause behind every flushed chunk Round-5 G5 finding: Node's flushStdio resumes both paused pipes when the background launcher exits, and the executor tracked the pause in its own flag rather than the stream's, so a writer that outlives its launcher pushed 230 MiB into the queue and climbed RSS by +200 MiB. The onOutput accounting now re-asserts the pause behind every delivered chunk, and applyPause is idempotent through the stream's isPaused state instead of the local flag — the stop-and-go the backpressure intended holds through the launcher exit, on both the Shell and the Monitor paths. The existing 17-MiB pipe pacing witness keeps its single pause call. * fix(managed-agent): mirror message.retracted from the H3 event vocabulary in the commit-side kind mirror * fix(managed-agent): validate every domain.committed payload, with or without a textual domain * test(integration): expect the monitor tool on the hosted shell profile The hosted shell profile now advertises `monitor` beside `run_shell_command`, so the fake-model driver's tool-name assertion — which the Hosted workspace tool-turn IT checks on every model call — must expect it. The file profile's expectation is unchanged: the declaration is shell-profile only. Fixes the single red IT in the MySQL fault-gates lane (HostedWorkspaceToolTurnIT.packagedHarnessUsesSavedWorkspacesThroughReal BrokerWorkerAndSqlStore), whose only difference was the extra 'monitor' entry. * fix(managed-agent): read the whole journal when deriving pending monitor wakes `readEvents()` is a bounded read: with no arguments it returns the first `defaultReadEvents` (100) events, capped at 256. Both consumers of the pending-input derivation called it bare, so on any Session whose log runs past that page — every real Session with tool calls — the derivation saw only the log's head: - the wake scheduler's `next()` never found a notification input, so a monitor wake was never delivered as a turn once the Session passed the page. The feature was dead in production while every rig stayed green, because the rigs commit a handful of events; - `settlePendingMonitorInputs` never found the owed inputs at close, so they stayed unsettled and the Session reopened as `hosted_turn_recovery_required` — exactly the wedge this close-path settle exists to prevent. Both now read the full committed prefix with `eventsInSequenceRange(1, committedSequence)`, the discipline every other journal scan in the hosted session already uses. The witness commits 110 observations so the notification lands beyond the default page, asserts the log really exceeds it, and was verified red under the inverse edit (the settle returns 0 and the input stays owed). * fix(managed-agent): stop a full task journal from wedging record commits The task event journal is a derived, bounded feed: its backlog bound refuses the next event past capacity rather than silently discarding it, and an unarchived output event pins the retention floor so the automatic pass cannot expire anything. The record commit called `appendStateChange` inline, so once a task's journal was pinned and full, that refusal propagated out of the extension-record commit — every later revision of that task failed, which would take the Shell and Monitor record lines with it. The record row is the authoritative state and is already written when the journal append runs, so a refusing journal now degrades the feed (warned with tenant, session, task, revision and the refusal code) and never the commit. The store still throws for callers that can apply backpressure — which is where the output producer's retry belongs. The witness drives the real wedge: it pins the floor with an unarchived output event, fills the journal to the bound, asserts the next append refuses with `managed_task_event_backlog_full`, then commits a record revision whose view changes and asserts it lands. Verified red under the inverse edit, where the commit dies with "The task's event journal holds 512 retained events". * refactor(managed-agent): drop the shell view's never-set error field The maintenance view declared an optional `error` object, the Java response validator carried a branch for it, and no producer on either side ever set one: the worker answers `running`, `exited` or `unknown`, and a route-level failure travels as an HTTP status with its own code body. A field that is declared and validated but never written is a dead switch, so it comes out of the TypeScript view and the Java field set — which now also refuses an answer that carries it, instead of silently accepting a shape nothing produces. * fix(managed-agent): let a replayed output advance pass instead of refusing The forward-only check refused any advance whose revision was not strictly newer, which included a redelivery of the very reference the record already holds. The publisher guards against that in memory, but the guard does not survive a worker restart, so a replayed advance could throw inside the background exit leg and wedge a Shell or Monitor that had already ended. Both funnels now treat an identical reference as the no-op the deep-equal skip already owns, and still refuse any other reference that is not strictly newer. For a Shell the replay commits one liveness step, since the live run line alternates by design; for a Monitor it commits nothing. Witnesses pin both, beside the existing refusal witness. * chore(managed-agent): keep the task-events lane report out of the PR diff A one-off construction report does not belong at the repository root: the project's directory rules put working artifacts under the ignored .qwen tree and keep the tracked docs/ for design and plans. The report is preserved at .qwen/investigations/2026-10-04-task-events-lane-report.md, and its Flyway numbering note is superseded by the V40 renumber the merge commit records. * merge: absorb review-round state and retarget onto the H3 stack * fix(managed-agent): quiet the wake pump's consume guard at a blocked owner runTurn's recovery-blocked path marks the Session blocked and returns 'settled' without the turn ever existing, so the input stays owed — the accurate answer the whole design gives for a parked Session. The pump's after-settle check read that as a consumption bug and threw, which both marked the already-blocked Session again and logged a spurious error on every such observation. It now stops at an accurate blocked state and still throws when the input simply did not move: that is the programming error the guard exists for. * fix(managed-agent): meld a monitor watch through one terminal step Round-6 found the shape the Monitor's final leg really had: prepare's open-then-advance could never survive the start order (output before a start receipt is a parse-level refusal), the manifest still advanced through the Shell funnel whatever the capture's recordDomain said (a Monitor never builds its own Shell record), and the finalize settled the record ahead of the loop's own exit chain, so a watch's last window died in denial between two async hops — with its swallow queued to silence it. Three changes, each with its own witness that goes red under the inverse edit: - The manifest advance reads the record's start receipt first and only advances past it; a missing record still fails loudly instead of the old "no record to revise" at whatever funnel happened to be there. - It routes by recordDomain: monitor captures land on the Monitor funnel, Shell captures on the Shell's. - The loop's onExit now returns a chain carrying the exit's own flush → settle in order, and the publisher awaits it rather than settling itself — a final window can no longer be dropped ahead of the settle, and a commit failure of the chain surfaces inside the finalize instead of returning an empty observation state silently. The end-to-end witness drives the production shape on a real authority, orchestrator and Publisher over HTTP: prepare ahead of attach, output advancing only after the receipt, then one terminal step holding observationSequence 1 with the settled mark behind it. * fix(managed-agent): settle the process row when the Runtime proves no start Round-6 found the permanent wedge: the `:process` row admits PREPARED ahead of dispatch, but an admitted-refusal answer (`{executionStatus: "not_started", capture: null}` — cgroup root missing, quota exceeded, and friends) settled only its invocation. The process was never registered by the Runtime, so every later shell-status can only be unknown, the exit-only settle path never applies, and every release carries the hold forever. A proven never-started is proof: when the invocation settles with executionStatus not_started, the sibling row now settles with the same not_started mark, its admitted refusal as the evidence. Busy computed after that never counts a process that could never have existed; an unproven-lost connection still wedges exactly as the declaration says. The witness drives the admitted refusal shape end-to-end: the process row lands SETTLED with not_started, and release passes without the busy answer. Verified red under the inverse edit, which reproduces the report's `expected: <SETTLED> but was: <PREPARED>` verbatim. * fix(managed-agent): attribute wake receipts to their own turn and stop re-driving it Two wake-path defects from the round-6 audit family, plus the monitor registry the maintenance route actually shares: - `recoverShellReceipts` exempted monitor inputs so thoroughly that a wake turn's receipts could never be attributed: any `tool.receipt` journaled under a wake turn would attach to the previous prompt or be skipped. Receipt attribution now advances on every accepted input, while the pending set itself still exempts monitor inputs as before — nothing reframes them as parked Turns. - A wake turn that started, journaled records and died was re-executed text-only on reopen, minting a second user record and leaving an unanswered first call in the model's history. `wakeHasPriorAttempt` names the condition simply; the pump now parks such turns accurately blocked for the recovery fleet instead of redriving. - The worker constructed its executor with no monitor registry, so the executor built a private one while the monitor maintenance routes got another empty one — every status and stop answered unknown. One registry is shared by executor and routes now, mirroring what the Shell path already does. * fix(managed-agent): read the task event floor after the events, not before Two stores-and-reads of an event page ran as separate autocommit queries: the retention floor was checked before `events.read`, so an expiry mid- read produced a 200 that silently skipped events under an unchanged nextCursor, the exact gap the contract's `409 cursor_expired` exists to answer. The replayable-events family already documents the sound order: events first, floor after — a monotoni…





What this PR does
This PR defines slice H0b of #12827, the shared record contract of stage H in the Managed Agent proposal #12380: the
managed-extension-record/1records that MCP, Hooks, background Shell and Monitor, child agents, Channels and automation will all use on the Managed path. It is a contract only. Nothing commits these records yet; H0c commits them with the Session authority and rebuilds the task list from them.The contract has six parts:
OperationGrantlets the owner of an accepted operation finish its planned phases without a model activation. It names the Session, operation, domain, revision, owner, Workspace generation, lease and a scope: the committed domain record that holds the plan (its kind must match the domain) and the phases it admits. It has no activation fields, so it can never pass for an activation. A Runtime gate may replace a grant only with a renewal that extends the lease or with a later revision that does not go back to an older Workspace generation.reserved…recovery_blocked), its physical execution (intent…outcome_unknown,corrupt) and the delivery or acceptance of its result, each with its allowed single steps. Delivery has two targets: a Channel (sending,partial,delivered) or a Session (accepting,accepted,consumed). The lines are tied together, so an unknown or corrupt execution keeps the runrecovery_blocked, a run ends only with an execution proven to have ended, and a Session accepts a result only after the run ended.outcome_unknown,execution_corrupt,runtime_lost,dispatch_unknown,handler_unavailable) and six quota reasons, each allowed only in the states it explains. A running or waiting run may carry a recovery reason, which is how a degraded run is recorded.executionCallIdoreffectId,dispatchIdanddeliveryIdidentify the physical call, the cross-Session dispatch and the external delivery.monitor_run. The record body covers section 10 of the reference design: the start call's arguments, limits, debounce window, start receipt, observation and notification watermarks, stop reason and output manifest. The domain joins the closed v1 index, as the storage design lists it among the 33 v1 names, but stays disabled for submission.One shared JSON schema and fixture file pins the contract, with 520 cases, each invalid case aimed at one rule. A Python implementation written from the design document, independent of both languages and kept outside the repository, labeled every case, and the generator stops when the two disagree. A TypeScript module in core and a Java validator in
managed-agent-server, the control plane that reads these records to build the task projection, replay every case, check every pair of states on every line, and pin the constants. The bilingual design document records the decisions, rules and open questions.Why it's needed
On the Managed path a Runtime can be reclaimed and a Harness replaced. The Legacy implementations of these capabilities recover from in-memory registries, PIDs and local files, and none of those survives. #12827 makes every asynchronous capability a durable resource plus a trigger intent, with stable identities and receipts. If each capability invented its own states, grant, reasons and identity rules, Java, qwen and the Runtime would read "unknown" and "settled" differently, and no single task projection could merge them. As with the attestation, Tool v2,
managed-context/1andmanaged-tool-result/1contracts, the shared shapes and fixtures are settled first, so that H0c and H1–H6 build against a fixed, reviewed target.Reviewer Test Plan
How to verify
monitor_runis still refused for submission.managed-agent-servercontract test and Checkstyle on JDK 21. The module needsqwencodeandruntime-brokerinstalled first, as its README says.Evidence (Before & After)
N/A: there is no user-visible change. Local results, on the branch rebased onto
mainat88d881491b:tscare clean for core. The contract suite passes 530 tests and the records suite 80; all 23 test files undersrc/managed-runtimepass (1443 tests).managed-agent-serverpasses 111 tests, 7 of them in the new contract test, and Checkstyle is clean.return falseand each equality, logical and relational operator of the module was mutated in turn; 218 of 219 mutants fail a test. The survivor turns>into>=right after the equal-revision branch has returned, so it is equivalent.requireand each earlyreturn falseof the validator was disabled in turn; 45 of 45 mutants fail a test.Tested on
Environment (optional)
Unit tests only. Linux: Node 22 and JDK 21. macOS and Windows were not run.
Risk & Scope
monitor_runto the closed v1 index rather than a v2 index (question 1 of feat(managed-agent): Stage H extension runtime, starting with H0 shared records and the task contract #12827);watch_failedas a Monitor stop reason, for the Legacy failures after start (non-zero exit, signal, binary output, stream error).domain.committednow parsesmonitor_run, while submission still refuses it; when H3 enables the domain, a Session that holds such records must keep readers from before this change out, for example by raising itsminimumReader.Design document: English · 中文
Linked Issues
Refs #12827 (H0b; H0a and H0c follow)
Refs #12380
中文说明
这个 PR 做了什么
本 PR 定义 #12827 的 H0b 切片,即 Managed Agent 提案 #12380 中 H 阶段的共用记录契约:
managed-extension-record/1。MCP、Hooks、后台 Shell 与 Monitor、子 Agent、Channels 和自动化在 Managed 路径上都将使用这些记录。本 PR 只定义契约,尚无组件提交这些记录;H0c 会随 Session authority 提交它们,并据此重建任务列表。契约包含六部分:
OperationGrant让已受理操作的 owner 在没有模型 activation 的情况下完成计划中的 phase。它写明 Session、操作、domain、修订号、owner、Workspace generation、租约和一个范围:保存计划的已提交 domain 记录(其 kind 必须与 domain 对应)以及它准入的 phase。它没有 activation 字段,因此永远不能冒充 activation。Runtime 门禁只能用延长租约的续租,或 Workspace generation 不回退的更晚修订来替换一个 grant。reserved…recovery_blocked)、其物理执行(intent…outcome_unknown、corrupt)以及结果的交付或接收,各有允许的单步迁移。交付有两个目标:Channel(sending、partial、delivered)或 Session(accepting、accepted、consumed)。三条线相互关联:未知或损坏的执行使运行保持recovery_blocked,运行只有在执行被证明结束时才能结束,Session 只有在运行结束后才能接收结果。outcome_unknown、execution_corrupt、runtime_lost、dispatch_unknown、handler_unavailable)和六个配额原因,每个只允许出现在它能解释的状态中。运行中或等待中的运行可以携带恢复原因,降级运行就是这样记录的。executionCallId或effectId、dispatchId和deliveryId分别标识物理调用、跨 Session 派发和外部交付。monitor_run。 记录正文覆盖参考设计第 10 节:启动调用的参数、限额、去抖窗口、启动回执、观测与通知水位、停止原因和输出 manifest。该 domain 加入封闭的 v1 索引(存储设计把它列在 33 个 v1 名称中),但仍不开放提交。一份共享的 JSON schema 和 fixture 文件固定整个契约,共 520 个用例,每个无效用例都针对一条规则。一个依据设计文档编写、独立于两种语言且放在仓库之外的 Python 实现为每个用例标注了结论,二者不一致时生成器停止。core 中的 TypeScript 模块和
managed-agent-server(读取这些记录以构建任务投影的控制面)中的 Java 校验器回放每个用例,检查每条线上的每一对状态,并固定常量。双语设计文档记录了各项决策、规则和开放问题。为什么需要
在 Managed 路径上,Runtime 可能被回收,Harness 也可能被替换。这些能力的 Legacy 实现依靠内存注册表、PID 和本地文件恢复,这些都保不住。#12827 让每项异步能力都成为"持久资源加触发意图",具备稳定身份和回执。如果每项能力各自定义状态、授权、原因和身份规则,Java、qwen 与 Runtime 对"unknown"和"settled"的理解就会不同,也无法用统一的任务投影把它们合并。与 attestation、Tool v2、
managed-context/1和managed-tool-result/1契约一样,先定下共用结构和 fixtures,让 H0c 与 H1–H6 面向一个固定且经过评审的目标来构建。评审测试计划
如何验证
monitor_run提交时仍被拒绝。managed-agent-server的契约测试和 Checkstyle。按该模块 README,需先安装qwencode与runtime-broker。证据(前后对比)
不适用:没有用户可见的变化。以下结果基于 rebase 到
main(88d881491b)之后的分支:tsc均无问题。契约测试通过 530 项,记录测试通过 80 项;src/managed-runtime下全部 23 个测试文件通过(1443 项)。managed-agent-server通过 111 项测试,其中 7 项来自新的契约测试;Checkstyle 无问题。return false以及每个相等、逻辑和关系运算符;219 个变异体中 218 个使测试失败。幸存者把相等修订分支返回之后的>改为>=,是等价变异。require和每个提前return false;45 个变异体全部使测试失败。测试平台
环境(可选)
仅单元测试。Linux:Node 22 与 JDK 21。macOS 与 Windows 未运行。
风险与范围
monitor_run加入封闭的 v1 索引,而不是 v2 索引(feat(managed-agent): Stage H extension runtime, starting with H0 shared records and the task contract #12827 的问题 1);watch_failed,对应 Legacy 在启动后的失败(非零退出码、信号、二进制输出、流错误)。domain.committed现在可以解析monitor_run,但提交仍被拒绝;H3 开放该 domain 时,含有这类记录的 Session 必须把本变更之前的 reader 挡在外面,例如提高其minimumReader。设计文档:English · 中文
关联 Issue
Refs #12827(H0b;H0a 与 H0c 随后)
Refs #12380