Skip to content

feat(managed-agent): persist Workspace session tool profiles - #13301

Merged
yiliang114 merged 2 commits into
mainfrom
codex/13271-persist-session-profile
Oct 4, 2026
Merged

yiliang114 merged 2 commits into
mainfrom
codex/13271-persist-session-profile

Conversation

@yiliang114

@yiliang114 yiliang114 commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Persist hosted-workspace-files/1 when a Workspace Session is created through the public API or WebShell. Creation, loading and recovery then use that saved choice. A bound Session with a missing or blank profile is refused before the Harness request.

This is the first implementation slice of #13271. Shell and /2 admission remain deferred.

Why it's needed

The control plane currently derives the file profile each time it attaches to a Session. Saving the choice now gives existing file Sessions a durable identity before later admission changes introduce other profiles. Idempotent creation retries preserve the original choice.

Reviewer Test Plan

How to verify

  1. Create a Workspace Session through each public and WebShell entry point. Confirm its stored profile is files/1, and retrying the same creation after registry changes returns the same Session and profile.
  2. Upgrade a database containing bound and unbound Sessions. Confirm bound rows receive files/1, existing unbound rows remain without a profile, and existing approval modes stay unchanged. An older writer must still be able to insert a bound Session after migration.
  3. Capture Harness creation, conflict recovery and cold recovery requests. Confirm they carry the stored profile. Missing or blank profiles on bound Sessions must prevent create/load calls; unbound Sessions must send no profile even when an old writer stored the SQL default.

Evidence (Before & After)

The Java module's full verification passed: 557 tests: 556 passed, 1 macOS-specific skip, 0 failures, 0 errors, including compilation, Checkstyle and SpotBugs. A negative control restored the old hardcoded profile and produced the expected assertion failure; restoring the implementation passed the full suite. The MariaDB CI follow-up corrected an old-schema fixture; all 52 non-Hosted integration cases now pass locally, with Checkstyle passing.

Browser verification passed against the real WebShell + Spring/H2 fixture: create a bound Session, reload the same Session and inspect desktop/mobile layouts. SQL confirms the stored files/1 profile after creation and reload. Test report; screenshot attachment pending. The fixture supplies test authentication and Workspace data with Harness execution disabled; this does not verify live Harness execution.

Live native verification now passed at c073f003bf16af29a85924c3275846365918f664: real Spring/H2/Flyway, Java Store/connector, packaged CLI Harness, Broker and worker completed initial and recovered read/write/read turns. Graceful Harness close plus Java attachment eviction/load preserved files/1; NULL, empty and whitespace profiles were refused before Harness requests. A mixed-classpath old-connector control accepted all six invalid-profile cases, confirming the new guard is discriminating. Native report records the 36 candidate and 36 control assertions, SQL/file effects and cleanup. It explicitly preserves a failed full before-classpath capture; only the previously frozen 1,928-file subset is compared byte-for-byte, including all 1,370 dist files. This was a new native API run, not a new browser or tmux run. Production/UI code is unchanged from the earlier browser revision; only the MariaDB test fixture differs.

Combined browser and live-runtime verification now passed at c073f003bf16af29a85924c3275846365918f664 using the committed Managed Workspace component host under Vite, with real Java/Store/Harness/Broker/worker execution. UI creation, reload/re-selection and graceful attachment recovery completed three read/write/read Turns and nine SETTLED executions with stored files/1. The run passed 36 evidence checks; its own complete before/after manifests match across 3,649 recorded files. The first intentionally held response crossed the fixture relay’s 90-second timeout and became visible only after reload; the normal second and third answers appeared live. Original screenshots are preserved, and the readable final capture changes only the fixture host body background; public attachment remains pending. This is candidate-only component-host evidence, with no production standalone-bundle, physical restart, vendor, Shell or /2 claim. Combined UI/runtime report.

Fresh local verification also passed 36/36 checks with the executed command and zero exit statuses preserved by tmux capture-pane. Real Java/Store/Harness/Broker/worker execution completed initial and recovered file turns, retained files/1, and refused all six invalid-profile create/load cases before Harness access. The tmux test report includes captured terminal text, independently checked file effects and matching before/after artifact manifests. The model is a deterministic loopback fixture; the external-model attempt was blocked by automatic approval and was not executed.

Tested on

OS Status
🍏 macOS ✅ Java module, browser and native Hosted verification
🪟 Windows ⚠️ Not tested
🐧 Linux ⚠️ Not tested

Environment (optional)

Java 21, Maven 3.9.11 and H2 in MySQL compatibility mode. The SQL migration was also verified against a disposable MariaDB 10.11 database, including all migrations through V34, the V35 upgrade, legacy rows and old-writer defaults. Native verification used Java 23 and Node 22.22 with an owned Spring/H2 fixture and live CLI/Broker/worker; it did not run the full Hosted Maven integration profile or a physical crash/reboot matrix.

Risk & Scope

  • Main risk or tradeoff: the SQL default supports older writers. Such a writer can also put files/1 on a new unbound row; the connector ignores it without a Workspace binding. All participating control planes must support persistence before admitting other profiles. The legacy unbound-row backfill runs in one UPDATE, so large deployments need to account for its table scan and lock duration during rollout.
  • Not validated / out of scope: MySQL 8 upgrade, full-stack Shell approval/FG6f, Shell opt-in, /2 admission and new approval behavior. The additional CI-aligned Qwen review remains unverified: all 10 dimension reports finished without a Critical finding, but the runner did not produce a final composed verdict. Local correctness and security reviews found no blocking issues.
  • Breaking changes / migration notes: additive nullable column; existing bound rows are pinned to files/1. V35 is provisional and must be checked against main before merge. No public request or response schema changes.

Design: English · 简体中文. Both versions include the same decisions, constraints, acceptance criteria and follow-up work.

Linked Issues

Part of #13271. The issue remains open for Shell admission and its prerequisite gates. The persisted field also supports the later admission work in #13166.

中文说明

本 PR 的改动

通过公开 API 或 WebShell 创建 Workspace Session 时,持久保存 hosted-workspace-files/1。创建、加载和恢复随后都使用该值。绑定会话的 profile 缺失或为空白时,在请求 Harness 前拒绝。

这是 #13271 的第一阶段实现。Shell 和 /2 准入仍留待后续完成。

为什么需要

控制面目前每次附着会话时都会重新推导文件 profile。现在保存该选择,可以在后续引入其他 profile 之前,为已有文件会话确定持久身份。幂等创建重试保留原选择。

评审验证计划

验证方式

  1. 分别通过公开入口和 WebShell 创建 Workspace Session。确认存储的 profile 为 files/1;注册信息变化后重试相同创建,仍返回同一会话和 profile。
  2. 升级含有绑定与未绑定会话的数据库。确认绑定行获得 files/1,已有未绑定行仍没有 profile,已有审批模式保持原值。旧版写入方在迁移后仍应能插入绑定会话。
  3. 捕获 Harness 创建、冲突恢复及冷恢复请求。确认请求携带存储的 profile。绑定会话缺少 profile 或值为空白时,必须阻止 create/load 调用;未绑定会话即使由旧版写入方保存了 SQL 默认值,也不能发送 profile。

前后对照证据

Java 模块完整验证通过:557 项测试:556 通过、1 项因 macOS 平台跳过,0 失败、0 错误,包括编译、Checkstyle 和 SpotBugs。反向对照恢复旧的硬编码 profile 后,出现预期断言失败;恢复实现后,完整测试通过。针对 MariaDB CI 的后续修正调整了旧库测试数据构造;全部 52 项非 Hosted 集成用例已在本地通过,Checkstyle 通过。

浏览器验证已通过真实 WebShell + Spring/H2 测试环境:创建绑定会话、刷新后加载同一会话,并检查桌面和移动端布局。SQL 确认创建及刷新后存储的 profile 都为 files/1。测试报告,截图待附上。测试环境提供测试身份和 Workspace 数据,并关闭 Harness 执行;本次不验证真实 Harness 执行。

c073f003bf16af29a85924c3275846365918f664 的真实原生验证现已通过:Spring/H2/Flyway、Java Store/connector、打包 CLI Harness、Broker 和 worker 完成首次及恢复后的真实读写读。先关闭 Harness 挂接,再清 Java 缓存并加载,仍使用存储的 files/1;NULL、空串及空白 profile 在请求 Harness 前被拒绝。只替换旧 connector 的混合 classpath 对照则放行全部六种非法 profile,确认新守卫能区分错误行为。原生报告记录 36 项候选及 36 项对照断言、SQL/文件效果和清理。完整 before classpath 捕获失败已明确保留;字节对比仅针对此前冻结的 1,928 个文件子集,其中包含全部 1,370 个 dist 文件。本次是新的原生 API 执行,没有重跑浏览器或 tmux;相对既有浏览器版本,生产/UI 代码不变,仅 MariaDB 测试数据构造不同。

c073f003bf16af29a85924c3275846365918f664 的浏览器与真实运行链路联合验证已完成:使用 Vite 提供仓库已有的 Managed Workspace 组件宿主页,后端实际经过 Java/Store/Harness/Broker/worker。UI 创建、刷新后重新选取同一会话、优雅分离后的挂接恢复共完成三次读写读 Turn、九次 SETTLED 执行,始终使用存储的 files/1。36 项证据检查通过,本轮独立的完整前后清单共 3,649 个文件且一致。首轮人为暂停跨过测试转发层的 90 秒超时,答案在刷新后显示;正常第二、第三轮答案实时可见。原图保留,可读的最终截图仅调整组件外的宿主背景,公开附件仍待上传授权。本次仅证明候选版本的组件宿主流程,不宣称生产独立 bundle、物理重启、真实厂商、Shell 或 /2 验收。联合 UI/运行链路报告。

新的本地验证也已通过 36/36 项检查,tmux capture-pane 保留了实际执行命令和全部为零的退出码。真实 Java/Store/Harness/Broker/worker 完成首次及恢复后的文件执行,保留 files/1,并在请求 Harness 前拒绝全部六种非法 profile 创建/加载情况。tmux 测试报告包含终端原文、独立核对的文件效果和一致的前后产物清单。模型为本地确定性测试服务;外部模型测试被自动审批阻止,未执行。

测试系统

系统 状态
🍏 macOS ✅ Java 模块、浏览器及原生 Hosted 验证
🪟 Windows ⚠️ 未测试
🐧 Linux ⚠️ 未测试

环境

Java 21、Maven 3.9.11,以及 MySQL 兼容模式的 H2。另在一次性 MariaDB 10.11 数据库中验证了 SQL 迁移,包括 V34 及之前的全部迁移、V35 升级、旧会话及旧版写入默认值。原生验证使用 Java 23、Node 22.22、自有 Spring/H2 环境及真实 CLI/Broker/worker;未运行完整 Hosted Maven 集成配置或物理崩溃/重启矩阵。

风险与范围

  • 主要风险或取舍:SQL 默认值兼容旧版写入方;旧版也可能在新的未绑定行写入 files/1,connector 在没有 Workspace 绑定时会忽略该值。准入其他 profile 前,所有参与的控制面都必须支持持久化。旧未绑定行的回填使用单条 UPDATE,大规模部署需要在发布时考虑扫描与持锁时间。
  • 未验证及范围外:MySQL 8 升级、完整链路 Shell 审批/FG6f、Shell 开关、/2 准入及新增审批行为。额外的 CI 对齐 Qwen 复审仍未验证:10 个维度的报告均已完成且没有 Critical,但工具未生成最终汇总结论。本地正确性与安全审查没有发现阻塞问题。
  • 兼容性及迁移说明:新增可空列,将已有绑定行固定为 files/1。V35 为暂定编号,合入前需对照 main 检查。公开请求和响应 schema 不变。

设计:English · 简体中文。两版的决策、约束、验收标准及后续工作一致。

关联 Issue

属于 #13271 的一部分。该 issue 保持打开,继续跟进 Shell 准入及其前置门槛。持久化字段也供 #13166 后续准入工作使用。

@yiliang114

yiliang114 commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator Author

Test report — 8da0715

Verified commit 8da0715edabc: 556 passed, 1 platform skip, 0 failures, 0 errors across 68 Java test suites. Compilation, Checkstyle and SpotBugs passed. The MariaDB migration smoke check also passed.

Scenario Observed result
Public API and WebShell Session creation HTTP admission tests confirm that bound Sessions persist files/1; idempotent retries preserve the Session and profile. New unbound Sessions store NULL.
Harness create, conflict fallback and restart recovery Captured requests carry the stored profile. The alternate /2 value is a mocked connector fixture that distinguishes persistence from the former constant; it does not enable /2 admission.
Missing bound profile NULL, empty and whitespace values are refused before Harness create/load calls.
Unbound Session with an old-writer default The captured load request omits toolProfile, even when the stored column contains files/1.
Legacy migration and old writers H2/Flyway and a disposable MariaDB 10.11 database both verified the V34 → V35 upgrade: existing bound rows receive files/1, existing unbound rows receive NULL, old-writer inserts retain the SQL default, and approval modes remain unchanged. The MariaDB test database was removed afterward.
Negative control Restoring the old hardcoded profile caused the expected regression failure: a stored files/2 value became files/1 on the request. Restoring the implementation passed the full suite.
Reproduction and result details

Local environment: macOS, Java 21, Maven 3.9.11, H2 in MySQL compatibility mode; separate MariaDB 10.11 migration smoke check.

From the repository root, with the sibling Java artifacts installed:

mvn -f packages/sdk-java/managed-agent-server/pom.xml verify
Full module:      557 tests, 0 failures, 0 errors, 1 skipped
Admission:        22 tests, 0 failures, 0 errors
Connector:        15 tests, 0 failures, 0 errors
Migration:         1 test,  0 failures, 0 errors
Negative control:  2 tests, 1 expected failure, 0 errors

The platform skip is RuntimeBrokerDefaultOnTest.defaultCombinationBootsWithTheYmlDefaultsBound, disabled on macOS. The negative-control failure is from the intentionally modified code, which was restored before the passing full-module verification.

Browser UI evidence

Added a real-browser check at the same commit: 1 Playwright test passed; 0 JavaScript page errors; all 11 recorded API responses were 2xx. The PR's ManagedAgentWebShell runs against its real Spring controllers and an isolated H2 database migrated through V35. The existing browser fixture supplies an authenticated test actor and Workspace; Harness execution is disabled. HTTP responses are not mocked.

  1. Create a Session bound to ws-default with services/./api: the server returns HTTP 202 and normalizes the directory to services/api.
  2. Reload the browser and select the saved Session: HTTP 200 loads the same Session and binding.
  3. Check the selected Session at desktop (1280×800) and mobile (390×844) viewport sizes.

The UI does not display tool_profile. A SQL check of the same Session (ded05e11-1d07-4ada-9207-7a7a211ff169) returned hosted-workspace-files/1 both after creation and after browser reload. These screenshots show the real create/load flow; the SQL check is the persistence evidence.

Desktop — create a Workspace Session

Create Workspace Session

Desktop — saved Session after browser reload

Reload saved Workspace Session

Mobile — same saved Session

Saved Workspace Session on mobile

Browser environment: macOS, Java 21, Vite 5.4.21, Playwright 1.61.1 Chromium. The fixture host supplies dark theme background/foreground and viewport sizing through the component's existing props. Product code and responses were unchanged; temporary fixture files and servers were cleaned up. This is after-only evidence, with no server-restart, model, Shell or approval-flow claim.

Not verified by this report: live Harness integration, MySQL 8, Shell approval/FG6f, Shell opt-in or /2 admission. The final model-review verdict is a separate incomplete check already disclosed in the PR description.

CI follow-up

The owner-assignment failure is now resolved on main by #13306. #13308 was closed as superseded after verifying that its remaining hosted-runner changes were unnecessary. At the latest check of head c073f003bf16, all 24 returned checks passed and 29 conditional checks were skipped; no failed or pending check remained.

The MariaDB CI failure came from an upgrade fixture calling the latest Session writer against a V31 schema. The fixture now seeds historical rows directly, then upgrades and verifies the preserved close receipt, files/1 backfill and retirement path. Production code is unchanged. Local verification: all 52 non-Hosted MariaDB integration cases passed (36 unaffected cases from the full run plus 16 retention cases rerun after the fixture correction), with Checkstyle passing. The original failure was reproduced locally before the fix. Disposable databases and credential files were removed.

The test-only follow-up is c073f003bf16; the production implementation and browser-tested UI remain unchanged.

Follow-up evidence at c073f003bf16: live Hosted execution and recovery and combined browser/runtime verification. Those reports describe their own deterministic-model and fixture boundaries; the screenshots above remain the earlier create/reload run with Harness execution disabled.

@yiliang114 yiliang114 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed c073f003bf16af29a85924c3275846365918f664 against main 5ddfacc9d4c18d6e85faeed18786ada134774541. No blocking finding in the bounded correctness, security, quality/performance and test-evidence review.

The review covered V35 and old-writer compatibility, the transactional profile write, idempotent replay, create/conflict-load/cold-recovery paths, and unchanged public/Web Shell admission. Bound sessions save their profile; unbound sessions omit it; a missing or blank bound profile fails before Harness access. V35 is unique on the checked main. The Java workflow for this exact head passed, including the migration uniqueness and daemon gates.

This pass was read-only: no tests or UI were rerun. The existing browser report uses a Spring/H2 fixture with Harness execution disabled, so it does not establish live Harness execution. The stored /2 test value checks persistence rather than new /2 admission. Shell and new profile admission remain outside this PR; it remains Draft.

@yiliang114

Copy link
Copy Markdown
Collaborator Author

Live Hosted files/1 execution and graceful recovery passed at c073f003bf16af29a85924c3275846365918f664, with the real Spring/H2/Flyway Store, Java connector, packaged CLI Harness, EmbeddedRuntimeBroker and local worker. The deterministic model ran on loopback; the observation relay forwarded actual Harness responses.

Check Candidate result
Bound Session and idempotent retry Public creation HTTP202; replay returns the same Session; SQL and actual Harness creation both use files/1.
Real initial tools read_file → write_file → read_file complete; the physical Workspace file matches a random marker absent from user input; three Broker executions are SETTLED.
Graceful recovery Harness DELETE returns 204; clear the Java attachment cache, then the production connector loads the journal with stored files/1 and returns 200.
Execution after recovery A new real read/write/read completes with the expected second file and six total SETTLED executions. Journal writer generation advances 1→2 and committed sequence 38→76. Neither proof file appears in the Harness decoy workspace.
Missing stored profile NULL, empty string and whitespace, each through create and journal-backed load: all six are refused by Java before any Harness request.
Unbound and public controls Unbound Session completes with NULL in SQL, no profile in the Harness request and no model tools. Both metadata.toolProfile overrides for Shell/1 and files/2 return public HTTP400/unsupported_feature.

The candidate passed 36 assertions. A separate control compiled only the connector from exact base 5ddfacc9d4c18d6e85faeed18786ada134774541 and kept the candidate Store, migration, API, SDK, Broker and CLI. All six invalid stored-profile cases then reached the real Harness with the old hardcoded files/1 and returned 200. This is a mixed-classpath discriminating control, not a complete base build; it passed its 36 expected-observation assertions. Each fixture made nine actual model requests.

The negative cases deliberately edit isolated SQL fixture rows. Their HTTP409 is the fixture wrapper around Java's Hosted Workspace Session tool profile is missing exception, not a claim about the public API's error contract. No Shell or /2 profile was admitted.

CLI SHA256: af7a319156308396c54adb749909f63def12a9c5d6b3f1b068afddb9617d5bbf. Source remains clean at c073. The earlier frozen 1,928-file subset, including all 1,370 dist files, matches byte-for-byte after this run. The new full before-classpath capture failed and is retained as failed; the complete after manifest contains 2,408 files and 121 JARs. No complete before-classpath comparison is claimed.

Independent cleanup confirmed all owned native processes and Java worker groups gone, all 12 observed ports closed and private configuration removed. Earlier setup/protocol failures remain separate evidence. This verifies graceful Harness close plus Java cache eviction/load; physical Host crash/restart, Broker reboot, MySQL 8, active Shell approval and real-model reliability remain outside this run.

The earlier browser report still covers creation/reload and desktop/mobile layouts with Harness execution disabled. Between that source revision 8da0715 and c073, only the MariaDB test fixture changed; production/UI code did not. This new report supplies live execution evidence separately and does not claim a new browser or tmux run. Public UI screenshot attachment remains pending.

@yiliang114

Copy link
Copy Markdown
Collaborator Author

Real browser + live Hosted files/1 verification passed within the scope below at exact head c073f003bf16af29a85924c3275846365918f664. This adds the combined UI/runtime evidence missing from the earlier native API report; no product source change was needed.

The browser opened the committed Managed Workspace fixture host, rendering the actual ManagedAgentWebShell with Workspace binding enabled under Vite. Its real Java provider called production WebShell routes through a loopback trusted-ingress/observation fixture. Spring/H2/Flyway, the Java Store/connector, packaged CLI Hosted Harness, Broker and local worker all executed normally. Browser automation did not substitute API or SSE responses. Only the model was deterministic and local.

Actual browser action Runtime and visible result
Create a bound Session and Send Public create returns 202; SQL and actual Harness creation use stored files/1. Real read/write/read produces the first Workspace proof and three SETTLED executions.
Reload, reselect the same Session, then Send The persisted first answer returns. A second real read/write/read uses the previous proof; its answer appears live and the UI reaches Completed. Same Session and Harness boot, writer generation 1.
Graceful private detach, then Send from the UI Actual Harness DELETE returns 204. After fixture-only Java attachment-cache eviction, the next public UI turn itself triggers real load, HTTP200, carrying stored files/1. A third read/write/read completes and displays its answer live; writer generation advances to 2.

Final SQL contains three COMPLETED Turns and nine SETTLED file executions, with journal committed sequence 132. All 12 model calls offered only read/write/edit. Random proof bytes were absent from the typed prompts; actual prior-file results reached the model, and all three physical proof files are correct inside the bound Workspace and absent from the Harness decoy. There was no process restart.

The first model response was intentionally held for an in-progress capture and crossed the fixture relay's 90-second SSE timeout. That live sample showed Completed without the assistant answer; SQL contained the full answer and browser reload restored it. This timing-affected sample is not a clean live-streaming pass or a proven product defect. The normal second and recovered third answers appeared live. Intermediate file-tool cards were not displayed and are not claimed.

Nine actual browser screenshots are retained locally. The original eight use the minimal host's white body with the component's transparent dark theme. The readable final capture changes only the fixture host body background to #101010; visible text was unchanged and original captures remain intact. I inspected the final capture: only isolated Session identity, test prompts and random proof values are visible, with no credentials or user content. Public screenshot attachment remains pending the previously blocked upload authorization; no screenshot-publication completion is claimed.

Both this run's own before/after manifests contain 3,649 identical files: all 1,370 dist files, explicit Java classpath files/JARs, frontend/SDK sources, fixture scripts and Node/Java executable bytes. Manifest SHA256: 587309d5dde264402f91bade626b23613d7f526e582a467099b0a12bd4b318d2; CLI SHA256: af7a319156308396c54adb749909f63def12a9c5d6b3f1b068afddb9617d5bbf. The candidate remained clean. The node_modules tree was not exhaustively hashed; this new comparison does not retroactively repair the older native run's missing full before capture.

All 36 evidence checks passed. Independent cleanup confirmed the Java, Harness, Vite, observed worker and runner processes are gone, all nine recorded loopback ports are closed, private fixture configuration is removed, and the browser task space is closed. The worker listener port was not separately sampled; its process/group absence was checked.

This is candidate-only component-host evidence, not browser Before/After, production standalone-bundle, external-vendor, MySQL 8, physical crash/restart, Shell, or /2 admission verification. The earlier mixed-classpath connector control remains API-only. Raw local traffic and process logs contain fixture credentials and were not published.

@yiliang114

Copy link
Copy Markdown
Collaborator Author

Fresh local execution at c073f003bf16af29a85924c3275846365918f664 passed 36/36 checks, with actual tmux capture-pane -p evidence. This run uses real Spring/H2/Flyway, the Java Store/connector, packaged CLI Hosted Harness, Broker and local worker. The model is a deterministic loopback fixture. tmux hosts the terminal commands; this is not an interactive CLI TUI or a new browser run.

Check Observed result
Create and retry a bound Session Public admission returns 202; SQL saves files/1; idempotent replay returns the same Session.
Initial file execution Actual read → write → read completes; the Workspace proof has the exact generated marker, with three SETTLED Broker executions and no decoy write.
Graceful attachment recovery Harness DELETE returns 204; fixture-only Java cache eviction followed by production connector load returns 200 and sends stored files/1.
Recovered execution Another read → write → read completes; the second physical proof matches. Six executions are SETTLED, writer generation advances 1→2, journal sequence 38→76, and SQL still contains files/1.
Missing profile guard NULL, empty and whitespace, each on create and journal-backed load: all six stop before any Harness request. The fixture control wrapper returns 409; this does not establish a public API error contract.
Unbound and admission controls The unbound Session stores NULL and omits the Harness profile. Public Shell/1 and files/2 overrides return 400/unsupported_feature.

All 269 main Java sources were freshly compiled from the candidate into isolated class directories; existing target class directories are absent from the runtime classpath. The successful before/after manifests contain 2,101 identical files, including all 1,370 dist files and the recorded Java classpath/fixture inputs. This is the recorded subset, not an exhaustive node_modules or host filesystem hash. Source remains clean. No product source change was needed.

Actual tmux capture excerpts

These are selected verbatim lines from the completed rendered terminal capture; complete step captures and the unabridged final capture are retained locally.

bash-3.2$ bash '/private/tmp/qwen-code-13271/tmp/pr13301-closeout-tmux-20261004-120228/run-local-in-tmux.sh'; TASK_EXIT=$?; printf 'TMUX_COMMAND_EXIT=%s\n' "$TASK_EXIT"
COMPILE qwencode: 110 source files, exit=0
COMPILE runtime-broker: 65 source files, exit=0
COMPILE managed-agent-server: 94 source files, exit=0
COMPILE isolated fixture: exit=0
BUILD_EXIT=0
PASS bound public admission202
PASS first turn completes
{"kind":"candidate","phase":"INITIAL_REAL_TOOL_TURN_COMPLETE","sessionId":"0c877ab8-20e5-43e8-80c1-7d9044e8346c"}
PASS real first file effect
PASS no decoy write
PASS snapshot HTTP200 afterFirst
PASS files1 persisted
PASS real broker executions settled
PASS actual journal exists
PASS idempotent replay same session
PASS graceful Harness detach preserves journal
PASS production connector load accepted
PASS load wire stored files1
PASS message admitted PR13301_RECOVERED
PASS recovered turn completes
PASS real recovered file effect
PASS snapshot HTTP200 afterRecovery
PASS profile remains files1
PASS six settled actual tool executions
{"kind":"candidate","phase":"REAL_WRITE_READ_AND_ATTACHMENT_RECOVERY_PASS","sessionId":"0c877ab8-20e5-43e8-80c1-7d9044e8346c"}
DETERMINISTIC_EXIT=0
IDENTITY_EXIT=0
RUN_COMPLETE deterministic=0
TMUX_COMMAND_EXIT=0

Local evidence: tmp/pr13301-closeout-tmux-20261004-120228/ in the candidate checkout; the primary transcript is tmux-readable-full.log, with separate command/compile and completed-chain captures. The two proof files were independently re-read and verified against the report. Cleanup independently confirmed the recorded Java/Harness processes gone, all six sampled ports refusing connections, temporary private configuration removed and the owned tmux Session removed.

Two setup failures were retained separately: the scratch compile initially selected Commons Lang 3.17.0 instead of the SDK's declared 3.20.0, then a preparation script had an indentation error. Both stopped before product tests; only the test setup was corrected.

The authenticated external-model attempt was rejected by automatic approval before execution, so no vendor-call result is claimed. This report covers real local product/file execution with nine deterministic model calls. Authentication/Workspace registration come from the isolated fixture; physical restart, MySQL 8, production deployment, Shell and /2 admission remain outside this run.

CLI SHA256: af7a319156308396c54adb749909f63def12a9c5d6b3f1b068afddb9617d5bbf. Recorded-manifest SHA256: d74b56573b72f1b83021f882640e06620ca9efda30c67fe387092c44cba31adb. Final capture SHA256: 8c8d93499cbf6eb9a9b79050ace1eb292586bafa15e53711ada9a2b8298753b7.

@yiliang114
yiliang114 marked this pull request as ready for review October 4, 2026 04:14

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

APPROVE at c073f003.

Critical-only scan over the production diff (the migration, the store, the record, the connector) plus the call chains that read them. No blocking issue found, and no prior blocking finding exists on this PR: there are no inline review threads, and the triage pass recorded no critical blocker.

The new IllegalStateException is not reachable, and the persisted value preserves today's behaviour

QwenHostedHarnessConnector.toolProfile (:377-385) replaced a hardcoded "hosted-workspace-files/1" with a DB-sourced value and added a throw when a Workspace session has a null/blank profile. I tried to find a path that reaches the throw and could not:

  • Existing rows — V35__managed_session_tool_profile.sql adds the column with DEFAULT 'hosted-workspace-files/1', so every pre-existing row is backfilled with exactly the string the code used to hardcode, then UPDATE ... SET tool_profile = NULL WHERE workspace_id IS NULL nulls precisely the rows for which the first guard already returns null. Behaviour is identical on both sides of the migration.
  • New rows — insertSession binds workspace == null ? null : "hosted-workspace-files/1", so tool_profile is non-null exactly when workspace is non-null, which is the only branch that can reach the throw.
  • Any writer that omits the column — the column DEFAULT supplies the value, so an omitted bind cannot produce null on a Workspace row.
  • The reduced constructor — SessionRecord's secondary constructor delegates workspace as null and toolProfile as null together, so a record built that way returns at the workspace() == null guard before the throw.

The invariant is therefore enforced, not merely asserted. This is the right direction for a value that will later vary per session.

Read-path column coverage and record arity both verified by reading, not inferred

The row mapper now calls result.getString("tool_profile"). All five queries that feed sessionMapper (ManagedAgentStore.java:1012, :1022, :1045, :1849, :2123) are SELECT * FROM managed_agent_session, so the column resolves on every path; the only projection query over that table (findMaterializationTargets, :1273) maps to MaterializationTarget and is unaffected.

Adding a component to a record is a compile-time hazard for every canonical-arity caller, so I enumerated them rather than trusting the diff. In src/main the only construction is the row mapper itself, which this PR updates. In src/test, five files construct SessionRecord and this PR updates four; the fifth, QwenHostedHarnessNewSessionRegressionTest.java, is untouched — I read it at this head and it calls the 13-argument secondary constructor (:49-50), whose parameter list this PR does not change (only its delegation gains one null). It still compiles. insertSession likewise stays balanced at 16 placeholders for 16 bound arguments.

CI gap, disclosed rather than relied on

At this head only housekeeping checks ran (assign, authorize, label, route, triage succeeded; 91 skipped; review-pr queued). No Java build or test job executed, so compilation and the migration test are not CI-verified here — which is why the arity and column-coverage checks above were done by reading source at this head instead of being inferred from a green build. Nothing in CI indicates a defect introduced by this PR; this is a coverage gap, not a failure.

Non-blocking: the migration version will need renumbering at merge time

main currently tops out at V34__managed_tool_output_collection.sql, so V35 is free there. However ten open PRs each add a different V35__*.sql to this same directory (#13210, #13217, #13247, #13289, #13260, #13301, #13325, #13336, #13354, #13355). Whichever lands second will collide and Flyway rejects duplicate versions at startup, so this file needs renumbering against whatever main holds when it merges. This is a merge-ordering action item common to all ten PRs, not a defect in this diff, and it does not gate approval.

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

APPROVE at c073f003. I found nothing blocking, and the part of this change that could quietly have been wrong is both handled and pinned by a test. No thread had been filed before this review, so this is an independent read rather than a confirmation.

The mixed-version window is the load-bearing part, and it holds

The migration adds the column with a SQL default and then clears it for unbound rows:

ALTER TABLE managed_agent_session
    ADD COLUMN tool_profile VARCHAR(64) DEFAULT 'hosted-workspace-files/1';
UPDATE managed_agent_session SET tool_profile = NULL WHERE workspace_id IS NULL;

The default is what lets an old writer keep inserting bound Sessions after the migration, which is the compatibility requirement the test plan names. The consequence worth checking is the awkward one: during a rolling upgrade an old writer inserting an unbound Session also gets that default, so the column can hold hosted-workspace-files/1 on a row that must not send a profile. toolProfile() is immune to that because it branches on the binding, not on the column:

if (session.workspace() == null) return null;
if (session.toolProfile() == null || session.toolProfile().isBlank()) throw new IllegalStateException(...);
return session.toolProfile();

And ManagedSessionToolProfileMigrationTest pins exactly that row — late-unbound is asserted to hold hosted-workspace-files/1 in the database while the connector still sends nothing. So the invariant is structural rather than incidental, and a future refactor that starts trusting the column instead of workspace() reddens a test instead of silently sending a profile for an unbound Session.

V35 is also the correct next version: main currently ends at V34__managed_tool_output_collection.sql, and Flyway migration version uniqueness is green at this head.

Two details that are right for non-obvious reasons

  • Hoisting toolProfile(session) to the first statement of load() is not cosmetic. In the previous form managedSessionStore(session) was assigned before toolProfile(session) was evaluated as an argument, so a refusal threw after the store connection had been obtained. Refusing first means the guard costs nothing when it fires.
  • The new writer states the profile explicitly instead of leaning on the column default — insertSession binds workspace == null ? null : "hosted-workspace-files/1" and lists tool_profile in the INSERT. That is what makes the default purely a compatibility device for old writers rather than part of the new path's semantics, and it means the IllegalStateException is unreachable through this code.

On that throw: I checked whether it is reachable at all, since it would surface as a 500 through ApiExceptionHandler's Exception.class mapping. It is not — the migration backfills every pre-existing bound row, and the new writer always sets the column for bound Sessions, so only an explicit NULL write could produce it. Keeping it as an invariant assertion is right; a 500 is the correct signal for a state the schema and the writer both promise cannot occur.

Adding toolProfile as a record component is compiler-enforced across every construction site, so the arity change needs no audit — the green Java lanes are the evidence.

Verification basis and one forward-looking note

All nine Java/DB lanes are green at this head: Java 11/17/21 on ubuntu, Java 21 on macos and windows, Real daemon E2E (Java 11), Runtime Broker and Managed Agent MariaDB (Java 21), Hosted process fault gates (MySQL 8.4, Java 21), and Flyway migration version uniqueness. CI is 22 pass / 0 fail / 35 skipped. Usual disclosure: I did not run the Java suite locally, so this rests on reading the code plus those lanes; the body's own numbers (557 tests, 556 passed, 1 macOS skip, plus a negative control that restored the hardcoded profile and produced the expected failure) are the author's and I have not reproduced them.

One thing to keep in view rather than fix here: the profile identifier now exists in two languages — once in insertSession and once as the migration's column default — and that agreement is protected only by the migration test. With /2 admission deferred to a later slice of #13271, a second profile value will land in both places, and the SQL side cannot be refactored to a constant once shipped. Worth deciding where the Java-side constant lives before that slice rather than after.

Vote effect

reviewRequests lists doudouOUC, LaZzyMan, qqqys, tanzhenxin and wenshao, and reviewDecision is REVIEW_REQUIRED. Those requests are not CODEOWNERS-derived: every changed file is under docs/design/ or packages/sdk-java/, and no CODEOWNERS rule covers either (the rules are /.github/CODEOWNERS, three release/security workflows, /packages/core/, /packages/cua-driver/ and /packages/mobile-mcp/), so they were requested explicitly. @qwen-code-ci-bot approved at 04:33 on this exact head and required_approving_review_count is 1, but an outstanding review request pins the decision regardless of how many approvals exist — so one of those five needs to submit, or the requests need clearing, before this merges.

@yiliang114
yiliang114 added this pull request to the merge queue Oct 4, 2026
Merged via the queue into main with commit 3229191 Oct 4, 2026
171 of 172 checks passed
wenshao pushed a commit that referenced this pull request Oct 4, 2026
…n added

Merging main brought #13301, which adds a toolProfile component to
SessionRecord and turns a null profile on a bound Session into an invalid
state: QwenHostedHarnessConnector throws "Hosted Workspace Session tool
profile is missing". The bound fixtures in the two retry-terminal suites
still used the 16-argument shape, so the module stopped compiling after
the merge.

The value is inert for these tests - both suites mock HarnessConnector, and
toolProfile() is read only inside the real connector - but a bound fixture
without a profile now describes a state production rejects, so it follows
main's own HarnessCoordinatorTest fixtures.
wenshao pushed a commit that referenced this pull request Oct 4, 2026
Main's #13301 persists Workspace session tool profiles and claimed
migration V35; this branch's four migrations move to V36-V39
(activation columns, event-type index, snapshot deferral marker,
journal sequence index) and SessionRecord carries both new fields
(approvalMode, toolProfile).
wenshao added a commit that referenced this pull request Oct 4, 2026
…the commit-time guarantee

#13301 landed V35__managed_session_tool_profile.sql on main, so this
branch's V35__managed_task_event.sql now shares its version: the merged
tree fails at startup with "Found more than one migration with version
35". Move the outbox migration to V36 and update the authority design
doc in both languages.

The commit-time validator mirrors the checks the authority applies to
every line whatever its kind (envelope, vocabularies, domain checks,
reserved ids, byte caps), not the per-kind payload schemas, subject
rules or commit-marker digests, which stay the authority's contract.
Reword the apply() Javadoc, the test comment and the design doc so they
no longer promise that no commit can brick the Session, and say what a
stored reserved id breaks: it collides with the id of that domain's next
Stage H record rather than failing the next open.

Pin the event id and operation id shape checks with two negative cases;
disabling either check previously left every test green.
wenshao added a commit that referenced this pull request Oct 4, 2026
main's #13301 claimed V35 for the session tool-profile migration while
this branch was in review; the two flyways collided on the classpath
(Found more than one migration with version 35).
yiliang114 pushed a commit to everyoneexe/qwen-code that referenced this pull request Oct 6, 2026
…LM#13300 (QwenLM#13355)

* docs(managed-agent): design H3 background Shell and Monitor runtime

* feat(core): add managed child_run (kind shell) record body for H3

* feat(managed-agent): mirror child_run shell record body in Java store

* test(core): commit and rebuild child_run shell chains through the session authority

* feat(managed-agent): close child_run reference closure and add stop-request draining

* feat(core): add named cgroup unit attach and managed child-run supervisor

* feat(cli): admit v3 background shell under the supervised worker registry

* feat(core): add open-ended shell stream capture with manifest revisions

* fix(core): type-narrow stream capture finalize and refused publish double

* fix(cli): close background shell type gaps after the main merge

* fix(managed-agent): answer the R1 review round on the child_run contract

* fix(managed-agent): close H0c critical follow-ups R3-1/R3-2/R3-3

PR 1 of #13300, fixing the three Critical findings from the round-3
review of #12855.

R3-1: the Broker-record execution mapping now reads the record's
dispatchGeneration: SETTLED/cancelled with generation 0 (the record's
own proof it was never claimed, per ToolExecutionRecord) maps to
not_started_proven instead of settled, removing the illegal
intent -> settled shape that left a Monitor cancelled before dispatch
with no committable settling revision. A shared brokerExecutionCases
row (settled-cancelled-unclaimed) replays it in both languages, the
TS divergence pin gains the second wire-reading divergence, and
Decision 10 of the authority design is reconciled with the H0b
record-contract doc in both language pairs.

R3-2: the store now runs the event-level half of the envelope checks
for every event line at commit time: closed envelope with an optional
subject, version, declared sequence, well-formed event id outside the
reserved <domain>:<n> namespace, this Session's closed key, valid time,
and kind from a mirrored EVENT_KINDS vocabulary; every domain.committed
payload is validated whether or not a body is registered (reusing the
pinned DOMAINS); unknown-subtype lines follow the scanner's "after the
Managed header" condition; a transaction carries at most one Stage H
record; and per-kind byte caps (maxEventBytes, maxCommitMarkerBytes)
are pinned in the shared limits contract and enforced per line kind.
All refusals answer the existing 409 so the commit rolls back.
Replay commits return before any of this runs, as before.

R3-3: task.updated rides its own outbox (managed_agent_task_event,
V35) written in the same commit transaction and keyed by a unique
source key, drained later by the task-events slice; managed_agent_event
stays turn and lifecycle events, so an announcement between two text
deltas can no longer split a message Part under the frozen projection
version. The route description is scoped to what the server does
(a Session being deleted announces nothing while its tasks stay
readable), and the design's replay-idempotency sentence is corrected.

Test helpers that wrote journals a real authority could not read back
(ActionJournal's reserved event ids, bare event/marker lines in the
publication and session-store integration suites) now write
well-formed lines. Mutations of every new check were verified red
against the suites.

* fix(managed-agent): repair the CI-only cgroup root and env-guard failures

How: wrap the delegated-root probes so a missing or unreadable root answers
the documented isolation error on Linux too (macOS refused at the platform
check first, which is why only CI saw the raw ENOENT), and document the
background Shell environment allowlist in the process.env guard.

Why: the Test lane on 4fb9f0a9c8 went red on exactly these two items; both
are this branch's own changes, not flake.

Test: hook-command-cgroup and process-env-guard suites green; tsc clean.

* feat(managed-agent): inject the child-run supervisor at worker boot

How: registerManagedContextRoutes builds a ManagedChildRunSupervisor from
the delegated cgroup root the Hook commands already use
(QWEN_MANAGED_HOOK_CGROUP_ROOT) and passes it as the executor's fifth
argument; a boot without the delegation keeps the executor's committed
refusal instead of failing to boot.

Why: executeV3Background landed in the previous increment with the
supervisor parameter unwired, so every real-stack background start answered
the committed no-supervisor refusal.

Test: new wiring case asserts the supervisor is injected exactly when the
root is set; managed-context-worker, process-env-guard and
managed-background-shell suites green (775); tsc clean.

* docs(managed-agent): pin the two-row ledger shape for background processes

How: the runtime-ownership section now says the ledger holds two rows —
the foreground-length start invocation settling with the handle, and a
second execution admitted at start that carries no model result and stays
non-terminal until physical exit evidence.

Why: the Java ledger writes a result exactly when an execution settles, so
a single row that is both handle-delivered and active cannot exist; the
two-row shape keeps the settled-if-and-only-if-result invariant while
preserving every cited behavior (Runtime holds, evidence-only settle).

* feat(managed-agent): commit child_run shell lines from the hosted side

How: HostedChildRunSession funnels every record write through one
serialized, replay-safe commitExtensionRecord path — admit first with the
start call's args as commandRef, dispatch and the set-once
managed-runtime-receipt after the physical start, outputRef advance-only,
and settlement only on proven exit evidence, a proven failure, or an
honored stop request. Live revisions alternate the run line (no
self-loops); the authority freeze refuses anything after a terminal
revision.

Why: the dual path puts product records on the hosted authority, but
nothing commits child_run from the hosted side yet — the worker-owned
registry intentionally does not touch records.

Test: four authority-level cases — full chain with projection, stop
request to cancelled, pre-start failure frozen on not_started_proven, and
serialization with deep-equal skips.

* feat(managed-agent): add the detached capture family to Tool v3 results

How: a background Shell start now settles success with a capture object
whose status is 'detached' — no reason, no manifest — because the live
output streams through the child_run record's growing manifest rather
than the result; the shared schema pins both invariants, the TS parser
mirrors them, the shared fixtures gain one valid and two invalid envelope
cases replayed on both sides, and the Java projector accepts the missing
manifest for unavailable or detached captures. Admission refusals settle
not_started with a null capture, the shape the durable unstarted family
already owns.

Why: every Session tool.receipt event lands in the Java delivery
projection, which requires a capture object and a committed-if-manifest
pairing for anything that started — a success result with a null capture
fails that projection and corrupts replay.

Test: core contract suite 519 (three new envelope cases), TS serve suites
33 + turn/harness 323 + background 6, Java publication contract and
projector suites 8; tsc clean on both packages.

* fix(managed-agent): answer the R2 review round on worker, capture and record lines

How — the four Criticals first: the stream capture publishes an ended
stream's revision only after its seal decision, so a sealed-or-incomplete
descriptor never changes again (R2-3); finalize's settle path degrades to
an unavailable capture instead of throwing past a writability failure, and
the last published revision stands like the worker-cut cap (R2-5); the
child_run body now refuses a settled or cancelled run that is not settled
execution, a failed run off its two ending lines, start_failed outside
not_started_proven, and a process-level failure without a settled
execution, on both languages with new witnesses replayed from the shared
fixtures (R2-6, R1-30); and the supervision suite skips win32 instead of
spawning a shell that cannot exist (R2-7).

How — the Suggestions land as: supervisor start now proves membership by
cgroup.procs with an fd-3 status channel, failing closed as isolation
(R2-37); the bounded EOF wait caps daemon-inherited pipes instead of
hanging, and a spawn-time error drains the entry (R2-1/R2-14); registry
hasHolds requires its session and setProcessResult carries the full result
shape (R2-13/R2-15); create() validates its caller-named unit like attach()
(R2-26); the executor checks the journal before any effect, mirrors the
closing/is_active re-check after prepare, gates the ninth live Shell as a
committed quota refusal, and keeps cgroup membership out of an empty
QWEN_MANAGED_HOOK_CGROUP_ROOT (R2-23/R2-19/R2-24/R2-18); the spawn uses
the configured shell and the session-context environment (R2-21/R2-20);
invalid fixture cases each pin their refusing clause on both languages
(R2-36); projection fixtures pin the draining precedence rows and the Java
replay reads stopRequested (R2-28/R2-38); the draining Javadoc states the
rule the code implements (R2-39); the duplicated ref helper is gone (R2-40);
the stream-capture suite is typed and gains the seal-before-publish,
page-cursor, late-failure and settle-degrade witnesses (R2-29/30/31/32);
and the design doc names the cgroup switch decision, the cursor's single
ordinal space, the unconditional host-scope evidence gate and the zh
stopped-object (R2-8/9/10/11/12).

Test: core suites 788, cli serve suites 776 + background 11, Java record/
projection/store suites 24; tsc clean.

* feat(managed-agent): thread the child-run orchestrator and the detached receipt family

How: HostedSession gains its session-scoped childRuns orchestrator next to
hooks/mcp, the tool turn receives it stored-ahead of the admission branch,
the reopen verifier and the recovery replay both accept the third durable
receipt family — a blocked delivery with a detached capture, beside the
complete and not-started ones — with the publication store correctly left
out of its proof, since its durable truth is the child_run record.

Why: the H3 background start settles with a handle whose output lives on
the record's growing manifest; without the third family, any Session
holding a background receipt could never reopen or be replayed.

Test: recovery-session suite 52 with two new detached-family cases;
harness-session and tool-turn suites green; tsc and lint clean.

* test(cli): type the background-shell capture double

How: drop the as-unknown cast on the FakeSink return and implement the
failCapture member the cast had been hiding (R2-29's cli location).

Why: an unchecked double can drift from the interface it claims to match —
and it already had.

Test: background-shell suite 10/10; tsc clean.

* feat(managed-agent): admit background Shell starts from the hosted tool turn

How: the turn replaces its two deliberate is_background refusals — exactly
when the Session owns its child_run orchestrator and the domain is
enabled — with orchestration that mirrors the shared line: the record
intent precedes every physical effect, dispatch_started lands the moment
the checkpoint commits, a settled detached handle attaches the physical
start to the record from the same facts and lands as the third durable
receipt family (blocked delivery with a null manifest, replay-validated
exactly like the not-started family), and a proven-unstarted refusal
settles start_failed on not_started_proven while riding the existing
unstarted family unchanged. Malformed is_background values keep their old
refusal, and the family stays out while the domain is disabled — the same
probe governs both paths.

Why: the background start is the first record-bearing, user-visible H3
effect, and without a hosted history family for the handle a Session that
ran one could never reopen or be replayed.

Test: three new authority-level turn cases — admitted detached flow with
admit/dispatch/attach order and the blocked receipt family, a
proven-unstarted refuse closing NOT_STARTED, and the disabled-domain
refusal preserving its exact text; turn suite 146, plus harness-session,
recovery, background and child-run suites 246; tsc clean.

* feat(managed-agent): admit background Tool v3 dispatches and acknowledge their detached settle

How: the v3 start admission flips from "foreground Shell only" to
"Shell only" — background dispatches now begin the same way; the poll
loop accepts the settled detached family through the status answer (the
only durable settle such a result can ever produce, since nothing is
published for a background handle), and the acknowledgement path
canonizes the settled handle envelope: a matching blocked receipt with a
null manifest and no history revision forwards exactly itself, without
consulting the publication receipt.

Why: without the detached branch, a background start settled a queued
handle only to mark the execution UNKNOWN after the 30-minute
foreground-shaped deadline — every background dispatch would wedge the
Runtime on its first call, and its acknowledgement would die on the
missing publication it was never going to have.

Test: new witness drives the detached start end-to-end — reservation,
v3 dispatch, settled handle via status, and ack with a null-manifest
blocked receipt, with the publication receipt path provably untouched;
service suite 139/139, module suite 600/600.

* feat(managed-agent): answer shell-status and shell-terminate from the worker registry

How: a new private maintenance route sibling to ManagedHookProtocol
(/internal/managed-runtime/v3/shells) backed by the process registry —
status answers running or unknown (never claimed without evidence),
terminate drains with the supervisor's own rules and answers exited with
its proof, and an end that cannot be proven is unknown with the hold
still registered. The unit name derives from the runtime invocation
identity through the shared shellUnitNameOf, so Java, the hosted turn
and the worker derive it identically; the executor takes the registry as
a constructor injection so the route and the journal share one.

Why: every later maintenance verb — recovery reconciliation, the ordered
close drain, the reconcile-settle — has to ask the only physical owner
these questions; answering them from anything but the registry would
invent truth.

Test: five worker-side cases — closed-key refusals, unknown for
unowned, running scoped to the registered Session, terminate-with-evidence
answering exited, and an unproven terminate answering unknown with the
hold intact; shell, context-worker, v3-route and child-run suites 784;
tsc clean.

* feat(managed-agent): admit the background process row and answer its maintenance protocol

How: the start dispatching of a background Shell admits a second ledger
execution beside the invocation row — the new process row carries no
model result and stays PREPARED (non-terminal) until physical proof, so
hasActiveBy* holds the Runtime exactly while the process lives. The
maintenance protocol mirrors ManagedHookProtocol on its own route:
shell-status and shell-terminate with a targetOperationId (both recovery
kinds), wire-validated on both transports and accepted through the same
workspace-generation and ownership gates. observeBackgroundProcess asks
the physical owner and settles the process row with proven evidence —
settle through the new repository verb settlePrepared, which mirrors the
PREPARED-requestCancel shape on both repositories; anything unproven
leaves the row active with its hold, never a claimed end.

Why: the ledger writes a result exactly when an execution settles, so a
single background execution carrying its handle at start cannot exist;
two rows keep the settled-if-and-only-if-result invariant while the
Runtime holds precisely, and the maintenance route gives every later
verb the worker's physical answer instead of an invented one.

Test: two new witnesses — a background start that admits the process row
(the invocation SETTLED, the process PREPARED, the Runtime busy until
status arrives, then exit evidence settling it and release succeeding),
and an unproven status keeping its hold; the runtime also gained the row
identity invariant fix and the settle-through-settlePrepared shape;
service suite 141/141, module 603/603.

* fix(managed-agent): answer maintenance views under the target identity

How: the shell-status/shell-terminate view now carries the target
operation's identity, exactly as the Hook lookup family does — the
requester asks about the process, never about the asking call, and the
Java wire validator's expectedId rule requires it.

Why: this repair was encoded in the route design but landed uncommitted
after ②e-a; the self-audit caught it before it could ride a later batch.

* feat(managed-agent): run the background exit leg through the record

How: the hosted publisher gains the background family — its registration
marker admits the open-ended capture through the session-level admission
(proven start on the record, never the model call), the open manifest
publishes at prepare, every awaited write or finish forwards the newest
manifest revision to the record's outputRef, and the exit finalizes with
the same evidence object settling the record. The client acknowledged
receipt keeps the detached-blocked shape; drain skips background stores
until the ordered close drain wires them; the register guard only binds
the foreground model-call family.

Why: the start handle already committed, so a background Shell's whole
physical truth is output and exit evidence — anything that reached the
record another way would let a second tool.receipt contradict its own
line, and an open-ended manifest published anywhere but the Session's
own store is an unreadable closure for the store-side commit.

Test: a full end-to-end authority case — admit-dispatch-attach, bytes
through the real rpc, outputRef advancing live, the final sealed manifest
reading the same exit (code 3, both streams sealed, manifest complete),
the record settled with exactly that evidence, and zero tool.receipt
events for the outcome; publisher, turn and child-run suites 174; tsc
clean.

* feat(managed-agent): close monitor_run revisions over their cited Session resources

The H0c closure rule — every ref a record cites must exist in the
Session's resource store at commit time — now covers monitor_run on both
stores: the authority reads commandRef, startReceiptRef, outputRef and
lastObservationRef through verifyExtensionResources, and the Java
applyRevision gains the same four-field branch beside child_run. This is
the precondition for enabling the monitor domains: without it a monitor
revision could anchor its chain to resources nobody committed, and the
replay would project a watch whose evidence was never verified.

Both fixture rigs stop citing fictional resource identities. The
authority test publishes its chain refs as real bodies once per harness
and mirrors each fixture-cited ref as a real same-kind placeholder of
the stated length, memoized so chain identity stays fixed across
revisions. The journal test harness does the same at its generic
request entry point, rewriting each cited ref to the placeholder's real
digest and appending the mirrored resources, but only for strictly
well-formed JSON monitor bodies — a body with trailing content keeps the
raw path so the store's grammar refusal still fires first over the
closure's.

Each side gains a witness that pins the closure itself: a monitor run
citing a resource outside its commit is refused with
managed_session_resource_missing, and each witness was observed red with
its branch disabled before landing.

* feat(managed-agent): gate monitor_run behind the managed-session/2 reader

The H3 design left the reader-gating mechanism to the implementation;
the journal settles it. Headers are immutable — the scanner refuses a
repeated managed_session_header_v1 line and any unknown subtype after
the header — so no transaction can legally raise a Session's
requirement in-journal, and a superseding header line would change the
journal format in both languages for a narrow mixed-version window.
New Sessions therefore stamp minimumReader managed-session/2 at
creation, and every path that commits a monitor_run record —
commitDomainRecord, commitExtensionRecord and the generic append guard
— additionally refuses when the Session's header names less. Readers
still accept requirements up to their own version, so pre-H3 Sessions
stay openable everywhere and simply can never receive monitor_run.

child_run is deliberately not header-gated: a managed-session/1 reader
already parses its name, and the asymmetric hazard lives on the server
side, which the established deploy-server-first sequencing covers.

The shared Stage H golden is regenerated for the v2 genesis bytes and
the Java integration replay passes against it unchanged. New witnesses:
new Sessions stamp v2, a v1 Session refuses monitor_run with the
precise error (observed red with the gate removed), and both header
sides of the version comparison are pinned in the records tests. The
design doc records the decision and the now-paid closure precondition
in both languages.

* feat(managed-agent): funnel monitor_run revisions through a hosted orchestrator

The mirror of HostedChildRunSession for the Monitor domain. The
managed-runtime worker will own the watch loop, but the dual path puts
every product record on the hosted authority, so HostedMonitorSession
commits the record line as watch facts arrive: revision 1 with the
intent before any side effect, dispatch onto a Runtime binding, the
set-once start receipt after the watch starts, observation revisions
only while attached with a watermark that never goes back, output
advancement only forward, and the terminal shapes — settled by
exited/max_events/idle_timeout, failed by a start failure on
not_started_proven or a settled watch failure, cancelled by a stop
request whose notification watermark may still advance after.

Writes serialize per Session and replay-safe by command id with
deep-equal skip, exactly like the Shell funnel. The suite drives the
whole line through a real authority: projection at admission and after
settle, max_events refused before its quota and accepted at it, a
second start receipt refused, an observation before attach refused,
and an idempotent output advance committing nothing. Quota
(enforcement), the debounce-floor observation loop with notification
composition, and rebuild-after-loss stay with the runner and route
increments that follow.

* feat(managed-agent): drive monitor observations through a debounced loop

The loop half of the Monitor runner, on top of the hosted funnel: an
admitted Monitor dispatches, starts an injected watch executor and runs
its own time. Stdout lines aggregate into one observation revision per
debounce window with a one-second floor, so the document's reopen bound
holds for any debounce a watch asks for; the run settles itself on the
contract's own terminal conditions — the observation quota, silence
past idle_timeout, a natural exit with its buffered lines flushed
first — or is stopped, which terminates the watch and lands cancelled.

The executor contract is typed for the worker's cgroup owner (onLine,
onExit after the start resolves, a terminate handle); the suite drives
the whole line through a real authority with a fake executor and a
manual clock: windowed aggregation and its floor, quota settlement and
watch termination exactly at max_events, idle settlement re-armed by
each accepted observation, exit flush order, start and mid-run failure
shapes, and a stop that ignores late lines. Every cross-boundary
settle — timer or executor callback — is counted by the loop's done
handle, which the close drain will await; window flushes count too, so
done never reports idle while an observation commit is in flight.
Notification composition and its wake stay with the next increment;
quota enforcement and rebuild-after-loss stay with the route slice.

* feat(managed-agent): notify from monitor observations in the same transaction

The H0c machinery now runs for Monitors: every accepted observation
commits its revision together with a notification input and the wake
the authority generates for it, and the run's notifiedThrough watermark
advances to exactly that observation's sequence. The loop composes the
input — monitor-1:notify:N ids, the window's joined lines as the
content a woken turn will read, an empty admission, no deadline — and
the funnel threads it through commitExtensionRecord, so a recovery
replay re-runs nothing the watermark already covers.

Dedupe semantics are pinned at both levels: the funnel advances the
watermark only when a notification rides the revision (an observation
without one keeps it where it was), and the loop suite shows the three
events — domain.committed, input.accepted, wake.requested carrying the
input's accepted id as its source — landing in one transaction with the
joined lines as its content. Consumption of that wake by the hosted
scheduler, and settling the input when no turn can take it, come next
(open question 6 of H0c).

* fix(managed-agent): keep a finished background Shell's receipt answerable

H7 of the real-stack rounds: the Broker settles its :process row only
when shell-status answers exited, but the worker's registry deleted an
entry the moment the Shell ended — a natural exit could never be
answered again, and the ledger row would hold the Runtime against
release and close forever. The registry now retains each completed
unit's receipt until the worker ends: the hold drops exactly as before,
but shell-status and shell-terminate answer a proven end from the
retained evidence, idempotently, for the owning Session scope only. An
end without evidence keeps nothing and answers unknown as before; the
existing live flows are untouched. Suite: natural exit keeps answering
exited with its evidence for both kinds after the hold dropped, the
wrong Session still hears unknown, and an evidence-less end stays
unknown. The close-drain/reconcile callers that consume this, and the
Java-side settle-before-busy on workspace release, land with the next
batch.

* fix(managed-agent): renumber the task-event outbox to V36 and narrow the commit-time guarantee

#13301 landed V35__managed_session_tool_profile.sql on main, so this
branch's V35__managed_task_event.sql now shares its version: the merged
tree fails at startup with "Found more than one migration with version
35". Move the outbox migration to V36 and update the authority design
doc in both languages.

The commit-time validator mirrors the checks the authority applies to
every line whatever its kind (envelope, vocabularies, domain checks,
reserved ids, byte caps), not the per-kind payload schemas, subject
rules or commit-marker digests, which stay the authority's contract.
Reword the apply() Javadoc, the test comment and the design doc so they
no longer promise that no commit can brick the Session, and say what a
stored reserved id breaks: it collides with the id of that domain's next
Stage H record rather than failing the next open.

Pin the event id and operation id shape checks with two negative cases;
disabling either check previously left every test green.

* fix(managed-agent): settle provable background exits before release's busy check

The second half of H7. The worker now keeps a finished Shell's evidence
answerable, so the Broker has someone to ask — but observeBackgroundProcess
still had no production caller, so a naturally exited Shell left its
:process row non-terminal and every release wedged on
runtime_session_busy. releaseSession now sweeps the background process
rows this Broker admitted for the Session before it computes busy: each
unproven row asks its physical owner once through the same
shell-status → settlePrepared leg observeBackgroundProcess already owns,
a proven exit settles with its evidence, and any lookup that fails or
cannot prove an end simply keeps the row so busy stays the accurate
answer. The sweep tracks per-context admissions because the
reconciliation scan deliberately excludes PREPARED rows, and mutates the
admission set under the context lock.

Witnesses: releaseSettlesAnExitedBackgroundProcessBeforeBusy proves an
exited row settles with its evidence and the release succeeds (observed
red with the sweep bypassed); the pre-existing
unprovenProcessStatusKeepsItsHold still proves an answer without proof
keeps the row and the busy refusal. Full module 604/604. The workspace
close drain still terminates shells in order rather than sweeping — its
ordered stop belongs to the close-drain slice, and these two verbs are
exactly what it will consume.

* fix(managed-agent): check the journaled receipt before attaching a background Shell's start

H6 of the real-stack rounds: acceptBackgroundShell called childRuns.attach
before its tool.receipt replay check. A retried accept — the leg recovering
pending results re-drives — would mint a second start receipt on the same
identity before reaching the replay path, and the record's set-once rule
would refuse it as a rerun shape, so a recoverable retry crashed as a
conflict. The lookup now runs first: a journaled receipt replays its
validated history with no attach at all, and a fresh accept attaches only
when this identity still lacks a start receipt — the record id IS the
execution identity, so an existing one is always our own crash-window
attach between record and journal, which the retry now tolerates instead
of minting a rerun. Witnesses: a retried accept answers from the journal
with exactly one attach and one tool.receipt line, and an
already-attached-but-unjournaled record completes its accept without
attaching again. The full hosted turn suite stays 148/148.

* fix(managed-agent): backpressure the background Shell output pipes

G5 of the real-stack rounds: the executor fed the bounded capture with a
fire-and-forget sink.write — every chunk queued behind asynchronous
publication, so a fast producer flooded memory (a 256 MiB run peaked near
315 MiB RSS at a 20 MiB/s store). The feed now counts bytes in flight and
pauses stdout and stderr when the drain falls sixteen MiB behind, resuming
at four MiB — enough to cover a store several seconds slower than the
process produces without stalling a steady producer against the floor.
Completion or rejection of a write both release their bytes, so a broken
capture can never leave the pipes paused. The witness drives a controlled
sink: seventeen MiB of chunks pauses a stream exactly once, draining past
the watermark but above it resumes nothing, and draining the rest resumes
exactly once. The outputRef-live-advance half of G5 belongs to the
turn-publisher's promotion to activation scope, sized with the close-drain
slice, and stays queued there.

* feat(managed-agent): promote the Shell publisher to Session scope

The ②d-out debt, and with it the live-advance half of G5. The Shell
publisher was a per-turn object: its server died with the turn, while
background Shell traffic flows across turns — between turns (or after a
close) the worker's retained descriptor pointed at a dead endpoint, so
no live output advance and no exit leg could ever complete past one
turn. The hosted Session now owns one long-lived publisher across
turns (the turn lazily structures it on the shell options and never
closes it; the Session's ordered close in this PR's fifth slice drains
it). The two small wire-level fixes that the promotion exposes:

- The worker's background execute stamps the background marker into its
  prepare request: the closed six-field wire capture has no such field,
  so a remote prepare used to arrive unmarked and die at admission; the
  executor derives the marker from the background input itself.
- register() takes the registering turn's prompt for the foreground
  admission check, so one instance admits the foreground of any turn of
  its Session while a foreground pretending another turn's identity is
  still refused; the background arm keeps its record-based admission.

Witnesses: a single instance admits a background watch plus the
foreground of two turns and refuses a misattributed one; the worker rig
pins the stamped marker on its prepare; the four touched suites stay
184/184 and both package typechecks pass (after a core dist rebuild for
the capture type's optional background marker). The publisher's own
descriptor installs idempotently per turn (identical descriptors are
accepted), so consecutive turns never fight for the worker slot.

* feat(managed-agent): drain a Session's background Shells ahead of release

The close-drain's missing legs. Until now a Session with a live
background Shell could never close: the worker's activation gate refused
release with managed_activation_conflict forever, and the CLI-side close
route skipped the background stores. The ordered sequence the design
asks for now exists end to end:

- The registry gains stopSession: terminate one Session's entries with
  the supervisor's evidence rules (TERM, escalate, empty-check), await
  each completion so its finalization — last manifest, sealed envelope,
  record settle through the session publisher — lands before the caller
  proceeds; an unproven stop keeps its hold for the caller to see.
- The worker's activation release gate drains before it refuses: active
  background work no longer wedges a close that can prove its stops,
  and anything still unproven keeps the exact same conflict.
- The hosted close route now closes the Session-scoped publisher right
  after the broker release returns — which by then has drained the
  Shells and settled their records through that publisher — and the
  drain closes every capture store, background families included.

Witnesses: a registry drain settles one Session's shells with evidence,
both kinds, while leaving another's holds untouched, and an unprovable
terminate under a session drain still preserves its hold; the 766-test
context-worker suite stays green through the release gate's new async
drain branch. What remains deliberately unproven locally: the
full-chain release-with-live-shell drain needs a real cgroup supervisor,
pinned as the Linux head acceptance instead of a mocked supervisor.

* docs(managed-agent): sync the H3 design's status line with landed slices

The status line still claimed publication and orchestration remained
design while most of them landed across this PR's batches; refresh both
language versions with the landed list and keep the remaining-design
list accurate: the Monitor physical side (route, rebuild, cgroup watch)
and wake consumer, both submission enablements, task events, and the
Linux physical acceptance.

* feat(managed-agent): spawn Monitor watches through a cgroup watcher

The physical half of the Monitor runner's executor interface, for the
worker: ManagedMonitorWatcher spawns the watch command under a unit
named for the execution identity, through the same supervisor (its
membership attestation unchanged), splits stdout into the observation
lines the loop buffers — boundaries respected, the Legacy partial-line
cap honored, the tail flushed as the final line at exit — and reports
the physical end exactly once, natural exit versus mid-run failure, so
the loop's settled and watch_failed shapes both hold. stderr rides only
the future output leg, never observations; terminate defers to the
supervisor's TERM-escalate-empty rules. The suite drives a supervisor
double: unit/executable/cwd derivation, chunk-boundary splits with the
final flush, the cap drop, spawn-error and membership-failure shapes,
and terminate delegation.

* fix(managed-agent): accept message.retracted at commit time

#13351 added the message.retracted event kind on main. The hosted text
delta stream writes it when a restarted model attempt replaces a prefix
it already published. The commit-time vocabulary this branch mirrors from
the authority still listed sixteen kinds, so after the merge with main
the store refused every retraction with 409 ("event.kind must be one
of [...]") and left the orphaned deltas in place.

Add the kind to ManagedExtensionRecords.EVENT_KINDS and to the shared
store fixture, so the parity pins in both languages agree again.

* feat(managed-agent): answer monitor status and stop from a worker registry

The monitor half of the maintenance route, mirroring the Shell's: a
worker-side registry of supervised watches that holds each Session's
Runtime while a watch runs, answers an end only with the supervisor's
own evidence, and retains a proven exit's receipt until the worker ends
so maintainers can learn the truth idempotently. The route sibling at
/internal/managed-runtime/v3/monitors serves monitor-status and
monitor-stop with the same closed wire shape, target-identity rule, and
unknown-but-never-unproven discipline as the shells' route; it joins
the worker's route list so the v2 envelope admits it. Registration
arrives with the Monitor start admission later — status/stop already
have consumers in the release drain and the rebuild slice that follows.
This also repairs the watcher test arity that broke the previous push's
build lanes (the mock's 2-arg onExit omission, seen as
error TS2554 (155,34) across every CI lane at that head).

* feat(managed-agent): admit monitor watches into the open-ended capture session

The γ3-α precondition for every later Monitor leg:
LocalShellStreamResultSession's prove-a-live-start admission reads the
capture family's record domain instead of being pinned to child_run —
same Session/binding/background-marker checks, same live-run gate,
same identity captureId, per-domain labels so a refusal names its own
family. The producer side names its domain explicitly (record presence
is never consulted where both domains could in principle compete), so
a monitor watch carries its monitor_run record into exactly the same
bounded stream the background Shell already rides. Witnesses: shell
admission unchanged; monitor admission succeeds under its domain;
wrong-domain proven start refuses; missing record and foreign binding
refuse with their own errors.

* feat(managed-agent): admit Monitor watches through the v3 executor

The workhorse of the Monitor leg. The worker's v3 executor gains its
is_monitor branch beside is_background: journaled admission first, then
refuses (never as a transport error), then the watch spawns through the
ManagedMonitorWatcher under a unit named for the execution identity,
streams stdout into the SAME bounded family (16/4 MiB backpressure on
the pipe, with an explicit refusal if the watcher reports no supervised
unit), and settles with its detached handle while a registry keeps the
Runtime hold. Three watch-level facts are done physically now:

- The monitor registry is a wrapper around the background Shell
  registry, never a copy: supervisors, EOF drains and publisher finishes
  stay in one place, while monitor holds and the four-watch quota bucket
  remain entirely their own. Its maintenance route answers run/exited/
  unknown with identical semantics to shells'.
- The execute path gates on the generated four-live-watch cap with
  `Session already runs 4 Monitor watches.` and additionally on a
  missing cgroup root and a missing capture service.
- Each prepare carries background+monitoring markers, so the hosted
  side of γ3-α knows the monitor_run domain to check.

Suite: 28 tests across the three touched files — the fork settles a
detached start under the right unit and marker, lines land as monitor
output, the fifth watch is refused, and supervisor-less boot refuses
with its dedicated text. The turn admission arm (funnel-driven
admit→dispatch→attach through HostedMonitorSession), the watcher→loop
observation feed, and rebuild still land in the following increments.

* fix(managed-agent): refuse a domain.committed without a textual domain

A domain.committed line whose payload.domain was missing or not a string
read as "no domain" and skipped requireDomainCommitted, so it committed
200 and the authority refused the Session at the next open. Decide that
a line is domain.committed from its kind alone, and take the domain from
the vocabulary check before building the expected record kind.

Also address the rest of the first /review round:
- The commit marker is the only line apply() measures. validateUtf8JsonLines
  already caps every line at MAX_EVENT_BYTES, which it now names.
- The outbox's MAX(sequence_id) stays a plain read. A locking read
  deadlocks concurrent first announcements of different Sessions on MySQL
  8.4 (85 of 480 in a probe). The comments now name the guarantee that
  actually holds: the Session store's journal-head lock is taken before
  the transaction's first plain read.
- Stop claiming that the Session event stream carries only turn and
  lifecycle events. V36 is not applied anywhere yet, so its comment can
  still change.
- Restore the shipped v1.19 OpenAPI sentence and record the move to the
  outbox as v1.30 (info.version 1.30.0).
- Reconcile both design docs: Decision 10's state table and the scope of
  its wire-only absolute, the goals, the public-contract bullet, the
  validation plan, the files list, and the restored caveat that the record
  contract does not catch an unprovable not_started_proven.
- Tests: refuse a non-textual and a missing domain, commit a
  message.retracted line, pin the final "events, then its commit marker"
  guard again, assert the empty outbox in the MySQL IT, and give the
  oversized-resource case bytes at its own sequence.

* feat(managed-agent): admit Monitor watches through the hosted tool turn

The turn-side arm of the Monitor leg. A call named monitor is now
admitted through its own three-state gate (requested/unavailable/ill-
formed, with the Shell-profile Monitor declaration now advertised),
implements the same record-first discipline as background Shells do —
the funnel admit commits monitor_run revision 1 before any side effect,
dispatchStarted follows the durable checkpoint, and the watch attaches
when its v3 execution settles. The granted publication, journal intent,
dispatch and accept forks are widened to carry is_monitor alongside
is_background, and acceptMonitor mirrors acceptBackgroundShell end to
end (not_started → start_failed with the unstarted family; detached →
same blocked receipt family with exactly one tool.receipt line, the H6
journal-verdict-before-attach order, and the blocked ack completing the
row). HostedSession gains a session-scoped HostedMonitorSession beside
childRuns; turn construction forwards it at both sites.

Two existing contract pins move deliberately: the Shell profile's
advertised list now names monitor (the harness declaration test names
it explicitly), and the three publisher-close witnesses now assert the
Session-scoped lifecycle settled by the earlier promotion — no close
at any turn ending (completed/error), the server still answering, close
exactly once at the Session's own delete. Suite: 3 new admission-arm
tests through a full hosted turn (funnel order, clamped defaults with
the Legacy monitor caps, start_failed shape, unavailable text), the
whole turn suite stays 151/151 and the harness session suite 180/180.
The observation fan-out (worker lines → hosted loop), quota witness at
the turn, and rebuild-after-runtime_lost land in the next increments.

* feat(managed-agent): fan a monitor watch's observations into the hosted loop

The observation channel of the Monitor leg. The session publisher now
marks monitor captures with recordDomain monitor_run through γ3-α, fans
each published stdout chunk through the Legacy line-split semantics —
boundaries kept, remainder held across chunks with the 4096-byte
partial cap — into a per-capture observer, and closes it with
onExit(failed) at finalization with the tail flushed first. The turn
gains resumeMonitorWatch: a fresh accept creates one HostedMonitorLoop
per Session report and resumes it from the already-attached record, fed
through a HostedMonitorRemoteExecutor whose start binds those observer
callbacks; a replay starts nothing twice, exactly like it never
re-attaches. Globally, loop timers unref so observation windows never
anchor the process.

Monitor capture finalize settles through the monitor funnel itself
(watch_failed when no evidence, exited via settleQuiet), beside the
shell's unchanged path. The publisher carries monitors next to
childRuns, and the turn's own construction forwards it. Suite: the
admission-arm tests now also pin one Session-level loop starting on a
fresh accept with the record-first flow intact; wide suites of 189
tests pass with cli tsc, prettier and eslint clean. Monitor output
before the observer registers stays durable in the open stream, with
observation starting at attach — noted in the design not as a flaw but
as the attach seam, a line crossing it already lands in the output
Artifact readers see.

* feat(managed-agent): rebuild a read-only Monitor after its Runtime is lost

The record side of the rebuild promise. HostMonitorSession gains the two
verbs the H3 lifecycle needs around runtime loss: blockedRuntimeLost
lands the run in recovery_blocked with the runtime_lost reason and the
lost binding still named, so degrading is what the projection shows
until something decides otherwise; rebuildFromRuntimeLost accepts only
the outcome_unknown/runtime_lost line the H0b rule allows, starts a
fresh watch under a newer Runtime generation with its own receipt, and
keeps the reason for honest projection while observation positions
never moves — the support case never restarts what the watermark holds.
monitorRebuildAllowed routes every candidate command through the Legacy
AST read-only check with its walk, so a side-effecting command stays
blocked accurately. Five witnesses cover the full rebuild chain
(recovery-block, rebuild with receipt remint and generation increase, a
rebuild refused from a live run and from an ended one, and the gate's
read-only/side-effect/empty answers). Binding-replacement detection and
re-drive of the new physical watch come with the recovery slice after
the wake consumer.

* feat(managed-agent): serve the task events routes with their SQL journal

H3 flips listSessionTaskEvents and queryWebShellTaskEvents from planned
to partial on both surfaces.

The Java control plane gains a bounded per-task event journal (V39): one
row per committed task event, written in the Session commit transaction,
with the per-task sequence allocated under the journal head lock in
commit order. Record revisions journal their task view changes at the
H0c announcement point. The durable retention floor lives in a per-task
cursor row, so it survives an empty retained set, restarts and
projection rebuilds; the Artifact visibility barrier pins it behind
unarchived output, and the backlog bound refuses past capacity instead
of silently discarding. Output events carry their per-stream segment
ordinals with overlap/gap refusals, and artifact references stop
fail-loud at the 100 bound. output_cursor/outputCursor lift with the
routes and the task views serve them.

The section 6.1 demonstrations run as store-level contract traffic in
the new journal contract test, the route probes and record gates cover
the flips, and the #12847 C15/C16 instances close alongside. WebShell
types regenerated from the bumped 1.30.0 contract.

* docs(managed-agent): add the task events lane's build report

Flyway numbering findings (V39 taken; V35-V38 claimed by in-flight PRs
on main, V15/V29 Java-only burns untouchable), implementation decisions
against the H3 design text, the CLI/daemon flow-surface finding (none
exists, regeneration is the only surface change), the full test ledger
and the owed items.

* feat(managed-agent): build the monitor wake envelope and the pending-input intake

Two support pieces for the H3 wake consumer (β2-b1a):

- monitorNotificationText wraps one due observation window in the exact
  Legacy task-notification envelope — task-id, optional tool-use-id,
  kind/status/event-count, summary and result — applying the same
  truncateNotificationLabel/stripDisplayControlChars/escapeXml pipeline
  the Legacy Monitor applies, so a turn learns nothing new.
- pendingSessionInputs derives the Session's still-pending inputs from
  the journal alone: an accepted input is consumed exactly when a turn
  settles under its turnId, order-insensitively, so a replayed admission
  after a restart is not re-driven.

* feat(managed-agent): deliver the Legacy notification envelope on monitor wakes

The loop's notification input now publishes the exact task-notification
envelope a Legacy Monitor's wake carried — task-id, tool-use-id, kind,
status, event-count, summary and the window's lines — so the turn the
wake raises reads nothing it has not already read. The description is
the tool call's, falling back to the watched command. A witness pins
the exact wire text of the first observation's input content.

* feat(managed-agent): run the monitor wake through an embedded scheduler

H3 closes the wake loop: an accepted observation already commits its
notification input and wake.requested in one transaction; this slice
makes that wake effective.

The pump re-derives the pending monitor inputs from the journal alone —
nothing is held in memory, so a restart re-derives the set. On an idle
Session the oldest notification runs as an ordinary text turn with the
Legacy envelope as its prompt, and the turn's settle consumes the input.
A busy Session queues in the journal exactly like channel and Goal
inputs; the retry loop redelivers. A Session that is parked, blocked or
mid-recovery keeps its reminders pending and reports them. On the close
path every pending notification settles cancelled without a model turn,
under the turn-result record's own idempotency key, so no wedged
notification ever holds a Session as hosted_turn_recovery_required at
its next open: the parked-turn scans now skip monitor inputs, which the
pump owns exclusively.

A wake turn that dies without settling re-drives on reopen — the input
is consumed exactly when a turn settles under its turnId, and a crashed
turn never settles — matching the queue semantics channel and Goal
inputs already honor.

* fix(managed-agent): keep every new Session on the managed-session/1 reader

Round-5 real-stack verification's cross-version matrix disproved the
v2 stamp's premise: monitor_run has been a parsed, known domain since
#12837 (v0.24.7), and a v1 reader opens a Session holding monitor_run
records without complaint. Stamping minimumReader: managed-session/2 on
every new Session — while monitor_run remains disabled — would make a
rollback or a mixed-version rollout lose access to every Session
created in between, and buys protection no deployed binary needs.

Revert the stamp to managed-session/1 and drop the per-domain header
refusal; the enablement list remains the single domain gate. The stamp
rises only when a change genuinely breaks an older reader mid-scan, and
only one release after a tolerant reader ships. The golden fixture's
genesis header regenerates to v1; the records and authority suites gain
v1-header admittance witnesses in place of the refusal cases; the
design's reader section now records the reversal with its evidence.

* fix(managed-agent): keep monitor output byte-true and advance only forward

Round-5 real-stack verification found three facts about the monitor
output path:

- A capture rebuilt from decoded lines cannot be byte-true: per-chunk
  toString split multi-byte runes into U+FFFD, and line-rounding dropped
  blank lines and the final unterminated line. The watcher now forwards
  raw stdout chunks to the capture on a new onChunk channel, so the
  durable Artifact reproduces the command's stdout exactly, while the
  observation lines keep the Legacy emit-path semantics (a StringDecoder
  holds a split rune until it completes; a blank line consumes no
  observation).
- The watcher registered its exit listener only after the supervisor's
  start resolved, so a fast watch could end unheard and the debounced
  loop never settled. A latched end with a check-phase sweep covers a
  watch that exited before its listeners attached, after the host is
  known to be listening.
- The funnels' advanceOutput promised "only ever to a newer revision"
  in prose and accepted an older manifest in code. Both hosted funnels
  now read the revision of both manifests and refuse a regression
  before anything commits; witnesses pin the refusal in place.

* feat(managed-agent): drain a busy background Shell at release, and answer the asked operation

Round-5 close-path findings (H8): with a running background Shell the
Broker refused busy before transport.release — yet that HTTP call is the
only route to the worker's drain — so a `tail -f` held session release
and workspace close forever, with nothing able to stop it.

The release path now distinguishes busy kinds: anything active that is
not one of the Session's own background process rows keeps the exact
busy refusal with zero worker hops, while proven-running rows trigger a
dedicated release call first whose 409 busy answer is swallowed (the
worker's route drains before refusing, and its maintenance routes are
not activation-gated, so the follow-up sweep still proves what the
drain ended). Rows settle from the drain's evidence and the complete
release converges on what is genuinely still running — one call when
the drain finishes in time, one retry when it does not. Two witnesses
pin the contract: a running Shell reaches the worker (the round-5 probe
found the refusal never did) and a non-background busy refuses without
a worker hop.

Also from the round: the shell maintenance validator computed the
expected operationId from targetOperationId but never compared it with
the answer's own operationId, and the sweep witness answered with the
process-row id where the real route echoes the call id. The validator
now refuses a view answering another operation, a new protocol suite
pins it, and the witnesses answer the shape the route really sends.

* fix(managed-agent): re-assert the output pause behind every flushed chunk

Round-5 G5 finding: Node's flushStdio resumes both paused pipes when
the background launcher exits, and the executor tracked the pause in
its own flag rather than the stream's, so a writer that outlives its
launcher pushed 230 MiB into the queue and climbed RSS by +200 MiB.

The onOutput accounting now re-asserts the pause behind every delivered
chunk, and applyPause is idempotent through the stream's isPaused state
instead of the local flag — the stop-and-go the backpressure intended
holds through the launcher exit, on both the Shell and the Monitor
paths. The existing 17-MiB pipe pacing witness keeps its single pause
call.

* fix(managed-agent): mirror message.retracted from the H3 event vocabulary in the commit-side kind mirror

* fix(managed-agent): validate every domain.committed payload, with or without a textual domain

* test(integration): expect the monitor tool on the hosted shell profile

The hosted shell profile now advertises `monitor` beside
`run_shell_command`, so the fake-model driver's tool-name assertion —
which the Hosted workspace tool-turn IT checks on every model call —
must expect it. The file profile's expectation is unchanged: the
declaration is shell-profile only.

Fixes the single red IT in the MySQL fault-gates lane
(HostedWorkspaceToolTurnIT.packagedHarnessUsesSavedWorkspacesThroughReal
BrokerWorkerAndSqlStore), whose only difference was the extra
'monitor' entry.

* fix(managed-agent): read the whole journal when deriving pending monitor wakes

`readEvents()` is a bounded read: with no arguments it returns the first
`defaultReadEvents` (100) events, capped at 256. Both consumers of the
pending-input derivation called it bare, so on any Session whose log runs
past that page — every real Session with tool calls — the derivation saw
only the log's head:

- the wake scheduler's `next()` never found a notification input, so a
  monitor wake was never delivered as a turn once the Session passed the
  page. The feature was dead in production while every rig stayed green,
  because the rigs commit a handful of events;
- `settlePendingMonitorInputs` never found the owed inputs at close, so
  they stayed unsettled and the Session reopened as
  `hosted_turn_recovery_required` — exactly the wedge this close-path
  settle exists to prevent.

Both now read the full committed prefix with
`eventsInSequenceRange(1, committedSequence)`, the discipline every other
journal scan in the hosted session already uses. The witness commits 110
observations so the notification lands beyond the default page, asserts
the log really exceeds it, and was verified red under the inverse edit
(the settle returns 0 and the input stays owed).

* fix(managed-agent): stop a full task journal from wedging record commits

The task event journal is a derived, bounded feed: its backlog bound
refuses the next event past capacity rather than silently discarding it,
and an unarchived output event pins the retention floor so the automatic
pass cannot expire anything. The record commit called `appendStateChange`
inline, so once a task's journal was pinned and full, that refusal
propagated out of the extension-record commit — every later revision of
that task failed, which would take the Shell and Monitor record lines
with it.

The record row is the authoritative state and is already written when the
journal append runs, so a refusing journal now degrades the feed (warned
with tenant, session, task, revision and the refusal code) and never the
commit. The store still throws for callers that can apply backpressure —
which is where the output producer's retry belongs.

The witness drives the real wedge: it pins the floor with an unarchived
output event, fills the journal to the bound, asserts the next append
refuses with `managed_task_event_backlog_full`, then commits a record
revision whose view changes and asserts it lands. Verified red under the
inverse edit, where the commit dies with "The task's event journal holds
512 retained events".

* refactor(managed-agent): drop the shell view's never-set error field

The maintenance view declared an optional `error` object, the Java
response validator carried a branch for it, and no producer on either
side ever set one: the worker answers `running`, `exited` or `unknown`,
and a route-level failure travels as an HTTP status with its own code
body. A field that is declared and validated but never written is a dead
switch, so it comes out of the TypeScript view and the Java field set —
which now also refuses an answer that carries it, instead of silently
accepting a shape nothing produces.

* fix(managed-agent): let a replayed output advance pass instead of refusing

The forward-only check refused any advance whose revision was not
strictly newer, which included a redelivery of the very reference the
record already holds. The publisher guards against that in memory, but
the guard does not survive a worker restart, so a replayed advance could
throw inside the background exit leg and wedge a Shell or Monitor that
had already ended.

Both funnels now treat an identical reference as the no-op the deep-equal
skip already owns, and still refuse any other reference that is not
strictly newer. For a Shell the replay commits one liveness step, since
the live run line alternates by design; for a Monitor it commits nothing.
Witnesses pin both, beside the existing refusal witness.

* chore(managed-agent): keep the task-events lane report out of the PR diff

A one-off construction report does not belong at the repository root:
the project's directory rules put working artifacts under the ignored
.qwen tree and keep the tracked docs/ for design and plans. The report
is preserved at .qwen/investigations/2026-10-04-task-events-lane-report.md,
and its Flyway numbering note is superseded by the V40 renumber the
merge commit records.

* merge: absorb review-round state and retarget onto the H3 stack

* fix(managed-agent): quiet the wake pump's consume guard at a blocked owner

runTurn's recovery-blocked path marks the Session blocked and returns
'settled' without the turn ever existing, so the input stays owed — the
accurate answer the whole design gives for a parked Session. The pump's
after-settle check read that as a consumption bug and threw, which both
marked the already-blocked Session again and logged a spurious error on
every such observation. It now stops at an accurate blocked state and
still throws when the input simply did not move: that is the programming
error the guard exists for.

* fix(managed-agent): meld a monitor watch through one terminal step

Round-6 found the shape the Monitor's final leg really had: prepare's
open-then-advance could never survive the start order (output before a
start receipt is a parse-level refusal), the manifest still advanced
through the Shell funnel whatever the capture's recordDomain said (a
Monitor never builds its own Shell record), and the finalize settled the
record ahead of the loop's own exit chain, so a watch's last window died
in denial between two async hops — with its swallow queued to silence it.

Three changes, each with its own witness that goes red under the inverse
edit:

- The manifest advance reads the record's start receipt first and only
  advances past it; a missing record still fails loudly instead of the
  old "no record to revise" at whatever funnel happened to be there.
- It routes by recordDomain: monitor captures land on the Monitor
  funnel, Shell captures on the Shell's.
- The loop's onExit now returns a chain carrying the exit's own
  flush → settle in order, and the publisher awaits it rather than
  settling itself — a final window can no longer be dropped ahead of
  the settle, and a commit failure of the chain surfaces inside the
  finalize instead of returning an empty observation state silently.

The end-to-end witness drives the production shape on a real authority,
orchestrator and Publisher over HTTP: prepare ahead of attach, output
advancing only after the receipt, then one terminal step holding
observationSequence 1 with the settled mark behind it.

* fix(managed-agent): settle the process row when the Runtime proves no start

Round-6 found the permanent wedge: the `:process` row admits PREPARED
ahead of dispatch, but an admitted-refusal answer
(`{executionStatus: "not_started", capture: null}` — cgroup root missing,
quota exceeded, and friends) settled only its invocation. The process
was never registered by the Runtime, so every later shell-status can
only be unknown, the exit-only settle path never applies, and every
release carries the hold forever.

A proven never-started is proof: when the invocation settles with
executionStatus not_started, the sibling row now settles with the same
not_started mark, its admitted refusal as the evidence. Busy computed
after that never counts a process that could never have existed; an
unproven-lost connection still wedges exactly as the declaration says.
The witness drives the admitted refusal shape end-to-end: the process
row lands SETTLED with not_started, and release passes without the busy
answer. Verified red under the inverse edit, which reproduces the
report's `expected: <SETTLED> but was: <PREPARED>` verbatim.

* fix(managed-agent): attribute wake receipts to their own turn and stop re-driving it

Two wake-path defects from the round-6 audit family, plus the monitor
registry the maintenance route actually shares:

- `recoverShellReceipts` exempted monitor inputs so thoroughly that a
  wake turn's receipts could never be attributed: any `tool.receipt`
  journaled under a wake turn would attach to the previous prompt or be
  skipped. Receipt attribution now advances on every accepted input,
  while the pending set itself still exempts monitor inputs as before —
  nothing reframes them as parked Turns.
- A wake turn that started, journaled records and died was re-executed
  text-only on reopen, minting a second user record and leaving an
  unanswered first call in the model's history. `wakeHasPriorAttempt`
  names the condition simply; the pump now parks such turns accurately
  blocked for the recovery fleet instead of redriving.
- The worker constructed its executor with no monitor registry, so the
  executor built a private one while the monitor maintenance routes got
  another empty one — every status and stop answered unknown. One
  registry is shared by executor and routes now, mirroring what the
  Shell path already does.

* fix(managed-agent): read the task event floor after the events, not before

Two stores-and-reads of an event page ran as separate autocommit queries:
the retention floor was checked before `events.read`, so an expiry mid-
read produced a 200 that silently skipped events under an unchanged
nextCursor, the exact gap the contract's `409 cursor_expired` exists to
answer. The replayable-events family already documents the sound order:
events first, floor after — a monotoni…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants