Repository navigation
feat(managed-agent): persist Workspace session tool profiles - #13301
Conversation
Test report — 8da0715Verified commit 8da0715edabc: 556 passed, 1 platform skip, 0 failures, 0 errors across 68 Java test suites. Compilation, Checkstyle and SpotBugs passed. The MariaDB migration smoke check also passed.
Reproduction and result detailsLocal environment: macOS, Java 21, Maven 3.9.11, H2 in MySQL compatibility mode; separate MariaDB 10.11 migration smoke check. From the repository root, with the sibling Java artifacts installed: mvn -f packages/sdk-java/managed-agent-server/pom.xml verifyThe platform skip is Browser UI evidenceAdded a real-browser check at the same commit: 1 Playwright test passed; 0 JavaScript page errors; all 11 recorded API responses were 2xx. The PR's ManagedAgentWebShell runs against its real Spring controllers and an isolated H2 database migrated through V35. The existing browser fixture supplies an authenticated test actor and Workspace; Harness execution is disabled. HTTP responses are not mocked.
The UI does not display Desktop — create a Workspace Session Desktop — saved Session after browser reload Browser environment: macOS, Java 21, Vite 5.4.21, Playwright 1.61.1 Chromium. The fixture host supplies dark theme background/foreground and viewport sizing through the component's existing props. Product code and responses were unchanged; temporary fixture files and servers were cleaned up. This is after-only evidence, with no server-restart, model, Shell or approval-flow claim. Not verified by this report: live Harness integration, MySQL 8, Shell approval/FG6f, Shell opt-in or CI follow-upThe owner-assignment failure is now resolved on main by #13306. #13308 was closed as superseded after verifying that its remaining hosted-runner changes were unnecessary. At the latest check of head The MariaDB CI failure came from an upgrade fixture calling the latest Session writer against a V31 schema. The fixture now seeds historical rows directly, then upgrades and verifies the preserved close receipt, files/1 backfill and retirement path. Production code is unchanged. Local verification: all 52 non-Hosted MariaDB integration cases passed (36 unaffected cases from the full run plus 16 retention cases rerun after the fixture correction), with Checkstyle passing. The original failure was reproduced locally before the fix. Disposable databases and credential files were removed. The test-only follow-up is c073f003bf16; the production implementation and browser-tested UI remain unchanged. Follow-up evidence at |
yiliang114
left a comment
There was a problem hiding this comment.
Reviewed c073f003bf16af29a85924c3275846365918f664 against main 5ddfacc9d4c18d6e85faeed18786ada134774541. No blocking finding in the bounded correctness, security, quality/performance and test-evidence review.
The review covered V35 and old-writer compatibility, the transactional profile write, idempotent replay, create/conflict-load/cold-recovery paths, and unchanged public/Web Shell admission. Bound sessions save their profile; unbound sessions omit it; a missing or blank bound profile fails before Harness access. V35 is unique on the checked main. The Java workflow for this exact head passed, including the migration uniqueness and daemon gates.
This pass was read-only: no tests or UI were rerun. The existing browser report uses a Spring/H2 fixture with Harness execution disabled, so it does not establish live Harness execution. The stored /2 test value checks persistence rather than new /2 admission. Shell and new profile admission remain outside this PR; it remains Draft.
|
Live Hosted files/1 execution and graceful recovery passed at
The candidate passed 36 assertions. A separate control compiled only the connector from exact base The negative cases deliberately edit isolated SQL fixture rows. Their HTTP409 is the fixture wrapper around Java's CLI SHA256: Independent cleanup confirmed all owned native processes and Java worker groups gone, all 12 observed ports closed and private configuration removed. Earlier setup/protocol failures remain separate evidence. This verifies graceful Harness close plus Java cache eviction/load; physical Host crash/restart, Broker reboot, MySQL 8, active Shell approval and real-model reliability remain outside this run. The earlier browser report still covers creation/reload and desktop/mobile layouts with Harness execution disabled. Between that source revision |
|
Real browser + live Hosted files/1 verification passed within the scope below at exact head The browser opened the committed Managed Workspace fixture host, rendering the actual
Final SQL contains three COMPLETED Turns and nine SETTLED file executions, with journal committed sequence 132. All 12 model calls offered only read/write/edit. Random proof bytes were absent from the typed prompts; actual prior-file results reached the model, and all three physical proof files are correct inside the bound Workspace and absent from the Harness decoy. There was no process restart. The first model response was intentionally held for an in-progress capture and crossed the fixture relay's 90-second SSE timeout. That live sample showed Completed without the assistant answer; SQL contained the full answer and browser reload restored it. This timing-affected sample is not a clean live-streaming pass or a proven product defect. The normal second and recovered third answers appeared live. Intermediate file-tool cards were not displayed and are not claimed. Nine actual browser screenshots are retained locally. The original eight use the minimal host's white body with the component's transparent dark theme. The readable final capture changes only the fixture host body background to Both this run's own before/after manifests contain 3,649 identical files: all 1,370 dist files, explicit Java classpath files/JARs, frontend/SDK sources, fixture scripts and Node/Java executable bytes. Manifest SHA256: All 36 evidence checks passed. Independent cleanup confirmed the Java, Harness, Vite, observed worker and runner processes are gone, all nine recorded loopback ports are closed, private fixture configuration is removed, and the browser task space is closed. The worker listener port was not separately sampled; its process/group absence was checked. This is candidate-only component-host evidence, not browser Before/After, production standalone-bundle, external-vendor, MySQL 8, physical crash/restart, Shell, or |
|
Fresh local execution at
All 269 main Java sources were freshly compiled from the candidate into isolated class directories; existing target class directories are absent from the runtime classpath. The successful before/after manifests contain 2,101 identical files, including all 1,370 dist files and the recorded Java classpath/fixture inputs. This is the recorded subset, not an exhaustive node_modules or host filesystem hash. Source remains clean. No product source change was needed. Actual tmux capture excerptsThese are selected verbatim lines from the completed rendered terminal capture; complete step captures and the unabridged final capture are retained locally. Local evidence: Two setup failures were retained separately: the scratch compile initially selected Commons Lang 3.17.0 instead of the SDK's declared 3.20.0, then a preparation script had an indentation error. Both stopped before product tests; only the test setup was corrected. The authenticated external-model attempt was rejected by automatic approval before execution, so no vendor-call result is claimed. This report covers real local product/file execution with nine deterministic model calls. Authentication/Workspace registration come from the isolated fixture; physical restart, MySQL 8, production deployment, Shell and /2 admission remain outside this run. CLI SHA256: |
qqqys
left a comment
There was a problem hiding this comment.
APPROVE at c073f003.
Critical-only scan over the production diff (the migration, the store, the record, the connector) plus the call chains that read them. No blocking issue found, and no prior blocking finding exists on this PR: there are no inline review threads, and the triage pass recorded no critical blocker.
The new IllegalStateException is not reachable, and the persisted value preserves today's behaviour
QwenHostedHarnessConnector.toolProfile (:377-385) replaced a hardcoded "hosted-workspace-files/1" with a DB-sourced value and added a throw when a Workspace session has a null/blank profile. I tried to find a path that reaches the throw and could not:
- Existing rows —
V35__managed_session_tool_profile.sqladds the column withDEFAULT 'hosted-workspace-files/1', so every pre-existing row is backfilled with exactly the string the code used to hardcode, thenUPDATE ... SET tool_profile = NULL WHERE workspace_id IS NULLnulls precisely the rows for which the first guard already returnsnull. Behaviour is identical on both sides of the migration. - New rows —
insertSessionbindsworkspace == null ? null : "hosted-workspace-files/1", sotool_profileis non-null exactly whenworkspaceis non-null, which is the only branch that can reach the throw. - Any writer that omits the column — the column
DEFAULTsupplies the value, so an omitted bind cannot produce null on a Workspace row. - The reduced constructor —
SessionRecord's secondary constructor delegatesworkspaceasnullandtoolProfileasnulltogether, so a record built that way returns at theworkspace() == nullguard before the throw.
The invariant is therefore enforced, not merely asserted. This is the right direction for a value that will later vary per session.
Read-path column coverage and record arity both verified by reading, not inferred
The row mapper now calls result.getString("tool_profile"). All five queries that feed sessionMapper (ManagedAgentStore.java:1012, :1022, :1045, :1849, :2123) are SELECT * FROM managed_agent_session, so the column resolves on every path; the only projection query over that table (findMaterializationTargets, :1273) maps to MaterializationTarget and is unaffected.
Adding a component to a record is a compile-time hazard for every canonical-arity caller, so I enumerated them rather than trusting the diff. In src/main the only construction is the row mapper itself, which this PR updates. In src/test, five files construct SessionRecord and this PR updates four; the fifth, QwenHostedHarnessNewSessionRegressionTest.java, is untouched — I read it at this head and it calls the 13-argument secondary constructor (:49-50), whose parameter list this PR does not change (only its delegation gains one null). It still compiles. insertSession likewise stays balanced at 16 placeholders for 16 bound arguments.
CI gap, disclosed rather than relied on
At this head only housekeeping checks ran (assign, authorize, label, route, triage succeeded; 91 skipped; review-pr queued). No Java build or test job executed, so compilation and the migration test are not CI-verified here — which is why the arity and column-coverage checks above were done by reading source at this head instead of being inferred from a green build. Nothing in CI indicates a defect introduced by this PR; this is a coverage gap, not a failure.
Non-blocking: the migration version will need renumbering at merge time
main currently tops out at V34__managed_tool_output_collection.sql, so V35 is free there. However ten open PRs each add a different V35__*.sql to this same directory (#13210, #13217, #13247, #13289, #13260, #13301, #13325, #13336, #13354, #13355). Whichever lands second will collide and Flyway rejects duplicate versions at startup, so this file needs renumbering against whatever main holds when it merges. This is a merge-ordering action item common to all ten PRs, not a defect in this diff, and it does not gate approval.
qwen-code-review-bot
left a comment
There was a problem hiding this comment.
APPROVE at c073f003. I found nothing blocking, and the part of this change that could quietly have been wrong is both handled and pinned by a test. No thread had been filed before this review, so this is an independent read rather than a confirmation.
The mixed-version window is the load-bearing part, and it holds
The migration adds the column with a SQL default and then clears it for unbound rows:
ALTER TABLE managed_agent_session
ADD COLUMN tool_profile VARCHAR(64) DEFAULT 'hosted-workspace-files/1';
UPDATE managed_agent_session SET tool_profile = NULL WHERE workspace_id IS NULL;The default is what lets an old writer keep inserting bound Sessions after the migration, which is the compatibility requirement the test plan names. The consequence worth checking is the awkward one: during a rolling upgrade an old writer inserting an unbound Session also gets that default, so the column can hold hosted-workspace-files/1 on a row that must not send a profile. toolProfile() is immune to that because it branches on the binding, not on the column:
if (session.workspace() == null) return null;
if (session.toolProfile() == null || session.toolProfile().isBlank()) throw new IllegalStateException(...);
return session.toolProfile();And ManagedSessionToolProfileMigrationTest pins exactly that row — late-unbound is asserted to hold hosted-workspace-files/1 in the database while the connector still sends nothing. So the invariant is structural rather than incidental, and a future refactor that starts trusting the column instead of workspace() reddens a test instead of silently sending a profile for an unbound Session.
V35 is also the correct next version: main currently ends at V34__managed_tool_output_collection.sql, and Flyway migration version uniqueness is green at this head.
Two details that are right for non-obvious reasons
- Hoisting
toolProfile(session)to the first statement ofload()is not cosmetic. In the previous formmanagedSessionStore(session)was assigned beforetoolProfile(session)was evaluated as an argument, so a refusal threw after the store connection had been obtained. Refusing first means the guard costs nothing when it fires. - The new writer states the profile explicitly instead of leaning on the column default —
insertSessionbindsworkspace == null ? null : "hosted-workspace-files/1"and liststool_profilein the INSERT. That is what makes the default purely a compatibility device for old writers rather than part of the new path's semantics, and it means theIllegalStateExceptionis unreachable through this code.
On that throw: I checked whether it is reachable at all, since it would surface as a 500 through ApiExceptionHandler's Exception.class mapping. It is not — the migration backfills every pre-existing bound row, and the new writer always sets the column for bound Sessions, so only an explicit NULL write could produce it. Keeping it as an invariant assertion is right; a 500 is the correct signal for a state the schema and the writer both promise cannot occur.
Adding toolProfile as a record component is compiler-enforced across every construction site, so the arity change needs no audit — the green Java lanes are the evidence.
Verification basis and one forward-looking note
All nine Java/DB lanes are green at this head: Java 11/17/21 on ubuntu, Java 21 on macos and windows, Real daemon E2E (Java 11), Runtime Broker and Managed Agent MariaDB (Java 21), Hosted process fault gates (MySQL 8.4, Java 21), and Flyway migration version uniqueness. CI is 22 pass / 0 fail / 35 skipped. Usual disclosure: I did not run the Java suite locally, so this rests on reading the code plus those lanes; the body's own numbers (557 tests, 556 passed, 1 macOS skip, plus a negative control that restored the hardcoded profile and produced the expected failure) are the author's and I have not reproduced them.
One thing to keep in view rather than fix here: the profile identifier now exists in two languages — once in insertSession and once as the migration's column default — and that agreement is protected only by the migration test. With /2 admission deferred to a later slice of #13271, a second profile value will land in both places, and the SQL side cannot be refactored to a constant once shipped. Worth deciding where the Java-side constant lives before that slice rather than after.
Vote effect
reviewRequests lists doudouOUC, LaZzyMan, qqqys, tanzhenxin and wenshao, and reviewDecision is REVIEW_REQUIRED. Those requests are not CODEOWNERS-derived: every changed file is under docs/design/ or packages/sdk-java/, and no CODEOWNERS rule covers either (the rules are /.github/CODEOWNERS, three release/security workflows, /packages/core/, /packages/cua-driver/ and /packages/mobile-mcp/), so they were requested explicitly. @qwen-code-ci-bot approved at 04:33 on this exact head and required_approving_review_count is 1, but an outstanding review request pins the decision regardless of how many approvals exist — so one of those five needs to submit, or the requests need clearing, before this merges.
…n added Merging main brought #13301, which adds a toolProfile component to SessionRecord and turns a null profile on a bound Session into an invalid state: QwenHostedHarnessConnector throws "Hosted Workspace Session tool profile is missing". The bound fixtures in the two retry-terminal suites still used the 16-argument shape, so the module stopped compiling after the merge. The value is inert for these tests - both suites mock HarnessConnector, and toolProfile() is read only inside the real connector - but a bound fixture without a profile now describes a state production rejects, so it follows main's own HarnessCoordinatorTest fixtures.
Main's #13301 persists Workspace session tool profiles and claimed migration V35; this branch's four migrations move to V36-V39 (activation columns, event-type index, snapshot deferral marker, journal sequence index) and SessionRecord carries both new fields (approvalMode, toolProfile).
…the commit-time guarantee #13301 landed V35__managed_session_tool_profile.sql on main, so this branch's V35__managed_task_event.sql now shares its version: the merged tree fails at startup with "Found more than one migration with version 35". Move the outbox migration to V36 and update the authority design doc in both languages. The commit-time validator mirrors the checks the authority applies to every line whatever its kind (envelope, vocabularies, domain checks, reserved ids, byte caps), not the per-kind payload schemas, subject rules or commit-marker digests, which stay the authority's contract. Reword the apply() Javadoc, the test comment and the design doc so they no longer promise that no commit can brick the Session, and say what a stored reserved id breaks: it collides with the id of that domain's next Stage H record rather than failing the next open. Pin the event id and operation id shape checks with two negative cases; disabling either check previously left every test green.
main's #13301 claimed V35 for the session tool-profile migration while this branch was in review; the two flyways collided on the classpath (Found more than one migration with version 35).
…LM#13300 (QwenLM#13355) * docs(managed-agent): design H3 background Shell and Monitor runtime * feat(core): add managed child_run (kind shell) record body for H3 * feat(managed-agent): mirror child_run shell record body in Java store * test(core): commit and rebuild child_run shell chains through the session authority * feat(managed-agent): close child_run reference closure and add stop-request draining * feat(core): add named cgroup unit attach and managed child-run supervisor * feat(cli): admit v3 background shell under the supervised worker registry * feat(core): add open-ended shell stream capture with manifest revisions * fix(core): type-narrow stream capture finalize and refused publish double * fix(cli): close background shell type gaps after the main merge * fix(managed-agent): answer the R1 review round on the child_run contract * fix(managed-agent): close H0c critical follow-ups R3-1/R3-2/R3-3 PR 1 of #13300, fixing the three Critical findings from the round-3 review of #12855. R3-1: the Broker-record execution mapping now reads the record's dispatchGeneration: SETTLED/cancelled with generation 0 (the record's own proof it was never claimed, per ToolExecutionRecord) maps to not_started_proven instead of settled, removing the illegal intent -> settled shape that left a Monitor cancelled before dispatch with no committable settling revision. A shared brokerExecutionCases row (settled-cancelled-unclaimed) replays it in both languages, the TS divergence pin gains the second wire-reading divergence, and Decision 10 of the authority design is reconciled with the H0b record-contract doc in both language pairs. R3-2: the store now runs the event-level half of the envelope checks for every event line at commit time: closed envelope with an optional subject, version, declared sequence, well-formed event id outside the reserved <domain>:<n> namespace, this Session's closed key, valid time, and kind from a mirrored EVENT_KINDS vocabulary; every domain.committed payload is validated whether or not a body is registered (reusing the pinned DOMAINS); unknown-subtype lines follow the scanner's "after the Managed header" condition; a transaction carries at most one Stage H record; and per-kind byte caps (maxEventBytes, maxCommitMarkerBytes) are pinned in the shared limits contract and enforced per line kind. All refusals answer the existing 409 so the commit rolls back. Replay commits return before any of this runs, as before. R3-3: task.updated rides its own outbox (managed_agent_task_event, V35) written in the same commit transaction and keyed by a unique source key, drained later by the task-events slice; managed_agent_event stays turn and lifecycle events, so an announcement between two text deltas can no longer split a message Part under the frozen projection version. The route description is scoped to what the server does (a Session being deleted announces nothing while its tasks stay readable), and the design's replay-idempotency sentence is corrected. Test helpers that wrote journals a real authority could not read back (ActionJournal's reserved event ids, bare event/marker lines in the publication and session-store integration suites) now write well-formed lines. Mutations of every new check were verified red against the suites. * fix(managed-agent): repair the CI-only cgroup root and env-guard failures How: wrap the delegated-root probes so a missing or unreadable root answers the documented isolation error on Linux too (macOS refused at the platform check first, which is why only CI saw the raw ENOENT), and document the background Shell environment allowlist in the process.env guard. Why: the Test lane on 4fb9f0a9c8 went red on exactly these two items; both are this branch's own changes, not flake. Test: hook-command-cgroup and process-env-guard suites green; tsc clean. * feat(managed-agent): inject the child-run supervisor at worker boot How: registerManagedContextRoutes builds a ManagedChildRunSupervisor from the delegated cgroup root the Hook commands already use (QWEN_MANAGED_HOOK_CGROUP_ROOT) and passes it as the executor's fifth argument; a boot without the delegation keeps the executor's committed refusal instead of failing to boot. Why: executeV3Background landed in the previous increment with the supervisor parameter unwired, so every real-stack background start answered the committed no-supervisor refusal. Test: new wiring case asserts the supervisor is injected exactly when the root is set; managed-context-worker, process-env-guard and managed-background-shell suites green (775); tsc clean. * docs(managed-agent): pin the two-row ledger shape for background processes How: the runtime-ownership section now says the ledger holds two rows — the foreground-length start invocation settling with the handle, and a second execution admitted at start that carries no model result and stays non-terminal until physical exit evidence. Why: the Java ledger writes a result exactly when an execution settles, so a single row that is both handle-delivered and active cannot exist; the two-row shape keeps the settled-if-and-only-if-result invariant while preserving every cited behavior (Runtime holds, evidence-only settle). * feat(managed-agent): commit child_run shell lines from the hosted side How: HostedChildRunSession funnels every record write through one serialized, replay-safe commitExtensionRecord path — admit first with the start call's args as commandRef, dispatch and the set-once managed-runtime-receipt after the physical start, outputRef advance-only, and settlement only on proven exit evidence, a proven failure, or an honored stop request. Live revisions alternate the run line (no self-loops); the authority freeze refuses anything after a terminal revision. Why: the dual path puts product records on the hosted authority, but nothing commits child_run from the hosted side yet — the worker-owned registry intentionally does not touch records. Test: four authority-level cases — full chain with projection, stop request to cancelled, pre-start failure frozen on not_started_proven, and serialization with deep-equal skips. * feat(managed-agent): add the detached capture family to Tool v3 results How: a background Shell start now settles success with a capture object whose status is 'detached' — no reason, no manifest — because the live output streams through the child_run record's growing manifest rather than the result; the shared schema pins both invariants, the TS parser mirrors them, the shared fixtures gain one valid and two invalid envelope cases replayed on both sides, and the Java projector accepts the missing manifest for unavailable or detached captures. Admission refusals settle not_started with a null capture, the shape the durable unstarted family already owns. Why: every Session tool.receipt event lands in the Java delivery projection, which requires a capture object and a committed-if-manifest pairing for anything that started — a success result with a null capture fails that projection and corrupts replay. Test: core contract suite 519 (three new envelope cases), TS serve suites 33 + turn/harness 323 + background 6, Java publication contract and projector suites 8; tsc clean on both packages. * fix(managed-agent): answer the R2 review round on worker, capture and record lines How — the four Criticals first: the stream capture publishes an ended stream's revision only after its seal decision, so a sealed-or-incomplete descriptor never changes again (R2-3); finalize's settle path degrades to an unavailable capture instead of throwing past a writability failure, and the last published revision stands like the worker-cut cap (R2-5); the child_run body now refuses a settled or cancelled run that is not settled execution, a failed run off its two ending lines, start_failed outside not_started_proven, and a process-level failure without a settled execution, on both languages with new witnesses replayed from the shared fixtures (R2-6, R1-30); and the supervision suite skips win32 instead of spawning a shell that cannot exist (R2-7). How — the Suggestions land as: supervisor start now proves membership by cgroup.procs with an fd-3 status channel, failing closed as isolation (R2-37); the bounded EOF wait caps daemon-inherited pipes instead of hanging, and a spawn-time error drains the entry (R2-1/R2-14); registry hasHolds requires its session and setProcessResult carries the full result shape (R2-13/R2-15); create() validates its caller-named unit like attach() (R2-26); the executor checks the journal before any effect, mirrors the closing/is_active re-check after prepare, gates the ninth live Shell as a committed quota refusal, and keeps cgroup membership out of an empty QWEN_MANAGED_HOOK_CGROUP_ROOT (R2-23/R2-19/R2-24/R2-18); the spawn uses the configured shell and the session-context environment (R2-21/R2-20); invalid fixture cases each pin their refusing clause on both languages (R2-36); projection fixtures pin the draining precedence rows and the Java replay reads stopRequested (R2-28/R2-38); the draining Javadoc states the rule the code implements (R2-39); the duplicated ref helper is gone (R2-40); the stream-capture suite is typed and gains the seal-before-publish, page-cursor, late-failure and settle-degrade witnesses (R2-29/30/31/32); and the design doc names the cgroup switch decision, the cursor's single ordinal space, the unconditional host-scope evidence gate and the zh stopped-object (R2-8/9/10/11/12). Test: core suites 788, cli serve suites 776 + background 11, Java record/ projection/store suites 24; tsc clean. * feat(managed-agent): thread the child-run orchestrator and the detached receipt family How: HostedSession gains its session-scoped childRuns orchestrator next to hooks/mcp, the tool turn receives it stored-ahead of the admission branch, the reopen verifier and the recovery replay both accept the third durable receipt family — a blocked delivery with a detached capture, beside the complete and not-started ones — with the publication store correctly left out of its proof, since its durable truth is the child_run record. Why: the H3 background start settles with a handle whose output lives on the record's growing manifest; without the third family, any Session holding a background receipt could never reopen or be replayed. Test: recovery-session suite 52 with two new detached-family cases; harness-session and tool-turn suites green; tsc and lint clean. * test(cli): type the background-shell capture double How: drop the as-unknown cast on the FakeSink return and implement the failCapture member the cast had been hiding (R2-29's cli location). Why: an unchecked double can drift from the interface it claims to match — and it already had. Test: background-shell suite 10/10; tsc clean. * feat(managed-agent): admit background Shell starts from the hosted tool turn How: the turn replaces its two deliberate is_background refusals — exactly when the Session owns its child_run orchestrator and the domain is enabled — with orchestration that mirrors the shared line: the record intent precedes every physical effect, dispatch_started lands the moment the checkpoint commits, a settled detached handle attaches the physical start to the record from the same facts and lands as the third durable receipt family (blocked delivery with a null manifest, replay-validated exactly like the not-started family), and a proven-unstarted refusal settles start_failed on not_started_proven while riding the existing unstarted family unchanged. Malformed is_background values keep their old refusal, and the family stays out while the domain is disabled — the same probe governs both paths. Why: the background start is the first record-bearing, user-visible H3 effect, and without a hosted history family for the handle a Session that ran one could never reopen or be replayed. Test: three new authority-level turn cases — admitted detached flow with admit/dispatch/attach order and the blocked receipt family, a proven-unstarted refuse closing NOT_STARTED, and the disabled-domain refusal preserving its exact text; turn suite 146, plus harness-session, recovery, background and child-run suites 246; tsc clean. * feat(managed-agent): admit background Tool v3 dispatches and acknowledge their detached settle How: the v3 start admission flips from "foreground Shell only" to "Shell only" — background dispatches now begin the same way; the poll loop accepts the settled detached family through the status answer (the only durable settle such a result can ever produce, since nothing is published for a background handle), and the acknowledgement path canonizes the settled handle envelope: a matching blocked receipt with a null manifest and no history revision forwards exactly itself, without consulting the publication receipt. Why: without the detached branch, a background start settled a queued handle only to mark the execution UNKNOWN after the 30-minute foreground-shaped deadline — every background dispatch would wedge the Runtime on its first call, and its acknowledgement would die on the missing publication it was never going to have. Test: new witness drives the detached start end-to-end — reservation, v3 dispatch, settled handle via status, and ack with a null-manifest blocked receipt, with the publication receipt path provably untouched; service suite 139/139, module suite 600/600. * feat(managed-agent): answer shell-status and shell-terminate from the worker registry How: a new private maintenance route sibling to ManagedHookProtocol (/internal/managed-runtime/v3/shells) backed by the process registry — status answers running or unknown (never claimed without evidence), terminate drains with the supervisor's own rules and answers exited with its proof, and an end that cannot be proven is unknown with the hold still registered. The unit name derives from the runtime invocation identity through the shared shellUnitNameOf, so Java, the hosted turn and the worker derive it identically; the executor takes the registry as a constructor injection so the route and the journal share one. Why: every later maintenance verb — recovery reconciliation, the ordered close drain, the reconcile-settle — has to ask the only physical owner these questions; answering them from anything but the registry would invent truth. Test: five worker-side cases — closed-key refusals, unknown for unowned, running scoped to the registered Session, terminate-with-evidence answering exited, and an unproven terminate answering unknown with the hold intact; shell, context-worker, v3-route and child-run suites 784; tsc clean. * feat(managed-agent): admit the background process row and answer its maintenance protocol How: the start dispatching of a background Shell admits a second ledger execution beside the invocation row — the new process row carries no model result and stays PREPARED (non-terminal) until physical proof, so hasActiveBy* holds the Runtime exactly while the process lives. The maintenance protocol mirrors ManagedHookProtocol on its own route: shell-status and shell-terminate with a targetOperationId (both recovery kinds), wire-validated on both transports and accepted through the same workspace-generation and ownership gates. observeBackgroundProcess asks the physical owner and settles the process row with proven evidence — settle through the new repository verb settlePrepared, which mirrors the PREPARED-requestCancel shape on both repositories; anything unproven leaves the row active with its hold, never a claimed end. Why: the ledger writes a result exactly when an execution settles, so a single background execution carrying its handle at start cannot exist; two rows keep the settled-if-and-only-if-result invariant while the Runtime holds precisely, and the maintenance route gives every later verb the worker's physical answer instead of an invented one. Test: two new witnesses — a background start that admits the process row (the invocation SETTLED, the process PREPARED, the Runtime busy until status arrives, then exit evidence settling it and release succeeding), and an unproven status keeping its hold; the runtime also gained the row identity invariant fix and the settle-through-settlePrepared shape; service suite 141/141, module 603/603. * fix(managed-agent): answer maintenance views under the target identity How: the shell-status/shell-terminate view now carries the target operation's identity, exactly as the Hook lookup family does — the requester asks about the process, never about the asking call, and the Java wire validator's expectedId rule requires it. Why: this repair was encoded in the route design but landed uncommitted after ②e-a; the self-audit caught it before it could ride a later batch. * feat(managed-agent): run the background exit leg through the record How: the hosted publisher gains the background family — its registration marker admits the open-ended capture through the session-level admission (proven start on the record, never the model call), the open manifest publishes at prepare, every awaited write or finish forwards the newest manifest revision to the record's outputRef, and the exit finalizes with the same evidence object settling the record. The client acknowledged receipt keeps the detached-blocked shape; drain skips background stores until the ordered close drain wires them; the register guard only binds the foreground model-call family. Why: the start handle already committed, so a background Shell's whole physical truth is output and exit evidence — anything that reached the record another way would let a second tool.receipt contradict its own line, and an open-ended manifest published anywhere but the Session's own store is an unreadable closure for the store-side commit. Test: a full end-to-end authority case — admit-dispatch-attach, bytes through the real rpc, outputRef advancing live, the final sealed manifest reading the same exit (code 3, both streams sealed, manifest complete), the record settled with exactly that evidence, and zero tool.receipt events for the outcome; publisher, turn and child-run suites 174; tsc clean. * feat(managed-agent): close monitor_run revisions over their cited Session resources The H0c closure rule — every ref a record cites must exist in the Session's resource store at commit time — now covers monitor_run on both stores: the authority reads commandRef, startReceiptRef, outputRef and lastObservationRef through verifyExtensionResources, and the Java applyRevision gains the same four-field branch beside child_run. This is the precondition for enabling the monitor domains: without it a monitor revision could anchor its chain to resources nobody committed, and the replay would project a watch whose evidence was never verified. Both fixture rigs stop citing fictional resource identities. The authority test publishes its chain refs as real bodies once per harness and mirrors each fixture-cited ref as a real same-kind placeholder of the stated length, memoized so chain identity stays fixed across revisions. The journal test harness does the same at its generic request entry point, rewriting each cited ref to the placeholder's real digest and appending the mirrored resources, but only for strictly well-formed JSON monitor bodies — a body with trailing content keeps the raw path so the store's grammar refusal still fires first over the closure's. Each side gains a witness that pins the closure itself: a monitor run citing a resource outside its commit is refused with managed_session_resource_missing, and each witness was observed red with its branch disabled before landing. * feat(managed-agent): gate monitor_run behind the managed-session/2 reader The H3 design left the reader-gating mechanism to the implementation; the journal settles it. Headers are immutable — the scanner refuses a repeated managed_session_header_v1 line and any unknown subtype after the header — so no transaction can legally raise a Session's requirement in-journal, and a superseding header line would change the journal format in both languages for a narrow mixed-version window. New Sessions therefore stamp minimumReader managed-session/2 at creation, and every path that commits a monitor_run record — commitDomainRecord, commitExtensionRecord and the generic append guard — additionally refuses when the Session's header names less. Readers still accept requirements up to their own version, so pre-H3 Sessions stay openable everywhere and simply can never receive monitor_run. child_run is deliberately not header-gated: a managed-session/1 reader already parses its name, and the asymmetric hazard lives on the server side, which the established deploy-server-first sequencing covers. The shared Stage H golden is regenerated for the v2 genesis bytes and the Java integration replay passes against it unchanged. New witnesses: new Sessions stamp v2, a v1 Session refuses monitor_run with the precise error (observed red with the gate removed), and both header sides of the version comparison are pinned in the records tests. The design doc records the decision and the now-paid closure precondition in both languages. * feat(managed-agent): funnel monitor_run revisions through a hosted orchestrator The mirror of HostedChildRunSession for the Monitor domain. The managed-runtime worker will own the watch loop, but the dual path puts every product record on the hosted authority, so HostedMonitorSession commits the record line as watch facts arrive: revision 1 with the intent before any side effect, dispatch onto a Runtime binding, the set-once start receipt after the watch starts, observation revisions only while attached with a watermark that never goes back, output advancement only forward, and the terminal shapes — settled by exited/max_events/idle_timeout, failed by a start failure on not_started_proven or a settled watch failure, cancelled by a stop request whose notification watermark may still advance after. Writes serialize per Session and replay-safe by command id with deep-equal skip, exactly like the Shell funnel. The suite drives the whole line through a real authority: projection at admission and after settle, max_events refused before its quota and accepted at it, a second start receipt refused, an observation before attach refused, and an idempotent output advance committing nothing. Quota (enforcement), the debounce-floor observation loop with notification composition, and rebuild-after-loss stay with the runner and route increments that follow. * feat(managed-agent): drive monitor observations through a debounced loop The loop half of the Monitor runner, on top of the hosted funnel: an admitted Monitor dispatches, starts an injected watch executor and runs its own time. Stdout lines aggregate into one observation revision per debounce window with a one-second floor, so the document's reopen bound holds for any debounce a watch asks for; the run settles itself on the contract's own terminal conditions — the observation quota, silence past idle_timeout, a natural exit with its buffered lines flushed first — or is stopped, which terminates the watch and lands cancelled. The executor contract is typed for the worker's cgroup owner (onLine, onExit after the start resolves, a terminate handle); the suite drives the whole line through a real authority with a fake executor and a manual clock: windowed aggregation and its floor, quota settlement and watch termination exactly at max_events, idle settlement re-armed by each accepted observation, exit flush order, start and mid-run failure shapes, and a stop that ignores late lines. Every cross-boundary settle — timer or executor callback — is counted by the loop's done handle, which the close drain will await; window flushes count too, so done never reports idle while an observation commit is in flight. Notification composition and its wake stay with the next increment; quota enforcement and rebuild-after-loss stay with the route slice. * feat(managed-agent): notify from monitor observations in the same transaction The H0c machinery now runs for Monitors: every accepted observation commits its revision together with a notification input and the wake the authority generates for it, and the run's notifiedThrough watermark advances to exactly that observation's sequence. The loop composes the input — monitor-1:notify:N ids, the window's joined lines as the content a woken turn will read, an empty admission, no deadline — and the funnel threads it through commitExtensionRecord, so a recovery replay re-runs nothing the watermark already covers. Dedupe semantics are pinned at both levels: the funnel advances the watermark only when a notification rides the revision (an observation without one keeps it where it was), and the loop suite shows the three events — domain.committed, input.accepted, wake.requested carrying the input's accepted id as its source — landing in one transaction with the joined lines as its content. Consumption of that wake by the hosted scheduler, and settling the input when no turn can take it, come next (open question 6 of H0c). * fix(managed-agent): keep a finished background Shell's receipt answerable H7 of the real-stack rounds: the Broker settles its :process row only when shell-status answers exited, but the worker's registry deleted an entry the moment the Shell ended — a natural exit could never be answered again, and the ledger row would hold the Runtime against release and close forever. The registry now retains each completed unit's receipt until the worker ends: the hold drops exactly as before, but shell-status and shell-terminate answer a proven end from the retained evidence, idempotently, for the owning Session scope only. An end without evidence keeps nothing and answers unknown as before; the existing live flows are untouched. Suite: natural exit keeps answering exited with its evidence for both kinds after the hold dropped, the wrong Session still hears unknown, and an evidence-less end stays unknown. The close-drain/reconcile callers that consume this, and the Java-side settle-before-busy on workspace release, land with the next batch. * fix(managed-agent): renumber the task-event outbox to V36 and narrow the commit-time guarantee #13301 landed V35__managed_session_tool_profile.sql on main, so this branch's V35__managed_task_event.sql now shares its version: the merged tree fails at startup with "Found more than one migration with version 35". Move the outbox migration to V36 and update the authority design doc in both languages. The commit-time validator mirrors the checks the authority applies to every line whatever its kind (envelope, vocabularies, domain checks, reserved ids, byte caps), not the per-kind payload schemas, subject rules or commit-marker digests, which stay the authority's contract. Reword the apply() Javadoc, the test comment and the design doc so they no longer promise that no commit can brick the Session, and say what a stored reserved id breaks: it collides with the id of that domain's next Stage H record rather than failing the next open. Pin the event id and operation id shape checks with two negative cases; disabling either check previously left every test green. * fix(managed-agent): settle provable background exits before release's busy check The second half of H7. The worker now keeps a finished Shell's evidence answerable, so the Broker has someone to ask — but observeBackgroundProcess still had no production caller, so a naturally exited Shell left its :process row non-terminal and every release wedged on runtime_session_busy. releaseSession now sweeps the background process rows this Broker admitted for the Session before it computes busy: each unproven row asks its physical owner once through the same shell-status → settlePrepared leg observeBackgroundProcess already owns, a proven exit settles with its evidence, and any lookup that fails or cannot prove an end simply keeps the row so busy stays the accurate answer. The sweep tracks per-context admissions because the reconciliation scan deliberately excludes PREPARED rows, and mutates the admission set under the context lock. Witnesses: releaseSettlesAnExitedBackgroundProcessBeforeBusy proves an exited row settles with its evidence and the release succeeds (observed red with the sweep bypassed); the pre-existing unprovenProcessStatusKeepsItsHold still proves an answer without proof keeps the row and the busy refusal. Full module 604/604. The workspace close drain still terminates shells in order rather than sweeping — its ordered stop belongs to the close-drain slice, and these two verbs are exactly what it will consume. * fix(managed-agent): check the journaled receipt before attaching a background Shell's start H6 of the real-stack rounds: acceptBackgroundShell called childRuns.attach before its tool.receipt replay check. A retried accept — the leg recovering pending results re-drives — would mint a second start receipt on the same identity before reaching the replay path, and the record's set-once rule would refuse it as a rerun shape, so a recoverable retry crashed as a conflict. The lookup now runs first: a journaled receipt replays its validated history with no attach at all, and a fresh accept attaches only when this identity still lacks a start receipt — the record id IS the execution identity, so an existing one is always our own crash-window attach between record and journal, which the retry now tolerates instead of minting a rerun. Witnesses: a retried accept answers from the journal with exactly one attach and one tool.receipt line, and an already-attached-but-unjournaled record completes its accept without attaching again. The full hosted turn suite stays 148/148. * fix(managed-agent): backpressure the background Shell output pipes G5 of the real-stack rounds: the executor fed the bounded capture with a fire-and-forget sink.write — every chunk queued behind asynchronous publication, so a fast producer flooded memory (a 256 MiB run peaked near 315 MiB RSS at a 20 MiB/s store). The feed now counts bytes in flight and pauses stdout and stderr when the drain falls sixteen MiB behind, resuming at four MiB — enough to cover a store several seconds slower than the process produces without stalling a steady producer against the floor. Completion or rejection of a write both release their bytes, so a broken capture can never leave the pipes paused. The witness drives a controlled sink: seventeen MiB of chunks pauses a stream exactly once, draining past the watermark but above it resumes nothing, and draining the rest resumes exactly once. The outputRef-live-advance half of G5 belongs to the turn-publisher's promotion to activation scope, sized with the close-drain slice, and stays queued there. * feat(managed-agent): promote the Shell publisher to Session scope The ②d-out debt, and with it the live-advance half of G5. The Shell publisher was a per-turn object: its server died with the turn, while background Shell traffic flows across turns — between turns (or after a close) the worker's retained descriptor pointed at a dead endpoint, so no live output advance and no exit leg could ever complete past one turn. The hosted Session now owns one long-lived publisher across turns (the turn lazily structures it on the shell options and never closes it; the Session's ordered close in this PR's fifth slice drains it). The two small wire-level fixes that the promotion exposes: - The worker's background execute stamps the background marker into its prepare request: the closed six-field wire capture has no such field, so a remote prepare used to arrive unmarked and die at admission; the executor derives the marker from the background input itself. - register() takes the registering turn's prompt for the foreground admission check, so one instance admits the foreground of any turn of its Session while a foreground pretending another turn's identity is still refused; the background arm keeps its record-based admission. Witnesses: a single instance admits a background watch plus the foreground of two turns and refuses a misattributed one; the worker rig pins the stamped marker on its prepare; the four touched suites stay 184/184 and both package typechecks pass (after a core dist rebuild for the capture type's optional background marker). The publisher's own descriptor installs idempotently per turn (identical descriptors are accepted), so consecutive turns never fight for the worker slot. * feat(managed-agent): drain a Session's background Shells ahead of release The close-drain's missing legs. Until now a Session with a live background Shell could never close: the worker's activation gate refused release with managed_activation_conflict forever, and the CLI-side close route skipped the background stores. The ordered sequence the design asks for now exists end to end: - The registry gains stopSession: terminate one Session's entries with the supervisor's evidence rules (TERM, escalate, empty-check), await each completion so its finalization — last manifest, sealed envelope, record settle through the session publisher — lands before the caller proceeds; an unproven stop keeps its hold for the caller to see. - The worker's activation release gate drains before it refuses: active background work no longer wedges a close that can prove its stops, and anything still unproven keeps the exact same conflict. - The hosted close route now closes the Session-scoped publisher right after the broker release returns — which by then has drained the Shells and settled their records through that publisher — and the drain closes every capture store, background families included. Witnesses: a registry drain settles one Session's shells with evidence, both kinds, while leaving another's holds untouched, and an unprovable terminate under a session drain still preserves its hold; the 766-test context-worker suite stays green through the release gate's new async drain branch. What remains deliberately unproven locally: the full-chain release-with-live-shell drain needs a real cgroup supervisor, pinned as the Linux head acceptance instead of a mocked supervisor. * docs(managed-agent): sync the H3 design's status line with landed slices The status line still claimed publication and orchestration remained design while most of them landed across this PR's batches; refresh both language versions with the landed list and keep the remaining-design list accurate: the Monitor physical side (route, rebuild, cgroup watch) and wake consumer, both submission enablements, task events, and the Linux physical acceptance. * feat(managed-agent): spawn Monitor watches through a cgroup watcher The physical half of the Monitor runner's executor interface, for the worker: ManagedMonitorWatcher spawns the watch command under a unit named for the execution identity, through the same supervisor (its membership attestation unchanged), splits stdout into the observation lines the loop buffers — boundaries respected, the Legacy partial-line cap honored, the tail flushed as the final line at exit — and reports the physical end exactly once, natural exit versus mid-run failure, so the loop's settled and watch_failed shapes both hold. stderr rides only the future output leg, never observations; terminate defers to the supervisor's TERM-escalate-empty rules. The suite drives a supervisor double: unit/executable/cwd derivation, chunk-boundary splits with the final flush, the cap drop, spawn-error and membership-failure shapes, and terminate delegation. * fix(managed-agent): accept message.retracted at commit time #13351 added the message.retracted event kind on main. The hosted text delta stream writes it when a restarted model attempt replaces a prefix it already published. The commit-time vocabulary this branch mirrors from the authority still listed sixteen kinds, so after the merge with main the store refused every retraction with 409 ("event.kind must be one of [...]") and left the orphaned deltas in place. Add the kind to ManagedExtensionRecords.EVENT_KINDS and to the shared store fixture, so the parity pins in both languages agree again. * feat(managed-agent): answer monitor status and stop from a worker registry The monitor half of the maintenance route, mirroring the Shell's: a worker-side registry of supervised watches that holds each Session's Runtime while a watch runs, answers an end only with the supervisor's own evidence, and retains a proven exit's receipt until the worker ends so maintainers can learn the truth idempotently. The route sibling at /internal/managed-runtime/v3/monitors serves monitor-status and monitor-stop with the same closed wire shape, target-identity rule, and unknown-but-never-unproven discipline as the shells' route; it joins the worker's route list so the v2 envelope admits it. Registration arrives with the Monitor start admission later — status/stop already have consumers in the release drain and the rebuild slice that follows. This also repairs the watcher test arity that broke the previous push's build lanes (the mock's 2-arg onExit omission, seen as error TS2554 (155,34) across every CI lane at that head). * feat(managed-agent): admit monitor watches into the open-ended capture session The γ3-α precondition for every later Monitor leg: LocalShellStreamResultSession's prove-a-live-start admission reads the capture family's record domain instead of being pinned to child_run — same Session/binding/background-marker checks, same live-run gate, same identity captureId, per-domain labels so a refusal names its own family. The producer side names its domain explicitly (record presence is never consulted where both domains could in principle compete), so a monitor watch carries its monitor_run record into exactly the same bounded stream the background Shell already rides. Witnesses: shell admission unchanged; monitor admission succeeds under its domain; wrong-domain proven start refuses; missing record and foreign binding refuse with their own errors. * feat(managed-agent): admit Monitor watches through the v3 executor The workhorse of the Monitor leg. The worker's v3 executor gains its is_monitor branch beside is_background: journaled admission first, then refuses (never as a transport error), then the watch spawns through the ManagedMonitorWatcher under a unit named for the execution identity, streams stdout into the SAME bounded family (16/4 MiB backpressure on the pipe, with an explicit refusal if the watcher reports no supervised unit), and settles with its detached handle while a registry keeps the Runtime hold. Three watch-level facts are done physically now: - The monitor registry is a wrapper around the background Shell registry, never a copy: supervisors, EOF drains and publisher finishes stay in one place, while monitor holds and the four-watch quota bucket remain entirely their own. Its maintenance route answers run/exited/ unknown with identical semantics to shells'. - The execute path gates on the generated four-live-watch cap with `Session already runs 4 Monitor watches.` and additionally on a missing cgroup root and a missing capture service. - Each prepare carries background+monitoring markers, so the hosted side of γ3-α knows the monitor_run domain to check. Suite: 28 tests across the three touched files — the fork settles a detached start under the right unit and marker, lines land as monitor output, the fifth watch is refused, and supervisor-less boot refuses with its dedicated text. The turn admission arm (funnel-driven admit→dispatch→attach through HostedMonitorSession), the watcher→loop observation feed, and rebuild still land in the following increments. * fix(managed-agent): refuse a domain.committed without a textual domain A domain.committed line whose payload.domain was missing or not a string read as "no domain" and skipped requireDomainCommitted, so it committed 200 and the authority refused the Session at the next open. Decide that a line is domain.committed from its kind alone, and take the domain from the vocabulary check before building the expected record kind. Also address the rest of the first /review round: - The commit marker is the only line apply() measures. validateUtf8JsonLines already caps every line at MAX_EVENT_BYTES, which it now names. - The outbox's MAX(sequence_id) stays a plain read. A locking read deadlocks concurrent first announcements of different Sessions on MySQL 8.4 (85 of 480 in a probe). The comments now name the guarantee that actually holds: the Session store's journal-head lock is taken before the transaction's first plain read. - Stop claiming that the Session event stream carries only turn and lifecycle events. V36 is not applied anywhere yet, so its comment can still change. - Restore the shipped v1.19 OpenAPI sentence and record the move to the outbox as v1.30 (info.version 1.30.0). - Reconcile both design docs: Decision 10's state table and the scope of its wire-only absolute, the goals, the public-contract bullet, the validation plan, the files list, and the restored caveat that the record contract does not catch an unprovable not_started_proven. - Tests: refuse a non-textual and a missing domain, commit a message.retracted line, pin the final "events, then its commit marker" guard again, assert the empty outbox in the MySQL IT, and give the oversized-resource case bytes at its own sequence. * feat(managed-agent): admit Monitor watches through the hosted tool turn The turn-side arm of the Monitor leg. A call named monitor is now admitted through its own three-state gate (requested/unavailable/ill- formed, with the Shell-profile Monitor declaration now advertised), implements the same record-first discipline as background Shells do — the funnel admit commits monitor_run revision 1 before any side effect, dispatchStarted follows the durable checkpoint, and the watch attaches when its v3 execution settles. The granted publication, journal intent, dispatch and accept forks are widened to carry is_monitor alongside is_background, and acceptMonitor mirrors acceptBackgroundShell end to end (not_started → start_failed with the unstarted family; detached → same blocked receipt family with exactly one tool.receipt line, the H6 journal-verdict-before-attach order, and the blocked ack completing the row). HostedSession gains a session-scoped HostedMonitorSession beside childRuns; turn construction forwards it at both sites. Two existing contract pins move deliberately: the Shell profile's advertised list now names monitor (the harness declaration test names it explicitly), and the three publisher-close witnesses now assert the Session-scoped lifecycle settled by the earlier promotion — no close at any turn ending (completed/error), the server still answering, close exactly once at the Session's own delete. Suite: 3 new admission-arm tests through a full hosted turn (funnel order, clamped defaults with the Legacy monitor caps, start_failed shape, unavailable text), the whole turn suite stays 151/151 and the harness session suite 180/180. The observation fan-out (worker lines → hosted loop), quota witness at the turn, and rebuild-after-runtime_lost land in the next increments. * feat(managed-agent): fan a monitor watch's observations into the hosted loop The observation channel of the Monitor leg. The session publisher now marks monitor captures with recordDomain monitor_run through γ3-α, fans each published stdout chunk through the Legacy line-split semantics — boundaries kept, remainder held across chunks with the 4096-byte partial cap — into a per-capture observer, and closes it with onExit(failed) at finalization with the tail flushed first. The turn gains resumeMonitorWatch: a fresh accept creates one HostedMonitorLoop per Session report and resumes it from the already-attached record, fed through a HostedMonitorRemoteExecutor whose start binds those observer callbacks; a replay starts nothing twice, exactly like it never re-attaches. Globally, loop timers unref so observation windows never anchor the process. Monitor capture finalize settles through the monitor funnel itself (watch_failed when no evidence, exited via settleQuiet), beside the shell's unchanged path. The publisher carries monitors next to childRuns, and the turn's own construction forwards it. Suite: the admission-arm tests now also pin one Session-level loop starting on a fresh accept with the record-first flow intact; wide suites of 189 tests pass with cli tsc, prettier and eslint clean. Monitor output before the observer registers stays durable in the open stream, with observation starting at attach — noted in the design not as a flaw but as the attach seam, a line crossing it already lands in the output Artifact readers see. * feat(managed-agent): rebuild a read-only Monitor after its Runtime is lost The record side of the rebuild promise. HostMonitorSession gains the two verbs the H3 lifecycle needs around runtime loss: blockedRuntimeLost lands the run in recovery_blocked with the runtime_lost reason and the lost binding still named, so degrading is what the projection shows until something decides otherwise; rebuildFromRuntimeLost accepts only the outcome_unknown/runtime_lost line the H0b rule allows, starts a fresh watch under a newer Runtime generation with its own receipt, and keeps the reason for honest projection while observation positions never moves — the support case never restarts what the watermark holds. monitorRebuildAllowed routes every candidate command through the Legacy AST read-only check with its walk, so a side-effecting command stays blocked accurately. Five witnesses cover the full rebuild chain (recovery-block, rebuild with receipt remint and generation increase, a rebuild refused from a live run and from an ended one, and the gate's read-only/side-effect/empty answers). Binding-replacement detection and re-drive of the new physical watch come with the recovery slice after the wake consumer. * feat(managed-agent): serve the task events routes with their SQL journal H3 flips listSessionTaskEvents and queryWebShellTaskEvents from planned to partial on both surfaces. The Java control plane gains a bounded per-task event journal (V39): one row per committed task event, written in the Session commit transaction, with the per-task sequence allocated under the journal head lock in commit order. Record revisions journal their task view changes at the H0c announcement point. The durable retention floor lives in a per-task cursor row, so it survives an empty retained set, restarts and projection rebuilds; the Artifact visibility barrier pins it behind unarchived output, and the backlog bound refuses past capacity instead of silently discarding. Output events carry their per-stream segment ordinals with overlap/gap refusals, and artifact references stop fail-loud at the 100 bound. output_cursor/outputCursor lift with the routes and the task views serve them. The section 6.1 demonstrations run as store-level contract traffic in the new journal contract test, the route probes and record gates cover the flips, and the #12847 C15/C16 instances close alongside. WebShell types regenerated from the bumped 1.30.0 contract. * docs(managed-agent): add the task events lane's build report Flyway numbering findings (V39 taken; V35-V38 claimed by in-flight PRs on main, V15/V29 Java-only burns untouchable), implementation decisions against the H3 design text, the CLI/daemon flow-surface finding (none exists, regeneration is the only surface change), the full test ledger and the owed items. * feat(managed-agent): build the monitor wake envelope and the pending-input intake Two support pieces for the H3 wake consumer (β2-b1a): - monitorNotificationText wraps one due observation window in the exact Legacy task-notification envelope — task-id, optional tool-use-id, kind/status/event-count, summary and result — applying the same truncateNotificationLabel/stripDisplayControlChars/escapeXml pipeline the Legacy Monitor applies, so a turn learns nothing new. - pendingSessionInputs derives the Session's still-pending inputs from the journal alone: an accepted input is consumed exactly when a turn settles under its turnId, order-insensitively, so a replayed admission after a restart is not re-driven. * feat(managed-agent): deliver the Legacy notification envelope on monitor wakes The loop's notification input now publishes the exact task-notification envelope a Legacy Monitor's wake carried — task-id, tool-use-id, kind, status, event-count, summary and the window's lines — so the turn the wake raises reads nothing it has not already read. The description is the tool call's, falling back to the watched command. A witness pins the exact wire text of the first observation's input content. * feat(managed-agent): run the monitor wake through an embedded scheduler H3 closes the wake loop: an accepted observation already commits its notification input and wake.requested in one transaction; this slice makes that wake effective. The pump re-derives the pending monitor inputs from the journal alone — nothing is held in memory, so a restart re-derives the set. On an idle Session the oldest notification runs as an ordinary text turn with the Legacy envelope as its prompt, and the turn's settle consumes the input. A busy Session queues in the journal exactly like channel and Goal inputs; the retry loop redelivers. A Session that is parked, blocked or mid-recovery keeps its reminders pending and reports them. On the close path every pending notification settles cancelled without a model turn, under the turn-result record's own idempotency key, so no wedged notification ever holds a Session as hosted_turn_recovery_required at its next open: the parked-turn scans now skip monitor inputs, which the pump owns exclusively. A wake turn that dies without settling re-drives on reopen — the input is consumed exactly when a turn settles under its turnId, and a crashed turn never settles — matching the queue semantics channel and Goal inputs already honor. * fix(managed-agent): keep every new Session on the managed-session/1 reader Round-5 real-stack verification's cross-version matrix disproved the v2 stamp's premise: monitor_run has been a parsed, known domain since #12837 (v0.24.7), and a v1 reader opens a Session holding monitor_run records without complaint. Stamping minimumReader: managed-session/2 on every new Session — while monitor_run remains disabled — would make a rollback or a mixed-version rollout lose access to every Session created in between, and buys protection no deployed binary needs. Revert the stamp to managed-session/1 and drop the per-domain header refusal; the enablement list remains the single domain gate. The stamp rises only when a change genuinely breaks an older reader mid-scan, and only one release after a tolerant reader ships. The golden fixture's genesis header regenerates to v1; the records and authority suites gain v1-header admittance witnesses in place of the refusal cases; the design's reader section now records the reversal with its evidence. * fix(managed-agent): keep monitor output byte-true and advance only forward Round-5 real-stack verification found three facts about the monitor output path: - A capture rebuilt from decoded lines cannot be byte-true: per-chunk toString split multi-byte runes into U+FFFD, and line-rounding dropped blank lines and the final unterminated line. The watcher now forwards raw stdout chunks to the capture on a new onChunk channel, so the durable Artifact reproduces the command's stdout exactly, while the observation lines keep the Legacy emit-path semantics (a StringDecoder holds a split rune until it completes; a blank line consumes no observation). - The watcher registered its exit listener only after the supervisor's start resolved, so a fast watch could end unheard and the debounced loop never settled. A latched end with a check-phase sweep covers a watch that exited before its listeners attached, after the host is known to be listening. - The funnels' advanceOutput promised "only ever to a newer revision" in prose and accepted an older manifest in code. Both hosted funnels now read the revision of both manifests and refuse a regression before anything commits; witnesses pin the refusal in place. * feat(managed-agent): drain a busy background Shell at release, and answer the asked operation Round-5 close-path findings (H8): with a running background Shell the Broker refused busy before transport.release — yet that HTTP call is the only route to the worker's drain — so a `tail -f` held session release and workspace close forever, with nothing able to stop it. The release path now distinguishes busy kinds: anything active that is not one of the Session's own background process rows keeps the exact busy refusal with zero worker hops, while proven-running rows trigger a dedicated release call first whose 409 busy answer is swallowed (the worker's route drains before refusing, and its maintenance routes are not activation-gated, so the follow-up sweep still proves what the drain ended). Rows settle from the drain's evidence and the complete release converges on what is genuinely still running — one call when the drain finishes in time, one retry when it does not. Two witnesses pin the contract: a running Shell reaches the worker (the round-5 probe found the refusal never did) and a non-background busy refuses without a worker hop. Also from the round: the shell maintenance validator computed the expected operationId from targetOperationId but never compared it with the answer's own operationId, and the sweep witness answered with the process-row id where the real route echoes the call id. The validator now refuses a view answering another operation, a new protocol suite pins it, and the witnesses answer the shape the route really sends. * fix(managed-agent): re-assert the output pause behind every flushed chunk Round-5 G5 finding: Node's flushStdio resumes both paused pipes when the background launcher exits, and the executor tracked the pause in its own flag rather than the stream's, so a writer that outlives its launcher pushed 230 MiB into the queue and climbed RSS by +200 MiB. The onOutput accounting now re-asserts the pause behind every delivered chunk, and applyPause is idempotent through the stream's isPaused state instead of the local flag — the stop-and-go the backpressure intended holds through the launcher exit, on both the Shell and the Monitor paths. The existing 17-MiB pipe pacing witness keeps its single pause call. * fix(managed-agent): mirror message.retracted from the H3 event vocabulary in the commit-side kind mirror * fix(managed-agent): validate every domain.committed payload, with or without a textual domain * test(integration): expect the monitor tool on the hosted shell profile The hosted shell profile now advertises `monitor` beside `run_shell_command`, so the fake-model driver's tool-name assertion — which the Hosted workspace tool-turn IT checks on every model call — must expect it. The file profile's expectation is unchanged: the declaration is shell-profile only. Fixes the single red IT in the MySQL fault-gates lane (HostedWorkspaceToolTurnIT.packagedHarnessUsesSavedWorkspacesThroughReal BrokerWorkerAndSqlStore), whose only difference was the extra 'monitor' entry. * fix(managed-agent): read the whole journal when deriving pending monitor wakes `readEvents()` is a bounded read: with no arguments it returns the first `defaultReadEvents` (100) events, capped at 256. Both consumers of the pending-input derivation called it bare, so on any Session whose log runs past that page — every real Session with tool calls — the derivation saw only the log's head: - the wake scheduler's `next()` never found a notification input, so a monitor wake was never delivered as a turn once the Session passed the page. The feature was dead in production while every rig stayed green, because the rigs commit a handful of events; - `settlePendingMonitorInputs` never found the owed inputs at close, so they stayed unsettled and the Session reopened as `hosted_turn_recovery_required` — exactly the wedge this close-path settle exists to prevent. Both now read the full committed prefix with `eventsInSequenceRange(1, committedSequence)`, the discipline every other journal scan in the hosted session already uses. The witness commits 110 observations so the notification lands beyond the default page, asserts the log really exceeds it, and was verified red under the inverse edit (the settle returns 0 and the input stays owed). * fix(managed-agent): stop a full task journal from wedging record commits The task event journal is a derived, bounded feed: its backlog bound refuses the next event past capacity rather than silently discarding it, and an unarchived output event pins the retention floor so the automatic pass cannot expire anything. The record commit called `appendStateChange` inline, so once a task's journal was pinned and full, that refusal propagated out of the extension-record commit — every later revision of that task failed, which would take the Shell and Monitor record lines with it. The record row is the authoritative state and is already written when the journal append runs, so a refusing journal now degrades the feed (warned with tenant, session, task, revision and the refusal code) and never the commit. The store still throws for callers that can apply backpressure — which is where the output producer's retry belongs. The witness drives the real wedge: it pins the floor with an unarchived output event, fills the journal to the bound, asserts the next append refuses with `managed_task_event_backlog_full`, then commits a record revision whose view changes and asserts it lands. Verified red under the inverse edit, where the commit dies with "The task's event journal holds 512 retained events". * refactor(managed-agent): drop the shell view's never-set error field The maintenance view declared an optional `error` object, the Java response validator carried a branch for it, and no producer on either side ever set one: the worker answers `running`, `exited` or `unknown`, and a route-level failure travels as an HTTP status with its own code body. A field that is declared and validated but never written is a dead switch, so it comes out of the TypeScript view and the Java field set — which now also refuses an answer that carries it, instead of silently accepting a shape nothing produces. * fix(managed-agent): let a replayed output advance pass instead of refusing The forward-only check refused any advance whose revision was not strictly newer, which included a redelivery of the very reference the record already holds. The publisher guards against that in memory, but the guard does not survive a worker restart, so a replayed advance could throw inside the background exit leg and wedge a Shell or Monitor that had already ended. Both funnels now treat an identical reference as the no-op the deep-equal skip already owns, and still refuse any other reference that is not strictly newer. For a Shell the replay commits one liveness step, since the live run line alternates by design; for a Monitor it commits nothing. Witnesses pin both, beside the existing refusal witness. * chore(managed-agent): keep the task-events lane report out of the PR diff A one-off construction report does not belong at the repository root: the project's directory rules put working artifacts under the ignored .qwen tree and keep the tracked docs/ for design and plans. The report is preserved at .qwen/investigations/2026-10-04-task-events-lane-report.md, and its Flyway numbering note is superseded by the V40 renumber the merge commit records. * merge: absorb review-round state and retarget onto the H3 stack * fix(managed-agent): quiet the wake pump's consume guard at a blocked owner runTurn's recovery-blocked path marks the Session blocked and returns 'settled' without the turn ever existing, so the input stays owed — the accurate answer the whole design gives for a parked Session. The pump's after-settle check read that as a consumption bug and threw, which both marked the already-blocked Session again and logged a spurious error on every such observation. It now stops at an accurate blocked state and still throws when the input simply did not move: that is the programming error the guard exists for. * fix(managed-agent): meld a monitor watch through one terminal step Round-6 found the shape the Monitor's final leg really had: prepare's open-then-advance could never survive the start order (output before a start receipt is a parse-level refusal), the manifest still advanced through the Shell funnel whatever the capture's recordDomain said (a Monitor never builds its own Shell record), and the finalize settled the record ahead of the loop's own exit chain, so a watch's last window died in denial between two async hops — with its swallow queued to silence it. Three changes, each with its own witness that goes red under the inverse edit: - The manifest advance reads the record's start receipt first and only advances past it; a missing record still fails loudly instead of the old "no record to revise" at whatever funnel happened to be there. - It routes by recordDomain: monitor captures land on the Monitor funnel, Shell captures on the Shell's. - The loop's onExit now returns a chain carrying the exit's own flush → settle in order, and the publisher awaits it rather than settling itself — a final window can no longer be dropped ahead of the settle, and a commit failure of the chain surfaces inside the finalize instead of returning an empty observation state silently. The end-to-end witness drives the production shape on a real authority, orchestrator and Publisher over HTTP: prepare ahead of attach, output advancing only after the receipt, then one terminal step holding observationSequence 1 with the settled mark behind it. * fix(managed-agent): settle the process row when the Runtime proves no start Round-6 found the permanent wedge: the `:process` row admits PREPARED ahead of dispatch, but an admitted-refusal answer (`{executionStatus: "not_started", capture: null}` — cgroup root missing, quota exceeded, and friends) settled only its invocation. The process was never registered by the Runtime, so every later shell-status can only be unknown, the exit-only settle path never applies, and every release carries the hold forever. A proven never-started is proof: when the invocation settles with executionStatus not_started, the sibling row now settles with the same not_started mark, its admitted refusal as the evidence. Busy computed after that never counts a process that could never have existed; an unproven-lost connection still wedges exactly as the declaration says. The witness drives the admitted refusal shape end-to-end: the process row lands SETTLED with not_started, and release passes without the busy answer. Verified red under the inverse edit, which reproduces the report's `expected: <SETTLED> but was: <PREPARED>` verbatim. * fix(managed-agent): attribute wake receipts to their own turn and stop re-driving it Two wake-path defects from the round-6 audit family, plus the monitor registry the maintenance route actually shares: - `recoverShellReceipts` exempted monitor inputs so thoroughly that a wake turn's receipts could never be attributed: any `tool.receipt` journaled under a wake turn would attach to the previous prompt or be skipped. Receipt attribution now advances on every accepted input, while the pending set itself still exempts monitor inputs as before — nothing reframes them as parked Turns. - A wake turn that started, journaled records and died was re-executed text-only on reopen, minting a second user record and leaving an unanswered first call in the model's history. `wakeHasPriorAttempt` names the condition simply; the pump now parks such turns accurately blocked for the recovery fleet instead of redriving. - The worker constructed its executor with no monitor registry, so the executor built a private one while the monitor maintenance routes got another empty one — every status and stop answered unknown. One registry is shared by executor and routes now, mirroring what the Shell path already does. * fix(managed-agent): read the task event floor after the events, not before Two stores-and-reads of an event page ran as separate autocommit queries: the retention floor was checked before `events.read`, so an expiry mid- read produced a 200 that silently skipped events under an unchanged nextCursor, the exact gap the contract's `409 cursor_expired` exists to answer. The replayable-events family already documents the sound order: events first, floor after — a monotoni…



What this PR does
Persist
hosted-workspace-files/1when a Workspace Session is created through the public API or WebShell. Creation, loading and recovery then use that saved choice. A bound Session with a missing or blank profile is refused before the Harness request.This is the first implementation slice of #13271. Shell and
/2admission remain deferred.Why it's needed
The control plane currently derives the file profile each time it attaches to a Session. Saving the choice now gives existing file Sessions a durable identity before later admission changes introduce other profiles. Idempotent creation retries preserve the original choice.
Reviewer Test Plan
How to verify
Evidence (Before & After)
The Java module's full verification passed: 557 tests: 556 passed, 1 macOS-specific skip, 0 failures, 0 errors, including compilation, Checkstyle and SpotBugs. A negative control restored the old hardcoded profile and produced the expected assertion failure; restoring the implementation passed the full suite. The MariaDB CI follow-up corrected an old-schema fixture; all 52 non-Hosted integration cases now pass locally, with Checkstyle passing.
Browser verification passed against the real WebShell + Spring/H2 fixture: create a bound Session, reload the same Session and inspect desktop/mobile layouts. SQL confirms the stored files/1 profile after creation and reload. Test report; screenshot attachment pending. The fixture supplies test authentication and Workspace data with Harness execution disabled; this does not verify live Harness execution.
Live native verification now passed at
c073f003bf16af29a85924c3275846365918f664: real Spring/H2/Flyway, Java Store/connector, packaged CLI Harness, Broker and worker completed initial and recovered read/write/read turns. Graceful Harness close plus Java attachment eviction/load preserved files/1; NULL, empty and whitespace profiles were refused before Harness requests. A mixed-classpath old-connector control accepted all six invalid-profile cases, confirming the new guard is discriminating. Native report records the 36 candidate and 36 control assertions, SQL/file effects and cleanup. It explicitly preserves a failed full before-classpath capture; only the previously frozen 1,928-file subset is compared byte-for-byte, including all 1,370 dist files. This was a new native API run, not a new browser or tmux run. Production/UI code is unchanged from the earlier browser revision; only the MariaDB test fixture differs.Combined browser and live-runtime verification now passed at
c073f003bf16af29a85924c3275846365918f664using the committed Managed Workspace component host under Vite, with real Java/Store/Harness/Broker/worker execution. UI creation, reload/re-selection and graceful attachment recovery completed three read/write/read Turns and nine SETTLED executions with stored files/1. The run passed 36 evidence checks; its own complete before/after manifests match across 3,649 recorded files. The first intentionally held response crossed the fixture relay’s 90-second timeout and became visible only after reload; the normal second and third answers appeared live. Original screenshots are preserved, and the readable final capture changes only the fixture host body background; public attachment remains pending. This is candidate-only component-host evidence, with no production standalone-bundle, physical restart, vendor, Shell or /2 claim. Combined UI/runtime report.Fresh local verification also passed 36/36 checks with the executed command and zero exit statuses preserved by tmux capture-pane. Real Java/Store/Harness/Broker/worker execution completed initial and recovered file turns, retained files/1, and refused all six invalid-profile create/load cases before Harness access. The tmux test report includes captured terminal text, independently checked file effects and matching before/after artifact manifests. The model is a deterministic loopback fixture; the external-model attempt was blocked by automatic approval and was not executed.
Tested on
Environment (optional)
Java 21, Maven 3.9.11 and H2 in MySQL compatibility mode. The SQL migration was also verified against a disposable MariaDB 10.11 database, including all migrations through V34, the V35 upgrade, legacy rows and old-writer defaults. Native verification used Java 23 and Node 22.22 with an owned Spring/H2 fixture and live CLI/Broker/worker; it did not run the full Hosted Maven integration profile or a physical crash/reboot matrix.
Risk & Scope
/2admission and new approval behavior. The additional CI-aligned Qwen review remains unverified: all 10 dimension reports finished without a Critical finding, but the runner did not produce a final composed verdict. Local correctness and security reviews found no blocking issues.Design: English · 简体中文. Both versions include the same decisions, constraints, acceptance criteria and follow-up work.
Linked Issues
Part of #13271. The issue remains open for Shell admission and its prerequisite gates. The persisted field also supports the later admission work in #13166.
中文说明
本 PR 的改动
通过公开 API 或 WebShell 创建 Workspace Session 时,持久保存
hosted-workspace-files/1。创建、加载和恢复随后都使用该值。绑定会话的 profile 缺失或为空白时,在请求 Harness 前拒绝。这是 #13271 的第一阶段实现。Shell 和
/2准入仍留待后续完成。为什么需要
控制面目前每次附着会话时都会重新推导文件 profile。现在保存该选择,可以在后续引入其他 profile 之前,为已有文件会话确定持久身份。幂等创建重试保留原选择。
评审验证计划
验证方式
前后对照证据
Java 模块完整验证通过:557 项测试:556 通过、1 项因 macOS 平台跳过,0 失败、0 错误,包括编译、Checkstyle 和 SpotBugs。反向对照恢复旧的硬编码 profile 后,出现预期断言失败;恢复实现后,完整测试通过。针对 MariaDB CI 的后续修正调整了旧库测试数据构造;全部 52 项非 Hosted 集成用例已在本地通过,Checkstyle 通过。
浏览器验证已通过真实 WebShell + Spring/H2 测试环境:创建绑定会话、刷新后加载同一会话,并检查桌面和移动端布局。SQL 确认创建及刷新后存储的 profile 都为 files/1。测试报告,截图待附上。测试环境提供测试身份和 Workspace 数据,并关闭 Harness 执行;本次不验证真实 Harness 执行。
c073f003bf16af29a85924c3275846365918f664的真实原生验证现已通过:Spring/H2/Flyway、Java Store/connector、打包 CLI Harness、Broker 和 worker 完成首次及恢复后的真实读写读。先关闭 Harness 挂接,再清 Java 缓存并加载,仍使用存储的 files/1;NULL、空串及空白 profile 在请求 Harness 前被拒绝。只替换旧 connector 的混合 classpath 对照则放行全部六种非法 profile,确认新守卫能区分错误行为。原生报告记录 36 项候选及 36 项对照断言、SQL/文件效果和清理。完整 before classpath 捕获失败已明确保留;字节对比仅针对此前冻结的 1,928 个文件子集,其中包含全部 1,370 个 dist 文件。本次是新的原生 API 执行,没有重跑浏览器或 tmux;相对既有浏览器版本,生产/UI 代码不变,仅 MariaDB 测试数据构造不同。c073f003bf16af29a85924c3275846365918f664的浏览器与真实运行链路联合验证已完成:使用 Vite 提供仓库已有的 Managed Workspace 组件宿主页,后端实际经过 Java/Store/Harness/Broker/worker。UI 创建、刷新后重新选取同一会话、优雅分离后的挂接恢复共完成三次读写读 Turn、九次 SETTLED 执行,始终使用存储的 files/1。36 项证据检查通过,本轮独立的完整前后清单共 3,649 个文件且一致。首轮人为暂停跨过测试转发层的 90 秒超时,答案在刷新后显示;正常第二、第三轮答案实时可见。原图保留,可读的最终截图仅调整组件外的宿主背景,公开附件仍待上传授权。本次仅证明候选版本的组件宿主流程,不宣称生产独立 bundle、物理重启、真实厂商、Shell 或 /2 验收。联合 UI/运行链路报告。新的本地验证也已通过 36/36 项检查,tmux capture-pane 保留了实际执行命令和全部为零的退出码。真实 Java/Store/Harness/Broker/worker 完成首次及恢复后的文件执行,保留 files/1,并在请求 Harness 前拒绝全部六种非法 profile 创建/加载情况。tmux 测试报告包含终端原文、独立核对的文件效果和一致的前后产物清单。模型为本地确定性测试服务;外部模型测试被自动审批阻止,未执行。
测试系统
环境
Java 21、Maven 3.9.11,以及 MySQL 兼容模式的 H2。另在一次性 MariaDB 10.11 数据库中验证了 SQL 迁移,包括 V34 及之前的全部迁移、V35 升级、旧会话及旧版写入默认值。原生验证使用 Java 23、Node 22.22、自有 Spring/H2 环境及真实 CLI/Broker/worker;未运行完整 Hosted Maven 集成配置或物理崩溃/重启矩阵。
风险与范围
/2准入及新增审批行为。额外的 CI 对齐 Qwen 复审仍未验证:10 个维度的报告均已完成且没有 Critical,但工具未生成最终汇总结论。本地正确性与安全审查没有发现阻塞问题。设计:English · 简体中文。两版的决策、约束、验收标准及后续工作一致。
关联 Issue
属于 #13271 的一部分。该 issue 保持打开,继续跟进 Shell 准入及其前置门槛。持久化字段也供 #13166 后续准入工作使用。