|
| 1 | +# Hosted Turn failover E2E and Harness Turn takeover (G1) |
| 2 | + |
| 3 | +[English](2026-09-30-hosted-turn-failover-e2e.md) | [简体中文](2026-09-30-hosted-turn-failover-e2e.zh-CN.md) |
| 4 | + |
| 5 | +Status: proposed. Implements slice G1 of #12952 (part of #12380), after G0 |
| 6 | +(#12955). |
| 7 | + |
| 8 | +## Problem and scope |
| 9 | + |
| 10 | +`--inflight-failover` and `--continuation-failover` in |
| 11 | +`scripts/run-managed-agent-server-e2e.ts` exit immediately with a |
| 12 | +`not-yet-enabled` throw whose stated reason is stale. G0 opened public |
| 13 | +admission for Workspace-bound file-tool Sessions, but running the modes against |
| 14 | +the packaged stack shows the deeper gap: **the Hosted Harness has no Turn |
| 15 | +takeover**. The Java coordinator and SDK already carry the full recovery |
| 16 | +contract — the `_meta.qwen.daemon.managedRuntimeRecovery` load payload, the |
| 17 | +`POST /session/:id/managed-runtime/continue` and `/cancel` routes, per-epoch |
| 18 | +output retraction — while the TypeScript Harness refuses a loaded Session with |
| 19 | +unsettled input (`hosted_turn_recovery_required`) and implements none of that |
| 20 | +contract. |
| 21 | + |
| 22 | +G1 therefore delivers: |
| 23 | + |
| 24 | +1. **Harness-side Turn takeover.** A replacement Harness loads a Session whose |
| 25 | + Turn parked at an `await_runtime` / `results_ready` checkpoint, reconciles |
| 26 | + and settles its Runtime executions without replay, reports the recovery |
| 27 | + snapshot, and continues the Turn on coordinator request. |
| 28 | +2. **Durable streamed text.** Assistant text commits to the journal as |
| 29 | + activation-scoped delta events while the model streams, so a published |
| 30 | + prefix survives the writer and a dead owner's prefix is retracted from the |
| 31 | + public projection by the existing epoch machinery. |
| 32 | +3. **Admission enabler.** The packaged server gains an explicit opt-in |
| 33 | + (`qwen.managed-agent.trusted-actor-header`, default empty) that stands in |
| 34 | + for the trusted gateway's `AuthenticatedTenantActor`, so the E2E can create |
| 35 | + Workspace-bound Sessions through the public route. Never enable it where |
| 36 | + untrusted clients can reach the server. |
| 37 | +4. **Ungated modes and CI.** Both modes lose the stale gate, drive their |
| 38 | + physical tool through `write_file` (the publicly admitted profile), run in |
| 39 | + the Hosted MySQL CI job, and the README describes them as runnable. |
| 40 | + |
| 41 | +Out of scope: public Shell profile selection, the G2 takeover reconciliation |
| 42 | +scan, G3 owner-affinity removal, and multi-Runtime takeover. Cancellation |
| 43 | +takeover is implemented for contract completeness but has no E2E mode. |
| 44 | + |
| 45 | +## Current state |
| 46 | + |
| 47 | +- The runner already contains the complete in-flight/continuation scenarios |
| 48 | + (held Broker `:start` proxy, fake model, process kills, home deletion, |
| 49 | + audits); they were gated by #12801 without ever passing. |
| 50 | +- `ManagedSessionJournalStore` writer fencing, `await_runtime` checkpoints and |
| 51 | + `resolveAwaitRuntime` exist in core; the `--session-failover` durable-owner |
| 52 | + proof passes. |
| 53 | +- Java side: `HostedHarnessClient` parses the recovery `_meta`, calls |
| 54 | + continue/cancel, and `HarnessCoordinator` retracts a dead owner's public |
| 55 | + deltas, binds the replacement Harness generation and records recovery |
| 56 | + admissions. |
| 57 | +- Missing (TS): recovery snapshot on load, continue/cancel routes, exactly-once |
| 58 | + re-dispatch of parked executions, and mid-Turn durable text deltas. |
| 59 | +- Missing (packaged server): any way to supply an actor principal, so public |
| 60 | + Workspace creation returns 401 outside in-JVM tests. |
| 61 | + |
| 62 | +## Decisions |
| 63 | + |
| 64 | +- **Load reports or completes the parked Runtime work; it never replays the |
| 65 | + model Prompt.** On load with unsettled input, a runnable `await_runtime` or |
| 66 | + `results_ready` checkpoint and a Workspace tool profile, the Harness resolves |
| 67 | + every in-progress execution against the Broker by its original |
| 68 | + `executionCallId`. A continuation load re-dispatches each parked execution |
| 69 | + through `execute` — the Broker's durable record makes that exactly-once — |
| 70 | + commits the tool result, and advances the checkpoint to `results_ready` |
| 71 | + before answering. A passive load (the coordinator's cancellation path) only |
| 72 | + reads execution status and reports `known`/`unknown` without dispatching. |
| 73 | + Executions the Broker cannot account for report `unknown`, the coordinator |
| 74 | + blocks the Turn as `managed_runtime_recovery_blocked`, and nothing replays. |
| 75 | +- **Continue runs the model from `results_ready`; cancel settles without new |
| 76 | + work.** `managed-runtime/continue` validates the prompt, checkpoint and |
| 77 | + activation identities, admits the continuation with a 200 receipt, then |
| 78 | + re-issues the model request with the journaled tool results and runs the |
| 79 | + normal tool-capable loop to a terminal record. `managed-runtime/cancel` |
| 80 | + best-effort cancels parked executions and settles the Turn as cancelled. |
| 81 | + Both refuse identity mismatches with 409. |
| 82 | +- **Assistant text streams as a new `message.delta` journal event.** The kind |
| 83 | + joins the closed v1 set with `harness` actor and activation subject, so a |
| 84 | + fenced former owner cannot append. Deltas carry the final message's |
| 85 | + pre-assigned `messageId`; `message.committed` records whose deltas already |
| 86 | + project do not project a second chunk, so from-scratch projections stay |
| 87 | + exact. Every model chunk commits immediately — a published prefix must be |
| 88 | + durable while its stream is still open — split only at the journal's text |
| 89 | + limit. |
| 90 | +- **Runtime takeover needs the W0e reclaim, so the modes are Linux-only |
| 91 | + today.** The replacement owner must retire the dead worker's Runtime binding |
| 92 | + before it can re-drive the parked execution. That reclaim exists only for |
| 93 | + the durable local-process provisioner, whose trusted host identity is |
| 94 | + Linux-only. The runner enables `durable-local-process` for these modes, |
| 95 | + creates the state directory owner-private, and refuses other platforms with |
| 96 | + an explicit error pointing at the Hosted MySQL CI job. This is the boundary |
| 97 | + #12766 already records; G1 does not change it. |
| 98 | +- **The E2E seeds Workspace admission as deployment data** (registry row, |
| 99 | + access grant, Broker mount) and enables `harness.workspace-files-enabled` |
| 100 | + plus the trusted-actor header on both Spring owners. The physical side effect |
| 101 | + is a fixed `write_file`; the exactly-once assertions ride the durable |
| 102 | + execution row, dispatch generation and model-request counts, not file bytes. |
| 103 | + `--session-failover` stays unbound and unchanged. |
| 104 | +- **Both modes join the `hosted-harness-mysql` job**, which installs the MySQL |
| 105 | + server binaries the runner needs for its private `mysqld`. |
| 106 | + |
| 107 | +## Changes and ownership |
| 108 | + |
| 109 | +| Layer | Change | Scope | |
| 110 | +| ------------ | ------------------------------------------------------------------------------- | ------------------------- | |
| 111 | +| Core journal | `message.delta` event kind (schema, harness actor, activation subject) | Managed Session log | |
| 112 | +| CLI Harness | Recovery snapshot + settlement on load; continue/cancel routes; delta streaming | Hosted Harness sessions | |
| 113 | +| CLI Broker | `status` read for passive reports | Workspace Broker | |
| 114 | +| Java API | `TrustedActorHeaderFilter` + property, default off | Deployment opt-in | |
| 115 | +| E2E runner | Ungate; Workspace seeding, mounts and actor wiring; `write_file` side effect | Local and CI verification | |
| 116 | +| CI workflow | MySQL binaries + both failover modes in `hosted-harness-mysql` | Hosted MySQL job | |
| 117 | +| README | Runnable modes; property risk note | Managed Agent Server docs | |
| 118 | + |
| 119 | +## Validation and acceptance |
| 120 | + |
| 121 | +Unit tests: the delta event schema round-trip; recovery meta shape and refusal |
| 122 | +paths; continue/cancel identity and phase gates; the actor filter. The G0 |
| 123 | +Hosted integration test and the Hosted MySQL suites keep passing. |
| 124 | + |
| 125 | +The E2E exit checks are the issue's own: |
| 126 | + |
| 127 | +- **In-flight:** the first Broker `:start` is held after the durable |
| 128 | + `await_runtime` checkpoint; both process trees die and their homes are |
| 129 | + deleted; the replacement reuses the original `executionCallId`, executes the |
| 130 | + physical tool exactly once, continues the Prompt without replay (one initial |
| 131 | + and one continuation model request), and commits one terminal event. |
| 132 | +- **Continuation:** the Harness dies after the first published text chunk; the |
| 133 | + tool executed once and settled beforehand; the replacement issues one further |
| 134 | + continuation; the public transcript shows only the replacement's answer (the |
| 135 | + retracted prefix is blank); the Turn has one terminal event. |
| 136 | + |
| 137 | +Deleting any assertion must fail its mode. `--session-failover` and the |
| 138 | +real-model check rerun as regressions. Build, typecheck, bundle, focused tests |
| 139 | +and two clean diff audits precede completion. |
| 140 | + |
| 141 | +## Boundaries and open questions |
| 142 | + |
| 143 | +G1 lives under #12952. The issue's slice text assumed the takeover machinery |
| 144 | +already existed; this design records that the Harness half is the bulk of the |
| 145 | +work. Whether delta streaming should later apply to no-tool Hosted Turns, and |
| 146 | +how multi-Runtime recovery composes with the G2 scan, stay open. The |
| 147 | +trusted-actor property is a local/E2E stand-in; a real gateway replaces it |
| 148 | +without touching admission logic. |
| 149 | + |
| 150 | +Known follow-ups from the maintainer's real-environment verification: |
| 151 | + |
| 152 | +- A takeover load whose reply is lost leaves the Turn stuck: the Harness has |
| 153 | + attached the Session and answers every later load with 409 |
| 154 | + `hosted_session_already_attached`, so the recovery snapshot cannot be |
| 155 | + fetched again. The takeover load needs to become idempotent. Note the load |
| 156 | + can legitimately take up to 120 s while the default coordinator |
| 157 | + `request-timeout` is 30 s. |
| 158 | +- A journal that contains `message.delta` events cannot be opened by a |
| 159 | + Harness of an older build (`managed_session_open_failed`). Readers of this |
| 160 | + build are fine; a rollback or a mixed fleet during a rolling deploy is not. |
| 161 | + Upgrade the fleet before enabling Hosted Workspace turns, or gate rollback. |
| 162 | +- A model fallback or retry that lands after the first streamed chunk fails |
| 163 | + the Turn terminally (`Hosted Harness cannot retract a published model |
| 164 | +attempt.`): once `message.delta` records are journaled, the partial attempt |
| 165 | + is public and cannot be withdrawn, so the Turn settles as `error` instead |
| 166 | + of discarding the partial attempt and retrying, which is what the |
| 167 | + pre-streaming behavior did. A transient provider capacity event mid-stream |
| 168 | + therefore fails the Turn permanently rather than being retried by the |
| 169 | + coordinator (a `turn_result` is terminal). Classifying this settlement as |
| 170 | + retryable for the coordinator is a follow-up. |
0 commit comments