Repository navigation
Track activeWork and background Agent recovery #8586
Description
Activity
Design change in layer 1, and why the plan is now five PRs
Review of #8588 surfaced a structural problem with the original layer-1 design, so the issue body above has been revised. Recording the reasoning here so the change is auditable.
What the original design said
Layer 1 combined three jobs in one mechanism: a per-Session boolean heartbeat that (a) fed activeWork, (b) gated Session retention, and (c) recycled the owning ACP channel after a missed heartbeat window.
Why that does not hold
The three consumers need opposite failure directions. For health reporting and for retention, "I did not hear anything" must mean assume busy, do not touch it. For a watchdog, the same silence meant destroy the channel. One signal cannot fail closed and fail open at the same time, and the destructive reading had the largest blast radius: recycling a channel kills every Session on that process, not the one that went quiet.
The watchdog was detecting the wrong thing anyway. A missed work report is not evidence of a dead process. A host suspend, a long synchronous stall, or a single dropped notification all present identically to a wedged child — and the genuine target (an Agent whose logic stops progressing while the process keeps answering) was never detectable this way. So it caught what it should not have owned and missed what it was for.
Retention was resting on an unreconcilable signal. Boolean transitions with no acknowledgement mean one lost message desyncs the daemon permanently in one direction. Because only the active state was re-asserted, a lost "now idle" message left the daemon pinned busy forever and armed the watchdog to kill the channel — a single dropped write escalating into destroyed work.
What replaces it
- Full snapshots at channel scope, derived not maintained. Every report carries the complete hold set for every Session on the channel, computed on demand from the owners of the work (the registry's unfinalized set, the notification queue, in-flight acceptance/continuation state). No acquire/release ledger, so a hold cannot leak past the work it names. A dropped report costs one interval and self-corrects in both directions; a Session absent from a snapshot is positive evidence the child released it.
- Confirmation before destruction. The cache decides when it is worth asking; it never authorizes a close. The child answers under its own close gate, which makes the check atomic — with the gate held no prompt is admitted and no automatic turn starts, so a hold cannot appear between the check and the teardown. This closes a TOCTOU that a read-only "are you idle?" query could not.
- No channel kill driven by work state. Transport liveness becomes its own layer with its own timeout, its own escalation, and a monotonic-clock check so suspend/resume does not read as a stall. That is the new PR 2, which is why the plan grew from four PRs to five.
- A trust grade on the health signal.
activeWorkReporting(full/partial/none) plusactiveWorkStaleMs. Without it a controller cannot distinguish "nothing is running" from "nobody told me what is running", and the old two-term busy rule silently trusted a degraded signal. The busy rule in the issue body now has three terms.
A concrete bug this fixed
Agent holds are keyed on hasUnfinalizedTasks(), not hasRunningTasks(). cancel() flips an Agent to cancelled and emits a status change immediately, but its terminal notification only arrives later from finalizeCancelled() or the 5s grace timer. Keying on "running" made the Session look idle for that whole window — and because the daemon closed detached Sessions the moment they read idle, a cancelled Agent's terminal notification could be stranded and its parent continuation never run. This is now an explicit acceptance criterion.
Honest limits
Health remains an observation cache, not a restart lease. A fresh, fully-graded, empty answer still only describes the sampling instant, and work can start immediately after. The three-term rule lowers the risk of a wrong restart substantially; it does not eliminate it. Anything requiring a hard guarantee needs a prepare-restart fence — stop new work admission, confirm the drain, then shut down — which stays out of scope here.
Nothing in layers 3–5 changes in substance; they are renumbered only because transport liveness became its own PR.
Background shell follow-up is now available as draft PR #9042.
This intentionally widens the layer-1 scope that the original issue body excluded: Session-managed background shells now contribute one bounded shell hold while the registry is running or the terminal notification is queued/in continuation. The v1 capability now negotiates categories explicitly, so a negotiated older child that omits shell reports partial coverage and is excluded from ordinary automatic cleanup, while completely unsupported children keep legacy behavior.
The released CLI baseline reproduced activePrompts: 0, activeWork: false, and full reporting while a background sleep 120 task was still running. The local build returned activeWork: true for the same state and retained it through terminal notification settlement. This PR references rather than closes #8586 because transport liveness, logical watchdogs, and runtime recycle remain outstanding.
Status cross-check against current main:
- Layer 1 is merged through feat(serve): Expose active work state #8588, feat(daemon): Track background shells in activeWork #9042, fix(daemon): Preserve sessions when active-work close is refused #9134, and fix(daemon): Bound conditional-close refusal holds #9820. fix(serve): bound and diagnose a session reclaim that can never succeed #11120 additionally bounds and diagnoses conditional-close probes that cannot settle, but fix(serve): a session doing cron, goal, monitor or history-mutation work can never be reclaimed #11118 identifies one remaining Layer-1 semantic mismatch: child-owned goal/cron/monitor/history activity can block settlement without appearing in the reported hold set. That needs the category/scope decision recorded there; it is not solved by retry backoff.
- Layer 2 is merged through feat(daemon): Add ACP channel transport liveness #9976.
- Layers 3–5 still have no implementation PRs linked here: logical background-Agent watchdogs, the reusable active/draining/dying runtime-generation primitive, and escalation for an Agent that ignores cooperative abort.
- serve: background shell output and wake notifications silently dropped when the session runtime recycles, wedging the session #11119 should not currently be treated as proof that a runtime-generation route drops every background-shell output chunk. Background shell stdout is intentionally file-backed and is not appended continuously to the transcript. fix(core): replay retained shell notifications after transient unbind #11132 only fixes same-registry terminal-notification redelivery after a transient callback absence. The confirmed part that belongs in this umbrella is that per-session live state still does not expose the existing shell/active-work hold, so background work can look idle in the session list.
Accordingly, this umbrella is correctly still open. The remaining implementation order stays: settle the Layer-1 scope mismatch, then Layer 3 watchdogs, Layer 4 generation draining, and Layer 5 non-cooperative escalation. The runtime-generation work should not be skipped in favor of a speculative shell-specific on-disk ledger.
Layer 1 follow-up is now implemented in PR #11265. It adds one negotiated session hold for goal/cron/history-mutation/Monitor-continuation work so the conditional-close authorization predicate matches the Session drain predicate without cancelling child-owned work. This closes the remaining Layer 1 semantic mismatch tracked in #11118 if merged. Layers 3–5 remain separate work; #11265 does not claim to implement watchdogs, runtime-generation draining, or non-cooperative Agent escalation.
Layer 3 is now implemented in PR #11270. It adds fixed model/control and per-tool progress deadlines to fresh, restored, and resident-continuation background Agents; approval and Monitor external-input waits pause the relevant deadline, and cooperative stalls settle once as TIMEOUT / failed without workflow-style retry. The PR is stacked on #11265. Layers 4–5 (runtime-generation draining and non-cooperative escalation) remain separate.
Layer 4 is now implemented in PR #11273. It reuses the existing ACP bridge ownership and retirement machinery, adds explicit active/draining/dying generation state, preserves each live Session on its owner generation, hard-caps the bridge at two OS-live generations, and returns retryable 503 runtime_recycling instead of creating a third. The PR is stacked on #11270. Layer 5 remains separate and will connect abort-ignoring Agent escalation to the recycle primitive and protected record-only terminal settlement.
The final watchdog-escalation layer is now in #11275, stacked on #11273. It handles an Agent that remains unsettled after the logical watchdog abort: record one failed terminal state, preserve the occupied physical slot until late settlement, persist/display the notification without starting another model turn, and recycle only the Session owner generation. With the dependency stack merged, this completes the implementation described in this issue. This does not claim the separate cross-generation background-shell route-loss case in #11119 is fixed.
PR #11270 is being closed to park its combined watchdog/runtime-generation implementation. The background-agent stall and recovery problem remains open here. #11206's collaboration watchdog is not a substitute for verifying ordinary background-agent recovery. Future fixes should use a concrete reproduction and a narrower scope; the closed PR retains its design and review history for reference.
What would you like to be added?
Add an explicit
activeWorkfact to deep daemon health and build the recovery path for background Agents that outlive their foreground prompt or stop making progress.The final behavior should cover five layers:
GET /health?deep=1reportsactiveWork, including accepted unsettled prompts, running background Agents, Session-managed background shells, and pending Agent or shell terminal notifications. It also reports how far that boolean can be trusted, so an unreported channel is never mistaken for an idle one. Supported ACP children publish channel-wide full snapshots of named holds on a negotiated cadence, and automatic cleanup confirms with the child before destroying a Session.TIMEOUT/failedwithout retry.503 runtime_recyclingwhile capacity is unavailable.The work should be delivered as five reviewable stacked PRs in that order. It intentionally does not add
activeBackgroundTasks,restartSafe, public timeout configuration, or a persistence migration.Scope of
activeWorkactiveWorkcovers accepted-but-unsettled prompts, running background Agents, Agent terminal notifications that are queued, awaiting acceptance, or being processed by a parent continuation, and Session-managed background shells while they are running, their terminal notifications are queued, or those notifications are driving a parent continuation.It deliberately excludes Monitors, workflows, cron, follow-up work, and processes detached from Session management. Those categories have no equivalent signal today, so a controller must not read
activeWork: falseas "nothing at all is running". Hold reports carry a category so this scope can be widened later by adding data rather than by silently changing what the boolean means.Why is this needed?
activePromptsonly represents foreground prompt execution. A prompt can launch background Agents and then settle, producingactivePrompts=0while useful work is still running. Restart controllers that rely on that count can therefore classify the daemon as idle and restart it, interrupting background work and leaving Sessions to recover from disk.Process liveness alone is also insufficient: an ACP child can stop responding, or an Agent can remain alive while its asynchronous model/tool logic makes no progress. These are three distinct failures — work retention, transport liveness, and logical progress — and conflating them is itself a source of bugs: inferring "this channel is dead" from "one Session stopped reporting" destroys every Session on that process, and a host suspend, a long event-loop stall, or a single dropped message all look identical to a stalled child. Each failure therefore gets its own mechanism.
The restart controller should fail closed and use:
The third term is required. Without it,
activeWork === falsecannot be distinguished from "no child ever reported", which is the one case where acting on the value is unsafe. Unknown health responses, probe failures, busy state, or insufficient idle grace must also prevent restart.Note that health remains an observation cache, not a restart lease: even a fresh, fully-graded, empty answer describes only the moment it was sampled. This rule substantially lowers the risk of a wrong restart but does not eliminate it; strict safety needs a prepare-restart fence that stops new work admission, confirms the drain, and only then shuts down.
Acceptance criteria
activePrompts=0whileactiveWorkremains true for running background Agents, Session-managed background shells, and their pending terminal continuations.activeWorkbecomes false only after the last relevant Agent or Session-managed shell terminal notification has fully settled.TIMEOUT/failedterminal result with no retry.activeBackgroundTasksorrestartSafeand does not add timeout configuration or migrate persisted data.Additional context
This issue is the umbrella tracker for the staged implementation. Each PR should remain behaviorally coherent at its merge point: