Skip to content

Track activeWork and background Agent recovery #8586

Description

@doudouOUC

What would you like to be added?

Add an explicit activeWork fact to deep daemon health and build the recovery path for background Agents that outlive their foreground prompt or stop making progress.

The final behavior should cover five layers:

  1. Deep health and ACP Session reporting: GET /health?deep=1 reports activeWork, including accepted unsettled prompts, running background Agents, Session-managed background shells, and pending Agent or shell terminal notifications. It also reports how far that boolean can be trusted, so an unreported channel is never mistaken for an idle one. Supported ACP children publish channel-wide full snapshots of named holds on a negotiated cadence, and automatic cleanup confirms with the child before destroying a Session.
  2. Transport and process liveness: a wedged ACP child, event loop, or transport is detected by a dedicated channel-level ping-pong with its own timeout and escalation policy, independent of any Session's work state.
  3. Logical background-Agent watchdogs: ordinary fresh and resumed Agents use fixed model/control and per-tool progress deadlines, while approval and external-input waits pause the relevant deadline. Cooperative stall aborts settle as TIMEOUT/failed without retry.
  4. Runtime generation draining: runtime recycle uses explicit active/draining/dying generations, preserves Session ownership, prevents a third live generation, and rejects new work with 503 runtime_recycling while capacity is unavailable.
  5. Unresponsive-Agent escalation: an Agent that ignores the stall abort retains its physical concurrency slot, emits a protected record-only terminal notification, and requests runtime recycle without allowing a late settle to overwrite its failed sidecar.

The work should be delivered as five reviewable stacked PRs in that order. It intentionally does not add activeBackgroundTasks, restartSafe, public timeout configuration, or a persistence migration.

Scope of activeWork

activeWork covers accepted-but-unsettled prompts, running background Agents, Agent terminal notifications that are queued, awaiting acceptance, or being processed by a parent continuation, and Session-managed background shells while they are running, their terminal notifications are queued, or those notifications are driving a parent continuation.

It deliberately excludes Monitors, workflows, cron, follow-up work, and processes detached from Session management. Those categories have no equivalent signal today, so a controller must not read activeWork: false as "nothing at all is running". Hold reports carry a category so this scope can be widened later by adding data rather than by silently changing what the boolean means.

Why is this needed?

activePrompts only represents foreground prompt execution. A prompt can launch background Agents and then settle, producing activePrompts=0 while useful work is still running. Restart controllers that rely on that count can therefore classify the daemon as idle and restart it, interrupting background work and leaving Sessions to recover from disk.

Process liveness alone is also insufficient: an ACP child can stop responding, or an Agent can remain alive while its asynchronous model/tool logic makes no progress. These are three distinct failures — work retention, transport liveness, and logical progress — and conflating them is itself a source of bugs: inferring "this channel is dead" from "one Session stopped reporting" destroys every Session on that process, and a host suspend, a long event-loop stall, or a single dropped message all look identical to a stalled child. Each failure therefore gets its own mechanism.

The restart controller should fail closed and use:

const busy =
  health.activePrompts > 0 ||
  health.activeWork ||
  health.activeWorkReporting !== 'full';

The third term is required. Without it, activeWork === false cannot be distinguished from "no child ever reported", which is the one case where acting on the value is unsafe. Unknown health responses, probe failures, busy state, or insufficient idle grace must also prevent restart.

Note that health remains an observation cache, not a restart lease: even a fresh, fully-graded, empty answer describes only the moment it was sampled. This rule substantially lowers the risk of a wrong restart but does not eliminate it; strict safety needs a prepare-restart fence that stops new work admission, confirms the drain, and only then shuts down.

Acceptance criteria

  • A foreground prompt may settle with activePrompts=0 while activeWork remains true for running background Agents, Session-managed background shells, and their pending terminal continuations.
  • activeWork becomes false only after the last relevant Agent or Session-managed shell terminal notification has fully settled.
  • An Agent that has been cancelled but has not yet emitted its terminal notification still counts as active work, so a detached Session is never reaped inside the cancel-to-finalize window.
  • Capability negotiation is backward compatible; unsupported ACP children keep the previous behavior and are reported as unvouched rather than as idle.
  • A dropped or reordered report is self-correcting: reports are complete snapshots, so the next one restores the truth in either direction without retransmit or acknowledgement.
  • A Session absent from a fresh snapshot is treated as released by the child.
  • No automatic cleanup destroys a Session on the strength of a cached report alone; the child confirms under its own close gate, and an unanswered confirmation is neither retried nor assumed.
  • Transport liveness never infers channel death from one Session's reporting, and is graded against a monotonic clock so host suspend does not present as a stall.
  • Cooperative logical stalls become one TIMEOUT/failed terminal result with no retry.
  • Approval and Monitor external-input waits do not cause false watchdog failures.
  • Abort-ignoring Agents cannot free their physical concurrency slot or overwrite the failed terminal state when they settle late.
  • Runtime draining preserves owner-generation routing and never creates more than two live generations.
  • Existing workflow stall/retry behavior remains unchanged.
  • The daemon does not expose activeBackgroundTasks or restartSafe and does not add timeout configuration or migrate persisted data.

Additional context

This issue is the umbrella tracker for the staged implementation. Each PR should remain behaviorally coherent at its merge point:

Activity

self-assigned this
on Aug 5, 2026

doudouOUC commented on Aug 6, 2026

@doudouOUC
CollaboratorAuthor

Design change in layer 1, and why the plan is now five PRs

Review of #8588 surfaced a structural problem with the original layer-1 design, so the issue body above has been revised. Recording the reasoning here so the change is auditable.

What the original design said

Layer 1 combined three jobs in one mechanism: a per-Session boolean heartbeat that (a) fed activeWork, (b) gated Session retention, and (c) recycled the owning ACP channel after a missed heartbeat window.

Why that does not hold

The three consumers need opposite failure directions. For health reporting and for retention, "I did not hear anything" must mean assume busy, do not touch it. For a watchdog, the same silence meant destroy the channel. One signal cannot fail closed and fail open at the same time, and the destructive reading had the largest blast radius: recycling a channel kills every Session on that process, not the one that went quiet.

The watchdog was detecting the wrong thing anyway. A missed work report is not evidence of a dead process. A host suspend, a long synchronous stall, or a single dropped notification all present identically to a wedged child — and the genuine target (an Agent whose logic stops progressing while the process keeps answering) was never detectable this way. So it caught what it should not have owned and missed what it was for.

Retention was resting on an unreconcilable signal. Boolean transitions with no acknowledgement mean one lost message desyncs the daemon permanently in one direction. Because only the active state was re-asserted, a lost "now idle" message left the daemon pinned busy forever and armed the watchdog to kill the channel — a single dropped write escalating into destroyed work.

What replaces it

  • Full snapshots at channel scope, derived not maintained. Every report carries the complete hold set for every Session on the channel, computed on demand from the owners of the work (the registry's unfinalized set, the notification queue, in-flight acceptance/continuation state). No acquire/release ledger, so a hold cannot leak past the work it names. A dropped report costs one interval and self-corrects in both directions; a Session absent from a snapshot is positive evidence the child released it.
  • Confirmation before destruction. The cache decides when it is worth asking; it never authorizes a close. The child answers under its own close gate, which makes the check atomic — with the gate held no prompt is admitted and no automatic turn starts, so a hold cannot appear between the check and the teardown. This closes a TOCTOU that a read-only "are you idle?" query could not.
  • No channel kill driven by work state. Transport liveness becomes its own layer with its own timeout, its own escalation, and a monotonic-clock check so suspend/resume does not read as a stall. That is the new PR 2, which is why the plan grew from four PRs to five.
  • A trust grade on the health signal. activeWorkReporting (full/partial/none) plus activeWorkStaleMs. Without it a controller cannot distinguish "nothing is running" from "nobody told me what is running", and the old two-term busy rule silently trusted a degraded signal. The busy rule in the issue body now has three terms.

A concrete bug this fixed

Agent holds are keyed on hasUnfinalizedTasks(), not hasRunningTasks(). cancel() flips an Agent to cancelled and emits a status change immediately, but its terminal notification only arrives later from finalizeCancelled() or the 5s grace timer. Keying on "running" made the Session look idle for that whole window — and because the daemon closed detached Sessions the moment they read idle, a cancelled Agent's terminal notification could be stranded and its parent continuation never run. This is now an explicit acceptance criterion.

Honest limits

Health remains an observation cache, not a restart lease. A fresh, fully-graded, empty answer still only describes the sampling instant, and work can start immediately after. The three-term rule lowers the risk of a wrong restart substantially; it does not eliminate it. Anything requiring a hard guarantee needs a prepare-restart fence — stop new work admission, confirm the drain, then shut down — which stays out of scope here.

Nothing in layers 3–5 changes in substance; they are renumbered only because transport liveness became its own PR.

doudouOUC commented on Aug 13, 2026

@doudouOUC
CollaboratorAuthor

Background shell follow-up is now available as draft PR #9042.

This intentionally widens the layer-1 scope that the original issue body excluded: Session-managed background shells now contribute one bounded shell hold while the registry is running or the terminal notification is queued/in continuation. The v1 capability now negotiates categories explicitly, so a negotiated older child that omits shell reports partial coverage and is excluded from ordinary automatic cleanup, while completely unsupported children keep legacy behavior.

The released CLI baseline reproduced activePrompts: 0, activeWork: false, and full reporting while a background sleep 120 task was still running. The local build returned activeWork: true for the same state and retained it through terminal notification settlement. This PR references rather than closes #8586 because transport liveness, logical watchdogs, and runtime recycle remain outstanding.

yiliang114 commented on Sep 7, 2026

@yiliang114
Collaborator

Status cross-check against current main:

Accordingly, this umbrella is correctly still open. The remaining implementation order stays: settle the Layer-1 scope mismatch, then Layer 3 watchdogs, Layer 4 generation draining, and Layer 5 non-cooperative escalation. The runtime-generation work should not be skipped in favor of a speculative shell-specific on-disk ledger.

yiliang114 commented on Sep 7, 2026

@yiliang114
Collaborator

Layer 1 follow-up is now implemented in PR #11265. It adds one negotiated session hold for goal/cron/history-mutation/Monitor-continuation work so the conditional-close authorization predicate matches the Session drain predicate without cancelling child-owned work. This closes the remaining Layer 1 semantic mismatch tracked in #11118 if merged. Layers 3–5 remain separate work; #11265 does not claim to implement watchdogs, runtime-generation draining, or non-cooperative Agent escalation.

yiliang114 commented on Sep 7, 2026

@yiliang114
Collaborator

Layer 3 is now implemented in PR #11270. It adds fixed model/control and per-tool progress deadlines to fresh, restored, and resident-continuation background Agents; approval and Monitor external-input waits pause the relevant deadline, and cooperative stalls settle once as TIMEOUT / failed without workflow-style retry. The PR is stacked on #11265. Layers 4–5 (runtime-generation draining and non-cooperative escalation) remain separate.

yiliang114 commented on Sep 7, 2026

@yiliang114
Collaborator

Layer 4 is now implemented in PR #11273. It reuses the existing ACP bridge ownership and retirement machinery, adds explicit active/draining/dying generation state, preserves each live Session on its owner generation, hard-caps the bridge at two OS-live generations, and returns retryable 503 runtime_recycling instead of creating a third. The PR is stacked on #11270. Layer 5 remains separate and will connect abort-ignoring Agent escalation to the recycle primitive and protected record-only terminal settlement.

yiliang114 commented on Sep 7, 2026

@yiliang114
Collaborator

The final watchdog-escalation layer is now in #11275, stacked on #11273. It handles an Agent that remains unsettled after the logical watchdog abort: record one failed terminal state, preserve the occupied physical slot until late settlement, persist/display the notification without starting another model turn, and recycle only the Session owner generation. With the dependency stack merged, this completes the implementation described in this issue. This does not claim the separate cross-generation background-shell route-loss case in #11119 is fixed.

yiliang114 commented on Sep 25, 2026

@yiliang114
Collaborator

PR #11270 is being closed to park its combined watchdog/runtime-generation implementation. The background-agent stall and recovery problem remains open here. #11206's collaboration watchdog is not a substitute for verifying ordinary background-agent recovery. Future fixes should use a concrete reproduction and a narrower scope; the closed PR retains its design and review history for reference.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions