Summary
Define what happens to long-running foreground and background shell processes when the managed server restarts.
Both shell process ownership and background job tracking are currently process-local. A restart interrupts that ownership, but foreground and background shells interact differently with Session restart continuity and need explicit, safe semantics.
Current behavior
Shell.Service owns running subprocess handles, output streams, timeout fibers, and waiters in one Location-scoped in-memory map. The process handle is scope-bound, and layer teardown clears the map and unblocks pending waiters.
Job.Service also tracks foreground/background state and completion through process-local maps, scopes, and Deferreds.
For a foreground shell:
- the tool call remains part of active Session execution;
- graceful managed shutdown suspends the active Session;
- teardown interrupts the shell/tool ownership chain;
- startup resumes the Session and reconciles the old running tool as interrupted rather than replaying it.
For a background shell:
- the shell tool has already returned a model-visible
running result;
- the Session may become idle while the subprocess continues;
- managed shutdown only snapshots active Session executions, so this Session may not receive a restart-continuation marker;
- teardown loses the shell and job registries, and the next server cannot observe the old PID, settle the shell, or deliver its completion through the Session inbox.
Output files may remain on disk, but they are not sufficient to establish process liveness, ownership, exit status, or completion.
Why automatic replay is unsafe
Shell commands have host-user filesystem, process, and network authority. A command interrupted after partial execution may already have produced side effects. Restarting the command from the beginning can duplicate writes, deployments, messages, payments, or other external actions.
Unexpected machine restart also guarantees that the original process is gone, while a managed-server-only restart could theoretically leave a detached process alive. The design should not infer either case from a stale PID without explicit ownership and identity guarantees.
Questions to decide
- On graceful managed-server restart, should OpenCode cancel running shells, wait for them, transfer/reattach ownership, or support more than one policy?
- Should foreground and background shells have different shutdown behavior?
- If a shell is interrupted, how is that terminal fact durably recorded and delivered to the Session?
- Should an idle Session with a background shell be included in restart continuation or receive a durable completion/interruption inbox entry independently?
- Can a surviving process be identified safely across server restart, including PID reuse and descendant process groups?
- What output remains readable after ownership is lost?
- What guarantees are possible on POSIX and Windows?
- Which semantics apply to graceful service restart, server crash, and machine restart?
- Should command replay ever be automatic, or only an explicit user/model decision after interruption is visible?
Required invariants
- Restart never silently reports an interrupted shell as still running.
- Restart never automatically repeats ambiguous shell side effects.
- Every shell that OpenCode stops during graceful shutdown receives a durable terminal/interrupted outcome.
- A background shell cannot disappear without a model-visible or user-visible terminal outcome.
- Completion or interruption is admitted durably before Session execution is scheduled.
- Stale PIDs cannot be adopted after PID reuse.
- Process-tree cleanup and orphan behavior are explicit per platform.
- Retained output accurately distinguishes complete, partial, truncated, and unavailable output.
Coverage
Cover at least:
- graceful restart during a foreground shell;
- graceful restart while an idle Session owns a background shell;
- shell exit racing with shutdown;
- restart after partial output;
- unexpected server death and machine restart;
- process-tree cleanup on POSIX and Windows;
- no duplicate command execution during Session continuation;
- durable interruption/completion delivery through the Session inbox.
Related
Summary
Define what happens to long-running foreground and background shell processes when the managed server restarts.
Both shell process ownership and background job tracking are currently process-local. A restart interrupts that ownership, but foreground and background shells interact differently with Session restart continuity and need explicit, safe semantics.
Current behavior
Shell.Serviceowns running subprocess handles, output streams, timeout fibers, and waiters in one Location-scoped in-memory map. The process handle is scope-bound, and layer teardown clears the map and unblocks pending waiters.Job.Servicealso tracks foreground/background state and completion through process-local maps, scopes, andDeferreds.For a foreground shell:
For a background shell:
runningresult;Output files may remain on disk, but they are not sufficient to establish process liveness, ownership, exit status, or completion.
Why automatic replay is unsafe
Shell commands have host-user filesystem, process, and network authority. A command interrupted after partial execution may already have produced side effects. Restarting the command from the beginning can duplicate writes, deployments, messages, payments, or other external actions.
Unexpected machine restart also guarantees that the original process is gone, while a managed-server-only restart could theoretically leave a detached process alive. The design should not infer either case from a stale PID without explicit ownership and identity guarantees.
Questions to decide
Required invariants
Coverage
Cover at least:
Related
specs/v2/session-restart-continuation.mdexplicitly excludes exactly-once provider or tool execution.