Skip to content

cli(win): 45s event-stream idle watchdog restarts the managed service, aborting all sessions/subagents #52049

Description

@wurenrumian

Description

On Windows the v2 managed background service (opencode serve --service) is repeatedly killed and respawned during normal use, aborting every in-flight session and background subagent ("failed"). It is not a crash — the client is killing it.

Every restart is preceded by the client log line event stream disconnected ... error="Event stream stalled", then ~6s later a new cli starting ... args=["serve","--service"].

The direct trigger is the client's event-stream idle watchdog (values read from the shipped 2.0.19 binary):

  • server sends : heartbeat on /api/event every 15s;
  • client resets a 45s idle timer on any received byte (idleTimeout ?? 45000);
  • after 45s with no bytes the client aborts with Event stream stalled, then treats the service as unresponsive and replaces it (Background service is unresponsive; recovery cannot preserve persistent terminals, Failed to shut down persistent terminals before service replacement) instead of reconnecting.

Measured correlation: server's last log line 08:47:12.749, stall logged 08:48:40.688 → 46.1s, i.e. the 45s watchdog, not the ~6s health probe.

This is the same class as #50594 (Windows: file-watcher re-registration storm on ~/.claude/skills → serve --service disappears with no error/dump) and #50591 (Linux: Event stream stalled + InterruptError on /api/event). The new information here is that the service is not dying on its own: the client kills it after a 45s heartbeat gap, so the "silent death" in #50594 is a client-side replacement, and the fix must also cover the watchdog margin / the replace-instead-of-reconnect policy.

Ruled out on this machine:

  • Not a crash — no opencode.exe APPCRASH in the Windows Application log; no crash dumps under %LOCALAPPDATA%\CrashDumps or WER (no RADAR_PRE_LEAK_64 either).
  • Not OOM — 23.7 GB RAM, 6.7 GB free, page file usage ~205 MB.
  • Not graceful shutdown — 2 shutting down lines in 12 days; the old process logs nothing before disappearing.
  • Not the subagents — they are victims of the replacement.

Frequency (opencode.log, one machine): Event stream stalled = 3× on 2026-09-28, 6× on 2026-09-29; serve --service starts = 7 and 6 on those days. It stops completely when idle (12+ min with zero restarts, liveness probe 0–2ms).

Suspected trigger for the ≥45s heartbeat gap (evidence below, hypothesis): concurrent cold-boot of many locations plus watcher re-registration churn, forming a feedback loop — replace service → all recently-active locations cold-boot → watcher/config/MCP churn blocks the loop → heartbeat gap ≥45s → watchdog replaces the service again.

Plugins

  • local ~/.config/opencode/plugins/orca-opencode-status.js
  • local ~/.config/opencode/plugins/win-notify/
  • no plugin array in opencode.json

(~/.claude/skills and ~/.agents/skills are both 44-entry directories and are re-subscribed ~950× each — see Additional context.)

OpenCode version

2.0.19 (channel=latest, npm global @opencode/cli)

Steps to reproduce

Intermittent, load-dependent; not reproducible on demand.

  1. On Windows with opencode 2.0.19 CLI/TUI, let the shared managed service run (~/.local/state/opencode/service.json → http://127.0.0.1:49372).

  2. Have several locations active (home directory plus several projects on other drives) and run subagent-heavy turns.

  3. Observe repeatedly in ~/.local/share/opencode/log/opencode.log:

    08:48:40  client  event stream disconnected ... error="Event stream stalled"
    08:48:46  cli     cli starting version=2.0.19 args=["serve","--service"]
    08:48:47  client  event stream connected
    
  4. In-flight work is aborted: InterruptError: All fibers interrupted without error ... http.url=/api/event http.status=200.

  5. Restarts recur every ~2.5–16 min while busy; zero restarts while idle.

Screenshot and/or share link

N/A — log excerpts, code snippets and counts are inline below.

Operating System

Windows 11 Home Single Language, build 26200

Terminal

Windows Terminal (WT_SESSION set, TERM=xterm-256color), shell C:\WINDOWS\system32\cmd.exe

Additional context

Relevant client code (from the shipped 2.0.19 binary):

// event stream idle watchdog
var yY = 2000, Pv = 1000, gY = 50, bY = 45000, EY = 20000;
function wv(t, n) {
  let u = n.idleTimeout ?? bY;              // 45000 ms
  ...
  let Ke = () => { x = Date.now(); ...; qe = setTimeout(() => te.abort(Error("Event stream stalled")), u) };
  // onActivity resets the timer on ANY received byte
}
// server SSE handler
let s = qK("15 seconds").pipe(wa(() => `: heartbeat\n\n`));
// SSE reader: any byte counts as activity, including the heartbeat comment
if (!y.done) n?.onActivity?.();

Watcher churn / location cold-boot evidence (hypothesis for the 45s gap):

stall 08:18:06 : watcher_subscribe=116  location_boot=8   spawns=40
stall 08:24:19 : watcher_subscribe=123  location_boot=9   spawns=12
stall 08:26:55 : watcher_subscribe=129  location_boot=10  spawns=6
path subscribes
C:\Users\<user>\.claude\skills 944
C:\Users\<user>\.agents\skills 923
C:\Users\<user>\.config\opencode\skills 250
C:\Users\<user> 162
C:\, C:\Users, D:\ (type=entries) 77 / 77 / 80

location services booted: C:\Users\<user> 73×, plus ~20 other project locations. This matches #50594 (.claude/skills storm) and #51852 (intermittent main-thread spin correlated with watcher churn).

Next step being tested: OPENCODE_DISABLE_FILEWATCHER=true (per #51852). Will report whether restarts stop under the same load.

Related issues: #50594, #50591, #51852, #51216, #37795, #48588, #36285.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions