Description
The web UI still loses live updates whenever the SSE event stream dies silently (tab suspended, network flap, proxy idle timeout) and only recovers after a hard refresh. This was already reported in #39030, #47258 and #45860, but there is an important detail those reports miss:
A fix for this exact failure mode was already merged — and the v2 client refactor removed it.
PR #47571 (fix(client): detect stalled event streams and resync on foreground, merged 2026-09-06) implemented a proper stall watchdog in packages/client/src/solid/connection.ts:
- every received chunk (including
server.heartbeat keepalives) pushed a stall deadline forward;
- an idle
setTimeout aborted the request with "Event stream stalled", triggering reconnect + resync;
visibilitychange and online events forced a resync;
- a
connected | connecting | reconnecting status was exposed to consumers.
On current dev (083ed266e, v1.18.33), packages/client/src/solid/ no longer exists, and the app's live event path is packages/app/src/context/server-sdk.tsx, which contains a naive reconnect loop with none of that liveness logic:
// packages/app/src/context/server-sdk.tsx (~L268-309)
while (!abort.signal.aborted && started && generation === active) {
attempt = new AbortController()
try {
const events = kind === "v1"
? (await eventSdk.global.event({ signal: attempt.signal })).stream
: eventApi.event.subscribe({ signal: attempt.signal })
for await (const event of events) { /* ... */ }
} catch (error) { /* logs, then falls through */ }
await wait(RECONNECT_DELAY_MS) // fixed 250ms — see also #48014
}
The server already emits server.heartbeat every 10s and sets X-Accel-Buffering: no (packages/opencode/src/server/routes/instance/httpapi/handlers/event.ts), so the keepalive exists on the wire — the client just never uses it to detect a dead stream. If the TCP connection half-opens (suspended tab, NAT/proxy drop without RST, cloudflared/reverse-proxy restart), for await never resolves and never rejects: the UI is stale forever.
Additionally, pagehide always calls stop(), but pageshow only restarts when event.persisted === true (resumeStreamAfterPageShow) — and an open streaming fetch makes the page ineligible for bfcache, so persisted is almost always false.
Expected behavior
Port the watchdog semantics that #47571 introduced into server-sdk.tsx (or wherever the app's event stream now lives):
// sketch: stall watchdog driven by the existing 10s server heartbeat
const STALL_TIMEOUT_MS = 3 * HEARTBEAT_INTERVAL_MS + 5_000
let watchdog: ReturnType<typeof setTimeout> | undefined
const arm = () => {
clearTimeout(watchdog)
watchdog = setTimeout(() => {
attempt?.abort(new Error("event stream stalled"))
// then resync: existing server.connected path already re-bootstraps
// active directories in server-sync.tsx
}, STALL_TIMEOUT_MS)
}
// inside the for-await loop: arm() on every event (heartbeats count)
plus a visibilitychange listener that aborts/resyncs when the page returns to the foreground after going quiet — restoring the behavior users had before the refactor.
Steps to reproduce
opencode web behind any proxy (or directly), open a session.
- Suspend the tab (Chrome:
chrome://discards, or switch apps on mobile / sleep the OS long enough for the TCP connection to drop without RST).
- Return to the tab.
- Trigger a new event server-side → the UI does not update.
Ctrl+F5 required.
OpenCode version
v1.18.32–v1.18.33 (dev @ 083ed266e)
Operating System
Observed on Linux server (Nginx + Cloudflare Tunnel in front); reproduction is client-side and platform-independent.
Related
#39030, #47258, #45860, #46733, #49861, #43519 (server-side: heartbeats reset chunkTimeout, so dead streams never expire either), #48014 / PR #48015 (missing reconnect backoff, adjacent), PR #47571 (the merged fix that was lost).
Related work (open issues this would resolve or interact with)
| Ref |
State |
What it reports / fixes |
| #39030 |
open issue |
Canonical repro: mobile/desktop tab returns from background → stream dead → manual refresh required. |
| #47258 |
open issue |
Same root cause, with precise pointers to server-sdk.tsx pagehide/pageshow + event.persisted. |
| #45860 |
open issue |
Session status goes stale after tab suspension (mobile). |
| #46733 |
open issue |
/global/event + /event deliver only server.connected + heartbeats — zero message events (1.18.25). |
| #49861 |
open issue |
SSE drops all events for worktree sessions (current dev). |
| #48771 |
open issue |
V2 streams stall without an SSE Accept header (AV/proxy interference). |
| #43519 |
open issue |
Server-side counterpart: chunkTimeout resets on heartbeats, so stalled streams never time out either. |
| #38458 |
open issue |
SSE closes mid-turn on opencode serve. |
| #48014 |
open issue |
Reconnect loop uses a fixed 250ms delay, no backoff — reconnect storms, self-inflicted 499s on re-bootstrap. |
| PR #47571 |
merged (lost) |
The original stall watchdog: idle abort, visibilitychange/online resync, connection status. File deleted in the v2 refactor — the regression this issue tracks. |
| PR #48015 |
open |
Backoff for event-stream reconnects (closes #48014) — same fix independently implemented downstream in the alltomatos/opencode fork (commit 6722d7a86). |
| PR #50375 / #51835 |
open (dup) |
Release SSE streams when a Bun client disconnects — server-side stream leak; plausible cause of listener-leak warnings under load. |
| PRs #44903, #39028, #29987, #26471, #19461, #10723 |
closed |
Prior fix attempts for the same failure class — this bug has regressed repeatedly; a watchdog (not another event listener) is the durable fix. |
Triage note for maintainers
This is the single highest-impact web reliability defect from the user's perspective: it produces daily "frozen" UIs that only Ctrl+F5 recovers, and it keeps regenerating new reports (#39030 → #47258 → mobile variants). The fix shape is already proven — it existed in packages/client/src/solid/connection.ts — and mostly needs porting to server-sdk.tsx plus coverage to prevent a third regression. Requesting prioritization together with #48014/PR #48015 (backoff) and #43519 (server-side stall reap) so client and server stop trusting half-open connections.
Description
The web UI still loses live updates whenever the SSE event stream dies silently (tab suspended, network flap, proxy idle timeout) and only recovers after a hard refresh. This was already reported in #39030, #47258 and #45860, but there is an important detail those reports miss:
A fix for this exact failure mode was already merged — and the v2 client refactor removed it.
PR #47571 (
fix(client): detect stalled event streams and resync on foreground, merged 2026-09-06) implemented a proper stall watchdog inpackages/client/src/solid/connection.ts:server.heartbeatkeepalives) pushed a stall deadline forward;setTimeoutaborted the request with"Event stream stalled", triggering reconnect + resync;visibilitychangeandonlineevents forced a resync;connected | connecting | reconnectingstatus was exposed to consumers.On current
dev(083ed266e, v1.18.33),packages/client/src/solid/no longer exists, and the app's live event path ispackages/app/src/context/server-sdk.tsx, which contains a naive reconnect loop with none of that liveness logic:The server already emits
server.heartbeatevery 10s and setsX-Accel-Buffering: no(packages/opencode/src/server/routes/instance/httpapi/handlers/event.ts), so the keepalive exists on the wire — the client just never uses it to detect a dead stream. If the TCP connection half-opens (suspended tab, NAT/proxy drop without RST,cloudflared/reverse-proxy restart),for awaitnever resolves and never rejects: the UI is stale forever.Additionally,
pagehidealways callsstop(), butpageshowonly restarts whenevent.persisted === true(resumeStreamAfterPageShow) — and an open streamingfetchmakes the page ineligible for bfcache, sopersistedis almost alwaysfalse.Expected behavior
Port the watchdog semantics that #47571 introduced into
server-sdk.tsx(or wherever the app's event stream now lives):plus a
visibilitychangelistener that aborts/resyncs when the page returns to the foreground after going quiet — restoring the behavior users had before the refactor.Steps to reproduce
opencode webbehind any proxy (or directly), open a session.chrome://discards, or switch apps on mobile / sleep the OS long enough for the TCP connection to drop without RST).Ctrl+F5required.OpenCode version
v1.18.32–v1.18.33 (
dev@083ed266e)Operating System
Observed on Linux server (Nginx + Cloudflare Tunnel in front); reproduction is client-side and platform-independent.
Related
#39030, #47258, #45860, #46733, #49861, #43519 (server-side: heartbeats reset
chunkTimeout, so dead streams never expire either), #48014 / PR #48015 (missing reconnect backoff, adjacent), PR #47571 (the merged fix that was lost).Related work (open issues this would resolve or interact with)
server-sdk.tsxpagehide/pageshow+event.persisted./global/event+/eventdeliver onlyserver.connected+ heartbeats — zero message events (1.18.25).Acceptheader (AV/proxy interference).chunkTimeoutresets on heartbeats, so stalled streams never time out either.opencode serve.visibilitychange/onlineresync, connection status. File deleted in the v2 refactor — the regression this issue tracks.6722d7a86).Triage note for maintainers
This is the single highest-impact web reliability defect from the user's perspective: it produces daily "frozen" UIs that only
Ctrl+F5recovers, and it keeps regenerating new reports (#39030 → #47258 → mobile variants). The fix shape is already proven — it existed inpackages/client/src/solid/connection.ts— and mostly needs porting toserver-sdk.tsxplus coverage to prevent a third regression. Requesting prioritization together with #48014/PR #48015 (backoff) and #43519 (server-side stall reap) so client and server stop trusting half-open connections.