Repository navigation
Conversation
wrapSSE re-armed the read timeout on any chunk, including SSE comment lines (`: keepalive`). Providers and load balancers that hold a stalled generation open with comment heartbeats kept the connection byte-warm forever: no event ever arrived, the timeout never fired, and the session hung with no error (anomalyco#43519). Replace the per-read timer with a stall deadline that only re-arms when a chunk contains the start of a real SSE field line (data:, event:, id:, retry:). A tiny line-head scanner tracks progress across chunk boundaries without parsing the stream. The deadline still fires the same typed, retryable timeout error (ResponseStreamError / "SSE read timed out") and still covers fully-silent connections. Applied identically to the v1 provider wrapper (packages/opencode) and the v2 SDK wrapper (packages/core/aisdk.ts). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
The following comment was made by an LLM, it may be inaccurate: |
|
cc @thdxr @adamdotdevin — requesting review when you get a chance. Fix details: The bug is subtle: The fix swaps the per-read timer for a single data-stall deadline that only re-arms when a chunk contains the start of a non-comment SSE line:
One semantic difference worth noting in review: while consumer backpressure pauses New regression test reproduces the keepalive-warm stall end-to-end (data event → Both |
|
Reviewed — this is the right fix for #43519 and the Consider — continuation bytes of a field line don't count as progress. The scanner only credits a line's head. A single FYI — event-level heartbeats still keep the deadline warm. Providers/LBs that heartbeat with real SSE events — Anthropic-style Verified the backpressure trade-off the author flagged: the deadline keeps running while the consumer pauses Verified both copies are identical ( Verdict: approve — closes the deterministic infinite-hang repro with minimal machinery and no contract change. cc @Hona @Brendonovich — could you review and approve? Provider-side counterpart to the client watchdog in #51871; together they close the silent-stall class on both SSE boundaries. |
Issue for this PR
Closes #43519. Related: #51857 (client-side SSE watchdog, PR #51871) — this is the provider-side counterpart; the same stall class exists on both SSE boundaries.
Type of change
What does this PR do?
wrapSSE(provider fetch wrapper forchunkTimeout) re-armed its read timeout on any chunk. SSE comment lines (: keepalive) carry nodata:event — the SSE spec explicitly allows them for keep-alive — so a provider or LB that holds a stalled generation open with comment heartbeats keeps the timer resetting forever. Result, per #43519:opencode runhangs indefinitely — no stdout event, no log line, no error, until an external watchdog kills it.The fix replaces the per-read timer with a stall deadline that only re-arms on SSE field progress:
sseFieldScannerinspects each chunk's line heads: a line starting with:(comment),\n/\r(blank) is not progress; any other line start (data:,event:,id:,retry:, any field) is.data:marker split between two reads still counts, and a comment split the same way still doesn't.ProviderError.ResponseStreamError("SSE read timed out")(v1) /Error("SSE read timed out")(core v2) — soSessionRetryclassification and provider retry behavior are unchanged.Applied identically to both copies of
wrapSSE:packages/opencode/src/provider/provider.ts(v1 runtime) andpackages/core/src/aisdk.ts(v2 SDK path).How did you verify your code works?
New regression test in
packages/opencode/test/provider/header-timeout.test.ts: a server sends one partial SSE event, then streams: keepalive\n\nevery 10ms forever. WithchunkTimeout: 100,result.fullStreamnow surfacesProviderError.ResponseStreamError("SSE read timed out") — previously it hung forever.bun typecheck(tsgo) clean;oxlinton touched files: 0 errors, no new warnings.Checklist