Skip to content

fix(app): back off event stream reconnects after repeated failures - #48015

Open
CannonRS wants to merge 1 commit into
anomalyco:devfrom
mQorva:event-reconnect-backoff
Open

CannonRS wants to merge 1 commit into
anomalyco:devfrom
mQorva:event-reconnect-backoff

Conversation

@CannonRS

@CannonRS CannonRS commented Sep 8, 2026

Copy link
Copy Markdown

Issue for this PR

Closes #48014

Type of change

  • Bug fix
  • New feature
  • Refactor / code improvement
  • Documentation

What does this PR do?

The global event stream (server-sdk.tsx) reconnected after every failure with
the same fixed RECONNECT_DELAY_MS delay. This adds an exponential backoff:

  • a consecutiveErrors counter resets on every successfully consumed event and
    increments on each catch;
  • the reconnect wait becomes
    min(5000, RECONNECT_DELAY_MS * 1.5 ** min(consecutiveErrors, 8)),
    so repeated failures back off from 250 ms up to a 5 s cap, while a healthy
    stream keeps the fast cadence.

No other behaviour in the loop changes.

How did you verify your code works?

  • bun typecheck in packages/app: clean.
  • The change is a 5-line delta restricted to the reconnect wait; the counter is
    scoped to the loop invocation and reset on any received event, so a recovered
    stream returns to the fast path immediately.

Screenshots / recordings

N/A — not a UI change.

Fork

Checklist

  • I have tested my changes locally
  • I have not included unrelated changes in this PR

@holny

holny commented Sep 11, 2026

Copy link
Copy Markdown

Been poking at reconnect behavior lately (see #48161 on v2 — client-side backoff, different layer, so no overlap here), so I read through this diff.

Main thing: the consecutiveErrors = 0 at the top of the for-await means a server that accepts the stream, emits a single event and then dies gets the full 250ms base delay again on every cycle, forever. That flapping pattern is basically where backoff earns its keep, so maybe reset the counter only after the stream has stayed healthy for a bit (say N seconds of events) instead of on first event.

Nit: with the 5s cap, the exponent clamp Math.min(consecutiveErrors, 8) can't actually engage at this base — 250ms x 1.5^7 is ~4.3s and anything beyond that hits the cap anyway, so the clamp is dead code unless the base ever drops under ~200ms.

#48014 looks assigned internally, so if this overlaps with planned work no worries — leaving notes in case they're useful.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Global event stream reconnects on a fixed 250ms timer with no backoff

2 participants