Skip to content

fix(mcp): raise Node's per-address connect attempt timeout - #53070

Open
jcoffi wants to merge 1 commit into
anomalyco:devfrom
jcoffi:fix/mcp-connect-attempt-timeout
Open

jcoffi wants to merge 1 commit into
anomalyco:devfrom
jcoffi:fix/mcp-connect-attempt-timeout

Conversation

@jcoffi

@jcoffi jcoffi commented Oct 4, 2026 •

Copy link
Copy Markdown

Issue for this PR

Closes #53053

Type of change

  • Bug fix
  • New feature
  • Refactor / code improvement
  • Documentation

What does this PR do?

Remote MCP servers never load when the endpoint's RTT is above roughly 250ms. My own host sits at ~270ms to an endpoint in us-west-2, which is an ordinary international hop, so this hits a lot of users. opencode reports the server as unavailable and its tools never register, even though the endpoint itself is healthy.

When a hostname resolves to more than one address, Node kills each attempt 250ms in (autoSelectFamilyAttemptTimeout). An endpoint whose TCP handshake needs longer never finishes connecting. Nothing here overrides that default.

This sets the value to 1000ms at module load in packages/opencode/src/mcp/index.ts:

const CONNECT_ATTEMPT_TIMEOUT = 1_000

net.setDefaultAutoSelectFamilyAttemptTimeout(CONNECT_ATTEMPT_TIMEOUT)

It's a stagger interval rather than a wait, so connects that succeed quickly aren't slowed down. I picked 1000ms because the cliff is near 300ms and the slowest connect I measured was 412ms.

Raising mcp.timeout won't help. It wraps client.connect() rather than the address attempts, and the connect dies well before then. Each server also pays that timeout twice, since opencode tries StreamableHTTP first and then SSE.

Only the desktop app is affected. Bun doesn't enforce the 250ms window, so the shipped CLI was never broken. There the MCP servers are child processes of Electron's node.mojom.NodeService, which is where this code runs.

How did you verify your code works?

I swept the value on a host with ~270ms RTT to mcp.deepwiki.com, which resolves 3 A and 3 AAAA records, with the addresses rotating between lookups:

250ms (default) -> no connect, attempt destroyed, no error surfaces
300ms           -> CONNECTED @901ms
400ms           -> CONNECTED @395ms
600ms           -> CONNECTED @294ms

Then I ran the patched side effect against the real endpoint on both runtimes opencode ships:

BUN   (1.4.2, matches the shipped CLI):        OK 400 @989ms / 451ms / 265ms
NODE  (matches the Electron desktop runtime):  OK 400 @869ms / 821ms / 317ms

Node failed all three probes before the change (ETIMEDOUT at 776ms and 752ms) and connected on all three after. The 400 is expected, since that POST has no MCP body; what matters is that the TCP and TLS handshakes complete.

I confirmed which process runs this code from the process tree, but I never watched a connect attempt inside it, so the timings above come from Node and Bun on my host.

Typechecking packages/opencode reports 49 errors with this change and 49 without it, identical apart from line numbers shifting by 12, so it adds none. A full bun typecheck doesn't come back clean in my clone either: @opencode-ai/protocol, @opencode-ai/client, @opencode-ai/server, and @opencode-ai/sdk-next all fail on Cannot find module 'effect'. That reproduces on pristine dev with this branch reverted, so it looks like missing workspace builds rather than anything from this change, but I couldn't prove that cleanly and CI hasn't run yet.

Screenshots / recordings

Not a UI change.

Checklist

  • I have tested my changes locally
  • I have not included unrelated changes in this PR

Node gives each address attempt in a multi-address connect 250ms before
killing it (autoSelectFamilyAttemptTimeout). Remote MCP endpoints whose TCP
handshake needs more than that never connect, so the server gets reported
as "server unavailable" and its tools never register, even though the
endpoint is healthy and reachable.

Measured on a host with ~270ms RTT to the endpoint (3 A + 3 AAAA records,
with the addresses rotating between lookups):

  250ms (default) -> no connect, attempt destroyed, no error surfaces
  300ms           -> CONNECTED @901ms
  400ms           -> CONNECTED @395ms
  600ms           -> CONNECTED @294ms

A config-level timeout can't help here: the connect dies well below
mcp.timeout ?? 30_000, which wraps client.connect() rather than the socket
itself, so raising it only makes the failure slower to surface.

The value is a stagger interval rather than a wait, so fast connects are
unaffected. 1000ms leaves headroom over the RTTs measured here without
adding perceptible delay. Bun and Node both implement the API; Bun doesn't
enforce the window, so this is a no-op there.

Fixes anomalyco#53053
@github-actions github-actions Bot added the needs:compliance This means the issue will auto-close after 2 hours. label Oct 4, 2026
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

The following comment was made by an LLM, it may be inaccurate:

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Thanks for updating your PR! It now meets our contributing guidelines. 👍

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Remote MCP servers fail to connect when endpoint RTT is above ~250ms (autoSelectFamilyAttemptTimeout too small)

1 participant