Skip to content

Remote MCP servers fail to connect when endpoint RTT is above ~250ms (autoSelectFamilyAttemptTimeout too small) #53053

Description

@jcoffi

Description

Remote MCP servers whose endpoint RTT exceeds ~250ms never load. On every session start opencode logs:

level=WARN run=7f94844c message="server unavailable" key=deepwiki type=remote status=failed

The server's healthy and reachable. curl completes a full MCP initialize handshake against the same URL and gets a valid result. The client's own TCP connect is what fails, so the tools never get registered. Startup drags too, since connectRemote tries StreamableHTTP and then SSE, and each attempt burns its own timeout before the warning shows up.

Plugins

None

OpenCode version

1.18.34 (the opencode deb package, @opencode-ai/desktop).

Steps to reproduce

  1. Be more than ~250ms RTT from the remote MCP endpoint. Much of the world is, for anything hosted in US or EU regions.
  2. Add a remote server: "deepwiki": { "type": "remote", "url": "https://mcp.deepwiki.com/mcp", "enabled": true }
  3. Start opencode and watch ~/.local/share/opencode/log/opencode.log for the line above.
  4. Reproduce outside opencode on the same host:
$ node -e "fetch('https://mcp.deepwiki.com/mcp',{method:'POST'}).catch(e=>console.log(e.code||e.cause?.code))"
UND_ERR_CONNECT_TIMEOUT

$ NODE_OPTIONS=--network-family-autoselection-attempt-timeout=2000 \
  node -e "fetch('https://mcp.deepwiki.com/mcp',{method:'POST'}).then(r=>console.log(r.status))"
400

Three runs fail without the flag, three succeed with it. The 400 is expected, since that POST carries no MCP body. The point is that the TCP and TLS handshakes now complete.

Screenshot and/or share link

Not applicable, there's no UI to capture. The log line above is the whole symptom.

Operating System

Ubuntu 24.04, x86_64

Terminal

n/a, opencode desktop app

Root cause

mcp.deepwiki.com returns 3 A and 3 AAAA records, and the addresses rotate between lookups. TCP connect to the A records measures 263-412ms from here, 3 samples each. Node's happy-eyeballs connect windows every address attempt at 250ms by default (autoSelectFamilyAttemptTimeout), so each attempt gets destroyed just before it finishes. Sweeping that value on this host with net.setDefaultAutoSelectFamilyAttemptTimeout():

250ms (default) -> no connect, attempt destroyed and no error surfaces
300ms           -> CONNECTED @901ms
400ms           -> CONNECTED @395ms
600ms           -> CONNECTED @294ms

Two failure signatures come out of this, depending on which addresses the pool hands back. Sometimes fetch fails in ~825ms with an AggregateError [ETIMEDOUT] listing every address. Other times fetch and net.connect hang with no error at all, which undici reports as UND_ERR_CONNECT_TIMEOUT after its 10s default. Instant failure and indefinite hang are the same bug.

For whoever picks this up: the warning comes from Effect.logWarning("server unavailable", ...) in create, and remote connects run through connectRemote in packages/opencode/src/mcp/index.ts. Nothing in that path touches autoSelectFamilyAttemptTimeout, and a search of the repo turns up zero references to it anywhere, so the 250ms default is inherited untouched. Config can't reach it either: connectRemote wraps each transport in withTimeout(client.connect(t), mcp.timeout ?? 30_000), which sits well above the point where the TCP connect dies. In practice the failure surfaces earlier than that, at undici's 10s connect timeout, once per transport.

I couldn't instrument the Electron process directly. The connect path was verified with Node's net and fetch on the same host, and the client log line matches.

Suggested fix

net.setDefaultAutoSelectFamilyAttemptTimeout(1000)

early in the process, before any MCP client is created. The cliff sits near 300ms and I measured up to 412ms, so 1000ms leaves headroom, and the value's a stagger interval rather than a wait, so fast connects aren't affected. Both Node and Bun implement the API. Retrying with family: 4 after a connect failure would cover the remaining cases.

Workaround I'm using for now that works: pin one endpoint IP in /etc/hosts. Single-address lookups skip the race entirely so the connect completes. But it's obviously fragile.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions