Repository navigation
Conversation
Node gives each address attempt in a multi-address connect 250ms before killing it (autoSelectFamilyAttemptTimeout). Remote MCP endpoints whose TCP handshake needs more than that never connect, so the server gets reported as "server unavailable" and its tools never register, even though the endpoint is healthy and reachable. Measured on a host with ~270ms RTT to the endpoint (3 A + 3 AAAA records, with the addresses rotating between lookups): 250ms (default) -> no connect, attempt destroyed, no error surfaces 300ms -> CONNECTED @901ms 400ms -> CONNECTED @395ms 600ms -> CONNECTED @294ms A config-level timeout can't help here: the connect dies well below mcp.timeout ?? 30_000, which wraps client.connect() rather than the socket itself, so raising it only makes the failure slower to surface. The value is a stagger interval rather than a wait, so fast connects are unaffected. 1000ms leaves headroom over the RTTs measured here without adding perceptible delay. Bun and Node both implement the API; Bun doesn't enforce the window, so this is a no-op there. Fixes anomalyco#53053
Contributor
|
The following comment was made by an LLM, it may be inaccurate: |
Contributor
|
Thanks for updating your PR! It now meets our contributing guidelines. 👍 |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue for this PR
Closes #53053
Type of change
What does this PR do?
Remote MCP servers never load when the endpoint's RTT is above roughly 250ms. My own host sits at ~270ms to an endpoint in us-west-2, which is an ordinary international hop, so this hits a lot of users. opencode reports the server as unavailable and its tools never register, even though the endpoint itself is healthy.
When a hostname resolves to more than one address, Node kills each attempt 250ms in (
autoSelectFamilyAttemptTimeout). An endpoint whose TCP handshake needs longer never finishes connecting. Nothing here overrides that default.This sets the value to 1000ms at module load in
packages/opencode/src/mcp/index.ts:It's a stagger interval rather than a wait, so connects that succeed quickly aren't slowed down. I picked 1000ms because the cliff is near 300ms and the slowest connect I measured was 412ms.
Raising
mcp.timeoutwon't help. It wrapsclient.connect()rather than the address attempts, and the connect dies well before then. Each server also pays that timeout twice, since opencode tries StreamableHTTP first and then SSE.Only the desktop app is affected. Bun doesn't enforce the 250ms window, so the shipped CLI was never broken. There the MCP servers are child processes of Electron's
node.mojom.NodeService, which is where this code runs.How did you verify your code works?
I swept the value on a host with ~270ms RTT to
mcp.deepwiki.com, which resolves 3 A and 3 AAAA records, with the addresses rotating between lookups:Then I ran the patched side effect against the real endpoint on both runtimes opencode ships:
Node failed all three probes before the change (ETIMEDOUT at 776ms and 752ms) and connected on all three after. The 400 is expected, since that POST has no MCP body; what matters is that the TCP and TLS handshakes complete.
I confirmed which process runs this code from the process tree, but I never watched a connect attempt inside it, so the timings above come from Node and Bun on my host.
Typechecking
packages/opencodereports 49 errors with this change and 49 without it, identical apart from line numbers shifting by 12, so it adds none. A fullbun typecheckdoesn't come back clean in my clone either:@opencode-ai/protocol,@opencode-ai/client,@opencode-ai/server, and@opencode-ai/sdk-nextall fail onCannot find module 'effect'. That reproduces on pristinedevwith this branch reverted, so it looks like missing workspace builds rather than anything from this change, but I couldn't prove that cleanly and CI hasn't run yet.Screenshots / recordings
Not a UI change.
Checklist