Skip to content

mcp: remote MCP servers that fail once are never retried until service restart #52237

Description

@duanhongyi

Summary

When a remote (HTTP) MCP server connection fails — e.g. after macOS sleep/wake kills long-lived connections — the background service keeps it in failed state indefinitely. Creating new sessions or restarting the TUI does not help, because MCP state lives in the background service (opencode serve --service), and there is neither a periodic retry nor a manual reconnect command. Only opencode service restart recovers.

Environment

  • opencode version: v2.0.20
  • OS: Darwin 27.0.0 arm64 (macOS, MacBook Air)
  • Terminal: ghostty (TERM=xterm-256color, COLORTERM=truecolor)
  • Shell: /bin/zsh
  • Install/channel: latest (npm global)
  • Active plugins: @omagents/omagents

Reproduction

  1. Configure a remote MCP server (e.g. { "type": "remote", "url": "https://mcp.context7.com/mcp" }) and start the TUI. Verify opencode mcp list shows it connected.
  2. Put the Mac to sleep for ~30+ minutes (close the lid). macOS runs DarkWake maintenance cycles every few minutes, which freeze user processes and kill/stall TCP connections.
  3. Wake the Mac. opencode mcp list now shows the remote servers as failed: Request timed out (or ECONNRESET), and their tools are gone from sessions.
  4. Create a new session, or quit and restart the TUI → servers stay failed.
  5. Run opencode service restart → all servers connect within seconds.

Expected Behavior

Failed MCP connections should recover without restarting the background service. Any of these would work:

  • periodic reconnect with backoff for servers in failed state,
  • retry on new session creation (per project/directory),
  • retry on system wake / network-change events,
  • a manual trigger, e.g. an opencode mcp reconnect [name] CLI subcommand or a TUI action.

Actual Behavior

Once marked failed, a remote server is never retried while the service keeps running. Observed: two remote servers (context7, websearch) failed at 13:05 local time and were still failed ~1 hour later, even though the network had fully recovered — both endpoints answered curl in <5s, and a different project directory initialized at 14:05 connected all 5 MCP servers within 6 seconds.

The failures correlate to the second with macOS DarkWake events from pmset -g log:

opencode log (UTC) pmset (local, UTC+8)
05:05:41 mcp connect failed DarkWake 13:05:41
05:10:54 mcp connect failed DarkWake 13:10:54
05:27:07 event stream stalled DarkWake 13:27:07
05:40:18 DarkWake 13:40:18
05:45:30 DarkWake 13:45:30
05:50:42 DarkWake 13:50:43
05:55:55 DarkWake 13:55:56

Additional Context

Typical log entries:

level=WARN message="mcp connect failed" server=context7 status.status=failed status.error="Request timed out" http.span=313211
level=WARN message="mcp http request failed" errors="[{... \"code\":\"ECONNRESET\" ...}]" durationMs=1252864 server=websearch

What seems to happen: an MCP request/reconnect attempt starts during a 10–60s DarkWake window, the machine goes back to sleep mid-request, the socket hangs across several sleep cycles, and ~313s later it fails with Request timed out or ECONNRESET. During sleep this repeats at every DarkWake (each stall triggers a retry that lands in the next DarkWake). After a real wake, the system is stable, so no retry is ever triggered again and the failed state sticks forever.

Local (stdio) MCP servers recover fine; only remote HTTP servers get stuck.

opencode mcp --help currently lists only list / add / auth / logout — there is no way to force a reconnect.

Workaround: opencode service restart (drops all attached TUI clients).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions