Skip to content

service: hardcoded default port 49374 locks out any second client on shared loopback (multi-user/WSL mirrored), surfacing only a generic timeout #50672

Description

@yveming

Summary

The managed background service binds a hardcoded default port (49374 for
latest/dev/beta/next channels). Service registration
(~/.local/state/opencode/service.json) is per-user, but the port is
machine-global on any shared loopback. Any second client on the same machine —
another Unix user, the Windows host under WSL2 networkingMode=mirrored, or a
--network host container — cannot bind the port. Instead of failing fast with
the real cause, the client spawns serve --service contenders for ~2 minutes
and then reports only Timed out waiting for the background service to start.
The actionable message exists only in the log file, and per the residual gap in
#41696 it can be swallowed entirely by contender overlap.

Environment

Both installations are v2.0.14, channel latest.

Environment A (holds the port):

  • opencode version: 2.0.14
  • OS: Linux 6.18.33.2-microsoft-standard-WSL2, networkingMode=Mirrored in .wslconfig
  • Terminal: TERM=xterm-256color
  • Shell: /bin/bash
  • Install/channel: latest
  • Active plugins: none found in config

Environment B (fails to start):

  • opencode version: 2.0.14
  • OS: Windows host of the WSL2 above (mirrored mode shares one loopback)
  • Terminal: cmd.exe (TERM unavailable: Windows console)
  • Shell: cmd
  • Install/channel: latest
  • Active plugins: none found in config

Reproduction

  1. In WSL2 with networkingMode=Mirrored, start opencode once and exit the
    TUI. The background service (opencode serve --service) persists, listening
    on 127.0.0.1:49374.
  2. On the Windows host (same loopback), run opencode.
  3. Wait ~2 minutes.

Same lockout without WSL: on a shared Linux box, run opencode as user A, exit,
then run opencode as user B.

Expected Behavior

Either:

  • the default service port is not a machine-global constant — pick an
    ephemeral/incremental port and record it in the registration file (it already
    stores url), or
  • at minimum, fail fast with the real cause and identify the occupying process.

Actual Behavior

Client spawns contenders in a loop, one failure every ~10s for ~2 minutes:

timestamp=2026-09-22T14:47:48.025Z level=ERROR message="cli process failed"
cause="Cause([Fail(Error: Managed service port 49374 on 127.0.0.1 is already in
use by another process. Configure another port with `opencode service set port
<port>` and start the service again. (cause: ServeError (cause: Error: Failed to
start server. Is port 49374 in use?)))])" args="[\"serve\",\"--service\"]" role=cli

(20+ identical repeats), then the only thing the TUI shows:

Error: Timed out waiting for the background service to start
    at service.ensure ...
    at cli.server-connection.resolve ...

Occupying process, from the WSL side:

$ ss -tlnp | grep 49374
LISTEN 0 512 127.0.0.1:49374 users:(("opencode",pid=12566,fd=10))

$ ps -p 12566
PID    STARTED              ELAPSED  CMD
12566  Tue Sep 22 21:46:33  01:14:39 /home/<user>/.opencode/bin/opencode serve --service

Additional Context

The port is a hardcoded constant (extracted from the v2.0.14 binary):

if (e==="latest"||e==="dev"||e==="beta"||e==="next") return 49374;
if (e==="local") return 49375;
return 1e4 + ...

All four release channels share 49374, for every user on the machine.

Collision matrix (fixed default port × shared loopback):

Scenario Collides
WSL2 mirrored ↔ Windows host yes (this report)
Two Unix users, same machine yes
--network host container ↔ host yes
-p 49374:... published container ↔ host yes
WSL2 NAT default / Docker bridge / full VM no

Multi-user impact is a silent lockout, not a data leak. Verified: API
returns 401 unauthenticated; service.json (registration + password) is 0600.
But user B's registration lives in B's home, so B cannot discover A's service;
B gets a bare timeout with no hint that another account is the cause, and the
remediation (opencode service set port) appears only in the log.

Warning on fix direction: #7629 proposes "probe the port and reuse the
existing service." That must not be applied to the default managed port across
users/environments. The service owns sessions, provider credentials,
permissions, and tool execution, so reuse keyed on port occupancy would
(a) require sharing the 0600 password across users or un-authenticating the
API, and (b) bind client A's in-flight work to client B's process lifecycle —
B logging out, wsl --shutdown, or a service auto-update restart (cf. #38567)
would kill A's running tasks with zero control on A's side. Reuse is only safe
when keyed on the client's own registration file (stale-service recovery),
never on port occupancy.

Related: #41696 + PR #41793 addressed surfacing stderr, but the follow-up
comment on #41696 documents a remaining contender-overlap gap that discards the
captured failure until the overall timeout — exactly what I observe on 2.0.14.
#19272 / #21170 / #46263 / #10357 cover adjacent symptom/solution ground; this
issue is the root cause: fixed machine-global default port + per-user
registration.

Workarounds: none applied on side B (user only retried). Side A unaffected.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions