Skip to content

MCP stdio servers leak pipe fds + orphan child processes → cumulative EMFILE ("Too many open files", os error 24) #26984

Description

@jacobcxdev

MCP stdio servers leak pipe fds + orphan child processes → cumulative EMFILE ("Too many open files", os error 24)

What version of Codex CLI is running?

codex-cli 0.137.0 (also reproduced symptom on long-running 0.12x sessions)

What subscription do you have?

Pro

Which model were you using?

n/a (process/transport bug, model-independent)

What platform is your computer?

Darwin 25.5.0 (macOS 26.5.1) arm64

What issue are you seeing?

Long-running sessions that use stdio MCP servers slowly accumulate open file descriptors in the codex parent process until they hit the per-process soft limit (256 on a default macOS login shell), at which point every subsequent open()/spawn() fails with Too many open files (os error 24).

This is almost certainly the underlying fd leak behind the now-closed #17806 ("Repeat Os { code: 24 } crashes"). That issue was closed via #12216, which only hardened the panic in AbsolutePathBuf::parent() so EMFILE no longer crashes the TUI — it did not address why fds are exhausted. The leak is still present.

Evidence (live system)

A single long-lived TUI session (codex resume, soft limit 256) sitting 3 fds under the ceiling:

$ lsof -p <codex_pid> | awk 'NR>1{print $5}' | sort | uniq -c | sort -rn
 186 PIPE        <-- the leak
  38 REG         (state_*.sqlite WAL/shm, rollout jsonl, logs_*.sqlite)
   8 KQUEUE
   7 unix
   5 IPv4
   ...
 253 total       (soft NOFILE = 256 -> next open() = EMFILE)

186 anonymous PIPE fds for a session with ~4–5 configured MCP servers (expected ≈3 pipes/server = ~15). The excess maps almost 1:1 to orphaned MCP child processes that the codex parent still holds read pipes to.

System-wide, orphaned stdio MCP servers far outnumber the live codex sessions that should own them:

$ pgrep -f slack-mcp-server  | wc -l   # 30
$ pgrep -f xcodebuildmcp     | wc -l   # 18
$ pgrep -f "codex app-server"| wc -l   #  9

These are spawned as npm exec <server> / npx, i.e. the actual server (node) is a grandchild under an npm wrapper.

Root-cause analysis (code, codex-rs @ 8f1aad5)

Cleanup of stdio MCP servers has three gaps that compound under the MCP refresh/rebuild churn.

1. Teardown is fire-and-forget — never awaits process exit

rmcp-client/src/stdio_server_launcher.rs LocalProcessTerminator::terminate (unix):

let should_escalate = terminate_process_group(pgid) /* killpg SIGTERM */ ...;
if should_escalate {
    spawn(move || {                 // detached std::thread
        sleep(PROCESS_GROUP_TERM_GRACE_PERIOD /* 2s */);
        kill_process_group(pgid);   // killpg SIGKILL
    });
}

RmcpClient::shutdown (rmcp-client/src/rmcp_client.rs:678) calls this and returns immediately — the SIGKILL lands ≥2 s later on a detached thread. Nothing awaits actual child exit or pipe EOF. If the runtime/parent is torn down (or the next refresh races) within that window, the SIGKILL thread can be lost, leaving the group alive.

2. Cancelled-startup path skips explicit terminate entirely

On MCP refresh, core/src/session/mcp.rs::refresh_mcp_servers_inner:

// :341  cancel old startup token
guard.cancel();
// :344  build a brand-new manager — spawns the ENTIRE MCP server set again
let (refreshed_manager, cancel_token) = McpConnectionManager::new(...).await;
// :377  swap in
let mut old_manager = std::mem::replace(&mut *manager, refreshed_manager);
// :379  shut down the old one
old_manager.shutdown().await;

For any old client still mid-handshake, AsyncManagedClient::shutdown (codex-mcp/src/rmcp_client.rs:240) resolves the shared startup future to Err(Cancelled):

match self.client().await {
    Ok(client) => client.client.shutdown().await,
    Err(StartupOutcomeError::Cancelled) => {}   // <-- terminate() never called
    Err(error) => warn!(...),
}

Cleanup then depends entirely on Drop for StdioServerProcessHandleInner, which only fires once the last Arc clone is dropped — and which itself uses the same fire-and-forget terminate from gap #1.

The same full-set spawn-then-shutdown churn also runs in the ephemeral connector-discovery manager (core/src/connectors.rs:273 … :359).

3. Group signalling misses npm/npx grandchildren

Spawn uses bare command.process_group(0) (stdio_server_launcher.rs:263) and signals pgid == direct-child PID. The direct child is the npm/npx wrapper; the real server is node. When the wrapper exits (or the server reparents) the original group no longer covers node, so killpg(SIGTERM/SIGKILL) misses it → orphaned node (the 30/18 above). Note utils/pty/src/process_group.rs already has detach_from_tty/setsid-style helpers, but the stdio MCP launcher does not use a wait-and-verify reap.

Net effect

Each refresh/rebuild cycle spawns a fresh server set and tears down the old one without awaiting death and without reliably reaping the npm grandchild. Every orphan keeps its stdout/stderr write ends open, so the codex-side read pipe + the detached stderr-reader task (stdio_server_launcher.rs:274) never reach EOF and never close. Pipe fds in the codex parent grow monotonically until EMFILE.

Suggested fixes

  1. Await actual exit on teardown. Make shutdown()/Drop waitpid/try_wait the group after SIGTERM→SIGKILL instead of fire-and-forget, so pipes are closed deterministically.
  2. Terminate on the cancelled-startup arm. In AsyncManagedClient::shutdown, call the process terminator on Err(Cancelled) rather than relying on Drop.
  3. Reap the whole group reliably. Either spawn the server binary directly (resolve npm exec/npx to the real entrypoint) or setsid + track the session and kill+wait the session, not just the direct child's PID-as-PGID.
  4. (Hardening) Bound the stderr-reader task to transport lifetime / close the read half on terminate so a lingering grandchild can't pin the pipe.

Workaround for users

Raise the per-process limit before launching (ulimit -n 65536, and/or sudo launchctl limit maxfiles 65536 524288 for GUI-launched servers) and periodically kill orphaned *-mcp-server / npx processes. This only delays the leak.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    CLIIssues related to the Codex CLIbugSomething isn't workingmcpIssues related to the use of model context protocol (MCP) servers

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions