MCP stdio servers leak pipe fds + orphan child processes → cumulative EMFILE ("Too many open files", os error 24)
What version of Codex CLI is running?
codex-cli 0.137.0 (also reproduced symptom on long-running 0.12x sessions)
What subscription do you have?
Pro
Which model were you using?
n/a (process/transport bug, model-independent)
What platform is your computer?
Darwin 25.5.0 (macOS 26.5.1) arm64
What issue are you seeing?
Long-running sessions that use stdio MCP servers slowly accumulate open file descriptors in the codex parent process until they hit the per-process soft limit (256 on a default macOS login shell), at which point every subsequent open()/spawn() fails with Too many open files (os error 24).
This is almost certainly the underlying fd leak behind the now-closed #17806 ("Repeat Os { code: 24 } crashes"). That issue was closed via #12216, which only hardened the panic in AbsolutePathBuf::parent() so EMFILE no longer crashes the TUI — it did not address why fds are exhausted. The leak is still present.
Evidence (live system)
A single long-lived TUI session (codex resume, soft limit 256) sitting 3 fds under the ceiling:
$ lsof -p <codex_pid> | awk 'NR>1{print $5}' | sort | uniq -c | sort -rn
186 PIPE <-- the leak
38 REG (state_*.sqlite WAL/shm, rollout jsonl, logs_*.sqlite)
8 KQUEUE
7 unix
5 IPv4
...
253 total (soft NOFILE = 256 -> next open() = EMFILE)
186 anonymous PIPE fds for a session with ~4–5 configured MCP servers (expected ≈3 pipes/server = ~15). The excess maps almost 1:1 to orphaned MCP child processes that the codex parent still holds read pipes to.
System-wide, orphaned stdio MCP servers far outnumber the live codex sessions that should own them:
$ pgrep -f slack-mcp-server | wc -l # 30
$ pgrep -f xcodebuildmcp | wc -l # 18
$ pgrep -f "codex app-server"| wc -l # 9
These are spawned as npm exec <server> / npx, i.e. the actual server (node) is a grandchild under an npm wrapper.
Root-cause analysis (code, codex-rs @ 8f1aad5)
Cleanup of stdio MCP servers has three gaps that compound under the MCP refresh/rebuild churn.
1. Teardown is fire-and-forget — never awaits process exit
rmcp-client/src/stdio_server_launcher.rs LocalProcessTerminator::terminate (unix):
let should_escalate = terminate_process_group(pgid) /* killpg SIGTERM */ ...;
if should_escalate {
spawn(move || { // detached std::thread
sleep(PROCESS_GROUP_TERM_GRACE_PERIOD /* 2s */);
kill_process_group(pgid); // killpg SIGKILL
});
}
RmcpClient::shutdown (rmcp-client/src/rmcp_client.rs:678) calls this and returns immediately — the SIGKILL lands ≥2 s later on a detached thread. Nothing awaits actual child exit or pipe EOF. If the runtime/parent is torn down (or the next refresh races) within that window, the SIGKILL thread can be lost, leaving the group alive.
2. Cancelled-startup path skips explicit terminate entirely
On MCP refresh, core/src/session/mcp.rs::refresh_mcp_servers_inner:
// :341 cancel old startup token
guard.cancel();
// :344 build a brand-new manager — spawns the ENTIRE MCP server set again
let (refreshed_manager, cancel_token) = McpConnectionManager::new(...).await;
// :377 swap in
let mut old_manager = std::mem::replace(&mut *manager, refreshed_manager);
// :379 shut down the old one
old_manager.shutdown().await;
For any old client still mid-handshake, AsyncManagedClient::shutdown (codex-mcp/src/rmcp_client.rs:240) resolves the shared startup future to Err(Cancelled):
match self.client().await {
Ok(client) => client.client.shutdown().await,
Err(StartupOutcomeError::Cancelled) => {} // <-- terminate() never called
Err(error) => warn!(...),
}
Cleanup then depends entirely on Drop for StdioServerProcessHandleInner, which only fires once the last Arc clone is dropped — and which itself uses the same fire-and-forget terminate from gap #1.
The same full-set spawn-then-shutdown churn also runs in the ephemeral connector-discovery manager (core/src/connectors.rs:273 … :359).
3. Group signalling misses npm/npx grandchildren
Spawn uses bare command.process_group(0) (stdio_server_launcher.rs:263) and signals pgid == direct-child PID. The direct child is the npm/npx wrapper; the real server is node. When the wrapper exits (or the server reparents) the original group no longer covers node, so killpg(SIGTERM/SIGKILL) misses it → orphaned node (the 30/18 above). Note utils/pty/src/process_group.rs already has detach_from_tty/setsid-style helpers, but the stdio MCP launcher does not use a wait-and-verify reap.
Net effect
Each refresh/rebuild cycle spawns a fresh server set and tears down the old one without awaiting death and without reliably reaping the npm grandchild. Every orphan keeps its stdout/stderr write ends open, so the codex-side read pipe + the detached stderr-reader task (stdio_server_launcher.rs:274) never reach EOF and never close. Pipe fds in the codex parent grow monotonically until EMFILE.
Suggested fixes
- Await actual exit on teardown. Make
shutdown()/Drop waitpid/try_wait the group after SIGTERM→SIGKILL instead of fire-and-forget, so pipes are closed deterministically.
- Terminate on the cancelled-startup arm. In
AsyncManagedClient::shutdown, call the process terminator on Err(Cancelled) rather than relying on Drop.
- Reap the whole group reliably. Either spawn the server binary directly (resolve
npm exec/npx to the real entrypoint) or setsid + track the session and kill+wait the session, not just the direct child's PID-as-PGID.
- (Hardening) Bound the stderr-reader task to transport lifetime / close the read half on terminate so a lingering grandchild can't pin the pipe.
Workaround for users
Raise the per-process limit before launching (ulimit -n 65536, and/or sudo launchctl limit maxfiles 65536 524288 for GUI-launched servers) and periodically kill orphaned *-mcp-server / npx processes. This only delays the leak.
Related
MCP stdio servers leak pipe fds + orphan child processes → cumulative EMFILE ("Too many open files", os error 24)
What version of Codex CLI is running?
codex-cli 0.137.0 (also reproduced symptom on long-running 0.12x sessions)
What subscription do you have?
Pro
Which model were you using?
n/a (process/transport bug, model-independent)
What platform is your computer?
Darwin 25.5.0 (macOS 26.5.1) arm64
What issue are you seeing?
Long-running sessions that use stdio MCP servers slowly accumulate open file descriptors in the codex parent process until they hit the per-process soft limit (256 on a default macOS login shell), at which point every subsequent
open()/spawn()fails withToo many open files (os error 24).This is almost certainly the underlying fd leak behind the now-closed #17806 ("Repeat
Os { code: 24 }crashes"). That issue was closed via #12216, which only hardened the panic inAbsolutePathBuf::parent()so EMFILE no longer crashes the TUI — it did not address why fds are exhausted. The leak is still present.Evidence (live system)
A single long-lived TUI session (
codex resume, soft limit 256) sitting 3 fds under the ceiling:186 anonymous PIPE fds for a session with ~4–5 configured MCP servers (expected ≈3 pipes/server = ~15). The excess maps almost 1:1 to orphaned MCP child processes that the codex parent still holds read pipes to.
System-wide, orphaned stdio MCP servers far outnumber the live codex sessions that should own them:
These are spawned as
npm exec <server>/npx, i.e. the actual server (node) is a grandchild under annpmwrapper.Root-cause analysis (code, codex-rs @ 8f1aad5)
Cleanup of stdio MCP servers has three gaps that compound under the MCP refresh/rebuild churn.
1. Teardown is fire-and-forget — never awaits process exit
rmcp-client/src/stdio_server_launcher.rsLocalProcessTerminator::terminate(unix):RmcpClient::shutdown(rmcp-client/src/rmcp_client.rs:678) calls this and returns immediately — the SIGKILL lands ≥2 s later on a detached thread. Nothing awaits actual child exit or pipe EOF. If the runtime/parent is torn down (or the next refresh races) within that window, the SIGKILL thread can be lost, leaving the group alive.2. Cancelled-startup path skips explicit terminate entirely
On MCP refresh,
core/src/session/mcp.rs::refresh_mcp_servers_inner:For any old client still mid-handshake,
AsyncManagedClient::shutdown(codex-mcp/src/rmcp_client.rs:240) resolves the shared startup future toErr(Cancelled):Cleanup then depends entirely on
Drop for StdioServerProcessHandleInner, which only fires once the lastArcclone is dropped — and which itself uses the same fire-and-forget terminate from gap #1.The same full-set spawn-then-shutdown churn also runs in the ephemeral connector-discovery manager (
core/src/connectors.rs:273…:359).3. Group signalling misses
npm/npxgrandchildrenSpawn uses bare
command.process_group(0)(stdio_server_launcher.rs:263) and signalspgid == direct-child PID. The direct child is thenpm/npxwrapper; the real server isnode. When the wrapper exits (or the server reparents) the original group no longer coversnode, sokillpg(SIGTERM/SIGKILL)misses it → orphaned node (the 30/18 above). Noteutils/pty/src/process_group.rsalready hasdetach_from_tty/setsid-style helpers, but the stdio MCP launcher does not use await-and-verify reap.Net effect
Each refresh/rebuild cycle spawns a fresh server set and tears down the old one without awaiting death and without reliably reaping the
npmgrandchild. Every orphan keeps its stdout/stderr write ends open, so the codex-side read pipe + the detached stderr-reader task (stdio_server_launcher.rs:274) never reach EOF and never close. Pipe fds in the codex parent grow monotonically until EMFILE.Suggested fixes
shutdown()/Dropwaitpid/try_waitthe group after SIGTERM→SIGKILL instead of fire-and-forget, so pipes are closed deterministically.AsyncManagedClient::shutdown, call the process terminator onErr(Cancelled)rather than relying onDrop.npm exec/npxto the real entrypoint) orsetsid+ track the session and kill+wait the session, not just the direct child's PID-as-PGID.Workaround for users
Raise the per-process limit before launching (
ulimit -n 65536, and/orsudo launchctl limit maxfiles 65536 524288for GUI-launched servers) and periodically kill orphaned*-mcp-server/npxprocesses. This only delays the leak.Related
Os { code: 24, kind: Uncategorized, message: "Too many open files" }crashes #17806 (closed) — same symptom; closed via panic-hardening only, leak not addressedAbsolutePathBuf::parent()syspolicyd), not MCP