Description
The TUI exhausts all system memory and is OOM-killed. Growth is linear at roughly
500MB/s-1GB/s with no GC sawtooth, reaching 24-28GB in under a minute, after which
the process is killed by the OOM killer.
I have not found a reliable trigger. It happens intermittently, sometimes
apparently on its own after the session has been open for a while. Resizing the
window (including fullscreen) often precedes an occurrence, but a resize on a
large session frequently does nothing, so I do not believe resize is a
sufficient cause. I have not established a pattern.
This is specific to v2. I hit it repeatedly on v2.0.18 and did not hit it on v1.
I have not gone back to v1 to re-verify, so treat that comparison as my
impression rather than a controlled test, but every occurrence I captured was on
the v2 build.
OpenCode version
v2.0.18 (stock release binary, official install)
Steps to reproduce
I cannot reproduce this on demand. There is no trigger I can identify or
reliably trigger, and I am not able to hand over a sequence of steps that
reliably brings it back. Sorry to whoever tries to follow this.
What I observe, offered as a pattern to correlate against rather than a recipe:
- Long session history is always present (~2150 messages in my case).
- Tool output streaming is usually in progress.
- The window freeze that precedes the growth is the only consistent early
signal I have noticed.
- Resizing, including fullscreen, often happens just before an occurrence. But
resizing a large session many times in a row does frequently nothing, so I do
not think resize is a sufficient cause. It may be incidental, or it may be
one of several contributing conditions.
- Once growth starts it is terminal: no recovery, OOM within about a minute.
What I actively ruled out, in a synthetic PTY harness:
- Resize alone: an empty session resized 89 times stayed flat at ~190MB.
- Resize at high rate on a long session: 12 resizes/second stayed bounded, peak
383MB, with the allocation sawtooth clearly visible and being reclaimed.
So neither resize alone nor a large transcript alone brings this on. I do not
know what the missing ingredient is.
Operating System
Arch Linux 7.2.7 (x64), Hyprland Wayland compositor, kitty terminal
Terminal
kitty, TERM=xterm-256color
Plugins
None that I'm aware of.
Mechanism
From gdb attached to a live process at 27GB. The main thread is the only busy
thread; every other thread is idle:
Thread 1 (main):
#0 mmap64 () from libc.so.6
#1 0x7f564a0c53ba ?? /tmp/.bun-1001-*.so
#2 0x7f564a0cebfb ?? /tmp/.bun-1001-*.so
#3 0x7f564a10b8a5 ?? /tmp/.bun-1001-*.so
#4 0x7f564a10a144 ?? /tmp/.bun-1001-*.so
#5 0x7f564be79e34 ?? ()
Thread 15 (HTTP Client): syscall <- idle
Thread 17-19 (Bun Pool): syscall <- idle
The main thread is in a synchronous allocation loop that never yields to the
event loop. That is the most informative part of this report: a loop that
starves the event loop cannot be reacting to user input, since input arrives
through the event loop. It has to be self-sustaining or driven by a timer.
A practical consequence: the built-in profilers cannot observe this.
SIGUSR1 (packages/cli/src/heap.ts, heap snapshot) and SIGPROF
(packages/cli/src/cpu-profile.ts, CPU profile) both dispatch and produce no
output, because their handlers run on the event loop that the loop never
reaches. gdb was the only tool that produced usable data. If you want to
profile this class of bug, the signal-based profilers will not work.
Candidate directions
Both of these are unconfirmed hypotheses, offered only to narrow the search.
-
packages/tui/src/ui/animation.ts has a module-level setInterval(tick, 16)
that self-perpetuates while any task remains in a module-level tasks Set.
Each surviving task calls setValue every tick, driving re-renders with no
user input. onCleanup(stop) is registered inside createAnimatable, so if
it is ever called outside a Solid owner the cleanup is a no-op and the task
can never be removed. Accumulated tasks would fit the observed
non-yielding-loop signature and the apparent randomness. I have not shown
the set actually grows unboundedly.
-
The tool renderers reprocess whole accumulated output per chunk, for example
stripAnsi(props.output?.trim()) in
packages/tui/src/routes/session/index.tsx. Over a long stream this is
quadratic in the output length. This would need streaming plus a long
transcript, which matches the fact that neither alone reproduces.
Known cost on the resize path
#37860 already identifies that collapseToolOutput scans complete tool output
via split("\n") and Array.from when only a short prefix is displayed, and
measures a 661.9ms -> 0.136ms improvement for the same fix shape. That is known
work and I am not proposing it again. Confirming only that the same code is
reachable from width-dependent memos in v2, so every resize re-runs the scan
across tool rows. Making that scan prefix-bounded moved peak RSS during re-layout
from 383MB to 148MB in my synthetic harness, but it did not prevent the OOM,
so I do not believe it is the cause of this issue.
Caveats
- I have no deterministic reproduction. Everything above comes from observing
live processes, plus a synthetic harness that demonstrates what does not
trigger it.
- I have not confirmed either candidate direction. The gdb evidence for a
non-yielding allocation loop is solid; the attribution to a specific file is
not.
- The v2-only observation is based on my own usage, not a controlled test.
Description
The TUI exhausts all system memory and is OOM-killed. Growth is linear at roughly
500MB/s-1GB/s with no GC sawtooth, reaching 24-28GB in under a minute, after which
the process is killed by the OOM killer.
I have not found a reliable trigger. It happens intermittently, sometimes
apparently on its own after the session has been open for a while. Resizing the
window (including fullscreen) often precedes an occurrence, but a resize on a
large session frequently does nothing, so I do not believe resize is a
sufficient cause. I have not established a pattern.
This is specific to v2. I hit it repeatedly on v2.0.18 and did not hit it on v1.
I have not gone back to v1 to re-verify, so treat that comparison as my
impression rather than a controlled test, but every occurrence I captured was on
the v2 build.
OpenCode version
v2.0.18 (stock release binary, official install)
Steps to reproduce
I cannot reproduce this on demand. There is no trigger I can identify or
reliably trigger, and I am not able to hand over a sequence of steps that
reliably brings it back. Sorry to whoever tries to follow this.
What I observe, offered as a pattern to correlate against rather than a recipe:
signal I have noticed.
resizing a large session many times in a row does frequently nothing, so I do
not think resize is a sufficient cause. It may be incidental, or it may be
one of several contributing conditions.
What I actively ruled out, in a synthetic PTY harness:
383MB, with the allocation sawtooth clearly visible and being reclaimed.
So neither resize alone nor a large transcript alone brings this on. I do not
know what the missing ingredient is.
Operating System
Arch Linux 7.2.7 (x64), Hyprland Wayland compositor, kitty terminal
Terminal
kitty, TERM=xterm-256color
Plugins
None that I'm aware of.
Mechanism
From gdb attached to a live process at 27GB. The main thread is the only busy
thread; every other thread is idle:
The main thread is in a synchronous allocation loop that never yields to the
event loop. That is the most informative part of this report: a loop that
starves the event loop cannot be reacting to user input, since input arrives
through the event loop. It has to be self-sustaining or driven by a timer.
A practical consequence: the built-in profilers cannot observe this.
SIGUSR1(packages/cli/src/heap.ts, heap snapshot) andSIGPROF(
packages/cli/src/cpu-profile.ts, CPU profile) both dispatch and produce nooutput, because their handlers run on the event loop that the loop never
reaches. gdb was the only tool that produced usable data. If you want to
profile this class of bug, the signal-based profilers will not work.
Candidate directions
Both of these are unconfirmed hypotheses, offered only to narrow the search.
packages/tui/src/ui/animation.tshas a module-levelsetInterval(tick, 16)that self-perpetuates while any task remains in a module-level
tasksSet.Each surviving task calls
setValueevery tick, driving re-renders with nouser input.
onCleanup(stop)is registered insidecreateAnimatable, so ifit is ever called outside a Solid owner the cleanup is a no-op and the task
can never be removed. Accumulated tasks would fit the observed
non-yielding-loop signature and the apparent randomness. I have not shown
the set actually grows unboundedly.
The tool renderers reprocess whole accumulated output per chunk, for example
stripAnsi(props.output?.trim())inpackages/tui/src/routes/session/index.tsx. Over a long stream this isquadratic in the output length. This would need streaming plus a long
transcript, which matches the fact that neither alone reproduces.
Known cost on the resize path
#37860already identifies thatcollapseToolOutputscans complete tool outputvia
split("\n")andArray.fromwhen only a short prefix is displayed, andmeasures a 661.9ms -> 0.136ms improvement for the same fix shape. That is known
work and I am not proposing it again. Confirming only that the same code is
reachable from width-dependent memos in v2, so every resize re-runs the scan
across tool rows. Making that scan prefix-bounded moved peak RSS during re-layout
from 383MB to 148MB in my synthetic harness, but it did not prevent the OOM,
so I do not believe it is the cause of this issue.
Caveats
live processes, plus a synthetic harness that demonstrates what does not
trigger it.
non-yielding allocation loop is solid; the attribution to a specific file is
not.