Skip to content

TUI OOM: intermittent 24-28GB memory exhaustion in v2, no trigger identified #51761

Description

@aasb13

Description

The TUI exhausts all system memory and is OOM-killed. Growth is linear at roughly
500MB/s-1GB/s with no GC sawtooth, reaching 24-28GB in under a minute, after which
the process is killed by the OOM killer.

I have not found a reliable trigger. It happens intermittently, sometimes
apparently on its own after the session has been open for a while. Resizing the
window (including fullscreen) often precedes an occurrence, but a resize on a
large session frequently does nothing, so I do not believe resize is a
sufficient cause. I have not established a pattern.

This is specific to v2. I hit it repeatedly on v2.0.18 and did not hit it on v1.
I have not gone back to v1 to re-verify, so treat that comparison as my
impression rather than a controlled test, but every occurrence I captured was on
the v2 build.

OpenCode version

v2.0.18 (stock release binary, official install)

Steps to reproduce

I cannot reproduce this on demand. There is no trigger I can identify or
reliably trigger, and I am not able to hand over a sequence of steps that
reliably brings it back. Sorry to whoever tries to follow this.

What I observe, offered as a pattern to correlate against rather than a recipe:

  • Long session history is always present (~2150 messages in my case).
  • Tool output streaming is usually in progress.
  • The window freeze that precedes the growth is the only consistent early
    signal I have noticed.
  • Resizing, including fullscreen, often happens just before an occurrence. But
    resizing a large session many times in a row does frequently nothing, so I do
    not think resize is a sufficient cause. It may be incidental, or it may be
    one of several contributing conditions.
  • Once growth starts it is terminal: no recovery, OOM within about a minute.

What I actively ruled out, in a synthetic PTY harness:

  • Resize alone: an empty session resized 89 times stayed flat at ~190MB.
  • Resize at high rate on a long session: 12 resizes/second stayed bounded, peak
    383MB, with the allocation sawtooth clearly visible and being reclaimed.

So neither resize alone nor a large transcript alone brings this on. I do not
know what the missing ingredient is.

Operating System

Arch Linux 7.2.7 (x64), Hyprland Wayland compositor, kitty terminal

Terminal

kitty, TERM=xterm-256color

Plugins

None that I'm aware of.

Mechanism

From gdb attached to a live process at 27GB. The main thread is the only busy
thread; every other thread is idle:

Thread 1 (main):
#0  mmap64 () from libc.so.6
#1  0x7f564a0c53ba ?? /tmp/.bun-1001-*.so
#2  0x7f564a0cebfb ?? /tmp/.bun-1001-*.so
#3  0x7f564a10b8a5 ?? /tmp/.bun-1001-*.so
#4  0x7f564a10a144 ?? /tmp/.bun-1001-*.so
#5  0x7f564be79e34 ?? ()

Thread 15 (HTTP Client): syscall    <- idle
Thread 17-19 (Bun Pool):    syscall    <- idle

The main thread is in a synchronous allocation loop that never yields to the
event loop. That is the most informative part of this report: a loop that
starves the event loop cannot be reacting to user input, since input arrives
through the event loop. It has to be self-sustaining or driven by a timer.

A practical consequence: the built-in profilers cannot observe this.
SIGUSR1 (packages/cli/src/heap.ts, heap snapshot) and SIGPROF
(packages/cli/src/cpu-profile.ts, CPU profile) both dispatch and produce no
output, because their handlers run on the event loop that the loop never
reaches. gdb was the only tool that produced usable data. If you want to
profile this class of bug, the signal-based profilers will not work.

Candidate directions

Both of these are unconfirmed hypotheses, offered only to narrow the search.

  1. packages/tui/src/ui/animation.ts has a module-level setInterval(tick, 16)
    that self-perpetuates while any task remains in a module-level tasks Set.
    Each surviving task calls setValue every tick, driving re-renders with no
    user input. onCleanup(stop) is registered inside createAnimatable, so if
    it is ever called outside a Solid owner the cleanup is a no-op and the task
    can never be removed. Accumulated tasks would fit the observed
    non-yielding-loop signature and the apparent randomness. I have not shown
    the set actually grows unboundedly.

  2. The tool renderers reprocess whole accumulated output per chunk, for example
    stripAnsi(props.output?.trim()) in
    packages/tui/src/routes/session/index.tsx. Over a long stream this is
    quadratic in the output length. This would need streaming plus a long
    transcript, which matches the fact that neither alone reproduces.

Known cost on the resize path

#37860 already identifies that collapseToolOutput scans complete tool output
via split("\n") and Array.from when only a short prefix is displayed, and
measures a 661.9ms -> 0.136ms improvement for the same fix shape. That is known
work and I am not proposing it again. Confirming only that the same code is
reachable from width-dependent memos in v2, so every resize re-runs the scan
across tool rows. Making that scan prefix-bounded moved peak RSS during re-layout
from 383MB to 148MB in my synthetic harness, but it did not prevent the OOM,
so I do not believe it is the cause of this issue.

Caveats

  • I have no deterministic reproduction. Everything above comes from observing
    live processes, plus a synthetic harness that demonstrates what does not
    trigger it.
  • I have not confirmed either candidate direction. The gdb evidence for a
    non-yielding allocation loop is solid; the attribution to a specific file is
    not.
  • The v2-only observation is based on my own usage, not a controlled test.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions