Skip to content

bug(processor): doom loop detection misses cross-message repetitions and has inverted filter order #25254

Description

@qz1543706741

Summary

Two bugs in the doom loop detection logic in packages/opencode/src/session/processor.ts allow infinite tool-call loops to go undetected.

Bug 1 — Detection scope limited to current message only

// current code
const parts = MessageV2.parts(ctx.assistantMessage.id)
const recentParts = parts.slice(-DOOM_LOOP_THRESHOLD)

MessageV2.parts(ctx.assistantMessage.id) returns only parts from the current assistant message. When the model repeats the same tool call across multiple turns (e.g. three separate messages each calling read_file with the same path), the doom loop is never detected because each individual message has fewer than DOOM_LOOP_THRESHOLD matching parts.

Expected behavior: Detection should span all messages since the last user turn, not just the current assistant message.

Bug 2 — Slice before filter inverts the logic

// current code
parts.slice(-DOOM_LOOP_THRESHOLD).every(part => part.type === "tool" && part.tool === value.toolName && ...)

Taking the last N parts first (which may include text or reasoning parts) and then checking whether all N match means that if any non-tool part appears in the tail, every returns false and the doom loop check silently passes — even if there are plenty of repeated tool calls.

Expected behavior: Filter matching tool parts first, then check whether the count reaches the threshold.

Proposed Fix

  • Use MessageV2.filterCompactedEffect(ctx.sessionID) to collect matching tool parts across message boundaries since the last user turn.
  • Reorder the logic: filter by tool + JSON(args) first, then slice to DOOM_LOOP_THRESHOLD.

Activity

added
coreAnything pertaining to core functionality of the application (opencode server stuff)
on May 1, 2026
removed
coreAnything pertaining to core functionality of the application (opencode server stuff)
on May 3, 2026

tim-mohrbach-ikigai commented on May 4, 2026

@tim-mohrbach-ikigai

Real-world reproduction with 1,827 repetitions, confirming both bugs exactly as described.

A sub-agent was tasked with auditing a large set of code changes. It tried to read a git diff but the output was being truncated, so it switched to writing to a temp file — but then looped on the exact same command 1,827 times (verified from the local SQLite part table):

# tool: bash
# description: "Get raw git diff using native git binary"
# same output returned every time

Both bugs were in play simultaneously:

Bug 1 (cross-message scope): The command repeated across many turns. Each individual assistant message had far fewer than DOOM_LOOP_THRESHOLD = 3 matching parts, so MessageV2.parts(ctx.assistantMessage.id) never accumulated enough to trigger.

Bug 2 (slice before filter): Reasoning and text parts in the tail of parts.slice(-3) caused every() to return false, silently bypassing detection even within a single message.

The agent ran for ~30 minutes before being manually interrupted.


PR #25255 fixes this correctly. I traced through the logic against the reproduction case:

  1. filterCompactedEffect(sessionID) collects all messages since the last compaction in chronological order
  2. findLastIndex(role === "user") anchors to the original user turn
  3. .slice(lastUserIndex + 1).flatMap(msg => msg.parts) spans all assistant turns after it — catching the cross-message repetitions Bug 1 missed
  4. .filter(tool === X && JSON(input) === JSON(args)) operates on matching parts only — so text/reasoning parts are invisible to the count, fixing Bug 2
  5. .slice(-DOOM_LOOP_THRESHOLD) + length === 3 check fires on the 3rd repetition ✓

The fix would have interrupted the loop after 3 turns instead of 1,827. Hoping this reproduction helps move #25255 toward merge.

isac322 commented on May 29, 2026

@isac322

This matches a Gemini case I reproduced where repeated tool calls happened across assistant messages, so current-message-only detection missed it. I published a small OpenCode plugin workaround here: https://github.com/isac322/opencode-gemini-dejavu / npm: opencode-gemini-dejavu

It does not replace a proper doom-loop detector fix, but it may be useful as a lightweight mitigation for repeated tool-call tails.

Keesan12 commented on Jun 5, 2026

@Keesan12

Once cross-message scope is fixed, I’d also emit a stable loop receipt from the detector: last_user_turn, fingerprint, count, action.

Headless/plugin environments need a machine-readable reason when the guard trips. Otherwise a loop stop looks too much like a normal tool failure, and supervisors end up retrying the same poisoned session instead of seeing “this run was stopped on purpose.”

JustSidus commented on Jun 12, 2026

@JustSidus

Opened PR #32089 to fix this. The root cause was two bugs in processor.ts: the doom loop check only searched within the current assistant message (not the full session), and it did slice(-3) before filtering, so any text/reasoning parts in the tail would cause the check to silently pass. Fix uses filterCompactedEffect to search all messages and filters first before counting.

github-actions commented on Aug 12, 2026

@github-actions
Contributor

To stay organized issues are automatically closed after 60 days of no activity. If the issue is still relevant please open a new one.

paolothomas72 commented on Sep 20, 2026

@paolothomas72

Still present on 1.18.31. processor.ts still does MessageV2.parts(ctx.assistantMessage.id) then parts.slice(-DOOM_LOOP_THRESHOLD) then every().

New measurements on top of the original report:

  1. Cross-turn / exact-JSON miss (production shape). Failed edit, identical failed edit, then read (or a third edit with a different oldString). No doom_loop permission. Count never reaches 3 identical payloads. steps unset → no iteration cap.

  2. Slice-before-filter miss inside one message. Five identical failed edits in a single assistant message. JSON event stream (opencode run --format json) had no doom_loop permission event. A non-tool part in the last-three window (step-start / text / reasoning) makes every() false. --auto was used, so this is “no permission event in the stream,” not a claim about a hidden TUI dialog.

  3. User-space backstops that do work (do not fix the detector):

    • A ~/.config/opencode/plugins/ hook that counts identical edits and failed edits on the same file across the session, throw on the 3rd (tool.execute.before). Live: 3rd identical failed edit aborted with that error; file unchanged.
    • Explicit agent.<name>.steps (no default in docs). Live steps: 3 → third iteration was text-only CRITICAL - MAXIMUM STEPS REACHED.
    • Plugin config hook can fill steps: 50 when omitted. MCP tools are not wrapped by tool.execute.*.

PRs #25255 / #32089 did not land in 1.18.31. Minimum that would have caught (1) and (2): filter matching tool+input first, count since last user turn, and give unset agents a default steps cap. A second rule for “same path + N consecutive failures” would catch the retry-then-read loop the JSON matcher will never see.

Happy to add a redacted JSONL of the 5× identical-edit stream if useful.

Field notes collecting this and related 1.18.x findings: #50073

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions