Skip to content

Subagent enters infinite retry loop on edit/write tool failure, causing excessive API costs #17169

Description

@cauboy

Description

When a subagent attempts to use the edit tool and it fails (e.g., due to invalid arguments), the subagent enters an infinite retry loop—continuously attempting the same failed operation without stopping. This causes excessive API usage and costs ($15+ per subagent invocation in some cases).

Steps to Reproduce:

  1. Invoke a subagent (e.g., call_omo_agent) with a task requiring the edit tool
  2. The edit tool fails (e.g., invalid parameters, schema error)
  3. The subagent retries the same failed call repeatedly
  4. No circuit breaker or max retry limit stops the loop

Expected Behavior:

  • After N failed attempts with the same error, the subagent should fail gracefully
  • Add a "giving up" message so users understand why the operation stopped
  • Prevent infinite retry loops that consume API credits

Impact:

  • Financial: $15+ per failed subagent call due to retry loops (this issue was costing me >$100 in total)
  • User Experience: Confusing, unclear what went wrong
  • Reliability: Subagent unreliable for edit operations

Requested Fix:

  • Add circuit breaker / max retry limit for subagent tool failures
  • Graceful failure after max retries with clear error message

Plugins

No response

OpenCode version

No response

Steps to reproduce

No response

Screenshot and/or share link

Image

Operating System

Sequoia 15.1.1

Terminal

iTerm2

Activity

added
coreAnything pertaining to core functionality of the application (opencode server stuff)
perfIndicates a performance issue or need for optimization
on Mar 12, 2026

github-actions commented on Mar 12, 2026

@github-actions
Contributor

This issue might be a duplicate of existing issues. Please check:

changed the title [-]Subagent enters infinite retry loop on edit tool failure, causing excessive API costs[/-] [+]Subagent enters infinite retry loop on edit/write tool failure, causing excessive API costs[/+] on Mar 12, 2026
removed
bugSomething isn't working
perfIndicates a performance issue or need for optimization
coreAnything pertaining to core functionality of the application (opencode server stuff)
on May 3, 2026

Keesan12 commented on May 12, 2026

@Keesan12

I’d split the fix into two layers: per-call retry policy at the tool boundary, and attempt-level admission at the loop boundary.

The first stops identical edit/write failures from hammering the same tool forever. The second decides whether the next attempt should run at all after N same-class failures. The important part is to emit a first-class stop reason, not just "gave up," so users can tell schema failure from budget failure from verifier failure. We ended up caring a lot about that distinction in MartinLoop because the expensive bug is usually not the single tool failure, it’s the ungoverned attempt that keeps going after the failure is already obvious.

Keesan12 commented on May 22, 2026

@Keesan12

The important retry boundary here is whether the tool failure is local to one action or evidence that the whole attempt should stop. A tiny stop receipt with failure class, verifier movement, and next-attempt admissibility would keep that line sharp.

Keesan12 commented on Jun 5, 2026

@Keesan12

This is exactly the kind of bug that makes people think the model is the expensive part when the real problem is the loop policy. The useful split is usually by failure class: transient, deterministic, permission, verifier, and unknown. Only one of those deserves automatic retries. The rest should stop early with a receipt so the parent agent does not keep buying the same mistake again. We ended up building MartinLoop around that exact boundary after seeing retry loops become the most expensive part of agent work.

JustSidus commented on Jun 12, 2026

@JustSidus

Related: opened PR #32089 which fixes the doom loop detection scope bug. That PR addresses the cross-message detection issue in processor.ts. The subagent retry loop issue you describe here is a separate problem that would need a different fix (circuit breaker at the subagent/task level, not the tool-call level).

github-actions commented on Aug 12, 2026

@github-actions
Contributor

To stay organized issues are automatically closed after 60 days of no activity. If the issue is still relevant please open a new one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions