Repository navigation
Subagent enters infinite retry loop on edit/write tool failure, causing excessive API costs #17169
Description
Activity
github-actions commented on Mar 12, 2026
This issue might be a duplicate of existing issues. Please check:
- Infinite retry loop when StreamIdleTimeoutError occurs during tool input generation #12234: Similar infinite retry loop when a tool fails repeatedly, also causing excessive API costs — though that issue focuses on stream timeouts, the underlying retry loop behavior and cost impact are very similar.
I’d split the fix into two layers: per-call retry policy at the tool boundary, and attempt-level admission at the loop boundary.
The first stops identical edit/write failures from hammering the same tool forever. The second decides whether the next attempt should run at all after N same-class failures. The important part is to emit a first-class stop reason, not just "gave up," so users can tell schema failure from budget failure from verifier failure. We ended up caring a lot about that distinction in MartinLoop because the expensive bug is usually not the single tool failure, it’s the ungoverned attempt that keeps going after the failure is already obvious.
The important retry boundary here is whether the tool failure is local to one action or evidence that the whole attempt should stop. A tiny stop receipt with failure class, verifier movement, and next-attempt admissibility would keep that line sharp.
This is exactly the kind of bug that makes people think the model is the expensive part when the real problem is the loop policy. The useful split is usually by failure class: transient, deterministic, permission, verifier, and unknown. Only one of those deserves automatic retries. The rest should stop early with a receipt so the parent agent does not keep buying the same mistake again. We ended up building MartinLoop around that exact boundary after seeing retry loops become the most expensive part of agent work.
Related: opened PR #32089 which fixes the doom loop detection scope bug. That PR addresses the cross-message detection issue in processor.ts. The subagent retry loop issue you describe here is a separate problem that would need a different fix (circuit breaker at the subagent/task level, not the tool-call level).
github-actions commented on Aug 12, 2026
To stay organized issues are automatically closed after 60 days of no activity. If the issue is still relevant please open a new one.
Description
When a subagent attempts to use the edit tool and it fails (e.g., due to invalid arguments), the subagent enters an infinite retry loop—continuously attempting the same failed operation without stopping. This causes excessive API usage and costs ($15+ per subagent invocation in some cases).
Steps to Reproduce:
Expected Behavior:
Impact:
Requested Fix:
Plugins
No response
OpenCode version
No response
Steps to reproduce
No response
Screenshot and/or share link
Operating System
Sequoia 15.1.1
Terminal
iTerm2