Repository navigation
[FEATURE]: Auto-continue when model hits output token limit (finish_reason: "length") #17471
Description
Activity
+1 with a headless data point. In unattended usage (no interactive user), this is not just an annoyance, it becomes a hard failure: there is no one to type "continue," so a turn that stops at the output cap returns no final answer and the caller records the whole step as failed. Several of those in a row can abort an unattended run.
Repro: a reasoning model behind an OpenAI-compatible provider, with a low effective per-step output budget. Turns that stayed under the budget produced a final answer and succeeded; turns that hit the ceiling produced no final answer. Raising the output cap reduced the failures, which is consistent with finish_reason: "length" ending the turn early.
Auto-continue on finish_reason: "length" (paired with honoring limit.output per #29363) would make opencode usable as a backend for unattended runs. Happy to test a build if one lands.
Update(2026-09-16): deepseek has fixed this issue on their end.
Additional data point from the oh-my-pi (omp) agent harness, where the same symptom is reported for opencode-go/deepseek-v4-flash (and the zen free tier): output stops mid-sentence with no warning, and retries are needed to finish a turn.
I verified what the omp client actually sends to the zen/go gateway (installed 17.2.9, by running its own request-resolver code):
- The opencode-go model entry ships
compat.maxTokensField: "max_tokens"andmaxTokens: 384000, so the wire body carries the correct field:max_tokens: 64000(omp clamps every openai-completions request to its global 64000OPENAI_MAX_OUTPUT_TOKENSunless the provider is whitelisted — GLM-5.2/ZAI and Moonshot K3 get their catalog caps lifted, deepseek does not).
So this is not a max_completion_tokens-vs-max_tokens field mismatch on the client side: the gateway receives max_tokens: 64000 and still truncates well below that (reports say after a few lines). That points at the gateway's own per-turn output budget — the OUTPUT_TOKEN_MAX (default 32K) cap described here — being applied to deepseek-v4-flash on zen/go with finish_reason: "length", and the client recording the turn as a normal completion (#40146) so nothing surfaces to the user.
Two asks, mirroring this thread:
- Confirm/raise the zen/go
OUTPUT_TOKEN_MAXfor deepseek-v4-flash (catalog advertises up to 384K output; a 32K cap truncates real work). - Surface
finish_reason: "length"to clients (warning + auto-continue), per this feature request — silent truncation breaks autonomous agents hard, as noted in the comments.
Filed #42900 before seeing this; closing it as a duplicate. Adding the data from it here since it comes from a variant neither this issue nor #18108 covers, and it strengthens the case for the prompt.ts fix.
The zero-output variant. Both issues here describe output that is truncated — content or tool-call JSON cut off mid-stream. We hit the same finish_reason: "length" with zero output tokens: the entire budget consumed by reasoning, no partial content, no tool call, nothing to continue from. Same code path, same fix, but worth knowing the turn can produce literally no output rather than a truncated one.
In headless opencode run this is worse than "waits for user input." The process exits, and to a process-liveness supervisor it is indistinguishable from success — ours logged "FINISHED clean" for two such deaths. One lost 59 uncommitted edits across 22 files about 50 minutes in. Frequency ~1 in 10 long headless runs.
Raising the output cap does not help, measured. After raising 8192 → 16384 → 32768, runs died at 8,191 / 16,383 / 31,999 reasoning tokens respectively, each with zero output — a runaway turn consumes essentially whatever budget it is given, so raising the cap only buys more wasted generation per death. Healthy steps in the same runs used 30–730 reasoning tokens, two orders of magnitude below the ceiling, so there is no gradual signal to alert on either.
The recovery path already exists and works, which supports P1 here: opencode run --session <id> "<continue instruction>" resumed with full prior context and completed the task with no re-verification. The loop just never takes that path itself.
Observed on 1.18.18, macOS, thinking-capable local models via LM Studio (OpenAI-compatible).
Investigated with Claude Code.
Real-world case
We observed this behavior on OpenCode 1.18.25 with an OpenAI-compatible local provider:
| Event | Result |
|---|---|
| Completed agent steps | 17 normal tool-driven steps |
| Final response duration | Approximately 13 minutes |
| Configured output limit | 16,384 tokens |
| Actual output | Exactly 16,384 tokens |
| Provider finish reason | length |
| OpenCode behavior | Logged exiting loop and left the unfinished task idle |
The model server stayed healthy, and the same session remained resumable through a manual Continue message. The interruption was specifically caused by OpenCode treating length as a terminal success instead of a truncation signal.
Why raising the limit is not enough
Increasing maxOutputTokens only postpones the same failure. It can also waste substantially more generation time before OpenCode silently stops.
Suggested behavior
When a response ends with finish_reason: "length", OpenCode should:
- Show a visible notice that the response reached its output limit.
- Add a concise synthetic continuation instruction.
- Continue within the same session, preserving conversation state and provider-side prompt-cache reuse.
- Limit consecutive automatic continuations to prevent runaway reasoning or repetition.
- If that safety limit is reached, stop with an explicit, actionable error instead of recording a normal completion.
A bounded automatic continuation would preserve autonomous task execution without creating an uncontrolled loop.
Another headless data point, and a mechanism I don't think is in this thread yet: whether a length finish ends the run depends on the auto-compaction threshold, which makes it fatal precisely when there is context to spare.
Terminal-Bench 4.0, OpenCode 1.18.31, OpenAI-compatible local provider, model limit context: 131072, output: 65536, effective cap 32,000 (#29363). Five truncations across three trials:
| context at truncation | >= usable | outcome |
|---|---|---|
| 42,975 | no | run ended |
| 48,595 | no | run ended |
| 99,918 | yes | continued |
| 124,651 | yes | continued |
| 66,729 | no | run ended |
usable = 131,072 - min(65,536, 32,000) = 99,072, so the threshold predicts all five.
The two that continued did so because compaction.isOverflow fired on the preceding step. The compaction inserts a new user message, so the lastAssistant.parentID === lastUser.id clause at prompt.ts:1115 no longer holds and the exit at :1113 is skipped. Nothing branches on length, as #40146 notes, so the rescue is incidental rather than intended.
This also qualifies the point above that raising the limit only postpones the failure. Here it also lowers the compaction threshold, since usable is context minus the cap. Raising the cap to 65,536 would move the threshold from 99,072 to 65,536, so truncations that still occur would be more likely to compact and continue instead of ending the run.
Cost in this run: three of eight scored trials ended this way, one of them after 6.9 hours of work, all recorded with no error.
Closing as a duplicate of #18108, which tracks how OpenCode handles a model hitting its output token limit (finishReason: "length"), including continuing automatically. Please follow and 👍 there.
mianabdulrehman1994-boop commented on Oct 2, 2026
Headless exit-0 on length truncation — recovery path works, automation missing
Confirmed: Headless opencode run exits 0 on finishReason: "length" (session loop treats it as success). Frequency ~1/10 long runs. One run lost 59 uncommitted edits across 22 files. Raising output cap doesn't help (runaway consumes whatever budget).
Recovery path verified: opencode run --session <id> "continue" resumes with full context and completes task. But resume_attempts=0 across all DBs — nothing automates this.
Auto-continue blocked by:
session.idlenot emitted in headless mode (0/879 build sessions)prompt.ts:1113clause (lastAssistant.parentID === lastUser.id) incidentally rescues via compaction, not by design- No watchdog / supervisor calls the recovery path
Deployed workaround:
- V2 stop-gate plugin on
session.idle(Desktop) + headless watchdog (LaunchAgent) for V1opencode runsessions - Watchdog uses
--attachfor V2 (avoids DB conflict) and proper XDG for V1 - Shared state file coordinates gate ↔ watchdog (inject count, NEEDS_USER flag)
Requested upstream:
- Make headless emit
session.idle(or equivalent) so continuation plugins work - Exit code: length-truncation should be non-zero (distinguish from clean completion)
- Expose
--session <id> "continue"as a documented recovery API
Related: #18108 (core truncation chain), #26063 (5min Undici timeout).
Feature hasn't been suggested before.
Describe the enhancement you want to request
When using models with large context windows (e.g., Claude Opus 4.6 with 1 million input context), the input context is rarely exhausted. However, the per-response output token cap (
OUTPUT_TOKEN_MAX, default 32K) frequently causes the model to stop mid-task withfinish_reason: "length". The user must then manually type "continue" to resume — which breaks the autonomous agent workflow.Current behavior:
The loop in
prompt.tsexits whenfinish_reasonis anything other than"tool-calls"or"unknown":When finish_reason is "length", the loop breaks and waits for user input.
Requested behavior:
When finish_reason is "length", it would be nice for OpenCode to automatically inject a synthetic "continue" user message and re-enter the loop — the same pattern already used for subtask summaries. The model's output was truncated, not intentionally finished, so the loop should continue.
Suggested implementation