Skip to content

[FEATURE]: Auto-continue when model hits output token limit (finish_reason: "length") #17471

Description

@stephenstubbs

Feature hasn't been suggested before.

  • I have verified this feature I'm about to request hasn't been suggested before.

Describe the enhancement you want to request

When using models with large context windows (e.g., Claude Opus 4.6 with 1 million input context), the input context is rarely exhausted. However, the per-response output token cap (OUTPUT_TOKEN_MAX, default 32K) frequently causes the model to stop mid-task with finish_reason: "length". The user must then manually type "continue" to resume — which breaks the autonomous agent workflow.
Current behavior:
The loop in prompt.ts exits when finish_reason is anything other than "tool-calls" or "unknown":

if (
  lastAssistant?.finish &&
  !["tool-calls", "unknown"].includes(lastAssistant.finish) &&
  lastUser.id < lastAssistant.id
) {
  break
}

When finish_reason is "length", the loop breaks and waits for user input.
Requested behavior:
When finish_reason is "length", it would be nice for OpenCode to automatically inject a synthetic "continue" user message and re-enter the loop — the same pattern already used for subtask summaries. The model's output was truncated, not intentionally finished, so the loop should continue.
Suggested implementation

  1. Add "length" to the continuation conditions in prompt.ts, or handle it separately
  2. When finish_reason === "length" is detected, inject a synthetic user message (e.g., "Continue from where you left off.") and continue the loop instead of breaking
  3. The infrastructure for synthetic user messages already exists (used for subtask command summaries)
  4. Consider adding a config option (compaction.auto_continue or similar) to make this opt-in/out
  • The MessageV2.OutputLengthError type is already defined in message-v2.ts but appears unused
  • The plugin system cannot solve this, as there is no hook between the finish reason being set and the loop exit decision

Activity

added
discussionUsed for feature requests, proposals, ideas, etc. Open discussion
on Mar 14, 2026
added
coreAnything pertaining to core functionality of the application (opencode server stuff)
on Mar 14, 2026
added a commit that references this issue on Apr 9, 2026
a266d02
removed
discussionUsed for feature requests, proposals, ideas, etc. Open discussion
coreAnything pertaining to core functionality of the application (opencode server stuff)
on May 3, 2026

ooiyeefei commented on Jun 27, 2026

@ooiyeefei

+1 with a headless data point. In unattended usage (no interactive user), this is not just an annoyance, it becomes a hard failure: there is no one to type "continue," so a turn that stops at the output cap returns no final answer and the caller records the whole step as failed. Several of those in a row can abort an unattended run.

Repro: a reasoning model behind an OpenAI-compatible provider, with a low effective per-step output budget. Turns that stayed under the budget produced a final answer and succeeded; turns that hit the ceiling produced no final answer. Raising the output cap reduced the failures, which is consistent with finish_reason: "length" ending the turn early.

Auto-continue on finish_reason: "length" (paired with honoring limit.output per #29363) would make opencode usable as a backend for unattended runs. Happy to test a build if one lands.

iacore commented on Aug 8, 2026

@iacore

Update(2026-09-16): deepseek has fixed this issue on their end.

Additional data point from the oh-my-pi (omp) agent harness, where the same symptom is reported for opencode-go/deepseek-v4-flash (and the zen free tier): output stops mid-sentence with no warning, and retries are needed to finish a turn.

I verified what the omp client actually sends to the zen/go gateway (installed 17.2.9, by running its own request-resolver code):

  • The opencode-go model entry ships compat.maxTokensField: "max_tokens" and maxTokens: 384000, so the wire body carries the correct field: max_tokens: 64000 (omp clamps every openai-completions request to its global 64000 OPENAI_MAX_OUTPUT_TOKENS unless the provider is whitelisted — GLM-5.2/ZAI and Moonshot K3 get their catalog caps lifted, deepseek does not).

So this is not a max_completion_tokens-vs-max_tokens field mismatch on the client side: the gateway receives max_tokens: 64000 and still truncates well below that (reports say after a few lines). That points at the gateway's own per-turn output budget — the OUTPUT_TOKEN_MAX (default 32K) cap described here — being applied to deepseek-v4-flash on zen/go with finish_reason: "length", and the client recording the turn as a normal completion (#40146) so nothing surfaces to the user.

Two asks, mirroring this thread:

  1. Confirm/raise the zen/go OUTPUT_TOKEN_MAX for deepseek-v4-flash (catalog advertises up to 384K output; a 32K cap truncates real work).
  2. Surface finish_reason: "length" to clients (warning + auto-continue), per this feature request — silent truncation breaks autonomous agents hard, as noted in the comments.

njhillson commented on Aug 16, 2026

@njhillson

Filed #42900 before seeing this; closing it as a duplicate. Adding the data from it here since it comes from a variant neither this issue nor #18108 covers, and it strengthens the case for the prompt.ts fix.

The zero-output variant. Both issues here describe output that is truncated — content or tool-call JSON cut off mid-stream. We hit the same finish_reason: "length" with zero output tokens: the entire budget consumed by reasoning, no partial content, no tool call, nothing to continue from. Same code path, same fix, but worth knowing the turn can produce literally no output rather than a truncated one.

In headless opencode run this is worse than "waits for user input." The process exits, and to a process-liveness supervisor it is indistinguishable from success — ours logged "FINISHED clean" for two such deaths. One lost 59 uncommitted edits across 22 files about 50 minutes in. Frequency ~1 in 10 long headless runs.

Raising the output cap does not help, measured. After raising 8192 → 16384 → 32768, runs died at 8,191 / 16,383 / 31,999 reasoning tokens respectively, each with zero output — a runaway turn consumes essentially whatever budget it is given, so raising the cap only buys more wasted generation per death. Healthy steps in the same runs used 30–730 reasoning tokens, two orders of magnitude below the ceiling, so there is no gradual signal to alert on either.

The recovery path already exists and works, which supports P1 here: opencode run --session <id> "<continue instruction>" resumed with full prior context and completed the task with no re-verification. The loop just never takes that path itself.

Observed on 1.18.18, macOS, thinking-capable local models via LM Studio (OpenAI-compatible).

Investigated with Claude Code.

BrunoCerberus commented on Aug 31, 2026

@BrunoCerberus

Real-world case

We observed this behavior on OpenCode 1.18.25 with an OpenAI-compatible local provider:

Event Result
Completed agent steps 17 normal tool-driven steps
Final response duration Approximately 13 minutes
Configured output limit 16,384 tokens
Actual output Exactly 16,384 tokens
Provider finish reason length
OpenCode behavior Logged exiting loop and left the unfinished task idle

The model server stayed healthy, and the same session remained resumable through a manual Continue message. The interruption was specifically caused by OpenCode treating length as a terminal success instead of a truncation signal.

Why raising the limit is not enough

Increasing maxOutputTokens only postpones the same failure. It can also waste substantially more generation time before OpenCode silently stops.

Suggested behavior

When a response ends with finish_reason: "length", OpenCode should:

  1. Show a visible notice that the response reached its output limit.
  2. Add a concise synthetic continuation instruction.
  3. Continue within the same session, preserving conversation state and provider-side prompt-cache reuse.
  4. Limit consecutive automatic continuations to prevent runaway reasoning or repetition.
  5. If that safety limit is reached, stop with an explicit, actionable error instead of recording a normal completion.

A bounded automatic continuation would preserve autonomous task execution without creating an uncontrolled loop.

beatakouchnir commented on Sep 23, 2026

@beatakouchnir

Another headless data point, and a mechanism I don't think is in this thread yet: whether a length finish ends the run depends on the auto-compaction threshold, which makes it fatal precisely when there is context to spare.

Terminal-Bench 4.0, OpenCode 1.18.31, OpenAI-compatible local provider, model limit context: 131072, output: 65536, effective cap 32,000 (#29363). Five truncations across three trials:

context at truncation >= usable outcome
42,975 no run ended
48,595 no run ended
99,918 yes continued
124,651 yes continued
66,729 no run ended

usable = 131,072 - min(65,536, 32,000) = 99,072, so the threshold predicts all five.

The two that continued did so because compaction.isOverflow fired on the preceding step. The compaction inserts a new user message, so the lastAssistant.parentID === lastUser.id clause at prompt.ts:1115 no longer holds and the exit at :1113 is skipped. Nothing branches on length, as #40146 notes, so the rescue is incidental rather than intended.

This also qualifies the point above that raising the limit only postpones the failure. Here it also lowers the compaction threshold, since usable is context minus the cap. Raising the cap to 65,536 would move the threshold from 99,072 to 65,536, so truncations that still occur would be more likely to compact and continue instead of ending the run.

Cost in this run: three of eight scored trials ended this way, one of them after 6.9 hours of work, all recorded with no error.

neriousy commented on Oct 1, 2026

@neriousy
Member

Closing as a duplicate of #18108, which tracks how OpenCode handles a model hitting its output token limit (finishReason: "length"), including continuing automatically. Please follow and 👍 there.

mianabdulrehman1994-boop commented on Oct 2, 2026

@mianabdulrehman1994-boop

Headless exit-0 on length truncation — recovery path works, automation missing

Confirmed: Headless opencode run exits 0 on finishReason: "length" (session loop treats it as success). Frequency ~1/10 long runs. One run lost 59 uncommitted edits across 22 files. Raising output cap doesn't help (runaway consumes whatever budget).

Recovery path verified: opencode run --session <id> "continue" resumes with full context and completes task. But resume_attempts=0 across all DBs — nothing automates this.

Auto-continue blocked by:

  • session.idle not emitted in headless mode (0/879 build sessions)
  • prompt.ts:1113 clause (lastAssistant.parentID === lastUser.id) incidentally rescues via compaction, not by design
  • No watchdog / supervisor calls the recovery path

Deployed workaround:

  • V2 stop-gate plugin on session.idle (Desktop) + headless watchdog (LaunchAgent) for V1 opencode run sessions
  • Watchdog uses --attach for V2 (avoids DB conflict) and proper XDG for V1
  • Shared state file coordinates gate ↔ watchdog (inject count, NEEDS_USER flag)

Requested upstream:

  1. Make headless emit session.idle (or equivalent) so continuation plugins work
  2. Exit code: length-truncation should be non-zero (distinguish from clean completion)
  3. Expose --session <id> "continue" as a documented recovery API

Related: #18108 (core truncation chain), #26063 (5min Undici timeout).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions