Skip to content

Bug: limit.output in config is silently capped at 32k; OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX is a poor workaround #29363

Description

@g199209

Summary

OpenCode silently caps per-step maxOutputTokens at 32,000 even when opencode.json sets a much larger limit.output (e.g. 384000 for DeepSeek, 128000 for GPT/Claude). The only documented escape hatch is an experimental environment variable, OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX, which users must discover via scattered GitHub threads.

This is surprising, breaks agent runs on reasoning-heavy models, and makes user configuration misleading.

Expected behavior

When a model defines explicit limits in config:

"deepseek-v4-flash": {
  "limit": {
    "context": 1000000,
    "output": 384000
  }
}

OpenCode should send max_output_tokens / max_tokens based on limit.output (possibly clamped to what the provider actually supports), not an undocumented global ceiling of 32k.

Actual behavior

ProviderTransform.maxOutputTokens() always applies a default cap:

export const OUTPUT_TOKEN_MAX = 32_000

export function maxOutputTokens(model: Provider.Model, outputTokenMax = OUTPUT_TOKEN_MAX): number {
  return Math.min(model.limit.output, outputTokenMax) || outputTokenMax
}

Unless OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX is set in the environment, min(384000, 32000) === 32000 is what gets sent to the API.

Real-world impact (agent benchmark)

On a coding task with reasoningEffort: "max" via @ai-sdk/openai-compatible:

  • A step finishes with reason: "length"
  • Token usage: reasoning: 32000, output: 0
  • The agent stops mid-task with no visible completion

The same model with reasoningEffort: "high" completes normally in the same step budget because reasoning does not consume the entire cap.

This is not a provider failure; logs show OpenCode requesting ~32k max output while config says 384k.

Why OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX feels like the wrong fix

  1. Misleading config — Users set limit.output in opencode.json believing it controls API behavior. It does not, unless they also export an env var whose name does not reference limit.output at all.

  2. Experimental / undiscoverable — The variable is buried under “experimental” flags in CLI docs. Issue max_tokens defaults to 32000 when using a custom provider #1735 was auto-closed; Custom provider (LM Studio) ignores limit.output config, hardcodes max_tokens to 32000 #20078 remains open with the same root cause.

  3. Global knob for per-model limits — Setting OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=1048576 is a shotgun fix across all models instead of respecting per-model limit.output.

  4. Still does not fix overflow semantics — overflow.ts uses context - maxOutputTokens() when limit.input is absent. Models with output close to context (e.g. Kimi K2.6 at 262144/262144) get usable === 0 once the env var is raised, causing immediate compaction/overflow unless users also add limit.input — another undocumented workaround.

Related issues

Suggested direction

  1. Respect configured limit.output by default — Use OUTPUT_TOKEN_MAX (32k) only as a fallback when limit.output is missing or zero, not as Math.min(limit.output, 32k) for every model. (PR fix(provider): respect configured output limit #24384 attempted this but was closed as “not the right fix”; the problem remains.)

  2. Optional global ceiling — If a safety cap is still desired, expose it in opencode.json (e.g. limits.maxOutputTokens) rather than a cryptic env var.

  3. Decouple overflow math from per-step output cap — usable should not use context - limit.output when output can equal the full window; consider limit.input defaults or context - reserved for shared-window models.

  4. Document the contract — Clarify in config schema/docs what limit.context, limit.input, and limit.output each control (API request vs compaction vs fallback).

Environment

  • OpenCode: v1.14.x (local install from anomalyco/opencode)
  • Custom provider: @ai-sdk/openai-compatible → OpenAI-compatible gateway
  • Models: DeepSeek V4 Flash/Pro, Kimi K2.6, GPT 5.5, etc. with explicit limit blocks in opencode.json

Workaround today (ugly)

export OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=1048576

Plus per-model limit.input for shared-window models — none of which should be required when limit.output is already configured.

Activity

github-actions commented on May 26, 2026

@github-actions
Contributor

This issue might be a duplicate of existing issues. Please check:

If your issue adds meaningful new context beyond what's already tracked in #20078 (e.g., the overflow/compaction math issue with shared-window models, or the OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX critique), it may be worth keeping open — but please coordinate with the maintainers to avoid duplicate effort.

divitkashyap commented on May 27, 2026

@divitkashyap

Investigated this with the goal of submitting a fix — there are two layers to the cap and it's worth being explicit about both before anyone opens a PR:

Layer 1 — config is dropped on the floor. When the model comes from models.dev, configured limit.output in opencode.json was being discarded entirely in provider/provider.ts. @sjawhar's #29354 (open since yesterday) fixes this — configured limits are now merged over the models.dev defaults. That PR closes #21564 but doesn't reference this one, even though it's a prerequisite.

Layer 2 — even with a correct model.limit.output, the transform re-caps at 32k. packages/opencode/src/provider/transform.ts:1261:

export function maxOutputTokens(model: Provider.Model, outputTokenMax = OUTPUT_TOKEN_MAX): number {
  // ...uses Math.min(model.limit.output, outputTokenMax)
}

OUTPUT_TOKEN_MAX = 32_000 (transform.ts:18). OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX overrides only that ceiling, not the Math.min. So even after #29354 merges limit.output: 384000 into the model, this function still returns 32000.

Even after #29354 lands, this issue would still reproduce.

Question for maintainers (@thdxr / @simonklee): is the 32k cap deliberately a safety net for providers that genuinely can't handle larger outputs, or is it vestigial? If it's a safety net, the right fix is probably to apply it only when model.limit.output is absent or zero (i.e. treat an explicit non-zero limit.output as authoritative). Happy to put up that follow-up to #29354 if there's interest.

rekram1-node commented on May 27, 2026

@rekram1-node
Collaborator

We aren't ignoring configs here, this is how many tokens we cap the agent to reply with for a given step of a turn, the case where it reaches limit prolly needs additional handling but making a larger value here is uncommon, we alreadyy have a pretty standard value and some other coding agents make it even smaller. Happy to make the configuration option more exposed tho

The limit.output tracks the maximum supported output tokens from a model, we don't want to always use that cause then u could potentially use up the entire context window in a single turn also it makes compaction math more annoying

lexlian commented on May 27, 2026

@lexlian

Submit a PR #29513 , and it fix the issue:

  • it makes maxOutputTokens() use limit.output directly when configured (> 0), fall back to OUTPUT_TOKEN_MAX only when no limit is set, and preserve OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX as an optional safety ceiling.

Could consider deprecating OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX in favor of a global limits.maxOutputTokens in opencode.json. This would make the safety ceiling discoverable and configurable without requiring an experimental env var. Per-model limit.output stays fully respected below the ceiling.

rekram1-node commented on May 27, 2026

@rekram1-node
Collaborator

did u read my message i literally said not to do that

g199209 commented on May 27, 2026

@g199209
Author

We aren't ignoring configs here, this is how many tokens we cap the agent to reply with for a given step of a turn, the case where it reaches limit prolly needs additional handling but making a larger value here is uncommon, we alreadyy have a pretty standard value and some other coding agents make it even smaller. Happy to make the configuration option more exposed tho

The limit.output tracks the maximum supported output tokens from a model, we don't want to always use that cause then u could potentially use up the entire context window in a single turn also it makes compaction math more annoying

deepseek v4 model has 1M context window, and it can generate very long chain of thought, much longer than 32k. Limit output to 32k making the thinking process truncated. This make the whole program unusable.

lexlian commented on May 27, 2026

@lexlian

did u read my message i literally said not to do that

@rekram1-node I didn't actually 'see' your message before I posted my previous message as I was using gh cli tool.

I spent some time trying to summarize all these and hopefully we all could find our conversation valuable for future decision:

1. What does modern LLM models support

Provider Latest Model Context Window Max Output Tokens
OpenAI GPT-5.5 / GPT-5.4 ~1,000,000 tokens 128,000
Anthropic Claude Opus 4.7 1,000,000 tokens 128,000 (sync) / 300,000 (Batch API)
Google Gemini 3.5 Flash ~1,000,000 tokens 65,536
DeepSeek DeepSeek V4 Pro 1,000,000 tokens 384,000

2. Claude Code and Codex CLI current mechanism

Claude Code lets the model cap drive output (from online tech doc):

  • CLAUDE_CODE_MAX_OUTPUT_TOKENS 'Set the maximum number of output tokens for most requests. Defaults and caps vary by model;' env var defaults to model's own max output (128k Opus 4.7, 64k Sonnet 4.6)
  • Sync API: up to model max, Batch API: up to 300k via output-300k-2026-03-24 beta header
  • Docs: code.claude.com/env-vars → platform.claude.com/models

OpenAI Codex CLI hardcodes 128k for all GPT-5 models (from github info):

  • max_output_tokens: 128_000 in codex-rs/core/src/openai_model_info.rs
  • Users hit stream disconnected before completion: reason: max_output_tokens (GitHub issue #14753)
  • Not user-configurable — feature requested in issue #20861
  • Source: codex-rs/openai_model_info.rs (mentioned by issue #3523)

3. Possible directions

Option 1 Maintain current behavior and allow global config to adjust max output tokens.

Option 2 Allow per-model max output tokens overrides.

Option 3 Allow per-profile or per-job configuration. As someone may want aggressively truncated tool call output for certain projects. Others might need larger windows for long run tasks like log analysis and self driven code generation, etc.

rekram1-node commented on May 27, 2026

@rekram1-node
Collaborator

max output tokens for a model, and the requested output tokens are 2 separate concepts

g199209 commented on May 28, 2026

@g199209
Author

max output tokens for a model, and the requested output tokens are 2 separate concepts

@rekram1-node Agreed that model capability and per-request output cap are different concepts — this issue isn't about the concept.

The concrete failure I'm reporting (reproducible):

  • DeepSeek V4 Pro, limit.output: 384000, reasoningEffort: "max"
  • Single step ends with finish_reason: "length", reasoning: 32000, output: 0
  • Agent loop hard-stops. No retry, no continuation, no surfaced error.

This is not "uncommon to want a larger value." As you can see in @lexlian 's previous comment, for any frontier reasoning model in 2026 — GPT-5.5 reasoning=max, Claude Opus 4.7 extended thinking, DeepSeek V4 long CoT — reasoning tokens routinely exceed 32k in a single step. Under the current cap that's 0 visible output, ~100% of the time, not an edge case. Comparable tools already account for this: Claude Code defaults CLAUDE_CODE_MAX_OUTPUT_TOKENS to the model's own max (128k Opus / 64k Sonnet); Codex CLI hardcodes 128k for GPT‑5.

Keeping your concept distinction intact, the asks are orthogonal to it:

  1. Expose the per-request cap in opencode.json (non-experimental). You said you'd be happy to — this is that.
  2. Default the cap from model metadata rather than a fixed 32k.
  3. Handle finish_reason: "length" gracefully instead of a hard stop.
  4. Decouple compaction math in overflow.ts from the per-step cap — this is the legitimate concern you raised, and it's solvable independently of (1)–(3).

None of these collapse the capability-vs-request-cap distinction.

Concrete question I'd like a direct answer on: is opencode intended to be usable with reasoningEffort: high/max on frontier models today? If the answer is no, please state that explicitly in the official docs — something like "opencode is not designed for frontier models with long-form output / extended reasoning; do not use it for such workloads." That at least sets correct expectations, instead of letting users configure limit.output and silently get it capped at 32k.

added a commit that references this issue on May 28, 2026
f3c9a1c
added a commit that references this issue on Jun 1, 2026
e755957

Mte90 commented on Jun 16, 2026

@Mte90

I am facing a similar issue, globally for various model of a specific provider where the limits are configured for every single one but automatically is capped to 32k.

Cleroth commented on Jul 1, 2026

@Cleroth

Regardless of whatever output max is set to... opencode should be handling this gracefully and continue rather than stopping dead like the turn is complete.

1 remaining item

axisrow commented on Jul 20, 2026

@axisrow

+1 with an independent headless reproduction on OpenCode 1.17.18 using zai-coding-plan/glm-5.1.

We run OpenCode through opencode serve as a benchmark backend. On a real coding task, we have 12 persisted runs: 5 successes and 7 failures. Every one of the 7 failures has the same signature:

  • no final assistant result;
  • no tool-created artifact;
  • non-zero provider usage;
  • output_tokens + reasoning_tokens = 32,000 exactly;
  • reasoning consumes almost the entire budget (31,933–31,987 tokens);
  • the session then reaches idle/terminal state without a usable result.

Persisted failure samples:

elapsed output reasoning completion
915.1s 33 31,967 32,000
357.6s 36 31,964 32,000
470.1s 61 31,939 32,000

We repeated the same task twice with --no-save; both failed again at exactly 32,000 completion tokens:

  • 457.1s: output 33 + reasoning 31,967;
  • 517.7s: output 29 + reasoning 31,971.

Control prompts on the same provider/model succeed quickly (hello_world in 21.9s, fast_sort in 19.4s), so this is not a general provider outage. We also saw no ECONNRESET, malformed SSE, or network error in these failures.

Our wrapper currently reports this as a hung POST /message, but that is misleading: the provider clearly returned usage. The actual failure is consistent with the 32k per-step cap being exhausted by reasoning and OpenCode ending the turn without exposing/recovering from finish_reason=length.

One important operational detail: successful and failed run durations overlap substantially (successes 515–1,077s; failures 358–915s), so a wall-clock watchdog cannot reliably identify this condition. Preserving and surfacing the terminal finish reason would let headless callers fail immediately and accurately; auto-continuation or an explicit actionable output-limit error would be even better.

Full local reproduction data and proposed downstream handling: axisrow/llm_benchmark#161

Related behavior also matches #17471 and #18108.

axisrow commented on Jul 20, 2026

@axisrow

Follow-up with a live A/B run and downstream mitigation on OpenCode 1.17.18 with zai-coding-plan/glm-5.1, using the same real coding prompt and cleared proxy variables:

mode elapsed input output reasoning aggregate completion result
default 32k ceiling 596.0s 45,650 9,029 26,298 35,327 success + artifact
OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=65536 959.9s 66,072 11,689 51,094 62,783 success + artifact

The default run happened to succeed this time, which is consistent with the stochastic behavior in the earlier sample: nine independent failures ended at exactly 32,000 completion tokens, while other runs succeeded.

With the 65,536 override, the OpenCode log shows the initial build step running for about 784 seconds before the first tool action, after which the agent wrote the artifact and completed. The 62,783 figure is aggregate completion usage for the full run, not proof that one step consumed all of it; nevertheless, this successful run is consistent with the higher ceiling avoiding the repeatedly observed 32k termination.

We also implemented downstream handling in the benchmark wrapper:

  • expose an explicit output-token ceiling;
  • preserve and surface terminal finish=length / MessageOutputLengthError;
  • stop waiting for a missing session.idle once terminal length is observed;
  • optional first-action watchdog.

A live watchdog smoke test with a 5s threshold exited this workload in 34.1s instead of waiting 10–20 minutes. The current floor is the existing 30s POST /message read timeout.

Details and test results: axisrow/llm_benchmark#161

diegonix commented on Jul 30, 2026

@diegonix

Confirmed on my side with the built-in ollama-cloud/glm-5.2 model as well.

opencode models ollama-cloud --verbose reports the model with:

  • limit.context = 976000
  • limit.output = 131072

So the model metadata is being loaded correctly.

However, without OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX, OpenCode still behaves as if the effective output cap is 32000. In other words, this is not just bad provider metadata or a bad custom config entry: the model is recognized with the correct limit.output, and the 32k ceiling is being applied later in the request path.

That makes this issue broader than any one provider/model family. I was able to reproduce the same pattern on a built-in OpenAI-compatible path (ollama-cloud) with GLM 5.2.

DEAN-Cherry commented on Aug 8, 2026

@DEAN-Cherry

A few things here don't add up.

The min() doesn't do what the rationale says.

we don't want to always use that cause then u could potentially use up the entire context window in a single turn

The code is Math.min(model.limit.output, 32000). It only caps models whose limit.output is above 32k. Anything below 32k is sent verbatim, which is exactly the "eat the whole window in one turn" case the cap is supposed to prevent. So 32k isn't a principled limit, it's just a constant sitting in the middle.

In practice it also kills the config. Every model in my setup has limit.output >= 64k (DeepSeek 384k, Kimi 262144, Claude 64k), so min() returns 32000 for all of them and limit.output does nothing. I set 384k, the API gets 32k. That is ignoring the config.

The context-window worry only applies to shared-window models.

DeepSeek at 1M/384k can't fill the window in one turn: 384k output still leaves 616k for input. The only models where output can eat the window are shared-window ones like Kimi 262144/262144, and there the cap protects nothing anyway. OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=1048576 removes it in one line, and then usable = context - maxOutputTokens() = 0 and every session compaction-loops immediately.

32k isn't a standard value in 2026.

The numbers are already in this thread (models + Claude Code/Codex, repro): GPT-5.5 128k, Opus 4.7 128k, Gemini 3.5 65k, DeepSeek 384k. Claude Code doesn't hardcode one number either. It uses a per-model (default, upper) pair (mainstream 32k/64k, opus/1M 64k/128k, legacy 4k/8k) that CLAUDE_CODE_MAX_OUTPUT_TOKENS can raise up to upper. So 32k there is an overridable default, not a ceiling. opencode is the only one applying a flat 32k that config can't touch.

"Uncommon" is wrong too: the repro above shows 7/12 runs on 1.17.18 dying at exactly 32,000 tokens with reasoning eating the whole budget and no result, on GLM-5.1, months after this was filed.

The real bug is the part you already flagged.

the case where it reaches limit prolly needs additional handling

There is no handling. Hit the cap and the turn ends silently: reason: "length", output: 0, no result, no error. You can't even catch it by timeout, since successes and failures overlap (358–1077s).

The fix is a per-model clamp to the provider's real per-request cap, plus decoupling the overflow math via limit.input, not a global 32k constant. Related: #20078, #17471, #18108.

DEAN-Cherry commented on Aug 8, 2026

@DEAN-Cherry

Worth noting: this looks already fixed on the v2 branch, which is probably why the dev-targeted PRs (#29513, #24384) weren't taken — the whole thing was rewritten there rather than patched.

On v2 the outbound Math.min(limit.output, 32000) cap is gone. There's no maxOutputTokens() helper capping the request anymore, and OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX no longer exists. Each protocol just sends the caller's maxTokens:

  • packages/ai/src/protocols/openai-chat.ts → max_tokens: generation?.maxTokens
  • packages/ai/src/protocols/open-responses.ts → max_output_tokens: generation?.maxTokens
  • packages/ai/src/protocols/gemini.ts → maxOutputTokens: generation?.maxTokens
  • packages/ai/src/protocols/bedrock-converse.ts → maxTokens: generation?.maxTokens
  • packages/ai/src/protocols/anthropic-messages.ts → max_tokens: generation?.maxTokens ?? (limit.output ?? 4096)

Anthropic keeps a limit.output fallback because its API requires max_tokens; the others omit the field when it isn't set, which lets the model use its own limit instead of being clamped. So limit.output is no longer silently overridden.

OUTPUT_TOKEN_MAX = 32_000 survives in exactly one place — packages/core/src/session/compaction.ts:25 — and it's only used to decide when to auto-compact, not in any outbound request. The overflow math is also decoupled from the per-step cap now (compaction.ts:354):

const output = Math.min(limits?.output ?? 0, OUTPUT_TOKEN_MAX)
const promptCeiling = Math.min(
  limits?.input === undefined ? Number.POSITIVE_INFINITY : limits.input - config.buffer,
  context - Math.max(output, config.buffer),
)

That closes the shared-window collapse from the original report: for a 262144/262144 model, output is clamped to 32k for the compaction estimate only, so promptCeiling = 262144 - max(32000, buffer) ≈ 230144 instead of 0 — no immediate compaction loop, and no need for a limit.input workaround.

The one thing still worth deciding is whether the compaction estimate should keep a fixed 32k assumption at all, or derive from the actual per-step budget — but the core bug (config-ignored 32k cap on the API request) is resolved on v2. Might be worth confirming and closing this against the v2 rewrite.

kevin-elser commented on Sep 5, 2026

@kevin-elser

Worth noting: this looks already fixed on the v2 branch, which is probably why the dev-targeted PRs (#29513, #24384) weren't taken — the whole thing was rewritten there rather than patched.

On v2 the outbound Math.min(limit.output, 32000) cap is gone. There's no maxOutputTokens() helper capping the request anymore, and OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX no longer exists. Each protocol just sends the caller's maxTokens:

* `packages/ai/src/protocols/openai-chat.ts` → `max_tokens: generation?.maxTokens`

* `packages/ai/src/protocols/open-responses.ts` → `max_output_tokens: generation?.maxTokens`

* `packages/ai/src/protocols/gemini.ts` → `maxOutputTokens: generation?.maxTokens`

* `packages/ai/src/protocols/bedrock-converse.ts` → `maxTokens: generation?.maxTokens`

* `packages/ai/src/protocols/anthropic-messages.ts` → `max_tokens: generation?.maxTokens ?? (limit.output ?? 4096)`

Anthropic keeps a limit.output fallback because its API requires max_tokens; the others omit the field when it isn't set, which lets the model use its own limit instead of being clamped. So limit.output is no longer silently overridden.

OUTPUT_TOKEN_MAX = 32_000 survives in exactly one place — packages/core/src/session/compaction.ts:25 — and it's only used to decide when to auto-compact, not in any outbound request. The overflow math is also decoupled from the per-step cap now (compaction.ts:354):

const output = Math.min(limits?.output ?? 0, OUTPUT_TOKEN_MAX)
const promptCeiling = Math.min(
limits?.input === undefined ? Number.POSITIVE_INFINITY : limits.input - config.buffer,
context - Math.max(output, config.buffer),
)

That closes the shared-window collapse from the original report: for a 262144/262144 model, output is clamped to 32k for the compaction estimate only, so promptCeiling = 262144 - max(32000, buffer) ≈ 230144 instead of 0 — no immediate compaction loop, and no need for a limit.input workaround.

The one thing still worth deciding is whether the compaction estimate should keep a fixed 32k assumption at all, or derive from the actual per-step budget — but the core bug (config-ignored 32k cap on the API request) is resolved on v2. Might be worth confirming and closing this against the v2 rewrite.

The concept that v1 fixes for a product in which v1 is the default shipped are pushed away for months because of a pending v2 merge is one of the single dumbest traps in open source development.

Until merged, v1 is the product. Intentionally avoiding fixing the product because you have a fix set in some amorphous time in the future is dumb.

joeym82956 commented on Sep 15, 2026

@joeym82956

Confirmed still present on v1.18.30 and v1.18.31 (v1.18.31 changelog is ACP/TUI/Copilot only; OUTPUT_TOKEN_MAX is unchanged).

// packages/opencode/src/provider/transform.ts
export const OUTPUT_TOKEN_MAX = 32_000

export function maxOutputTokens(model: Provider.Model, outputTokenMax = OUTPUT_TOKEN_MAX): number {
  return Math.min(model.limit.output, outputTokenMax) || outputTokenMax
}

session/llm/request.ts then sends that as maxOutputTokens. Anthropic budget variants are also clamped with OUTPUT_TOKEN_MAX - 1.

What we see

Same silent 32k request cap on reasoning-heavy turns, including models whose catalog already advertises far more:

Model Advertised limit.output What OpenCode actually requests
anthropic/claude-fable-5-1 (Max) 128000 32000
anthropic/claude-opus-5 128000 32000
openai/gpt-5.6-sol 128000 32000
DeepSeek V4-class via @ai-sdk/openai-compatible 64k–384k 32000

On Fable 5.1 Max and DeepSeek thinking, a step can finish reason: "length" with reasoning consuming the entire 32k budget and little or no answer/tool call. Title/summary agents that set a small explicit cap are fine; the bug is the global clamp on ordinary/thinking turns.

OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX still works as a shotgun, but it is global and does not honor per-model limit.output.

Please use configured limit.output (and any smaller explicit request cap) instead of Math.min(limit.output, 32000) for every model.

dakotad74 commented on Sep 17, 2026

@dakotad74

Independent confirmation on v1.18.31 (stable) — same provider class, three models, reasoning-heavy sub-agent turns.

Provider is @ai-sdk/openai-compatible (custom gateway). Declared limits come from the provider catalog, not from a local limit override:

model limit.context limit.output
deepseek-v4-flash 1,000,000 384,000
glm5.3-flash 1,000,000 131,072
qwen3.8-flash 262,144 131,072

Every long-thinking turn ended finish_reason: length with no assistant text, at a shared ceiling across three different models:

model reasoning chars finish emitted text
deepseek-v4-flash 59,991–59,999 length none
qwen3.8-flash 59,996–59,997 length none
glm5.3-flash 59,994 length none

A turn on the same model that stayed under the ceiling (56,178 reasoning chars) finished stop with normal output, so the budget — not the model — is what ends the turn. (We measured characters from the session store, not tokens; ~60k chars is consistent with a 32,000-token cap for this tokenizer.)

This is still the exact code path in 1.18.31, not only 1.14.x. Extracted from the released binary:

OUTPUT_TOKEN_MAX:()=>M7}); ... var M7=32000 ...
function by($,Z=M7){return Math.min($.limit.output,Z)||Z}
outputTokenMax:G("OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX")
maxOutputTokens(e.model, e.flags.outputTokenMax)

With the env var unset, all three models receive min(limit.output, 32000) === 32000 regardless of the 384k / 131k declared in config.

The documented workaround does work: raising OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX above the declared limit makes each model receive its own limit.output. On the overflow caveat mentioned above, all three of our models declare limit.context well above limit.output, so the context - maxOutputTokens() path stays positive (616k / 869k / 131k) and we did not need limit.input.

The suggested direction still stands: use the constant as a fallback (limit.output || OUTPUT_TOKEN_MAX) instead of Math.min(limit.output, OUTPUT_TOKEN_MAX), and the env var becomes unnecessary.


Correction / update (2026-09-17, measured later the same day).

We instrumented our own outbound requests with a local logging proxy (OPENCODE_CONFIG_CONTENT overriding the provider baseURL) and can now separate two things we had previously conflated:

  • The clamp is real and the env var lifts it. On 1.18.31 with OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=1048576, the request body carries max_tokens: 384000 for the title, primary, sub-agent and post-tool steps. Long turns do complete: we observed a sub-agent turn of 59,013 output tokens ending finish: stop, and a review-shaped request of 38,222 completion tokens. So Math.min(limit.output, OUTPUT_TOKEN_MAX) is bypassed exactly as described.

  • The shared ~59.9k-char ceiling is not explained by the opencode clamp alone. The same truncation (finish_reason: length, reasoning ~59,994 chars, no assistant text, no usage block) reproduces when calling the provider's /v1/chat/completions directly, with opencode out of the loop, using max_tokens: 384000. The opencode 32k clamp and this provider-side ceiling coincide at ~32k tokens, which is why session data alone could not tell them apart.

We therefore withdraw the "so the budget - not the model - is what ends the turn" attribution for the residual case, and we have moved the remaining report to the provider. The suggested direction here still stands on its own: limit.output || OUTPUT_TOKEN_MAX is strictly better than Math.min(limit.output, OUTPUT_TOKEN_MAX), and the env var becomes unnecessary.

We also withdraw the char-to-token estimate: measured ratios ranged from ~1.85 to ~3.5 chars/token across turns, so ~60k chars is not a reliable proxy for a 32,000-token cap.

benjamin-bonneton commented on Sep 20, 2026

@benjamin-bonneton

Same issue when using : ollama launch opencode --model glm-5.3-flash:cloud

dotCipher commented on Oct 1, 2026

@dotCipher

Any updates on this? I'm running into the same problem

simonklee commented on Oct 1, 2026

@simonklee
Member

Yes, this is fixed in latest v2 release

kevin-elser commented on Oct 1, 2026

@kevin-elser

simonklee commented on Oct 1, 2026

@simonklee
Member

Cool so merge v2 as main because as is you're refusing to fix things in v1
for months because "it's fixed in v2".
…

We don’t really use main. V2 is the default version on our website for some time already.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions