Repository navigation
Bug: limit.output in config is silently capped at 32k; OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX is a poor workaround #29363
Description
Activity
github-actions commented on May 26, 2026
This issue might be a duplicate of existing issues. Please check:
- Custom provider (LM Studio) ignores limit.output config, hardcodes max_tokens to 32000 #20078: Same root cause —
limit.outputconfig ignored,max_tokenshardcoded to 32000 for custom providers (open, assigned)
If your issue adds meaningful new context beyond what's already tracked in #20078 (e.g., the overflow/compaction math issue with shared-window models, or the OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX critique), it may be worth keeping open — but please coordinate with the maintainers to avoid duplicate effort.
Investigated this with the goal of submitting a fix — there are two layers to the cap and it's worth being explicit about both before anyone opens a PR:
Layer 1 — config is dropped on the floor. When the model comes from models.dev, configured limit.output in opencode.json was being discarded entirely in provider/provider.ts. @sjawhar's #29354 (open since yesterday) fixes this — configured limits are now merged over the models.dev defaults. That PR closes #21564 but doesn't reference this one, even though it's a prerequisite.
Layer 2 — even with a correct model.limit.output, the transform re-caps at 32k. packages/opencode/src/provider/transform.ts:1261:
export function maxOutputTokens(model: Provider.Model, outputTokenMax = OUTPUT_TOKEN_MAX): number {
// ...uses Math.min(model.limit.output, outputTokenMax)
}OUTPUT_TOKEN_MAX = 32_000 (transform.ts:18). OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX overrides only that ceiling, not the Math.min. So even after #29354 merges limit.output: 384000 into the model, this function still returns 32000.
Even after #29354 lands, this issue would still reproduce.
Question for maintainers (@thdxr / @simonklee): is the 32k cap deliberately a safety net for providers that genuinely can't handle larger outputs, or is it vestigial? If it's a safety net, the right fix is probably to apply it only when model.limit.output is absent or zero (i.e. treat an explicit non-zero limit.output as authoritative). Happy to put up that follow-up to #29354 if there's interest.
We aren't ignoring configs here, this is how many tokens we cap the agent to reply with for a given step of a turn, the case where it reaches limit prolly needs additional handling but making a larger value here is uncommon, we alreadyy have a pretty standard value and some other coding agents make it even smaller. Happy to make the configuration option more exposed tho
The limit.output tracks the maximum supported output tokens from a model, we don't want to always use that cause then u could potentially use up the entire context window in a single turn also it makes compaction math more annoying
Submit a PR #29513 , and it fix the issue:
- it makes maxOutputTokens() use limit.output directly when configured (> 0), fall back to OUTPUT_TOKEN_MAX only when no limit is set, and preserve OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX as an optional safety ceiling.
Could consider deprecating OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX in favor of a global limits.maxOutputTokens in opencode.json. This would make the safety ceiling discoverable and configurable without requiring an experimental env var. Per-model limit.output stays fully respected below the ceiling.
did u read my message i literally said not to do that
We aren't ignoring configs here, this is how many tokens we cap the agent to reply with for a given step of a turn, the case where it reaches limit prolly needs additional handling but making a larger value here is uncommon, we alreadyy have a pretty standard value and some other coding agents make it even smaller. Happy to make the configuration option more exposed tho
The limit.output tracks the maximum supported output tokens from a model, we don't want to always use that cause then u could potentially use up the entire context window in a single turn also it makes compaction math more annoying
deepseek v4 model has 1M context window, and it can generate very long chain of thought, much longer than 32k. Limit output to 32k making the thinking process truncated. This make the whole program unusable.
did u read my message i literally said not to do that
@rekram1-node I didn't actually 'see' your message before I posted my previous message as I was using gh cli tool.
I spent some time trying to summarize all these and hopefully we all could find our conversation valuable for future decision:
1. What does modern LLM models support
| Provider | Latest Model | Context Window | Max Output Tokens |
|---|---|---|---|
| OpenAI | GPT-5.5 / GPT-5.4 | ~1,000,000 tokens | 128,000 |
| Anthropic | Claude Opus 4.7 | 1,000,000 tokens | 128,000 (sync) / 300,000 (Batch API) |
| Gemini 3.5 Flash | ~1,000,000 tokens | 65,536 | |
| DeepSeek | DeepSeek V4 Pro | 1,000,000 tokens | 384,000 |
2. Claude Code and Codex CLI current mechanism
Claude Code lets the model cap drive output (from online tech doc):
CLAUDE_CODE_MAX_OUTPUT_TOKENS'Set the maximum number of output tokens for most requests. Defaults and caps vary by model;' env var defaults to model's own max output (128k Opus 4.7, 64k Sonnet 4.6)- Sync API: up to model max, Batch API: up to 300k via
output-300k-2026-03-24beta header - Docs: code.claude.com/env-vars → platform.claude.com/models
OpenAI Codex CLI hardcodes 128k for all GPT-5 models (from github info):
max_output_tokens: 128_000incodex-rs/core/src/openai_model_info.rs- Users hit
stream disconnected before completion: reason: max_output_tokens(GitHub issue #14753) - Not user-configurable — feature requested in issue #20861
- Source: codex-rs/openai_model_info.rs (mentioned by issue #3523)
3. Possible directions
Option 1 Maintain current behavior and allow global config to adjust max output tokens.
Option 2 Allow per-model max output tokens overrides.
Option 3 Allow per-profile or per-job configuration. As someone may want aggressively truncated tool call output for certain projects. Others might need larger windows for long run tasks like log analysis and self driven code generation, etc.
max output tokens for a model, and the requested output tokens are 2 separate concepts
max output tokens for a model, and the requested output tokens are 2 separate concepts
@rekram1-node Agreed that model capability and per-request output cap are different concepts — this issue isn't about the concept.
The concrete failure I'm reporting (reproducible):
- DeepSeek V4 Pro,
limit.output: 384000,reasoningEffort: "max" - Single step ends with
finish_reason: "length",reasoning: 32000,output: 0 - Agent loop hard-stops. No retry, no continuation, no surfaced error.
This is not "uncommon to want a larger value." As you can see in @lexlian 's previous comment, for any frontier reasoning model in 2026 — GPT-5.5 reasoning=max, Claude Opus 4.7 extended thinking, DeepSeek V4 long CoT — reasoning tokens routinely exceed 32k in a single step. Under the current cap that's 0 visible output, ~100% of the time, not an edge case. Comparable tools already account for this: Claude Code defaults CLAUDE_CODE_MAX_OUTPUT_TOKENS to the model's own max (128k Opus / 64k Sonnet); Codex CLI hardcodes 128k for GPT‑5.
Keeping your concept distinction intact, the asks are orthogonal to it:
- Expose the per-request cap in
opencode.json(non-experimental). You said you'd be happy to — this is that. - Default the cap from model metadata rather than a fixed 32k.
- Handle
finish_reason: "length"gracefully instead of a hard stop. - Decouple compaction math in
overflow.tsfrom the per-step cap — this is the legitimate concern you raised, and it's solvable independently of (1)–(3).
None of these collapse the capability-vs-request-cap distinction.
Concrete question I'd like a direct answer on: is opencode intended to be usable with reasoningEffort: high/max on frontier models today? If the answer is no, please state that explicitly in the official docs — something like "opencode is not designed for frontier models with long-form output / extended reasoning; do not use it for such workloads." That at least sets correct expectations, instead of letting users configure limit.output and silently get it capped at 32k.
I am facing a similar issue, globally for various model of a specific provider where the limits are configured for every single one but automatically is capped to 32k.
Regardless of whatever output max is set to... opencode should be handling this gracefully and continue rather than stopping dead like the turn is complete.
1 remaining item
+1 with an independent headless reproduction on OpenCode 1.17.18 using zai-coding-plan/glm-5.1.
We run OpenCode through opencode serve as a benchmark backend. On a real coding task, we have 12 persisted runs: 5 successes and 7 failures. Every one of the 7 failures has the same signature:
- no final assistant result;
- no tool-created artifact;
- non-zero provider usage;
output_tokens + reasoning_tokens = 32,000exactly;- reasoning consumes almost the entire budget (31,933–31,987 tokens);
- the session then reaches idle/terminal state without a usable result.
Persisted failure samples:
| elapsed | output | reasoning | completion |
|---|---|---|---|
| 915.1s | 33 | 31,967 | 32,000 |
| 357.6s | 36 | 31,964 | 32,000 |
| 470.1s | 61 | 31,939 | 32,000 |
We repeated the same task twice with --no-save; both failed again at exactly 32,000 completion tokens:
- 457.1s: output 33 + reasoning 31,967;
- 517.7s: output 29 + reasoning 31,971.
Control prompts on the same provider/model succeed quickly (hello_world in 21.9s, fast_sort in 19.4s), so this is not a general provider outage. We also saw no ECONNRESET, malformed SSE, or network error in these failures.
Our wrapper currently reports this as a hung POST /message, but that is misleading: the provider clearly returned usage. The actual failure is consistent with the 32k per-step cap being exhausted by reasoning and OpenCode ending the turn without exposing/recovering from finish_reason=length.
One important operational detail: successful and failed run durations overlap substantially (successes 515–1,077s; failures 358–915s), so a wall-clock watchdog cannot reliably identify this condition. Preserving and surfacing the terminal finish reason would let headless callers fail immediately and accurately; auto-continuation or an explicit actionable output-limit error would be even better.
Full local reproduction data and proposed downstream handling: axisrow/llm_benchmark#161
Follow-up with a live A/B run and downstream mitigation on OpenCode 1.17.18 with zai-coding-plan/glm-5.1, using the same real coding prompt and cleared proxy variables:
| mode | elapsed | input | output | reasoning | aggregate completion | result |
|---|---|---|---|---|---|---|
| default 32k ceiling | 596.0s | 45,650 | 9,029 | 26,298 | 35,327 | success + artifact |
OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=65536 |
959.9s | 66,072 | 11,689 | 51,094 | 62,783 | success + artifact |
The default run happened to succeed this time, which is consistent with the stochastic behavior in the earlier sample: nine independent failures ended at exactly 32,000 completion tokens, while other runs succeeded.
With the 65,536 override, the OpenCode log shows the initial build step running for about 784 seconds before the first tool action, after which the agent wrote the artifact and completed. The 62,783 figure is aggregate completion usage for the full run, not proof that one step consumed all of it; nevertheless, this successful run is consistent with the higher ceiling avoiding the repeatedly observed 32k termination.
We also implemented downstream handling in the benchmark wrapper:
- expose an explicit output-token ceiling;
- preserve and surface terminal
finish=length/MessageOutputLengthError; - stop waiting for a missing
session.idleonce terminal length is observed; - optional first-action watchdog.
A live watchdog smoke test with a 5s threshold exited this workload in 34.1s instead of waiting 10–20 minutes. The current floor is the existing 30s POST /message read timeout.
Details and test results: axisrow/llm_benchmark#161
Confirmed on my side with the built-in ollama-cloud/glm-5.2 model as well.
opencode models ollama-cloud --verbose reports the model with:
limit.context = 976000limit.output = 131072
So the model metadata is being loaded correctly.
However, without OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX, OpenCode still behaves as if the effective output cap is 32000. In other words, this is not just bad provider metadata or a bad custom config entry: the model is recognized with the correct limit.output, and the 32k ceiling is being applied later in the request path.
That makes this issue broader than any one provider/model family. I was able to reproduce the same pattern on a built-in OpenAI-compatible path (ollama-cloud) with GLM 5.2.
A few things here don't add up.
The min() doesn't do what the rationale says.
we don't want to always use that cause then u could potentially use up the entire context window in a single turn
The code is Math.min(model.limit.output, 32000). It only caps models whose limit.output is above 32k. Anything below 32k is sent verbatim, which is exactly the "eat the whole window in one turn" case the cap is supposed to prevent. So 32k isn't a principled limit, it's just a constant sitting in the middle.
In practice it also kills the config. Every model in my setup has limit.output >= 64k (DeepSeek 384k, Kimi 262144, Claude 64k), so min() returns 32000 for all of them and limit.output does nothing. I set 384k, the API gets 32k. That is ignoring the config.
The context-window worry only applies to shared-window models.
DeepSeek at 1M/384k can't fill the window in one turn: 384k output still leaves 616k for input. The only models where output can eat the window are shared-window ones like Kimi 262144/262144, and there the cap protects nothing anyway. OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=1048576 removes it in one line, and then usable = context - maxOutputTokens() = 0 and every session compaction-loops immediately.
32k isn't a standard value in 2026.
The numbers are already in this thread (models + Claude Code/Codex, repro): GPT-5.5 128k, Opus 4.7 128k, Gemini 3.5 65k, DeepSeek 384k. Claude Code doesn't hardcode one number either. It uses a per-model (default, upper) pair (mainstream 32k/64k, opus/1M 64k/128k, legacy 4k/8k) that CLAUDE_CODE_MAX_OUTPUT_TOKENS can raise up to upper. So 32k there is an overridable default, not a ceiling. opencode is the only one applying a flat 32k that config can't touch.
"Uncommon" is wrong too: the repro above shows 7/12 runs on 1.17.18 dying at exactly 32,000 tokens with reasoning eating the whole budget and no result, on GLM-5.1, months after this was filed.
The real bug is the part you already flagged.
the case where it reaches limit prolly needs additional handling
There is no handling. Hit the cap and the turn ends silently: reason: "length", output: 0, no result, no error. You can't even catch it by timeout, since successes and failures overlap (358–1077s).
The fix is a per-model clamp to the provider's real per-request cap, plus decoupling the overflow math via limit.input, not a global 32k constant. Related: #20078, #17471, #18108.
Worth noting: this looks already fixed on the v2 branch, which is probably why the dev-targeted PRs (#29513, #24384) weren't taken — the whole thing was rewritten there rather than patched.
On v2 the outbound Math.min(limit.output, 32000) cap is gone. There's no maxOutputTokens() helper capping the request anymore, and OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX no longer exists. Each protocol just sends the caller's maxTokens:
packages/ai/src/protocols/openai-chat.ts→max_tokens: generation?.maxTokenspackages/ai/src/protocols/open-responses.ts→max_output_tokens: generation?.maxTokenspackages/ai/src/protocols/gemini.ts→maxOutputTokens: generation?.maxTokenspackages/ai/src/protocols/bedrock-converse.ts→maxTokens: generation?.maxTokenspackages/ai/src/protocols/anthropic-messages.ts→max_tokens: generation?.maxTokens ?? (limit.output ?? 4096)
Anthropic keeps a limit.output fallback because its API requires max_tokens; the others omit the field when it isn't set, which lets the model use its own limit instead of being clamped. So limit.output is no longer silently overridden.
OUTPUT_TOKEN_MAX = 32_000 survives in exactly one place — packages/core/src/session/compaction.ts:25 — and it's only used to decide when to auto-compact, not in any outbound request. The overflow math is also decoupled from the per-step cap now (compaction.ts:354):
const output = Math.min(limits?.output ?? 0, OUTPUT_TOKEN_MAX)
const promptCeiling = Math.min(
limits?.input === undefined ? Number.POSITIVE_INFINITY : limits.input - config.buffer,
context - Math.max(output, config.buffer),
)That closes the shared-window collapse from the original report: for a 262144/262144 model, output is clamped to 32k for the compaction estimate only, so promptCeiling = 262144 - max(32000, buffer) ≈ 230144 instead of 0 — no immediate compaction loop, and no need for a limit.input workaround.
The one thing still worth deciding is whether the compaction estimate should keep a fixed 32k assumption at all, or derive from the actual per-step budget — but the core bug (config-ignored 32k cap on the API request) is resolved on v2. Might be worth confirming and closing this against the v2 rewrite.
Worth noting: this looks already fixed on the
v2branch, which is probably why thedev-targeted PRs (#29513, #24384) weren't taken — the whole thing was rewritten there rather than patched.On v2 the outbound
Math.min(limit.output, 32000)cap is gone. There's nomaxOutputTokens()helper capping the request anymore, andOPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAXno longer exists. Each protocol just sends the caller'smaxTokens:* `packages/ai/src/protocols/openai-chat.ts` → `max_tokens: generation?.maxTokens` * `packages/ai/src/protocols/open-responses.ts` → `max_output_tokens: generation?.maxTokens` * `packages/ai/src/protocols/gemini.ts` → `maxOutputTokens: generation?.maxTokens` * `packages/ai/src/protocols/bedrock-converse.ts` → `maxTokens: generation?.maxTokens` * `packages/ai/src/protocols/anthropic-messages.ts` → `max_tokens: generation?.maxTokens ?? (limit.output ?? 4096)`Anthropic keeps a
limit.outputfallback because its API requiresmax_tokens; the others omit the field when it isn't set, which lets the model use its own limit instead of being clamped. Solimit.outputis no longer silently overridden.
OUTPUT_TOKEN_MAX = 32_000survives in exactly one place —packages/core/src/session/compaction.ts:25— and it's only used to decide when to auto-compact, not in any outbound request. The overflow math is also decoupled from the per-step cap now (compaction.ts:354):const output = Math.min(limits?.output ?? 0, OUTPUT_TOKEN_MAX)
const promptCeiling = Math.min(
limits?.input === undefined ? Number.POSITIVE_INFINITY : limits.input - config.buffer,
context - Math.max(output, config.buffer),
)That closes the shared-window collapse from the original report: for a 262144/262144 model,
outputis clamped to 32k for the compaction estimate only, sopromptCeiling = 262144 - max(32000, buffer) ≈ 230144instead of 0 — no immediate compaction loop, and no need for alimit.inputworkaround.The one thing still worth deciding is whether the compaction estimate should keep a fixed 32k assumption at all, or derive from the actual per-step budget — but the core bug (config-ignored 32k cap on the API request) is resolved on v2. Might be worth confirming and closing this against the v2 rewrite.
The concept that v1 fixes for a product in which v1 is the default shipped are pushed away for months because of a pending v2 merge is one of the single dumbest traps in open source development.
Until merged, v1 is the product. Intentionally avoiding fixing the product because you have a fix set in some amorphous time in the future is dumb.
Confirmed still present on v1.18.30 and v1.18.31 (v1.18.31 changelog is ACP/TUI/Copilot only; OUTPUT_TOKEN_MAX is unchanged).
// packages/opencode/src/provider/transform.ts
export const OUTPUT_TOKEN_MAX = 32_000
export function maxOutputTokens(model: Provider.Model, outputTokenMax = OUTPUT_TOKEN_MAX): number {
return Math.min(model.limit.output, outputTokenMax) || outputTokenMax
}session/llm/request.ts then sends that as maxOutputTokens. Anthropic budget variants are also clamped with OUTPUT_TOKEN_MAX - 1.
What we see
Same silent 32k request cap on reasoning-heavy turns, including models whose catalog already advertises far more:
| Model | Advertised limit.output |
What OpenCode actually requests |
|---|---|---|
anthropic/claude-fable-5-1 (Max) |
128000 | 32000 |
anthropic/claude-opus-5 |
128000 | 32000 |
openai/gpt-5.6-sol |
128000 | 32000 |
DeepSeek V4-class via @ai-sdk/openai-compatible |
64k–384k | 32000 |
On Fable 5.1 Max and DeepSeek thinking, a step can finish reason: "length" with reasoning consuming the entire 32k budget and little or no answer/tool call. Title/summary agents that set a small explicit cap are fine; the bug is the global clamp on ordinary/thinking turns.
OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX still works as a shotgun, but it is global and does not honor per-model limit.output.
Please use configured limit.output (and any smaller explicit request cap) instead of Math.min(limit.output, 32000) for every model.
Independent confirmation on v1.18.31 (stable) — same provider class, three models, reasoning-heavy sub-agent turns.
Provider is @ai-sdk/openai-compatible (custom gateway). Declared limits come from the provider catalog, not from a local limit override:
| model | limit.context |
limit.output |
|---|---|---|
| deepseek-v4-flash | 1,000,000 | 384,000 |
| glm5.3-flash | 1,000,000 | 131,072 |
| qwen3.8-flash | 262,144 | 131,072 |
Every long-thinking turn ended finish_reason: length with no assistant text, at a shared ceiling across three different models:
| model | reasoning chars | finish | emitted text |
|---|---|---|---|
| deepseek-v4-flash | 59,991–59,999 | length | none |
| qwen3.8-flash | 59,996–59,997 | length | none |
| glm5.3-flash | 59,994 | length | none |
A turn on the same model that stayed under the ceiling (56,178 reasoning chars) finished stop with normal output, so the budget — not the model — is what ends the turn. (We measured characters from the session store, not tokens; ~60k chars is consistent with a 32,000-token cap for this tokenizer.)
This is still the exact code path in 1.18.31, not only 1.14.x. Extracted from the released binary:
OUTPUT_TOKEN_MAX:()=>M7}); ... var M7=32000 ...
function by($,Z=M7){return Math.min($.limit.output,Z)||Z}
outputTokenMax:G("OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX")
maxOutputTokens(e.model, e.flags.outputTokenMax)
With the env var unset, all three models receive min(limit.output, 32000) === 32000 regardless of the 384k / 131k declared in config.
The documented workaround does work: raising OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX above the declared limit makes each model receive its own limit.output. On the overflow caveat mentioned above, all three of our models declare limit.context well above limit.output, so the context - maxOutputTokens() path stays positive (616k / 869k / 131k) and we did not need limit.input.
The suggested direction still stands: use the constant as a fallback (limit.output || OUTPUT_TOKEN_MAX) instead of Math.min(limit.output, OUTPUT_TOKEN_MAX), and the env var becomes unnecessary.
Correction / update (2026-09-17, measured later the same day).
We instrumented our own outbound requests with a local logging proxy (OPENCODE_CONFIG_CONTENT overriding the provider baseURL) and can now separate two things we had previously conflated:
-
The clamp is real and the env var lifts it. On 1.18.31 with
OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=1048576, the request body carriesmax_tokens: 384000for the title, primary, sub-agent and post-tool steps. Long turns do complete: we observed a sub-agent turn of 59,013 output tokens endingfinish: stop, and a review-shaped request of 38,222 completion tokens. SoMath.min(limit.output, OUTPUT_TOKEN_MAX)is bypassed exactly as described. -
The shared ~59.9k-char ceiling is not explained by the opencode clamp alone. The same truncation (
finish_reason: length, reasoning ~59,994 chars, no assistant text, nousageblock) reproduces when calling the provider's/v1/chat/completionsdirectly, with opencode out of the loop, usingmax_tokens: 384000. The opencode 32k clamp and this provider-side ceiling coincide at ~32k tokens, which is why session data alone could not tell them apart.
We therefore withdraw the "so the budget - not the model - is what ends the turn" attribution for the residual case, and we have moved the remaining report to the provider. The suggested direction here still stands on its own: limit.output || OUTPUT_TOKEN_MAX is strictly better than Math.min(limit.output, OUTPUT_TOKEN_MAX), and the env var becomes unnecessary.
We also withdraw the char-to-token estimate: measured ratios ranged from ~1.85 to ~3.5 chars/token across turns, so ~60k chars is not a reliable proxy for a 32,000-token cap.
Same issue when using : ollama launch opencode --model glm-5.3-flash:cloud
Any updates on this? I'm running into the same problem
Yes, this is fixed in latest v2 release
Cool so merge v2 as main because as is you're refusing to fix things in v1
for months because "it's fixed in v2".
…
We don’t really use main. V2 is the default version on our website for some time already.
Summary
OpenCode silently caps per-step
maxOutputTokensat 32,000 even whenopencode.jsonsets a much largerlimit.output(e.g. 384000 for DeepSeek, 128000 for GPT/Claude). The only documented escape hatch is an experimental environment variable,OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX, which users must discover via scattered GitHub threads.This is surprising, breaks agent runs on reasoning-heavy models, and makes user configuration misleading.
Expected behavior
When a model defines explicit limits in config:
OpenCode should send
max_output_tokens/max_tokensbased onlimit.output(possibly clamped to what the provider actually supports), not an undocumented global ceiling of 32k.Actual behavior
ProviderTransform.maxOutputTokens()always applies a default cap:Unless
OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAXis set in the environment,min(384000, 32000) === 32000is what gets sent to the API.Real-world impact (agent benchmark)
On a coding task with
reasoningEffort: "max"via@ai-sdk/openai-compatible:reason: "length"reasoning: 32000,output: 0The same model with
reasoningEffort: "high"completes normally in the same step budget because reasoning does not consume the entire cap.This is not a provider failure; logs show OpenCode requesting ~32k max output while config says 384k.
Why
OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAXfeels like the wrong fixMisleading config — Users set
limit.outputinopencode.jsonbelieving it controls API behavior. It does not, unless they also export an env var whose name does not referencelimit.outputat all.Experimental / undiscoverable — The variable is buried under “experimental” flags in CLI docs. Issue max_tokens defaults to 32000 when using a custom provider #1735 was auto-closed; Custom provider (LM Studio) ignores limit.output config, hardcodes max_tokens to 32000 #20078 remains open with the same root cause.
Global knob for per-model limits — Setting
OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=1048576is a shotgun fix across all models instead of respecting per-modellimit.output.Still does not fix overflow semantics —
overflow.tsusescontext - maxOutputTokens()whenlimit.inputis absent. Models withoutputclose tocontext(e.g. Kimi K2.6 at 262144/262144) getusable === 0once the env var is raised, causing immediate compaction/overflow unless users also addlimit.input— another undocumented workaround.Related issues
limit.outputignored, hardcodes 32000 (open)Suggested direction
Respect configured
limit.outputby default — UseOUTPUT_TOKEN_MAX(32k) only as a fallback whenlimit.outputis missing or zero, not asMath.min(limit.output, 32k)for every model. (PR fix(provider): respect configured output limit #24384 attempted this but was closed as “not the right fix”; the problem remains.)Optional global ceiling — If a safety cap is still desired, expose it in
opencode.json(e.g.limits.maxOutputTokens) rather than a cryptic env var.Decouple overflow math from per-step output cap —
usableshould not usecontext - limit.outputwhenoutputcan equal the full window; considerlimit.inputdefaults orcontext - reservedfor shared-window models.Document the contract — Clarify in config schema/docs what
limit.context,limit.input, andlimit.outputeach control (API request vs compaction vs fallback).Environment
anomalyco/opencode)@ai-sdk/openai-compatible→ OpenAI-compatible gatewaylimitblocks inopencode.jsonWorkaround today (ugly)
export OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX=1048576Plus per-model
limit.inputfor shared-window models — none of which should be required whenlimit.outputis already configured.