Repository navigation
Conversation
|
The following comment was made by an LLM, it may be inaccurate: Based on the search results, I found several related PRs that address similar functionality: Potential Related PRs
These PRs address related concerns around provider fallback chains, transient error handling, and retry logic, though they appear to be separate implementations in different components. PR #26292 appears to be the consolidated, comprehensive implementation of this functionality. |
|
Thanks for updating your PR! It now meets our contributing guidelines. 👍 |
9526a6b to
bd85c37
Compare
0d07b43 to
e54c5c1
Compare
|
Update: Added a structured Design section to this PR that covers the full data flow, cooldown duration table, key architectural decisions, config contract, and explicit scope boundaries. Goal: make it easy to validate the approach at a glance. For anyone following #7602 — this is the most complete implementation of native fallback to date:
Running in production for 1+ week, all tests passing. Would love a review from the core team — happy to iterate on the design if there are concerns about the approach. |
34e3b42 to
a7d8ea5
Compare
7dda3d0 to
cc4902c
Compare
|
This is a hard +1 from me and my entire team. There are rate limits everywhere you look now. What happens when Opencode Go hits limits? You just wait it out or go manually edit the file. At least having some sort of fallback in the event of rate limits, downtime, or errors so that the work continues in unexpected circumstances is essential in production. |
|
This is much needed. |
|
🙏 |
|
After a month of rebasing my PR with almost guaranteed conflicts, I gave up and switched to oh-my-pi (https://omp.sh) - it supports canonical model names with "provider order". It does the job and maybe better than opencode. I'll keep that PR open just in case the team ever look at it. I understand they have a thousands PR now, so it's probably lost. |
|
I'd really like to have this one... |
|
+1 we need this |
|
We need this feature! |
|
this is outside the scope of this project. For this use case, use a dedicated proxy: https://github.com/router-for-me/CLIProxyAPI |
What is the definition of project scope here? What you suggest is to use another project/tool to just fulfilling a missing feature of this project? |
|
much needed.. |
|
whats stopping from merging this? |
Agents can now configure a fallback model chain (agents.<name>.fallback in config, or the older top-level agent: format). When the selected model fails after RequestExecutor's transport retry budget is spent (429/5xx/529), or with a quota/transport/model-not-found failure, the V2 session runner switches to the next untried fallback and retries the turn -- rather than erroring the whole turn out to the user. Ported from real, working upstream implementations rather than written from scratch (per the standing "reuse upstream PR work" directive for this fork): anomalyco#35188 already targeted this repo's V2 session runner (unlike anomalyco#49125/anomalyco#26292/anomalyco#27939/anomalyco#42424, which patch the legacy V1 processor.ts our own AGENTS.md says new orchestration code must not bridge through) and is the primary source for the config schema shape and general approach. The retry-exhaustion classifier's names/shape come from anomalyco#49125; its logic reuses RequestExecutor's existing LLMError.retryable flag as the single source of truth for "is this retryable" instead of re-deriving a second status-code list. Model-switch attribution is recorded via the session's existing SessionEvent.ModelSwitched event rather than mutating an in-flight input object in place -- a lesson taken directly from anomalyco#27939's writeup of a real bug in its own V1-era approach. New packages/core/src/session/runner/fallback.ts (~40 lines: a classifier + a candidate resolver) and a new ContinueAfterFallback turn-transition in packages/core/src/session/runner/llm.ts, following the exact pattern already used for post-compaction turn continuation. Config plumbing threads the new field through packages/schema/src/ agent.ts, packages/core/src/config/agent.ts (+ plugin and v1 config variants). Regenerated the client and legacy JS SDKs. Known, honestly-scoped gaps: the switch is sticky for the rest of the session (no cooldown/return-to-primary); no mid-stream failover once assistant output has started; only errors marked retryable trigger a switch, so some in-stream provider errors won't while an equivalent HTTP status will; fallback is per-agent only, not per-model or global; a fallback model uses its default variant; no TUI/docs yet. New packages/core/test/session-runner-fallback.test.ts (9 tests) exercises real RequestExecutor retry counts against a scripted fake HTTP client -- exact call sequences, session projection correctness (model-switched message inserted, no leaked failed assistant message, correct final session model), chain-walking past an unresolvable model, non-retryable failures correctly not switching, and chain exhaustion still surfacing the real terminal error. Full packages/core suite: 1164 pass (only the pre-existing, unrelated pty flake). bun typecheck clean across all 30 packages. Co-Authored-By: Claude Sonnet 5 <[email protected]>
The entire project is about being able to use whatever provider you want, correct? How on earth is being able to assign provider fallbacks not in scope? I'd love any sort of explanation. |
Issue for this PR
Closes #7602
Type of change
What does this PR do?
Adds a configurable fallback chain so that when a provider returns a transient error (rate limit, overload, 5xx), OpenCode automatically retries on the next model in the chain instead of failing the session.
{ "model": "anthropic/claude-sonnet-4-20250514", "fallbacks": ["openai/gpt-4.1", "deepseek/deepseek-v4"], "cooldown_seconds": 300 }fallbackscan be set at the top level or per-agent.cooldown_secondsdefaults to 300 — after a retryable failure, that provider/model is skipped for the cooldown duration so you don't wait on retries to an overloaded provider.Why built-in instead of a proxy: cheaper providers are unreliable, and routing through LiteLLM degrades tool-call quality. When a provider gets overloaded, falling through immediately is faster than retrying the same one.
Design
Data flow
Cooldown durations
cooldown_seconds(default 300)retry-afterheaderOn success, the winning provider's cooldown is cleared so it's immediately available next request.
Key decisions
Cooldown over session state — We track what's failed, not what's succeeded. No "sticky" fallback — once a provider recovers, traffic routes back naturally.
Stream-level error detection — Providers can return HTTP 200 with an error in the stream body. We wrap
fullStreamto throw on{ type: "error" }chunks, triggering fallback just like a connection error.Quota limits trigger fallback (6h cooldown) — A quota-limited provider is unavailable, not the session. Fallback keeps the session alive. The 6h cooldown prevents hammering a capped provider.
No dedup in the chain — The same model can appear twice. After falling through once, the primary may be worth retrying with fresh context. Cooldown handles this naturally — if primary is still on cooldown, it's skipped; if it's cleared, trying it again is valid.
Model attribution updates on fallback — When a fallback succeeds,
usedFallbackpropagates back so events, logs, and billing reflect the actual provider that handled the request.Config contract
New fields:
fallbacks(array ofprovider/modelstrings),cooldown_seconds(positive int, default 300).Not in scope
cooldown_secondsglobally for now)How did you verify your code works?
CooldownManager(put/get/clear/expiry) and config validation (fallbacks array and cooldown_seconds)bun typecheckpasses for all 12 packages in the monorepoScreenshots / recordings
N/A — no UI changes visible in screenshots (toast is a runtime notification)
Checklist
Comparison with related PRs
Reviewed #24369, #26192, #24013, #18443, and the closed #13189:
resolveFallbackChainutility — minor convenience we can add later.Key differences in our approach: