Skip to content

feat(processor): add model fallback chain when retries are exhausted - #42424

Closed
herjarsa wants to merge 5 commits into
anomalyco:devfrom
herjarsa:feat/model-fallback
Closed

herjarsa wants to merge 5 commits into
anomalyco:devfrom
herjarsa:feat/model-fallback

Conversation

@herjarsa

Copy link
Copy Markdown

Issue for this PR

Closes #10287

Type of change

  • New feature

What does this PR do?

Adds automatic model fallback chain when the primary model fails after all retries are exhausted.

When a model encounters an error that retries cannot resolve, the processor now checks if a fallbackModels chain was configured (e.g. via fallback_model in agent config). If so, it resolves the next model in the chain and re-attempts the stream with that model.

Error classification for fallback:

  • Fallback triggered: retryable APIError (rate limits, 5xx, transient errors)
  • No fallback: AuthError (bad credentials), AbortedError (user cancelled), non-retryable APIError (permanent failures), ContextOverflowError (triggers compaction instead)

Implementation details:

  • Added fallbackModels field to MessageV2.User schema
  • Added resolveFallbackChain helper to Provider service (returns {model, remaining})
  • Added outer while loop in SessionProcessor.process() to iterate through fallback models
  • Modified halt() to classify errors and set ctx.shouldFallback flag when appropriate
  • Reset error state when switching to a fallback model

How did you verify your code works?

  • bun typecheck passes with zero errors in modified files
  • Added integration tests in test/session/fallback.test.ts:
    • Auth error (401) → returns "stop" without fallback
    • Non-retryable error (400) → returns "stop" without fallback

Checklist

  • I have tested my changes locally
  • I have not included unrelated changes in this PR

@github-actions

Copy link
Copy Markdown
Contributor

The following comment was made by an LLM, it may be inaccurate:

Found one potentially related PR:

PR #26292: feat(opencode): add LLM provider fallback chain
#26292

Why it might be related: This PR implements fallback chain functionality for LLM providers. While the current PR (#42424) adds model fallback at the processor level when retries are exhausted, PR #26292 addresses fallback at the provider level. These could be related features or might address overlapping concerns around fallback handling.

All other search results returned the current PR (#42424) itself, which confirms this is a unique implementation. No other duplicate PRs were identified.

@github-actions

Copy link
Copy Markdown
Contributor

Automated PR Cleanup

Thank you for contributing to opencode.

Due to the high volume of PRs from users and AI agents, we periodically close older PRs using automated criteria so maintainers can focus review time on the most active and community-supported contributions.

This PR was closed because it matched the following cleanup criteria:

  • The PR was created more than 1 month ago
  • The PR had fewer than 2 positive reactions
  • Positive reactions are counted as thumbs-up, heart, celebration, or rocket reactions on the PR

PRs created within the last month are not affected by this cleanup.

If you believe this PR was closed incorrectly, or if you are still actively working on it, please leave a comment explaining why it should be reopened. A maintainer can review and reopen it if appropriate.

Thanks again for taking the time to contribute.

@github-actions github-actions Bot closed this Sep 14, 2026
hunterchristian added a commit to chipp-ai/opencode that referenced this pull request Sep 23, 2026
Agents can now configure a fallback model chain (agents.<name>.fallback
in config, or the older top-level agent: format). When the selected
model fails after RequestExecutor's transport retry budget is spent
(429/5xx/529), or with a quota/transport/model-not-found failure, the
V2 session runner switches to the next untried fallback and retries
the turn -- rather than erroring the whole turn out to the user.

Ported from real, working upstream implementations rather than written
from scratch (per the standing "reuse upstream PR work" directive for
this fork): anomalyco#35188 already targeted this repo's V2
session runner (unlike anomalyco#49125/anomalyco#26292/anomalyco#27939/anomalyco#42424, which patch the
legacy V1 processor.ts our own AGENTS.md says new orchestration code
must not bridge through) and is the primary source for the config
schema shape and general approach. The retry-exhaustion classifier's
names/shape come from anomalyco#49125; its logic reuses RequestExecutor's
existing LLMError.retryable flag as the single source of truth for
"is this retryable" instead of re-deriving a second status-code list.
Model-switch attribution is recorded via the session's existing
SessionEvent.ModelSwitched event rather than mutating an in-flight
input object in place -- a lesson taken directly from anomalyco#27939's
writeup of a real bug in its own V1-era approach.

New packages/core/src/session/runner/fallback.ts (~40 lines: a
classifier + a candidate resolver) and a new ContinueAfterFallback
turn-transition in packages/core/src/session/runner/llm.ts, following
the exact pattern already used for post-compaction turn continuation.
Config plumbing threads the new field through packages/schema/src/
agent.ts, packages/core/src/config/agent.ts (+ plugin and v1 config
variants). Regenerated the client and legacy JS SDKs.

Known, honestly-scoped gaps: the switch is sticky for the rest of the
session (no cooldown/return-to-primary); no mid-stream failover once
assistant output has started; only errors marked retryable trigger a
switch, so some in-stream provider errors won't while an equivalent
HTTP status will; fallback is per-agent only, not per-model or
global; a fallback model uses its default variant; no TUI/docs yet.

New packages/core/test/session-runner-fallback.test.ts (9 tests)
exercises real RequestExecutor retry counts against a scripted fake
HTTP client -- exact call sequences, session projection correctness
(model-switched message inserted, no leaked failed assistant message,
correct final session model), chain-walking past an unresolvable
model, non-retryable failures correctly not switching, and chain
exhaustion still surfacing the real terminal error. Full packages/core
suite: 1164 pass (only the pre-existing, unrelated pty flake). bun
typecheck clean across all 30 packages.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Critical bug revert undo messages

1 participant