Repository navigation
feat(core): implement models fallback - #35188
ReStranger wants to merge 3 commits into
Conversation
|
Hey! Your PR title Please update it to start with one of:
Where See CONTRIBUTING.md for details. |
|
The following comment was made by an LLM, it may be inaccurate: I found a potentially related PR: #26292 - This PR appears to address a similar feature for implementing fallback behavior with LLM providers. While PR #35188 focuses on specifying fallback models for agents through configuration, PR #26292 implements an LLM provider fallback chain. These may be related or overlapping approaches to the same problem. You should review PR #26292 to ensure there isn't duplicate work or if it should be merged before this PR. |
|
Thanks for updating your PR! It now meets our contributing guidelines. 👍 |
|
Audit hygiene: the body still says |
|
This is exactly the missing piece in my workflow. I use different providers for different subagents, and when one provider hits its quota, the parent agent can stall while waiting for that subagent instead of continuing with another model. I hope this gets serious consideration for v2. Thank you for working on it @ReStranger. |
Agents can now configure a fallback model chain (agents.<name>.fallback in config, or the older top-level agent: format). When the selected model fails after RequestExecutor's transport retry budget is spent (429/5xx/529), or with a quota/transport/model-not-found failure, the V2 session runner switches to the next untried fallback and retries the turn -- rather than erroring the whole turn out to the user. Ported from real, working upstream implementations rather than written from scratch (per the standing "reuse upstream PR work" directive for this fork): anomalyco#35188 already targeted this repo's V2 session runner (unlike anomalyco#49125/anomalyco#26292/anomalyco#27939/anomalyco#42424, which patch the legacy V1 processor.ts our own AGENTS.md says new orchestration code must not bridge through) and is the primary source for the config schema shape and general approach. The retry-exhaustion classifier's names/shape come from anomalyco#49125; its logic reuses RequestExecutor's existing LLMError.retryable flag as the single source of truth for "is this retryable" instead of re-deriving a second status-code list. Model-switch attribution is recorded via the session's existing SessionEvent.ModelSwitched event rather than mutating an in-flight input object in place -- a lesson taken directly from anomalyco#27939's writeup of a real bug in its own V1-era approach. New packages/core/src/session/runner/fallback.ts (~40 lines: a classifier + a candidate resolver) and a new ContinueAfterFallback turn-transition in packages/core/src/session/runner/llm.ts, following the exact pattern already used for post-compaction turn continuation. Config plumbing threads the new field through packages/schema/src/ agent.ts, packages/core/src/config/agent.ts (+ plugin and v1 config variants). Regenerated the client and legacy JS SDKs. Known, honestly-scoped gaps: the switch is sticky for the rest of the session (no cooldown/return-to-primary); no mid-stream failover once assistant output has started; only errors marked retryable trigger a switch, so some in-stream provider errors won't while an equivalent HTTP status will; fallback is per-agent only, not per-model or global; a fallback model uses its default variant; no TUI/docs yet. New packages/core/test/session-runner-fallback.test.ts (9 tests) exercises real RequestExecutor retry counts against a scripted fake HTTP client -- exact call sequences, session projection correctness (model-switched message inserted, no leaked failed assistant message, correct final session model), chain-walking past an unresolvable model, non-retryable failures correctly not switching, and chain exhaustion still surfacing the real terminal error. Full packages/core suite: 1164 pass (only the pre-existing, unrelated pty flake). bun typecheck clean across all 30 packages. Co-Authored-By: Claude Sonnet 5 <[email protected]>
Issue for this PR
Closes #123456
Type of change
What does this PR do?
Adds the ability to specify flallback models for agents
How did you verify your code works?
Use this config:
And just call the explore agent
Screenshots / recordings
Checklist