Skip to content

Reasoning effort tiers are still hardcoded / manually declared — expose them from the models.dev catalog like limits and modalities (#11959 follow-up) #13393

Description

@tyunn

Summary

#11959 landed a models.dev-backed catalog for context windows, output limits, and input modalities (merged 2026-10-02, after the 0.24.7 stable). Reasoning effort tiers did not make it into scope (the PR description says so explicitly). Today they come from four places:

  1. the built-in provider registry (ALL_PROVIDERS[].models[].capabilities.reasoning — inline entries for qwen/glm models),
  2. hardcoded per-family tables (getGptReasoningCapabilities for gpt-* models, anthropicSupportedEffortTiers for claude),
  3. per-surface wire adapters, and
  4. manual capabilities.reasoning declarations in user config (parseModelReasoningCapabilities).

For any model not covered by any of those, /effort offers the full default tier ladder and clampReasoningEffort clamps to whatever the current surface supports — with no UI-visible signal (at most a one-time debug-log warning), so the user never learns what the model actually accepts.

Meanwhile the data to fix this is already in the catalog that was just merged:

  • models.dev api.json exposes reasoning_options per model: {type: "effort", values: [...]} (3,865 entries), {type: "budget_tokens", min and/or max} (599), and {type: "toggle"} (on/off only, 1,396). Spot-checked 2026-10-04: anthropic/claude-opus-5 → ["low","medium","high","xhigh","max"], matching the native Anthropic API exactly.
  • Provider APIs increasingly advertise this natively:
    • Anthropic GET /v1/models → capabilities.effort with per-level support (low/medium/high/max/xhigh, each CapabilitySupport{supported}) — the richest source, needs no catalog at all;
    • OpenRouter GET /v1/models → reasoning: {supported_efforts, default_effort, mandatory} plus supported_parameters (includes reasoning_effort);
    • Ollama POST /api/show → thinking: {values: [...], default}.

The base OpenAI-compatible GET /v1/models schema has no capability fields (verified against the official openai-openapi spec), so catalog+native-API discovery is the only workable route for non-first-party endpoints.

Proposal

Extend the #11959 precedence chain to reasoning capabilities (the built-in provider registry is the natural insertion point — the inherited slot in validateReasoningCapabilities):

explicit user config (capabilities.reasoning)
  > curated tables + per-surface wire adapters (unchanged)
  > models.dev catalog (new)

Mapping sketch:

models.dev reasoning_options existing qwen-code concept
type: "effort", values efforts list for /effort
type: "budget_tokens", min/max thinking-budget surface (no /effort ladder)
type: "toggle" toggleOnly: true (already representable)

Open question — vocabulary bridge: the catalog uses minimal (e.g. openai/gpt-5 → ["minimal","low","medium","high"]), while qwen-code tiers are low..max plus thinking-off. Either drop out-of-vocabulary values or add the tier.

Expected benefit

  • /effort shows the real tier list for catalog-covered models without user config;
  • fewer silent clamps (clampReasoningEffort operates on real data instead of surface defaults);
  • unknown/aliased ids keep today's behavior (catalog miss → fallback), so no regression path;
  • closes the reasoning-effort portion of the umbrella request Use API-backed model metadata for limits and capabilities #8558, which asks for reasoning support and the supported controls, such as effort tiers, toggles, and token budgets.

Caveats (from probing real endpoints)

  • The catalog lists model capability, not endpoint semantics. Per-provider wire adapters must stay authoritative where they collapse tiers (e.g. DeepSeek maps low/medium → high), hence the precedence order above.
  • models.dev carries no mandatory/can-disable signal. The hardcoded gpt tables encode thinkingMandatory (e.g. gpt-5); a catalog-derived capability for a mandatory-reasoning model outside the tables would let /effort offer disabling thinking, which the API would reject. Catalog-derived capabilities should either default thinking to always-on or leave canDisable adapter-owned.
  • Same model id exposed through different endpoints (direct / Azure / OpenRouter) can advertise different effort sets; the feat(core): resolve model limits and modalities from a models.dev catalog #11959 conflict-omission rule (drop the field when providers disagree) should be inherited for effort values.
  • Anthropic-surface models could bypass the catalog entirely via capabilities.effort from GET /v1/models.

Related

Environment

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions