Skip to content

Use API-backed model metadata for limits and capabilities #8558

Description

@DragonnZhang

What would you like to be added?

Introduce a resolved model metadata layer that uses exact provider-and-model entries from models.dev, with room for provider-native metadata in the future, to supply model defaults and capabilities that are currently maintained in code.

The metadata should cover:

  • Context, maximum input, and maximum output token limits.
  • Reasoning support and the supported controls, such as effort tiers, toggles, and token budgets.
  • Support for temperature, tool calling, structured output, interleaved reasoning, and input/output modalities.
  • Display-only information such as model status, release date, knowledge cutoff, and token cost.

Capability facts must remain separate from request policy. For example, a model's maximum output limit must not automatically become the requested max_tokens, and reasoning: true must not automatically enable reasoning. Explicit user or provider configuration should take precedence over provider-native metadata, then models.dev metadata, then a small legacy fallback, and finally conservative defaults.

The integration should use exact provider and model matching, validate remote data locally, remain non-fatal when metadata is unavailable, and never inject catalog-provided request bodies, headers, credentials, or routing endpoints into inference requests.

Why is this needed?

Qwen Code currently maintains model context limits, output limits, modality rules, reasoning behavior, and provider presets across model-name regular expressions and per-provider tables. These values drift as new model versions and aliases are released, require a Qwen Code release to correct, and can resolve differently depending on the configuration path.

Incorrect limits affect context reporting, compaction thresholds, and output clamping. Incorrect reasoning or temperature assumptions can produce rejected requests or unexpected latency and cost, while assuming tool support can select a model that cannot run the normal coding-agent loop.

models.dev already publishes normalized limits and capability metadata for thousands of provider-specific model entries. Using it as a default metadata source would remove most per-model maintenance while retaining explicit configuration and protocol-specific adapters for provider wire formats and exceptional behavior.

Additional context

Draft PR #8529 demonstrates the approach for input modalities only. A broader implementation should first generalize that catalog into a typed model metadata object, then adopt limits, runtime capabilities, and display metadata in separate reviewable steps.

The cold-start path should be addressed before critical behavior depends on remote metadata. A bundled snapshot or another non-blocking fallback should let the CLI start without waiting for models.dev, while disk caching and background refresh keep the data current.

OpenCode already converts models.dev records into first-class model objects containing capabilities, context/input/output limits, cost, and lifecycle status, with user configuration layered on top. Hermes Agent similarly uses a typed models.dev record as its primary metadata source while retaining narrow fallbacks.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions