Skip to content

gpt-5.6-sol/terra/luna hardcode use_responses_lite / multi_agent_version, causing 400s on Azure (and likely any non-ChatGPT-backend) model_provider #31882

Description

@tianwei-mo

What version of Codex CLI is running?

codex-cli 0.144.0 (codex --version)

What subscription do you have?

Azure OpenAI, API-key auth (env_key in model_providers.azure). No ChatGPT account/subscription involved in this setup.

Which model were you using?

gpt-5.6-sol (also affects gpt-5.6-terra and gpt-5.6-luna per the bundled catalog — see root cause below)

What platform is your computer?

Linux 6.17.0-1019-aws x86_64 x86_64

What terminal emulator and version are you using (if applicable)?

VS Code integrated terminal (TERM_PROGRAM=vscode)

Codex doctor report

{
  "schemaVersion": 1,
  "overallStatus": "ok",
  "codexVersion": "0.144.0",
  "checks": {
    "auth.credentials": {
      "status": "ok",
      "summary": "auth is provided by the active model provider",
      "details": {
        "auth storage mode": "File",
        "model provider requires OpenAI auth": "false",
        "provider auth env var": "AZURE_OPENAI_KEY (present)"
      }
    },
    "config.load": {
      "status": "ok",
      "summary": "config loaded",
      "details": {
        "model": "gpt-5.6-sol",
        "model provider": "azure"
      }
    },
    "network.provider_reachability": {
      "status": "ok",
      "summary": "active provider endpoints are reachable over HTTP",
      "details": {
        "azure API base URL": "https://<redacted>.cognitiveservices.azure.com/openai/<redacted> reachable (HTTP 404)",
        "reachability mode": "provider auth"
      }
    }
  }
}

(Trimmed to the relevant checks; full report available on request.)

What issue are you seeing?

Two distinct 400s from the Azure endpoint, in sequence (the second only surfaces once the first is worked around):

1. Responses-Lite transport rejected:

{
  "error": {
    "message": "X-OpenAI-Internal-Codex-Responses-Lite only supports function tools, custom tools, and client-executed tool search.",
    "type": "invalid_request_error",
    "param": "tools",
    "code": "unsupported_value"
  }
}

2. After locally forcing use_responses_lite: false (see workaround below), a second failure surfaces:

{
  "error": {
    "message": "Invalid Value: 'tools'. Namespace 'collaboration' is reserved for encrypted tool use by this model.",
    "type": "invalid_request_error",
    "param": "tools",
    "code": null
  }
}

Both errors come from the provider (Azure), not from Codex — Codex is constructing a request shape the provider was never told to expect, and the provider correctly rejects it. Reproduced independently by a second engineer on the same endpoint/key, so this isn't a one-off local misconfiguration.

Root cause: codex-rs/models-manager/models.json is compiled directly into the Codex binary via include_str! (codex-rs/models-manager/src/lib.rs):

pub fn bundled_models_response() -> ... {
    serde_json::from_str(include_str!("../models.json"))
}

As of the models.json update that landed today (commit 3380969a29, "Update models.json (#31684)"), the three new gpt-5.6-* entries ship with:

"tool_mode": "code_mode_only",
"multi_agent_version": "v2",   // v1 for gpt-5.6-luna
"use_responses_lite": true,

These flags are unconditional per-model constants — nothing in codex-rs/core/src/client.rs gates them on the active provider:

add_responses_lite_header(&mut extra_headers, model_info.use_responses_lite);

For providers that are not the Codex/ChatGPT-hosted backend, codex-rs/models-manager/src/manager.rs::should_refresh_models() returns false:

async fn should_refresh_models(&self) -> bool {
    self.endpoint_client.uses_codex_backend().await || self.endpoint_client.has_command_auth()
}

so these providers never fetch a remote catalog and never get a chance to override these flags — they run with the bundled, ChatGPT-backend-shaped metadata as-is. There's already a provider/auth signal available elsewhere in the codebase for exactly this kind of gating (e.g. AuthManager::current_auth_uses_codex_backend, used to filter the model picker by auth mode) — it just isn't consulted before deciding to send X-OpenAI-Internal-Codex-Responses-Lite or the collaboration tool namespace.

Both X-OpenAI-Internal-Codex-Responses-Lite and the collaboration reserved tool namespace read as internal transport/session-state optimizations specific to the OpenAI-hosted Codex backend (compact tool wire format, multi-agent session state) — not general Responses-API capabilities that third-party-hosted deployments of the same model would be expected to implement.

Verified workaround: model_catalog_json (documented as "Optional path to a JSON model catalog (applied on startup only)") lets a StaticModelsManager fully replace the bundled catalog (codex-rs/model-provider/src/provider.rs::models_manager()). Setting "use_responses_lite": false and "multi_agent_version": null for the three gpt-5.6-* entries in a local catalog file, then pointing model_catalog_json at it in config.toml, resolves both errors — verified with plain chat and real tool calls (codex exec "Run the shell command ...") succeeding end-to-end against Azure. This is a local mitigation, not a fix — it will drift out of sync the next time models.json is updated upstream.

What steps can reproduce the bug?

codex exec -m gpt-5.6-sol --skip-git-repo-check "Reply with exactly: OK"

With model_provider = "azure" configured against an Azure OpenAI deployment named gpt-5.6-sol, wire_api = "responses", API-key auth (no ChatGPT sign-in).

What is the expected behavior?

gpt-5.6-sol should work against any configured model_provider the same way gpt-5.5 and gpt-5.4 currently do, or Codex should not send ChatGPT-backend-only request shapes (Responses-Lite header, collaboration tool namespace) to providers that never advertised support for them.

Suggested fix: gate use_responses_lite (and other ChatGPT-backend-only behavior baked into ModelInfo, e.g. multi_agent_version, tool_mode: code_mode_only) on whether the active provider/auth actually talks to the Codex/ChatGPT-hosted backend — the same signal already used for model-picker filtering (AuthManager::current_auth_uses_codex_backend / endpoint_client.uses_codex_backend()) — rather than applying them unconditionally per model slug regardless of model_provider.

Additional information

This looks related to, but distinct from, the existing cluster of "gpt-5.5 routed into Responses-Lite" reports, which are all reported from ChatGPT-account sign-in flows (unexpected server-side routing): #31150, #30403, #30238, #30912, #31705, #31717.

This report is filed from a plain API-key / custom model_provider (Azure) setup with no ChatGPT account involved, which points at the same flag (use_responses_lite) but a different root cause: the bundled catalog is applied unconditionally to all providers, not just ChatGPT-backend ones. Also possibly related to #31843 (gpt-5.6 gets no tools in read-only sandbox), which may share the same "gpt-5.6 metadata wasn't validated against non-default configurations before shipping" theme.

Suggested labels: azure, custom-model, CLI.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    CLIIssues related to the Codex CLIazureIssues related to the Azure-hosted OpenAI modelsbugSomething isn't workingcustom-modelIssues related to custom model providers (including local models)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions