Skip to content

GPT-6 Astra: repeated "Selected model is at capacity" on Pro 20x — rollout logs show server_overloaded with 0% usage used #43700

Description

@FranklinChip

What version of the Codex App are you using (From “About Codex” dialog)?

Codex Desktop (bundled codex-cli 0.153.4)

What subscription do you have?

ChatGPT Pro (20x)

What platform is your computer?

Darwin 25.5.0 arm64 (Apple Silicon Mac, macOS)

What issue are you seeing?

Codex Desktop repeatedly terminates turns with:

Selected model is at capacity. Please try a different model.

Model: gpt-6-astra. This has been recurring since Aug 28 and escalated sharply on Sep 7–8.

Local rollout JSONL logs classify all of these terminations as:

"codex_error_info": "server_overloaded"

while the rate-limit snapshots in the same logs show the account is nowhere near any limit:

plan_type: pro
used_percent: 0.0
rate_limit_reached_type: null

Measured from my local rollout logs:

Date (UTC) Occurrences Notes
Aug 28 39 banner lines / 30 server_overloaded single day
Sep 7 65 server_overloaded across 5 sessions worst session: 27 terminal errors between 07:14–12:27 UTC
Sep 8 6 server_overloaded between 03:02–04:38 UTC today

Important detail: turns are not rejected upfront — they are admitted, run for a while, then killed mid-flight. In the Sep 7 session the killed turns ran on average 85 s before termination (max 275 s), losing in-progress work on long-running tasks.

Retrying succeeds intermittently, then fails again in bursts. Switching models and starting new sessions does not reliably recover.

What steps can reproduce the bug?

  1. Sign in with a ChatGPT Pro account (auth via ChatGPT, not API key) in Codex Desktop 0.153.4 on macOS arm64.
  2. Select gpt-6-astra.
  3. Start normal coding tasks.
  4. Turns intermittently terminate with the capacity banner; the rollout log records codex_error_info: "server_overloaded" while used_percent stays at 0.0.

Session IDs for server-side correlation:

  • 01a07ab7-4da3-7cf0-b711-638991ba5df0 (Sep 7, 27 terminal errors, 07:14–12:27 UTC)
  • 01a07ef7-1841-70e0-9344-a8d33ad34714 (Sep 8, 6 terminal errors, 03:02–04:38 UTC)

What is the expected behavior?

  1. The UI message should distinguish server-side capacity / admission-control failures from per-account quota exhaustion — "Selected model is at capacity. Please try a different model." reads like quota guidance, but my quota was at 0% used and switching models does not help.
  2. Long-running turns should be retried with bounded backoff within the same turn (or resume from a preserved state) instead of being killed mid-flight after minutes of work.
  3. status.openai.com showed "fully operational" with 100% Codex uptime during these windows; Codex-side regional serving degradation is not reflected there.

Additional information

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    appIssues related to the Codex desktop appbugSomething isn't workingconnectivityIssues involving networking or endpoint connectivity problems (disconnections)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions