Skip to content

Spark compaction fails with Responses-Lite and hides upstream 400 #1154

Description

@bcdonadio

Observed behavior

timeout 30 lcm compact --all repeatedly fails through codex-process with a generic CLI compatibility message. A supplied proxy report identifies an upstream HTTP 400 rejection of the Spark/Responses-Lite combination. LCM hides that rejection behind HTTP 502 and advice to upgrade Codex.

Reproduced on 2026-09-07:

Found 29 uncompacted conversations (626.4k tokens)

conv #29 (151 msgs, 37.6k tokens) FAILED (Codex CLI rejected the compaction request (exit 1; diagnostic output omitted): provider codex-process, model "gpt-5.3-codex-spark", reasoning effort "default/omitted", fast mode false. Upgrade the Codex CLI or choose a supported model and control combination.)
[1/29] ... 12.8s
conv #19 (202 msgs, 37.3k tokens) FAILED (same diagnostic)
[2/29] ... 14.4s
conv #30 (44 msgs, 24.4k tokens) FAILED (The operation was aborted)
[3/29] ... 0.3s
Batch compact complete.

The shell returned 124. The third failure is cancellation from the requested 30-second timeout; the first two provider exits happened before cancellation. No successful compaction was reported in this run. Project paths and conversation identifiers are omitted.

A separate call to the installed package's createCodexProcessSummarizer, using model: "gpt-5.3-codex-spark", fastMode: false, and a short synthetic greeting-summary prompt, reproduced the exact generic compatibility error without reading or compacting stored conversations. A spawn wrapper captured stderr privately. Sanitized relevant output:

OpenAI Codex v0.153.4
model: gpt-5.3-codex-spark
provider: lcm_compaction
ERROR: unexpected status 502 Bad Gateway: codex responses gateway request failed, url: http://127.0.0.1:<port>/<redacted-capability>/responses

A preceding model-catalog refresh also received 404 from the private gateway. Its causal role in selection of the Lite protocol is unproven.

Proxy evidence supplied by the reporter

The supplied record is dated 2026-09-07T04:37:31.445Z. It is supporting evidence from a separate request, not a proven one-to-one correlation to either of the two batch failures or the synthetic probe.

{
  "model": "gpt-5.3-codex-spark",
  "inboundProtocol": "responses",
  "inboundTransport": "http",
  "upstreamTransport": "websocket",
  "status": 400,
  "durationMs": 800,
  "errorCode": "invalid_request_error",
  "upstreamError": "This model is not supported when using X-OpenAI-Internal-Codex-Responses-Lite.",
  "requestedEffort": "low",
  "effectiveEffort": "low",
  "configuredServiceTier": "default",
  "upstreamRequestAccepted": false,
  "websocketHandshakeStatus": 101,
  "errorParam": "model",
  "toolDefinitionCount": 0,
  "toolCallCount": 0,
  "streamingRequested": true,
  "usageStatus": "unreported",
  "proxyVersion": "2.45.0"
}

This rules out a connection/handshake failure for that supplied request: the WebSocket handshake succeeded, then upstream rejected the request before accepting model work. It does not establish account quota exhaustion. The source report has redactionApplied: true and captureTruncated: true; raw outbound headers are not included. Account identifiers, capability URLs, request/session IDs, and user content are intentionally excluded from this issue.

Root cause and source trace

The diagnostic loss is confirmed by source at checkout a824d63d31876fc1fcca960f32a621b4392cf615 and by the installed runtime probe:

  1. src/llm/codex-responses-gateway.ts:63,200-204,313-335,710-718 accepts and forwards x-openai-internal-codex-responses-lite. With Lite enabled, it builds additional_tools plus the immutable user summary prompt, rather than the standard top-level tools: [] shape.
  2. src/llm/codex-responses-gateway.ts:734-737 categorizes only HTTP 429 (usage) and 401 (authentication). Every non-success response, including an upstream 400 model/protocol rejection, becomes a generic 502. The underlying structured rejection is not preserved as a safe failure category.
  3. src/llm/codex-process.ts:489-505 sees the nonzero Codex exit. Neither the gateway category nor stderr identifies the original 400, so it falls through to createProcessCompatibilityError.
  4. src/llm/process-utils.ts:37-52 renders the upgrade/model-controls advice, obscuring the specific model/protocol incompatibility.

The supplied upstream report establishes a Spark/Lite incompatibility for that request. The origin of the Lite selection remains to be isolated: Codex bundled model metadata, effective catalog/configuration under LCM's isolated invocation, or proxy behavior. LCM's --ignore-user-config invocation and a gateway that rejects /models deserve a targeted capability-resolution check. Do not assume that merely upgrading Codex fixes this, or that deleting the header alone is valid: the request body also changes with Lite mode.

Expected behavior

  • The configured Spark compactor should select a supported transport/body contract or fail with an accurate, safely categorized model/protocol error.
  • Preserve a bounded structured upstream failure category through the gateway and CLI without exposing raw diagnostics, credentials, prompts, or capability URLs.
  • Reserve CLI upgrade advice for evidence of CLI incompatibility.
  • Treat external timeout cancellation separately from provider rejection.

How to reproduce

  1. Use the installed LCM CLI with a managed daemon, llm.provider = "codex-process", and llm.model = "gpt-5.3-codex-spark", through the affected Codex/proxy combination.
  2. Have eligible stored conversations.
  3. Run timeout 30 lcm compact --all.
  4. Compare pre-timeout provider failures with the proxy's upstream error, rather than interpreting exit 124 as the root cause.
  5. For a smaller probe, import the installed dist/src/llm/codex-process.js and invoke createCodexProcessSummarizer({ model: "gpt-5.3-codex-spark", fastMode: false, timeoutMs: 25000 }) with synthetic text. This still makes a real provider request, but requires no stored conversation.

Regression coverage to add

  • A deterministic fake upstream returning HTTP 400 with the supplied model/Lite error; assert a safe useful error reaches the public summarizer and batch CLI.
  • A model-capability case for Spark through the actual isolated Codex invocation, including catalog-fetch failure/fallback. Validate header and body together.
  • Preserve redaction, no-tools isolation, and cancellation/drain behavior.

Existing tests cover recognized stderr phrases and gateway 401/429 classifications; arbitrary errors currently assert the generic compatibility fallback. No production files were changed and no fix is claimed.

Environment

  • Agent: Codex investigation; provider CLI codex-cli 0.153.4
  • Connector: LCM CLI -> managed daemon -> codex-process -> private Responses gateway -> OpenCodex proxy
  • OS: Fedora Linux 44, x86_64
  • Installed LCM: 1.4.2 (independent global package, not this checkout)
  • Node: v25.9.0
  • Proxy report: OpenCodex 2.45.0 / Bun 1.4.0
  • Configured provider request timeout: 600000 ms; requested outer command timeout: 30 seconds
  • Related historical issue: Codex process hides usage-limit failures as compatibility errors #780 fixed usage-limit classification; this is the remaining upstream 400/model-protocol diagnostic path, not a demonstrated recurrence of quota exhaustion.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    Fields

    Priority

    High

    Effort

    None yet

    Projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions