Skip to content

[BUG] Headless claude -p (entrypoint sdk-cli) uses ~1.8× more of the 5-hour window per token than interactive cli, same workload #97074

Description

@qvarker

Preflight Checklist

  • I have searched existing issues and this hasn't been reported yet
  • This is a single bug report (please file separate reports for different bugs)
  • I am using the latest version of Claude Code

What's Wrong?

On a Max subscription, the same workload uses about 1.8× more of the 5-hour window when it runs headless (claude -p --input-format stream-json, recorded as entrypoint: "sdk-cli" in the transcripts) than when it runs in the interactive TUI (entrypoint: "cli"). Token usage is the same in both modes.

The Agent SDK help article (https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan) says claude -p "still draw[s] from your subscription's usage limits", and I could not find anything saying it is weighted differently.

After moving the same work from the TUI to claude -p, we hit the 5h limit several times a day with no change in workload, and we paid for extra usage (rejected + isUsingOverage: true) that the same work in the TUI would not have needed.

What Should Happen?

The same tokens on the same model should use the same share of the 5-hour window, whether they come from claude -p or from the interactive TUI. If headless usage is intentionally weighted differently, it should be documented.

Error Messages/Logs

Controlled A/B test, 2026-09-25 10:27–10:41 UTC, Claude Code 2.1.282, claude-opus-5, effort high

Phase        | Test spend (API-equiv.) | Other spend        | five_hour utilization | $ per 1% of 5h
-------------|-------------------------|--------------------|-----------------------|---------------
headless     | $12.74                  | $0.40 (headless)   | 52% -> 60%            | ~1.6 (1.5-1.9)
interactive  | $12.91                  | $1.12 (headless)   | 60% -> 65%            | ~3.0 (2.4-3.9)

Ranges come from the 1% resolution of the utilization reading; they do not overlap.

Historical (same account, all Opus 5, API-equivalent spend per 1% of each 5h window that peaked above 25%):
2026-09-18..22  cli      2.1.267/2.1.270  3.1-4.1 $ per 1%  (8 windows)
2026-09-23..24  sdk-cli  2.1.267          1.6-1.9 $ per 1%  (5 windows)

Steps to Reproduce

  1. Create a directory with ~250K tokens of plain text split into files small enough for Read (I used 18 files of ~45K characters of source code).
  2. Prepare a fixed prompt list: turn 1 asks Claude to read all files in parallel and reply only "read"; turns 2–41 are short questions about the content answered in one sentence, without tools.
  3. Headless phase: run
    claude -p --input-format stream-json --output-format stream-json --verbose --permission-prompt-tool stdio --replay-user-messages --append-system-prompt "" --dangerously-skip-permissions --model claude-opus-5
    and send the prompts one by one as stream-json user messages, waiting for each result. Record rate_limit_event.rate_limit_info.unifiedWindows.five_hour.utilization.
  4. Interactive phase: run claude --model claude-opus-5 --dangerously-skip-permissions in tmux, send the same prompts with tmux send-keys, detect turn end with a Stop hook, and read rate_limits.five_hour.used_percentage from the statusline JSON.
  5. Before and after each phase, run a tiny probe (claude -p --model claude-haiku-4-5 --output-format stream-json --verbose "reply: ok") to read utilization the same way.
  6. Compute each phase's token cost from the session transcripts (message.usage, deduplicated by message id + request id) at list API prices, and subtract any other activity on the account during the phase (measured from the transcripts of every machine on the account).
  7. Compare the $ per 1% of the 5-hour window: headless comes out ~1.8× more expensive for the same tokens.

Checks already done:

  • Not the auto-mode classifier: both phases used --dangerously-skip-permissions, and 40 of 41 turns made no tool calls.
  • Not a larger harness in headless: the first request of a headless session is slightly smaller than an interactive one (median ~66K vs ~88K input tokens on 2.1.267).
  • Not missing transcript entries: headless result.total_cost_usd matches the transcript totals.
  • Not usage from other machines: measured and subtracted.

Claude Model

Opus

Is this a regression?

I don't know

Last Working Version

No response

Claude Code Version

2.1.282 (Claude Code) — also observed on 2.1.267

Platform

Anthropic API

Operating System

Ubuntu/Debian Linux

Terminal/Shell

Non-interactive/CI environment

Additional Information

Related:

I can share the test scripts and the per-turn measurements if that helps.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions