Preflight Checklist
What's Wrong?
On a Max subscription, the same workload uses about 1.8× more of the 5-hour window when it runs headless (claude -p --input-format stream-json, recorded as entrypoint: "sdk-cli" in the transcripts) than when it runs in the interactive TUI (entrypoint: "cli"). Token usage is the same in both modes.
The Agent SDK help article (https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan) says claude -p "still draw[s] from your subscription's usage limits", and I could not find anything saying it is weighted differently.
After moving the same work from the TUI to claude -p, we hit the 5h limit several times a day with no change in workload, and we paid for extra usage (rejected + isUsingOverage: true) that the same work in the TUI would not have needed.
What Should Happen?
The same tokens on the same model should use the same share of the 5-hour window, whether they come from claude -p or from the interactive TUI. If headless usage is intentionally weighted differently, it should be documented.
Error Messages/Logs
Controlled A/B test, 2026-09-25 10:27–10:41 UTC, Claude Code 2.1.282, claude-opus-5, effort high
Phase | Test spend (API-equiv.) | Other spend | five_hour utilization | $ per 1% of 5h
-------------|-------------------------|--------------------|-----------------------|---------------
headless | $12.74 | $0.40 (headless) | 52% -> 60% | ~1.6 (1.5-1.9)
interactive | $12.91 | $1.12 (headless) | 60% -> 65% | ~3.0 (2.4-3.9)
Ranges come from the 1% resolution of the utilization reading; they do not overlap.
Historical (same account, all Opus 5, API-equivalent spend per 1% of each 5h window that peaked above 25%):
2026-09-18..22 cli 2.1.267/2.1.270 3.1-4.1 $ per 1% (8 windows)
2026-09-23..24 sdk-cli 2.1.267 1.6-1.9 $ per 1% (5 windows)
Steps to Reproduce
- Create a directory with ~250K tokens of plain text split into files small enough for
Read (I used 18 files of ~45K characters of source code).
- Prepare a fixed prompt list: turn 1 asks Claude to read all files in parallel and reply only "read"; turns 2–41 are short questions about the content answered in one sentence, without tools.
- Headless phase: run
claude -p --input-format stream-json --output-format stream-json --verbose --permission-prompt-tool stdio --replay-user-messages --append-system-prompt "" --dangerously-skip-permissions --model claude-opus-5
and send the prompts one by one as stream-json user messages, waiting for each result. Record rate_limit_event.rate_limit_info.unifiedWindows.five_hour.utilization.
- Interactive phase: run
claude --model claude-opus-5 --dangerously-skip-permissions in tmux, send the same prompts with tmux send-keys, detect turn end with a Stop hook, and read rate_limits.five_hour.used_percentage from the statusline JSON.
- Before and after each phase, run a tiny probe (
claude -p --model claude-haiku-4-5 --output-format stream-json --verbose "reply: ok") to read utilization the same way.
- Compute each phase's token cost from the session transcripts (
message.usage, deduplicated by message id + request id) at list API prices, and subtract any other activity on the account during the phase (measured from the transcripts of every machine on the account).
- Compare the $ per 1% of the 5-hour window: headless comes out ~1.8× more expensive for the same tokens.
Checks already done:
- Not the auto-mode classifier: both phases used --dangerously-skip-permissions, and 40 of 41 turns made no tool calls.
- Not a larger harness in headless: the first request of a headless session is slightly smaller than an interactive one (median ~66K vs ~88K input tokens on 2.1.267).
- Not missing transcript entries: headless
result.total_cost_usd matches the transcript totals.
- Not usage from other machines: measured and subtracted.
Claude Model
Opus
Is this a regression?
I don't know
Last Working Version
No response
Claude Code Version
2.1.282 (Claude Code) — also observed on 2.1.267
Platform
Anthropic API
Operating System
Ubuntu/Debian Linux
Terminal/Shell
Non-interactive/CI environment
Additional Information
Related:
I can share the test scripts and the per-turn measurements if that helps.
Preflight Checklist
What's Wrong?
On a Max subscription, the same workload uses about 1.8× more of the 5-hour window when it runs headless (
claude -p --input-format stream-json, recorded asentrypoint: "sdk-cli"in the transcripts) than when it runs in the interactive TUI (entrypoint: "cli"). Token usage is the same in both modes.The Agent SDK help article (https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan) says
claude -p"still draw[s] from your subscription's usage limits", and I could not find anything saying it is weighted differently.After moving the same work from the TUI to
claude -p, we hit the 5h limit several times a day with no change in workload, and we paid for extra usage (rejected+isUsingOverage: true) that the same work in the TUI would not have needed.What Should Happen?
The same tokens on the same model should use the same share of the 5-hour window, whether they come from
claude -por from the interactive TUI. If headless usage is intentionally weighted differently, it should be documented.Error Messages/Logs
Steps to Reproduce
Read(I used 18 files of ~45K characters of source code).claude -p --input-format stream-json --output-format stream-json --verbose --permission-prompt-tool stdio --replay-user-messages --append-system-prompt "" --dangerously-skip-permissions --model claude-opus-5
and send the prompts one by one as stream-json user messages, waiting for each
result. Recordrate_limit_event.rate_limit_info.unifiedWindows.five_hour.utilization.claude --model claude-opus-5 --dangerously-skip-permissionsin tmux, send the same prompts withtmux send-keys, detect turn end with a Stop hook, and readrate_limits.five_hour.used_percentagefrom the statusline JSON.claude -p --model claude-haiku-4-5 --output-format stream-json --verbose "reply: ok") to read utilization the same way.message.usage, deduplicated by message id + request id) at list API prices, and subtract any other activity on the account during the phase (measured from the transcripts of every machine on the account).Checks already done:
result.total_cost_usdmatches the transcript totals.Claude Model
Opus
Is this a regression?
I don't know
Last Working Version
No response
Claude Code Version
2.1.282 (Claude Code) — also observed on 2.1.267
Platform
Anthropic API
Operating System
Ubuntu/Debian Linux
Terminal/Shell
Non-interactive/CI environment
Additional Information
Related:
claude -pdraws from #81692 (which usage poolclaude -pdraws from; closed)I can share the test scripts and the per-turn measurements if that helps.