Repository navigation
fix(proxy): withhold the primer from offloaded results and stop counting them as context savings - #495
Merged
Conversation
…ing them as context savings Claude Code saves an MCP result over its limit (25,000 tokens by default, MAX_MCP_OUTPUT_TOKENS to change it) to a file and gives the model only the path. The lazy primer could attach to such a result and never be read, and the ledger counted its saving as context saved. On the live ledger about 20% of the reported saving was in these results. - The primer skips a result that would be offloaded, checked before the claim, so a router's shared latch stays armed (text block and #463 typed wrapper). - Each ledger row of an affected result is tagged offload: "out" | "raw"; on the context basis both save nothing. Older rows read as before. - over_limit is bounded (tokenizes at most limit*8 chars) and fail-open (tiktoken raises on <|endoftext|>, which would kill the reader thread). Also graduates the v0.44.0 CHANGELOG section.
inth3shadows
added a commit
that referenced
this pull request
Oct 6, 2026
…ts token count (#496) (#497) Fixes #496. ## Why Since 0.44.1 (#495), terse holds the lazy primer back from a result Claude Code will offload to a file. But it only treated a result as offloaded above 25,000 cl100k tokens. Claude Code also offloads by length: | | chars | Claude Code | |---|---|---| | largest MCP result shown inline | 49,034 | 2.1.284 | | smallest offloaded ("Output too large (50KB)") | 51,246 | 2.1.287 | So a result of about 13k–25k tokens was offloaded with the primer still attached. The model sees an offloaded result only as a short preview of its start, and the primer, at block 0, filled that preview. In the pre-registered Opus 5.5 check (2026-10-02), a 57 KB `list_principles` result was offloaded in both arms. Without terse, the preview held the answer row and Opus answered in 2 turns. Through terse, Opus grepped the file 2–3 more times, so that task cost 88% more per success. That one task turned the Opus result from about −2.8% into +4.2%, and the pre-registered check failed. ## What `OFFLOAD_DEFAULT_CHARS = 50_000`, checked first in `over_limit`. A result whose text is longer than that counts as offloaded, whatever its token count. Both the primer guard and the ledger's `offload` tag use `over_limit`, so both are fixed, and no tokenizer pass ever sees more than 50,000 chars. The token limit (`MAX_MCP_OUTPUT_TOKENS`, default 25,000) is unchanged. The CHANGELOG's v0.44.1 section is graduated and an `[Unreleased]` entry added. ## Tests - Test-first: 3 new tests failed on the old code before the fix: - a 60,852-char / ~17.8k-token result counts as offloaded, and a 46,770-char one does not; - the primer skips a 56,750-char compressed result at the default limits, then attaches to a smaller one; - the ledger tags that result `offload: "out"`. - The tests clear any inherited `MAX_MCP_OUTPUT_TOKENS`, since the suite also runs inside Claude Code. - The bounded-tokenizer test now also checks that a 2M-char text is never tokenized. - Mutation check: with the char cutoff removed, 4 tests fail. - Full suite: 2885 passed; ruff and mypy clean. ## Not covered - Per-tool `_meta["anthropic/maxResultSizeChars"]`: no live server sets it. - Whether `MAX_MCP_OUTPUT_TOKENS` also moves the char cutoff: unmeasured.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Claude Code saves an MCP result over its limit to a file and gives the model only the path. The limit is 25,000 tokens by default, set with
MAX_MCP_OUTPUT_TOKENS(docs: https://code.claude.com/docs/en/mcp, "MCP output limits and warnings"). Transcripts put the real cutoff between 39,269 chars (inline) and 104,414 chars (offloaded). This caused two problems:secret.list_credentials, were over the limit raw but brought under it by terse (298,248 tokens).What
offload_limit(client_name): onlyclaude-code, which is the only client documented to offload. It readsMAX_MCP_OUTPUT_TOKENS, else uses 25,000.over_limit(text, limit):limitbytes;limit * 8characters;_claim_primer, so a router's shared latch stays armed for the next result._emit_statsruns after the lock is released and classifies the whole result over what the client reads: the typed field if there is one, otherwise the text blocks joined. Every row of an affected result is taggedoffload: "out"(terse's output was offloaded) or"raw"(only the raw result would have been). The tag is passed to the writer only when set, so older writers keep working."out"rows count as (0, 0)."raw"rows count as (out, out): terse's inline copy is paid, and the raw path's short file pointer is not treated as the baseline.terse stats --jsonkeeps its shape, and older rows read as before.[Unreleased]entry is added.Review
A domain-briefed reviewer found:
<|endoftext|>in input, and the primer guard ran outside any handler, so that would have killed the proxy's reader thread.over_limitis now fail-open._local_lock, and a server whose results are all huge would pay that on every response while the primer was pending. It is now capped atlimit * 8characters.Cleared: guard ordering and the shared latch; reverting to the hold at the typed-wrapper site; the classification edge cases;
context_tokensnever going negative; the client name reaching each peer on both the cold and warm initialize paths.Tests
tests/test_offload.py:out,rawand none;<|endoftext|>input;Not covered
_meta["anthropic/maxResultSizeChars"]: no live server sets it.