Skip to content

fix(proxy): withhold the primer from offloaded results and stop counting them as context savings - #495

Merged
inth3shadows merged 1 commit into
mainfrom
feat/offload-aware-ledger
Sep 30, 2026
Merged

inth3shadows merged 1 commit into
mainfrom
feat/offload-aware-ledger

Conversation

@inth3shadows

Copy link
Copy Markdown
Owner

Why

Claude Code saves an MCP result over its limit to a file and gives the model only the path. The limit is 25,000 tokens by default, set with MAX_MCP_OUTPUT_TOKENS (docs: https://code.claude.com/docs/en/mcp, "MCP output limits and warnings"). Transcripts put the real cutoff between 39,269 chars (inline) and 104,414 chars (offloaded). This caused two problems:

  • Primer: the lazy primer attaches to a session's first compressed result. If that result was over the limit, the primer went into the file, the model might never read it, and every later result went out compressed and unexplained.
  • Ledger: on the live ledger, about 20% of the reported saving was in these results. 48 results were over the limit even after terse (382,799 tokens that never reached context). Another 17, all secret.list_credentials, were over the limit raw but brought under it by terse (298,248 tokens).

What

  • offload_limit(client_name): only claude-code, which is the only client documented to offload. It reads MAX_MCP_OUTPUT_TOKENS, else uses 25,000. over_limit(text, limit):
    • skips tokenizing a text no longer than limit bytes;
    • tokenizes at most limit * 8 characters;
    • is fail-open.
  • Primer guard: at both attach sites (the text block and the Router never compresses structuredContent results in a session with no text-result tool: primer_hold never releases #463 typed wrapper), the primer is not attached to a result that would be over the limit with the primer counted in. The check runs before _claim_primer, so a router's shared latch stays armed for the next result.
  • Ledger: _emit_stats runs after the lock is released and classifies the whole result over what the client reads: the typed field if there is one, otherwise the text blocks joined. Every row of an affected result is tagged offload: "out" (terse's output was offloaded) or "raw" (only the raw result would have been). The tag is passed to the writer only when set, so older writers keep working.
  • Context basis (terse stats publishes a WIRE-basis saving; for a client that discards the mirror block the absolute tokens (and turns_covered) are ~1.9x overstated #420): "out" rows count as (0, 0). "raw" rows count as (out, out): terse's inline copy is paid, and the raw path's short file pointer is not treated as the baseline. terse stats --json keeps its shape, and older rows read as before.
  • The CHANGELOG v0.44.0 section is graduated, and an [Unreleased] entry is added.

Review

A domain-briefed reviewer found:

  • High: tiktoken raises on <|endoftext|> in input, and the primer guard ran outside any handler, so that would have killed the proxy's reader thread. over_limit is now fail-open.
  • Moderate: the tokenizer pass had no bound. A 2.4 MB payload took 558 ms, sometimes under _local_lock, and a server whose results are all huge would pay that on every response while the primer was pending. It is now capped at limit * 8 characters.
  • Smaller: the typed-wrapper check serialized before reading the limit, and the stats label now names offloaded results.

Cleared: guard ordering and the shared latch; reverting to the hold at the typed-wrapper site; the classification edge cases; context_tokens never going negative; the client name reaching each peer on both the cold and warm initialize paths.

Tests

  • 20 new tests in tests/test_offload.py:
    • the limit and its env override;
    • token rather than byte counting;
    • the primer skipping an over-limit result and attaching to the next, including when only the primer pushes it over;
    • a client that doesn't offload is unaffected;
    • a router's shared latch is not spent;
    • the typed wrapper;
    • ledger tags out, raw and none;
    • old seven-argument writers;
    • context basis;
    • <|endoftext|> input;
    • the bounded tokenizer pass.
  • Mutation checks: each test fails with its guard removed (text guard, typed-wrapper guard, out/raw classification, the context-basis zero, the fail-open catch).
  • Full suite: 2882 passed; ruff and mypy clean.

Not covered

  • Per-tool _meta["anthropic/maxResultSizeChars"]: no live server sets it.
  • Rewriting old rows: the ledger doesn't record which client wrote them.
  • cl100k is not Claude Code's own tokenizer, so counts near the limit are approximate.

…ing them as context savings

Claude Code saves an MCP result over its limit (25,000 tokens by default,
MAX_MCP_OUTPUT_TOKENS to change it) to a file and gives the model only the
path. The lazy primer could attach to such a result and never be read, and
the ledger counted its saving as context saved. On the live ledger about 20%
of the reported saving was in these results.

- The primer skips a result that would be offloaded, checked before the
  claim, so a router's shared latch stays armed (text block and #463 typed
  wrapper).
- Each ledger row of an affected result is tagged offload: "out" | "raw";
  on the context basis both save nothing. Older rows read as before.
- over_limit is bounded (tokenizes at most limit*8 chars) and fail-open
  (tiktoken raises on <|endoftext|>, which would kill the reader thread).

Also graduates the v0.44.0 CHANGELOG section.
@inth3shadows
inth3shadows merged commit 069059f into main Sep 30, 2026
7 checks passed
@inth3shadows
inth3shadows deleted the feat/offload-aware-ledger branch September 30, 2026 21:45
inth3shadows added a commit that referenced this pull request Oct 6, 2026
…ts token count (#496) (#497)

Fixes #496.

## Why
Since 0.44.1 (#495), terse holds the lazy primer back from a result
Claude Code will offload to a file. But it only treated a result as
offloaded above 25,000 cl100k tokens. Claude Code also offloads by
length:

| | chars | Claude Code |
|---|---|---|
| largest MCP result shown inline | 49,034 | 2.1.284 |
| smallest offloaded ("Output too large (50KB)") | 51,246 | 2.1.287 |

So a result of about 13k–25k tokens was offloaded with the primer still
attached. The model sees an offloaded result only as a short preview of
its start, and the primer, at block 0, filled that preview.

In the pre-registered Opus 5.5 check (2026-10-02), a 57 KB
`list_principles` result was offloaded in both arms. Without terse, the
preview held the answer row and Opus answered in 2 turns. Through terse,
Opus grepped the file 2–3 more times, so that task cost 88% more per
success. That one task turned the Opus result from about −2.8% into
+4.2%, and the pre-registered check failed.

## What
`OFFLOAD_DEFAULT_CHARS = 50_000`, checked first in `over_limit`. A
result whose text is longer than that counts as offloaded, whatever its
token count. Both the primer guard and the ledger's `offload` tag use
`over_limit`, so both are fixed, and no tokenizer pass ever sees more
than 50,000 chars. The token limit (`MAX_MCP_OUTPUT_TOKENS`, default
25,000) is unchanged.

The CHANGELOG's v0.44.1 section is graduated and an `[Unreleased]` entry
added.

## Tests
- Test-first: 3 new tests failed on the old code before the fix:
- a 60,852-char / ~17.8k-token result counts as offloaded, and a
46,770-char one does not;
- the primer skips a 56,750-char compressed result at the default
limits, then attaches to a smaller one;
  - the ledger tags that result `offload: "out"`.
- The tests clear any inherited `MAX_MCP_OUTPUT_TOKENS`, since the suite
also runs inside Claude Code.
- The bounded-tokenizer test now also checks that a 2M-char text is
never tokenized.
- Mutation check: with the char cutoff removed, 4 tests fail.
- Full suite: 2885 passed; ruff and mypy clean.

## Not covered
- Per-tool `_meta["anthropic/maxResultSizeChars"]`: no live server sets
it.
- Whether `MAX_MCP_OUTPUT_TOKENS` also moves the char cutoff:
unmeasured.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant