Skip to content

Tags: inth3shadows/terse

Tags

v0.44.2

Toggle v0.44.2's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(proxy): treat a result over 50,000 chars as offloaded, whatever i…

…ts token count (#496) (#497)

Fixes #496.

## Why
Since 0.44.1 (#495), terse holds the lazy primer back from a result
Claude Code will offload to a file. But it only treated a result as
offloaded above 25,000 cl100k tokens. Claude Code also offloads by
length:

| | chars | Claude Code |
|---|---|---|
| largest MCP result shown inline | 49,034 | 2.1.284 |
| smallest offloaded ("Output too large (50KB)") | 51,246 | 2.1.287 |

So a result of about 13k–25k tokens was offloaded with the primer still
attached. The model sees an offloaded result only as a short preview of
its start, and the primer, at block 0, filled that preview.

In the pre-registered Opus 5.5 check (2026-10-02), a 57 KB
`list_principles` result was offloaded in both arms. Without terse, the
preview held the answer row and Opus answered in 2 turns. Through terse,
Opus grepped the file 2–3 more times, so that task cost 88% more per
success. That one task turned the Opus result from about −2.8% into
+4.2%, and the pre-registered check failed.

## What
`OFFLOAD_DEFAULT_CHARS = 50_000`, checked first in `over_limit`. A
result whose text is longer than that counts as offloaded, whatever its
token count. Both the primer guard and the ledger's `offload` tag use
`over_limit`, so both are fixed, and no tokenizer pass ever sees more
than 50,000 chars. The token limit (`MAX_MCP_OUTPUT_TOKENS`, default
25,000) is unchanged.

The CHANGELOG's v0.44.1 section is graduated and an `[Unreleased]` entry
added.

## Tests
- Test-first: 3 new tests failed on the old code before the fix:
- a 60,852-char / ~17.8k-token result counts as offloaded, and a
46,770-char one does not;
- the primer skips a 56,750-char compressed result at the default
limits, then attaches to a smaller one;
  - the ledger tags that result `offload: "out"`.
- The tests clear any inherited `MAX_MCP_OUTPUT_TOKENS`, since the suite
also runs inside Claude Code.
- The bounded-tokenizer test now also checks that a 2M-char text is
never tokenized.
- Mutation check: with the char cutoff removed, 4 tests fail.
- Full suite: 2885 passed; ruff and mypy clean.

## Not covered
- Per-tool `_meta["anthropic/maxResultSizeChars"]`: no live server sets
it.
- Whether `MAX_MCP_OUTPUT_TOKENS` also moves the char cutoff:
unmeasured.

v0.44.1

Toggle v0.44.1's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(proxy): withhold the primer from offloaded results and stop count…

…ing them as context savings (#495)

## Why
Claude Code saves an MCP result over its limit to a file and gives the
model only the path. The limit is 25,000 tokens by default, set with
`MAX_MCP_OUTPUT_TOKENS` (docs: https://code.claude.com/docs/en/mcp, "MCP
output limits and warnings"). Transcripts put the real cutoff between
39,269 chars (inline) and 104,414 chars (offloaded). This caused two
problems:
- **Primer:** the lazy primer attaches to a session's first compressed
result. If that result was over the limit, the primer went into the
file, the model might never read it, and every later result went out
compressed and unexplained.
- **Ledger:** on the live ledger, about 20% of the reported saving was
in these results. 48 results were over the limit even after terse
(382,799 tokens that never reached context). Another 17, all
`secret.list_credentials`, were over the limit raw but brought under it
by terse (298,248 tokens).

## What
- `offload_limit(client_name)`: only `claude-code`, which is the only
client documented to offload. It reads `MAX_MCP_OUTPUT_TOKENS`, else
uses 25,000. `over_limit(text, limit)`:
  - skips tokenizing a text no longer than `limit` bytes;
  - tokenizes at most `limit * 8` characters;
  - is fail-open.
- **Primer guard:** at both attach sites (the text block and the #463
typed wrapper), the primer is not attached to a result that would be
over the limit with the primer counted in. The check runs before
`_claim_primer`, so a router's shared latch stays armed for the next
result.
- **Ledger:** `_emit_stats` runs after the lock is released and
classifies the whole result over what the client reads: the typed field
if there is one, otherwise the text blocks joined. Every row of an
affected result is tagged `offload: "out"` (terse's output was
offloaded) or `"raw"` (only the raw result would have been). The tag is
passed to the writer only when set, so older writers keep working.
- **Context basis (#420):** `"out"` rows count as (0, 0). `"raw"` rows
count as (out, out): terse's inline copy is paid, and the raw path's
short file pointer is not treated as the baseline. `terse stats --json`
keeps its shape, and older rows read as before.
- The CHANGELOG v0.44.0 section is graduated, and an `[Unreleased]`
entry is added.

## Review
A domain-briefed reviewer found:
- **High:** tiktoken raises on `<|endoftext|>` in input, and the primer
guard ran outside any handler, so that would have killed the proxy's
reader thread. `over_limit` is now fail-open.
- **Moderate:** the tokenizer pass had no bound. A 2.4 MB payload took
558 ms, sometimes under `_local_lock`, and a server whose results are
all huge would pay that on every response while the primer was pending.
It is now capped at `limit * 8` characters.
- **Smaller:** the typed-wrapper check serialized before reading the
limit, and the stats label now names offloaded results.

Cleared: guard ordering and the shared latch; reverting to the hold at
the typed-wrapper site; the classification edge cases; `context_tokens`
never going negative; the client name reaching each peer on both the
cold and warm initialize paths.

## Tests
- 20 new tests in `tests/test_offload.py`:
  - the limit and its env override;
  - token rather than byte counting;
- the primer skipping an over-limit result and attaching to the next,
including when only the primer pushes it over;
  - a client that doesn't offload is unaffected;
  - a router's shared latch is not spent;
  - the typed wrapper;
  - ledger tags `out`, `raw` and none;
  - old seven-argument writers;
  - context basis;
  - `<|endoftext|>` input;
  - the bounded tokenizer pass.
- Mutation checks: each test fails with its guard removed (text guard,
typed-wrapper guard, out/raw classification, the context-basis zero, the
fail-open catch).
- Full suite: 2882 passed; ruff and mypy clean.

## Not covered
- Per-tool `_meta["anthropic/maxResultSizeChars"]`: no live server sets
it.
- Rewriting old rows: the ledger doesn't record which client wrote them.
- cl100k is not Claude Code's own tokenizer, so counts near the limit
are approximate.

v0.44.0

Toggle v0.44.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat(multiproxy): declared snapshot_markers so new worktrees start wa…

…rm (#479) (#494)

## Why
#479's known limit: the router's saved replies (#270) were keyed by
launch directory whenever a peer inherits the router's cwd. The first
session in every new directory, which is most per-session worktrees,
blocked on the slowest peer's `initialize` (4.2s measured, plus a
`tools_changed` cache miss for headless runs).

## What
- **`snapshot_markers` (new, optional, per peer in the peers file).** A
list of relative paths declaring that the peer's replies depend on the
directory only through the nearest match at or above it, resolved
through symlinks. `[]` means not at all. When every cwd-inheriting peer
declares markers, the snapshot scope is those resolved matches.
Otherwise it is the cwd, exactly as before: an undeclared config keys
and fingerprints byte-identically to `main`, so existing snapshots stay
valid. A malformed value, or one on a `url` peer or a peer with its own
`cwd`, is a config error.
- **Prune:** saving a snapshot deletes sibling snapshots that have been
neither saved nor served for 30 days. `load` touches the file it serves,
so an in-use snapshot whose replies never change is not pruned.
- **Warm/cold on the ledger:** each `router_session` row carries
`snapshot: "warm"|"cold"`. It is ledger-only; `terse stats --json` is a
pinned contract, and counting older rows as not-warm would read unknown
as cold.
- **`install-mcp`:** carries a peer's markers over when it re-folds that
peer, but only onto a `command` peer without its own `cwd`.
- USAGE and CHANGELOG `[Unreleased]` document the field.

## Rejected along the way (in this PR's history)
Keying by `git rev-parse --git-common-dir` (c540714). Review and a live
check showed that 9 repos mix worktrees with and without a `.codegraph`
index (Claude Code's agent worktrees have none; terse has 14 of 19
without one), and the container dir and `.bare` resolved to the same
key. One key per repo served the wrong replies on every switch. A
hardcoded ".codegraph walking up" bit was also rejected: it puts one
peer's behaviour into a generic router.

## Review
Three review passes: a broad correctness pass and a domain-briefed pass
on c540714+0c5bb53, then a domain pass on the redesign. Found: the
repo-keying premise (both reviewers), prune deleting an in-use snapshot
(both), a `url` peer keeping markers on re-fold, which would make the
router refuse the fleet (fixed in 58b161f), and test nits. Cleared:
warm/cold attribution; prune races; the undeclared-config fingerprint
matching `main` (computed on both versions); the scope agreeing with
`codegraph-gate` for `$HOME`, tableless indexes, a `.codegraph` file,
broken symlinks and nested indexes.

## Tests
- New: shared marker → shared snapshot and warm start (symlinked
worktree + subdirectory); a different or missing marker stays separate;
no marker anywhere is shared; one undeclared peer keeps per-directory;
undeclared fingerprint equals the pre-field formula; malformed markers
and markers alongside `cwd` rejected; prune keeps its own and young
files; a served stale snapshot survives prune; warm/cold rows;
install-mcp carry-over, including `url` and `cwd`.
- Mutation checks: each behaviour's test fails under the mutated code
(markers ignored, no symlink resolve, no touch, no carry-over, url
carry-over, unstable fingerprint, no-op prune, always-cold).
- Full suite: 2862 passed; ruff and mypy clean.

## Live rollout (separate, Protected Change)
Takes effect only after the peers file declares markers: codegraph
`[".codegraph"]`, runecho `[]`, kb `[]`.

v0.43.0

Toggle v0.43.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat(proxy): --primer always|never|auto (#325 Phase 1) (#488)

Refs #325 (Phase 1 of 2: the switch. Phase 2, the cost-per-task harness
arm comparing always / never / auto, is not in this PR.)

## What

- `terse proxy --primer {always,never,auto}`, standalone and `--config`
router. Default `always` is today's lazy attach, so no existing install
changes.
- `never`: no primer, ever (also skips the eager `initialize` path).
Under a router the `structuredContent` hold is lifted, since no primer
is coming.
- `auto`: attaches lazily, but only on the first result whose compressed
form carries `subcols`, `absent_cols`/`sentinel_cols`, an object- or
array-valued alias, or a diff, text diff, embedded JSON or dropped
marker. Flat tables and scalar-only aliases go out without it. Once
attached, the session stays latched as today. Router: one mode on the
shared `PrimerLatch`. The #463 typed-wrap path releases an auto-declined
result compressed rather than reverting it to the raw hold.
- Ledger: `never` and `auto` declines are written once per session as
`attached: false` primer rows with a new `reason` (`never` / `auto` /
`structured`; older suppression rows default to `structured`).
`aggregate` keys on the reason, and `--json` `primers[]` rows now carry
`reason`.
- `primer_liability` reads the baked `--primer` (via `parse_proxy_opts`,
terse's side of `--` only). `never` gives a zero with the new
`primer_source: "configured"` unless an attach was recorded. `auto` uses
the existing recorded, measured-zero or estimate path, and when it falls
back to the estimate the report calls it an upper bound. Each liability
server row gets `primer_mode`.

## The shape signal: deviation from the plan

The plan named alias density and table width. Measured over #249's
corpora (codec only, no model), neither separates stress from real,
because the real payloads score higher on both (gh_pulls: 36 cols,
density 0.42; stress.heavy_alias: 4 cols, 0.33). Instead, `auto` uses an
allowlist taken from #249's per-question evidence: every regression was
`deref` on a `subcols` row or on an object alias with absent keys, and
flat or wide or heavily aliased tables were read at raw parity. The
constants are marked provisional in `proxy.py`.

What this means in practice: on the real corpus, `auto` still attaches
on 6 of 9 payloads. It declines only on flat tables with scalar aliases
(gh_labels, gh_commits_flat; gh_repo_single is vacuous). On stress it
declines wide_table, long_table, heavy_alias and mixed_realistic, and
attaches on nested_records, object_alias and the synthetic shape. Phase
2 should measure whether that is worth anything.

## Not done
- `install-mcp` does not bake `--primer`, and a re-wrap drops a
hand-added one.
- POSITIONING.md economics rewrite is deferred until Phase 2 has
numbers, per the plan.

## Verification
`uv run pytest -q`: 2804 passed (33 new in `tests/test_primer_mode.py`).
`ruff check .` and `mypy` are clean. No model or benchmark runs.

v0.42.1

Toggle v0.42.1's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(multiproxy): answer initialize from a persisted snapshot (#479)

Fixes #270.

## Problem
The router broadcast `initialize` and replied only when the slowest peer
answered: 4.2s end to end on the live fleet (kb 4.3s, runecho 3.8s,
codegraph 0.3s; [per-peer
timings](#270 (comment))).
Headless clients (`claude --print`, SDK, CI) send their first request
before then, so terse's tools arrive on turn 2 and the API reports
`cache_miss_reason: tools_changed` ([headless
evidence](#270 (comment));
43% of one session's cache writes were never read).

## Design
- **Snapshot.** After a complete session the router saves each peer's
**raw** `initialize` and list replies (`tools/list`, plus `prompts/list`
/ `resources/list` / `resources/templates/list` if the client asked for
them) to `$XDG_STATE_HOME/terse/router-snapshots/<sha256(resolved config
path [+ launch dir])[:16]>.json`. The file is written atomically and is
0600. It carries `peers_fingerprint`, a sha256 over the parsed
`downstreams[]` specs plus the router's cwd whenever any peer inherits
it (codegraph serves zero tools outside an indexed repo, so a snapshot
must never cross directories). It also stores the client
`protocolVersion` the replies were negotiated for; only the digest is
stored, because headers and env can hold credentials.
- **Fast path.** If the fingerprint matches and the client requests the
same `protocolVersion`, `initialize` and the list methods are answered
from the snapshot straight away. The stored replies go through the same
live merge code, so naming, `terse.retrieve`, primer mode,
`serverInfo.version` and the routing table all come from the running
process. The peers start in the background.
- **Calls to a peer that isn't ready.** They wait in that peer's
existing FIFO sender queue, behind its own `initialize`, with the same
routed-call timeout as before. There is no new buffer.
- **Reconcile.** Once every peer has answered `initialize` **and** the
client has sent `notifications/initialized`, the router re-lists as its
own background broadcasts. It installs the live tables (a reserved seq
means live listings always win the seq guard) and compares per-peer
replies. If anything differs it sends one `notifications/*/list_changed`
per changed list and updates the file; if nothing differs it sends
nothing. The fast `initialize` advertises `listChanged` for tools and
for every prompts/resources capability the merged result declares, and
only on that path; a notification is sent only for an advertised
capability. The live `initialize` (instructions/capabilities) is
compared with the one served too: MCP cannot re-send it mid-session, so
a difference is logged to stderr and persisted for the next session.
- **Fallback.** With no snapshot, an edited config, a different launch
directory, a different client `protocolVersion`, or a corrupt or
unreadable file, the old blocking behaviour runs unchanged and a
snapshot is written afterwards. A listing in which any peer timed out is
never saved.
- The lazy primer (#212/#451/#463), per-session retract (#252) and the
late-reply/timeout paths are untouched; the full existing suite passes
as-is.

## Known limitation
The snapshot is keyed per launch directory, so the first headless
session in a **new** directory still takes the blocking path. With
per-session git worktrees (`claudew`), that means most new worktrees.
Snapshot files are also never pruned. Possible follow-ups, not done
here: key by the git common dir where codegraph's index location allows
it, and add an age-based prune.

## Tests (`tests/test_router_snapshot.py`, real subprocess peers over a
pipe; `fake_mcp_server.py` gains `FAKE_INIT_DELAY` / `FAKE_TOOLS` /
`FAKE_DIE_IF`)
1. With a snapshot, `initialize` and `tools/list` reply in under 0.5s
while a peer takes 1.5s to initialize.
2. A `tools/call` to the not-yet-ready peer is delivered once that peer
is up, with no error.
3. Snapshot differs from live: exactly one `tools/list_changed`, the
file is updated, and the stale name then gets -32601. Identical: no
notification, and the file's bytes and mtime are unchanged.
4. No snapshot: blocking as today, no `listChanged` claim, and a
snapshot is written with the right fingerprint.
5. Edited peers config: the old snapshot is ignored and rewritten.
6. Garbage, wrong-shape, or unreadable (directory) snapshot: falls back
to blocking without crashing. Also a unit test covering each shape
`load()` rejects.
7. A peer that fails to start: calls to it still error with -32001,
calls to the other peer work, `list_changed` is sent, and the good
snapshot is not overwritten.

Review follow-ups (11ae025):
- Two launch directories keep separate snapshots, and a snapshot from
one directory is not served in another.
- A difference only in `initialize` is logged and persisted, with no
`list_changed`.
- A different client `protocolVersion` takes the blocking path and
re-keys the snapshot.
- Two instruction-bearing peers that swap `initialize` arrival order are
not reported as a change (603a886: the served and live views are both
merged in config order).
- `listChanged` is advertised only for declared capabilities, and a
changed resources list that was never declared is not announced (unit
test).

`uv run pytest -q`: 2724 passed · `ruff check .`: clean · `mypy`: clean.

v0.42.0

Toggle v0.42.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
feat(proxy): record the dropped block's index on retrieve ledger rows (…

…#252) (#473)

Refs #252 (step 3a: collect position data for auto-tuning keep_first /
retract_after).

## What
- `lossy.apply_text_drops` records `(tool, path, index)` per handle;
`index` = position among qualifying blocks (>= `min`), **counting blocks
kept by `keep_first`**, so a retrieve at index i means keep_first i+1
would have kept it. JSON-field drops record `None`.
- Duplicates: identical spans share one content-addressed handle; within
a result the **first position** wins (`setdefault`). Across results the
latest drop still overwrites, as before.
- Proxy `_drop_origin` widens to `(server, tool, path, index)`;
`_note_drop_origins` tolerates old 2-tuples. `answer_retrieve` passes
`index=` to the writer only when set, so 5-arg writers keep working.
- `build_retrieve_record` / `build_retrieve_writer` gain optional
`index`; key omitted when None. Aggregation ignores it; old rows read
unchanged. Still payload-free.
- Per-session retract (#471) keys on `(origin[1], origin[2])`,
unaffected; covered by a test.

## Tests
New `tests/test_retrieve_block_index.py` (11 tests, written first, 8
failed before the change). `uv run pytest -q`: 2720 passed. `ruff check
.` and `mypy` clean.

v0.41.5

Toggle v0.41.5's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(codeceval): keep the gateway model's compliance line beside a cli…

…: model (#480)

A run mixing a text-channel (`cli:`) model with a gateway model dropped
the #432 tool-call compliance line from SAFE rows: the `cli:` model's
rows carry no `*_calls` counters (compliance doesn't apply to it), so
the "every model has a full reading" check failed for both arms.

Text-channel models are now excluded from the compliance quote; a partly
counted *tool* model still withholds it, as before. Model-naming uses
the measured-model count, so a single gateway model's rate names nobody.

New test
`test_a_text_channel_model_does_not_withhold_the_gateway_models_rate`
fails on main (SAFE row with no compliance line) and passes here. Full
suite 2710 passed; ruff and mypy clean.

Refs #450

v0.41.4

Toggle v0.41.4's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(codeceval): honest empty diff reports; size cli: requests by chan…

…nel (#476)

## Summary

- **#266 (a)**: `_build_diff_style_report` now names one of three empty
states: no model configured / no same-tool pairs / N pairs that
generated no question. The old single hint is gone. New
`fluency.diff_pairs` / `fluency.text_diff_pairs` (pulled out of the
harnesses, so the pairing logic is unchanged) give the CLI the pair
count. Part (b), committing a corpus, is being closed as obsolete.
- **#450 items covered:**
- [x] `request_tokens` counts the tool-channel instruction and tool defs
for `cli:` models. Now it follows the answer channel through a shared
`_channel_instruction`, which `_codec_turn` also uses. The `claude -p`
preamble part is **not** covered: still not counted, and the docstring
says so.
- [x] `_codec_complete` fallback branch: added a direct test (the
`_codec_answered` fallback was already pinned by
`test_R7_legacy_rows_fall_back_to_trials_minus_fails`).
- [x] Stale docs: the `min_paired` comment (report.py) now names
`_CODEC_MIN_QUESTIONS`, and the `_codec_turn` docstring now notes the
`claude -p` preamble.
  - [x] Pre-flight refusal message branches on `answer_channel`.
- [x] "Two tests can't fail": **already fixed by #449.**
Mutation-checked:
`test_the_primer_counts_toward_the_terse_arms_input_limit` (drives
`run_codec_fluency`) and `test_an_UNSAFE_row_names_lost_calls_too` both
fail when their code is removed. No change needed.
- Not touched: `_TEXT_INSTRUCTION` duplication (pending decision), and
the mixed cli:/gateway SAFE-row compliance line.

## Test plan

- `uv run ruff check .`: clean. `uv run mypy`: clean.
- `uv run pytest -q`: 2718 passed.

Refs #450
Fixes #266

v0.41.3

Toggle v0.41.3's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(stats): contest a label two entries each bake as --server-name (#478

)

Fixes #426.

Two distinct entries sharing a `--server-name` label (no router) each
banked the label's full blocks/savings and reported KEEP — #285's double
count with the router removed.

**Rule (owner decision 2026-09-26 + review):** in `_contested_labels`, a
label is contested when it has more than one ledger-writing writer AND
at least one writer CHOSE it — a router peer name or an explicit
`--server-name` (`ledger_identity_explicit`). This covers
explicit+explicit and explicit+guess. Labels every writer merely guesses
(same binary) keep #285's behaviour;
`test_ambiguity_needs_a_LAUNCHER_not_merely_a_shared_label` is
unaffected. Existing `_precedence_winner` and `_writes_ledger_rows`
gates are unchanged. Reuses the existing contested machinery: both
entries go dark and the report prints the duplicate-label line and
per-label remedy.

The `router/live duplicate label` verdict reason (closed `--json`
vocabulary, not renamed) now also covers declared-name collisions with
no router; noted in the CHANGELOG.

Diff is local to `_contested_labels` to limit conflicts with #474/#477.

## Tests
- New:
`test_two_entries_baking_the_same_explicit_server_name_are_contested`
(issue repro, failed before fix),
`test_one_explicit_name_beside_a_guess_of_the_same_label_is_contested`
- `uv run pytest -q`: 2711 passed
- `uv run ruff check .`, `uv run mypy`: clean

v0.41.2

Toggle v0.41.2's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
fix(stats): rank rejected project .mcp.json entries as not launched (#…

…477)

test(stats): isolate CLAUDE_CONFIG_DIR for scanner tests; note --setting-sources gap

Refs #448

fix(stats): only a rejected project entry is absent; pending still launches

Review of #477: Claude Code loads pending project-scoped servers without
asking in claude -p, Agent SDK and cloud sessions (and under bypass mode),
so a pending entry is a live writer there. Only disabledMcpjsonServers
blocks a server in every mode. Pending rows keep their real state and only
record approval=pending; rejected rows become unapproved. User settings
honour CLAUDE_CONFIG_DIR. Documents the enable-all OR-merge and ignored
workspace trust as false-approved (safe) approximations.

Refs #448

fix(stats): rank unapproved project .mcp.json entries as not launched

A project-scope server Claude Code has not approved (pending or rejected)
sits in .mcp.json but never runs, yet the precedence rule treated it as the
running definition and dropped the user/local entry that actually runs.
scan_scopes now reads enabledMcpjsonServers / disabledMcpjsonServers /
enableAllProjectMcpServers from the ~/.claude.json project block and the
user/project settings files, marks unapproved entries as state
"unapproved" (with an approval field), and _ABSENT_FROM_SCOPE includes it.