Repository navigation
Tags: inth3shadows/terse
Tags
fix(proxy): treat a result over 50,000 chars as offloaded, whatever i… …ts token count (#496) (#497) Fixes #496. ## Why Since 0.44.1 (#495), terse holds the lazy primer back from a result Claude Code will offload to a file. But it only treated a result as offloaded above 25,000 cl100k tokens. Claude Code also offloads by length: | | chars | Claude Code | |---|---|---| | largest MCP result shown inline | 49,034 | 2.1.284 | | smallest offloaded ("Output too large (50KB)") | 51,246 | 2.1.287 | So a result of about 13k–25k tokens was offloaded with the primer still attached. The model sees an offloaded result only as a short preview of its start, and the primer, at block 0, filled that preview. In the pre-registered Opus 5.5 check (2026-10-02), a 57 KB `list_principles` result was offloaded in both arms. Without terse, the preview held the answer row and Opus answered in 2 turns. Through terse, Opus grepped the file 2–3 more times, so that task cost 88% more per success. That one task turned the Opus result from about −2.8% into +4.2%, and the pre-registered check failed. ## What `OFFLOAD_DEFAULT_CHARS = 50_000`, checked first in `over_limit`. A result whose text is longer than that counts as offloaded, whatever its token count. Both the primer guard and the ledger's `offload` tag use `over_limit`, so both are fixed, and no tokenizer pass ever sees more than 50,000 chars. The token limit (`MAX_MCP_OUTPUT_TOKENS`, default 25,000) is unchanged. The CHANGELOG's v0.44.1 section is graduated and an `[Unreleased]` entry added. ## Tests - Test-first: 3 new tests failed on the old code before the fix: - a 60,852-char / ~17.8k-token result counts as offloaded, and a 46,770-char one does not; - the primer skips a 56,750-char compressed result at the default limits, then attaches to a smaller one; - the ledger tags that result `offload: "out"`. - The tests clear any inherited `MAX_MCP_OUTPUT_TOKENS`, since the suite also runs inside Claude Code. - The bounded-tokenizer test now also checks that a 2M-char text is never tokenized. - Mutation check: with the char cutoff removed, 4 tests fail. - Full suite: 2885 passed; ruff and mypy clean. ## Not covered - Per-tool `_meta["anthropic/maxResultSizeChars"]`: no live server sets it. - Whether `MAX_MCP_OUTPUT_TOKENS` also moves the char cutoff: unmeasured.
fix(proxy): withhold the primer from offloaded results and stop count… …ing them as context savings (#495) ## Why Claude Code saves an MCP result over its limit to a file and gives the model only the path. The limit is 25,000 tokens by default, set with `MAX_MCP_OUTPUT_TOKENS` (docs: https://code.claude.com/docs/en/mcp, "MCP output limits and warnings"). Transcripts put the real cutoff between 39,269 chars (inline) and 104,414 chars (offloaded). This caused two problems: - **Primer:** the lazy primer attaches to a session's first compressed result. If that result was over the limit, the primer went into the file, the model might never read it, and every later result went out compressed and unexplained. - **Ledger:** on the live ledger, about 20% of the reported saving was in these results. 48 results were over the limit even after terse (382,799 tokens that never reached context). Another 17, all `secret.list_credentials`, were over the limit raw but brought under it by terse (298,248 tokens). ## What - `offload_limit(client_name)`: only `claude-code`, which is the only client documented to offload. It reads `MAX_MCP_OUTPUT_TOKENS`, else uses 25,000. `over_limit(text, limit)`: - skips tokenizing a text no longer than `limit` bytes; - tokenizes at most `limit * 8` characters; - is fail-open. - **Primer guard:** at both attach sites (the text block and the #463 typed wrapper), the primer is not attached to a result that would be over the limit with the primer counted in. The check runs before `_claim_primer`, so a router's shared latch stays armed for the next result. - **Ledger:** `_emit_stats` runs after the lock is released and classifies the whole result over what the client reads: the typed field if there is one, otherwise the text blocks joined. Every row of an affected result is tagged `offload: "out"` (terse's output was offloaded) or `"raw"` (only the raw result would have been). The tag is passed to the writer only when set, so older writers keep working. - **Context basis (#420):** `"out"` rows count as (0, 0). `"raw"` rows count as (out, out): terse's inline copy is paid, and the raw path's short file pointer is not treated as the baseline. `terse stats --json` keeps its shape, and older rows read as before. - The CHANGELOG v0.44.0 section is graduated, and an `[Unreleased]` entry is added. ## Review A domain-briefed reviewer found: - **High:** tiktoken raises on `<|endoftext|>` in input, and the primer guard ran outside any handler, so that would have killed the proxy's reader thread. `over_limit` is now fail-open. - **Moderate:** the tokenizer pass had no bound. A 2.4 MB payload took 558 ms, sometimes under `_local_lock`, and a server whose results are all huge would pay that on every response while the primer was pending. It is now capped at `limit * 8` characters. - **Smaller:** the typed-wrapper check serialized before reading the limit, and the stats label now names offloaded results. Cleared: guard ordering and the shared latch; reverting to the hold at the typed-wrapper site; the classification edge cases; `context_tokens` never going negative; the client name reaching each peer on both the cold and warm initialize paths. ## Tests - 20 new tests in `tests/test_offload.py`: - the limit and its env override; - token rather than byte counting; - the primer skipping an over-limit result and attaching to the next, including when only the primer pushes it over; - a client that doesn't offload is unaffected; - a router's shared latch is not spent; - the typed wrapper; - ledger tags `out`, `raw` and none; - old seven-argument writers; - context basis; - `<|endoftext|>` input; - the bounded tokenizer pass. - Mutation checks: each test fails with its guard removed (text guard, typed-wrapper guard, out/raw classification, the context-basis zero, the fail-open catch). - Full suite: 2882 passed; ruff and mypy clean. ## Not covered - Per-tool `_meta["anthropic/maxResultSizeChars"]`: no live server sets it. - Rewriting old rows: the ledger doesn't record which client wrote them. - cl100k is not Claude Code's own tokenizer, so counts near the limit are approximate.
feat(multiproxy): declared snapshot_markers so new worktrees start wa… …rm (#479) (#494) ## Why #479's known limit: the router's saved replies (#270) were keyed by launch directory whenever a peer inherits the router's cwd. The first session in every new directory, which is most per-session worktrees, blocked on the slowest peer's `initialize` (4.2s measured, plus a `tools_changed` cache miss for headless runs). ## What - **`snapshot_markers` (new, optional, per peer in the peers file).** A list of relative paths declaring that the peer's replies depend on the directory only through the nearest match at or above it, resolved through symlinks. `[]` means not at all. When every cwd-inheriting peer declares markers, the snapshot scope is those resolved matches. Otherwise it is the cwd, exactly as before: an undeclared config keys and fingerprints byte-identically to `main`, so existing snapshots stay valid. A malformed value, or one on a `url` peer or a peer with its own `cwd`, is a config error. - **Prune:** saving a snapshot deletes sibling snapshots that have been neither saved nor served for 30 days. `load` touches the file it serves, so an in-use snapshot whose replies never change is not pruned. - **Warm/cold on the ledger:** each `router_session` row carries `snapshot: "warm"|"cold"`. It is ledger-only; `terse stats --json` is a pinned contract, and counting older rows as not-warm would read unknown as cold. - **`install-mcp`:** carries a peer's markers over when it re-folds that peer, but only onto a `command` peer without its own `cwd`. - USAGE and CHANGELOG `[Unreleased]` document the field. ## Rejected along the way (in this PR's history) Keying by `git rev-parse --git-common-dir` (c540714). Review and a live check showed that 9 repos mix worktrees with and without a `.codegraph` index (Claude Code's agent worktrees have none; terse has 14 of 19 without one), and the container dir and `.bare` resolved to the same key. One key per repo served the wrong replies on every switch. A hardcoded ".codegraph walking up" bit was also rejected: it puts one peer's behaviour into a generic router. ## Review Three review passes: a broad correctness pass and a domain-briefed pass on c540714+0c5bb53, then a domain pass on the redesign. Found: the repo-keying premise (both reviewers), prune deleting an in-use snapshot (both), a `url` peer keeping markers on re-fold, which would make the router refuse the fleet (fixed in 58b161f), and test nits. Cleared: warm/cold attribution; prune races; the undeclared-config fingerprint matching `main` (computed on both versions); the scope agreeing with `codegraph-gate` for `$HOME`, tableless indexes, a `.codegraph` file, broken symlinks and nested indexes. ## Tests - New: shared marker → shared snapshot and warm start (symlinked worktree + subdirectory); a different or missing marker stays separate; no marker anywhere is shared; one undeclared peer keeps per-directory; undeclared fingerprint equals the pre-field formula; malformed markers and markers alongside `cwd` rejected; prune keeps its own and young files; a served stale snapshot survives prune; warm/cold rows; install-mcp carry-over, including `url` and `cwd`. - Mutation checks: each behaviour's test fails under the mutated code (markers ignored, no symlink resolve, no touch, no carry-over, url carry-over, unstable fingerprint, no-op prune, always-cold). - Full suite: 2862 passed; ruff and mypy clean. ## Live rollout (separate, Protected Change) Takes effect only after the peers file declares markers: codegraph `[".codegraph"]`, runecho `[]`, kb `[]`.
feat(proxy): --primer always|never|auto (#325 Phase 1) (#488) Refs #325 (Phase 1 of 2: the switch. Phase 2, the cost-per-task harness arm comparing always / never / auto, is not in this PR.) ## What - `terse proxy --primer {always,never,auto}`, standalone and `--config` router. Default `always` is today's lazy attach, so no existing install changes. - `never`: no primer, ever (also skips the eager `initialize` path). Under a router the `structuredContent` hold is lifted, since no primer is coming. - `auto`: attaches lazily, but only on the first result whose compressed form carries `subcols`, `absent_cols`/`sentinel_cols`, an object- or array-valued alias, or a diff, text diff, embedded JSON or dropped marker. Flat tables and scalar-only aliases go out without it. Once attached, the session stays latched as today. Router: one mode on the shared `PrimerLatch`. The #463 typed-wrap path releases an auto-declined result compressed rather than reverting it to the raw hold. - Ledger: `never` and `auto` declines are written once per session as `attached: false` primer rows with a new `reason` (`never` / `auto` / `structured`; older suppression rows default to `structured`). `aggregate` keys on the reason, and `--json` `primers[]` rows now carry `reason`. - `primer_liability` reads the baked `--primer` (via `parse_proxy_opts`, terse's side of `--` only). `never` gives a zero with the new `primer_source: "configured"` unless an attach was recorded. `auto` uses the existing recorded, measured-zero or estimate path, and when it falls back to the estimate the report calls it an upper bound. Each liability server row gets `primer_mode`. ## The shape signal: deviation from the plan The plan named alias density and table width. Measured over #249's corpora (codec only, no model), neither separates stress from real, because the real payloads score higher on both (gh_pulls: 36 cols, density 0.42; stress.heavy_alias: 4 cols, 0.33). Instead, `auto` uses an allowlist taken from #249's per-question evidence: every regression was `deref` on a `subcols` row or on an object alias with absent keys, and flat or wide or heavily aliased tables were read at raw parity. The constants are marked provisional in `proxy.py`. What this means in practice: on the real corpus, `auto` still attaches on 6 of 9 payloads. It declines only on flat tables with scalar aliases (gh_labels, gh_commits_flat; gh_repo_single is vacuous). On stress it declines wide_table, long_table, heavy_alias and mixed_realistic, and attaches on nested_records, object_alias and the synthetic shape. Phase 2 should measure whether that is worth anything. ## Not done - `install-mcp` does not bake `--primer`, and a re-wrap drops a hand-added one. - POSITIONING.md economics rewrite is deferred until Phase 2 has numbers, per the plan. ## Verification `uv run pytest -q`: 2804 passed (33 new in `tests/test_primer_mode.py`). `ruff check .` and `mypy` are clean. No model or benchmark runs.
fix(multiproxy): answer initialize from a persisted snapshot (#479) Fixes #270. ## Problem The router broadcast `initialize` and replied only when the slowest peer answered: 4.2s end to end on the live fleet (kb 4.3s, runecho 3.8s, codegraph 0.3s; [per-peer timings](#270 (comment))). Headless clients (`claude --print`, SDK, CI) send their first request before then, so terse's tools arrive on turn 2 and the API reports `cache_miss_reason: tools_changed` ([headless evidence](#270 (comment)); 43% of one session's cache writes were never read). ## Design - **Snapshot.** After a complete session the router saves each peer's **raw** `initialize` and list replies (`tools/list`, plus `prompts/list` / `resources/list` / `resources/templates/list` if the client asked for them) to `$XDG_STATE_HOME/terse/router-snapshots/<sha256(resolved config path [+ launch dir])[:16]>.json`. The file is written atomically and is 0600. It carries `peers_fingerprint`, a sha256 over the parsed `downstreams[]` specs plus the router's cwd whenever any peer inherits it (codegraph serves zero tools outside an indexed repo, so a snapshot must never cross directories). It also stores the client `protocolVersion` the replies were negotiated for; only the digest is stored, because headers and env can hold credentials. - **Fast path.** If the fingerprint matches and the client requests the same `protocolVersion`, `initialize` and the list methods are answered from the snapshot straight away. The stored replies go through the same live merge code, so naming, `terse.retrieve`, primer mode, `serverInfo.version` and the routing table all come from the running process. The peers start in the background. - **Calls to a peer that isn't ready.** They wait in that peer's existing FIFO sender queue, behind its own `initialize`, with the same routed-call timeout as before. There is no new buffer. - **Reconcile.** Once every peer has answered `initialize` **and** the client has sent `notifications/initialized`, the router re-lists as its own background broadcasts. It installs the live tables (a reserved seq means live listings always win the seq guard) and compares per-peer replies. If anything differs it sends one `notifications/*/list_changed` per changed list and updates the file; if nothing differs it sends nothing. The fast `initialize` advertises `listChanged` for tools and for every prompts/resources capability the merged result declares, and only on that path; a notification is sent only for an advertised capability. The live `initialize` (instructions/capabilities) is compared with the one served too: MCP cannot re-send it mid-session, so a difference is logged to stderr and persisted for the next session. - **Fallback.** With no snapshot, an edited config, a different launch directory, a different client `protocolVersion`, or a corrupt or unreadable file, the old blocking behaviour runs unchanged and a snapshot is written afterwards. A listing in which any peer timed out is never saved. - The lazy primer (#212/#451/#463), per-session retract (#252) and the late-reply/timeout paths are untouched; the full existing suite passes as-is. ## Known limitation The snapshot is keyed per launch directory, so the first headless session in a **new** directory still takes the blocking path. With per-session git worktrees (`claudew`), that means most new worktrees. Snapshot files are also never pruned. Possible follow-ups, not done here: key by the git common dir where codegraph's index location allows it, and add an age-based prune. ## Tests (`tests/test_router_snapshot.py`, real subprocess peers over a pipe; `fake_mcp_server.py` gains `FAKE_INIT_DELAY` / `FAKE_TOOLS` / `FAKE_DIE_IF`) 1. With a snapshot, `initialize` and `tools/list` reply in under 0.5s while a peer takes 1.5s to initialize. 2. A `tools/call` to the not-yet-ready peer is delivered once that peer is up, with no error. 3. Snapshot differs from live: exactly one `tools/list_changed`, the file is updated, and the stale name then gets -32601. Identical: no notification, and the file's bytes and mtime are unchanged. 4. No snapshot: blocking as today, no `listChanged` claim, and a snapshot is written with the right fingerprint. 5. Edited peers config: the old snapshot is ignored and rewritten. 6. Garbage, wrong-shape, or unreadable (directory) snapshot: falls back to blocking without crashing. Also a unit test covering each shape `load()` rejects. 7. A peer that fails to start: calls to it still error with -32001, calls to the other peer work, `list_changed` is sent, and the good snapshot is not overwritten. Review follow-ups (11ae025): - Two launch directories keep separate snapshots, and a snapshot from one directory is not served in another. - A difference only in `initialize` is logged and persisted, with no `list_changed`. - A different client `protocolVersion` takes the blocking path and re-keys the snapshot. - Two instruction-bearing peers that swap `initialize` arrival order are not reported as a change (603a886: the served and live views are both merged in config order). - `listChanged` is advertised only for declared capabilities, and a changed resources list that was never declared is not announced (unit test). `uv run pytest -q`: 2724 passed · `ruff check .`: clean · `mypy`: clean.
feat(proxy): record the dropped block's index on retrieve ledger rows (… …#252) (#473) Refs #252 (step 3a: collect position data for auto-tuning keep_first / retract_after). ## What - `lossy.apply_text_drops` records `(tool, path, index)` per handle; `index` = position among qualifying blocks (>= `min`), **counting blocks kept by `keep_first`**, so a retrieve at index i means keep_first i+1 would have kept it. JSON-field drops record `None`. - Duplicates: identical spans share one content-addressed handle; within a result the **first position** wins (`setdefault`). Across results the latest drop still overwrites, as before. - Proxy `_drop_origin` widens to `(server, tool, path, index)`; `_note_drop_origins` tolerates old 2-tuples. `answer_retrieve` passes `index=` to the writer only when set, so 5-arg writers keep working. - `build_retrieve_record` / `build_retrieve_writer` gain optional `index`; key omitted when None. Aggregation ignores it; old rows read unchanged. Still payload-free. - Per-session retract (#471) keys on `(origin[1], origin[2])`, unaffected; covered by a test. ## Tests New `tests/test_retrieve_block_index.py` (11 tests, written first, 8 failed before the change). `uv run pytest -q`: 2720 passed. `ruff check .` and `mypy` clean.
fix(codeceval): keep the gateway model's compliance line beside a cli… …: model (#480) A run mixing a text-channel (`cli:`) model with a gateway model dropped the #432 tool-call compliance line from SAFE rows: the `cli:` model's rows carry no `*_calls` counters (compliance doesn't apply to it), so the "every model has a full reading" check failed for both arms. Text-channel models are now excluded from the compliance quote; a partly counted *tool* model still withholds it, as before. Model-naming uses the measured-model count, so a single gateway model's rate names nobody. New test `test_a_text_channel_model_does_not_withhold_the_gateway_models_rate` fails on main (SAFE row with no compliance line) and passes here. Full suite 2710 passed; ruff and mypy clean. Refs #450
fix(codeceval): honest empty diff reports; size cli: requests by chan… …nel (#476) ## Summary - **#266 (a)**: `_build_diff_style_report` now names one of three empty states: no model configured / no same-tool pairs / N pairs that generated no question. The old single hint is gone. New `fluency.diff_pairs` / `fluency.text_diff_pairs` (pulled out of the harnesses, so the pairing logic is unchanged) give the CLI the pair count. Part (b), committing a corpus, is being closed as obsolete. - **#450 items covered:** - [x] `request_tokens` counts the tool-channel instruction and tool defs for `cli:` models. Now it follows the answer channel through a shared `_channel_instruction`, which `_codec_turn` also uses. The `claude -p` preamble part is **not** covered: still not counted, and the docstring says so. - [x] `_codec_complete` fallback branch: added a direct test (the `_codec_answered` fallback was already pinned by `test_R7_legacy_rows_fall_back_to_trials_minus_fails`). - [x] Stale docs: the `min_paired` comment (report.py) now names `_CODEC_MIN_QUESTIONS`, and the `_codec_turn` docstring now notes the `claude -p` preamble. - [x] Pre-flight refusal message branches on `answer_channel`. - [x] "Two tests can't fail": **already fixed by #449.** Mutation-checked: `test_the_primer_counts_toward_the_terse_arms_input_limit` (drives `run_codec_fluency`) and `test_an_UNSAFE_row_names_lost_calls_too` both fail when their code is removed. No change needed. - Not touched: `_TEXT_INSTRUCTION` duplication (pending decision), and the mixed cli:/gateway SAFE-row compliance line. ## Test plan - `uv run ruff check .`: clean. `uv run mypy`: clean. - `uv run pytest -q`: 2718 passed. Refs #450 Fixes #266
fix(stats): contest a label two entries each bake as --server-name (#478 ) Fixes #426. Two distinct entries sharing a `--server-name` label (no router) each banked the label's full blocks/savings and reported KEEP — #285's double count with the router removed. **Rule (owner decision 2026-09-26 + review):** in `_contested_labels`, a label is contested when it has more than one ledger-writing writer AND at least one writer CHOSE it — a router peer name or an explicit `--server-name` (`ledger_identity_explicit`). This covers explicit+explicit and explicit+guess. Labels every writer merely guesses (same binary) keep #285's behaviour; `test_ambiguity_needs_a_LAUNCHER_not_merely_a_shared_label` is unaffected. Existing `_precedence_winner` and `_writes_ledger_rows` gates are unchanged. Reuses the existing contested machinery: both entries go dark and the report prints the duplicate-label line and per-label remedy. The `router/live duplicate label` verdict reason (closed `--json` vocabulary, not renamed) now also covers declared-name collisions with no router; noted in the CHANGELOG. Diff is local to `_contested_labels` to limit conflicts with #474/#477. ## Tests - New: `test_two_entries_baking_the_same_explicit_server_name_are_contested` (issue repro, failed before fix), `test_one_explicit_name_beside_a_guess_of_the_same_label_is_contested` - `uv run pytest -q`: 2711 passed - `uv run ruff check .`, `uv run mypy`: clean
fix(stats): rank rejected project .mcp.json entries as not launched (#… …477) test(stats): isolate CLAUDE_CONFIG_DIR for scanner tests; note --setting-sources gap Refs #448 fix(stats): only a rejected project entry is absent; pending still launches Review of #477: Claude Code loads pending project-scoped servers without asking in claude -p, Agent SDK and cloud sessions (and under bypass mode), so a pending entry is a live writer there. Only disabledMcpjsonServers blocks a server in every mode. Pending rows keep their real state and only record approval=pending; rejected rows become unapproved. User settings honour CLAUDE_CONFIG_DIR. Documents the enable-all OR-merge and ignored workspace trust as false-approved (safe) approximations. Refs #448 fix(stats): rank unapproved project .mcp.json entries as not launched A project-scope server Claude Code has not approved (pending or rejected) sits in .mcp.json but never runs, yet the precedence rule treated it as the running definition and dropped the user/local entry that actually runs. scan_scopes now reads enabledMcpjsonServers / disabledMcpjsonServers / enableAllProjectMcpServers from the ~/.claude.json project block and the user/project settings files, marks unapproved entries as state "unapproved" (with an approval field), and _ABSENT_FROM_SCOPE includes it.
PreviousNext