Program plan: ~/.claude/plans/terse-defensibility-program.md, Phase 0b. Highest-ROI item in the repo.
Why
The entire docs/POSITIONING.md economic model exists to price a 248–555 cl100k token/session tax whose only job is teaching the model to read terse's wire form. If the model doesn't need to be taught, the tax — and most of that document — is optional.
There is already evidence it may not:
Neither of those was run as a deliberate "can we drop the primer entirely on a capable model" experiment across a frontier panel.
What to do
Add a terse-no-primer arm to the existing fluency harness (src/terse/fluency/harnesses.py already runs raw / terse / terse+primer / terse+inline) across a frontier panel — Haiku 4.5, Sonnet 5, Opus 5, Gemini 2.5 Flash — on the stress corpus, paired-scored, same discipline as the existing arms.
Decision rule (pre-register before looking)
- Accuracy holds across the panel → ship
--primer=auto|always|never (auto = never for known-capable models). Consequences are large and must be followed through: every UNWRAP/TUNE verdict in terse stats --recommend flips, the 248/312/555 arithmetic in POSITIONING.md becomes a legacy-mode footnote, and the diff tier's central objection (its 190-token paragraph, 34–43% of a router primer) largely evaporates.
- Accuracy drops on any panel model → keep the primer and attack its size instead: 155 of the 248 baseline tokens are the table paragraph. Measure a shortened primer against the same panel.
Blocks
This likely retroactively decides #0a (the diff hit-rate re-measurement), so run it first or read 0a net of whichever primer regime this lands on.
Program plan:
~/.claude/plans/terse-defensibility-program.md, Phase 0b. Highest-ROI item in the repo.Why
The entire
docs/POSITIONING.mdeconomic model exists to price a 248–555 cl100k token/session tax whose only job is teaching the model to read terse's wire form. If the model doesn't need to be taught, the tax — and most of that document — is optional.There is already evidence it may not:
tersearm (no primer anywhere) at 96% (glm-5.2) and 92% (deepseek-v4-flash) against 100% raw.Neither of those was run as a deliberate "can we drop the primer entirely on a capable model" experiment across a frontier panel.
What to do
Add a
terse-no-primer arm to the existing fluency harness (src/terse/fluency/harnesses.pyalready runs raw / terse / terse+primer / terse+inline) across a frontier panel — Haiku 4.5, Sonnet 5, Opus 5, Gemini 2.5 Flash — on the stress corpus, paired-scored, same discipline as the existing arms.Decision rule (pre-register before looking)
--primer=auto|always|never(auto= never for known-capable models). Consequences are large and must be followed through: everyUNWRAP/TUNEverdict interse stats --recommendflips, the 248/312/555 arithmetic in POSITIONING.md becomes a legacy-mode footnote, and the diff tier's central objection (its 190-token paragraph, 34–43% of a router primer) largely evaporates.Blocks
This likely retroactively decides #0a (the diff hit-rate re-measurement), so run it first or read 0a net of whichever primer regime this lands on.