Skip to content
Open
Show file tree
Hide file tree
Changes from 1 commit
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
8416804
feat(pricing): 1/10 — price_turn is the one cost rule for adapters an…
uipreliga Sep 16, 2026
644e26e
feat(streaming): 2/10 — TurnEmitter, Window, TimingBasis and coder_ev…
uipreliga Sep 17, 2026
ac69e40
feat(agents): 3/10 — communicate returns a TurnOutcome; the orchestra…
uipreliga Sep 17, 2026
5b6e12a
feat(agents): 4/10 — port Pi onto TurnEmitter; the CLI never inherits…
uipreliga Sep 17, 2026
111d36e
feat(agents): 5/10 — SubprocessJsonlAgent under Pi and OpenCode; Open…
uipreliga Sep 17, 2026
f62c170
feat(agents): 6/10 — port Antigravity onto TurnEmitter; the user's pr…
uipreliga Sep 17, 2026
68958d9
feat(agents): 7/10 — port Codex onto TurnEmitter; tag its sub-agent e…
uipreliga Sep 17, 2026
0be2ed3
feat(agents): 8/10 — port Claude Code onto TurnEmitter; tag its sub-a…
uipreliga Sep 17, 2026
726a961
feat(lint): 9/10 — CE072 keeps adapters on the emitter; retire CE059-…
uipreliga Sep 17, 2026
853a38c
test(agents): 10/10 — live verification; the Claude harness live case…
uipreliga Sep 17, 2026
eaabe82
fix: code review fixes for turn-emitter-and-ports
uipreliga Sep 17, 2026
7a42ee1
test(tasks): the timeout fixtures sleep through python3 so Claude Cod…
uipreliga Sep 17, 2026
28c057b
docs(harness): defer six candidates from the turn-emitter plan
uipreliga Sep 17, 2026
3221e30
feat(run-limits): 1/5 — run_limits.max_turns field and the model-turn…
uipreliga Sep 17, 2026
8c722bf
feat(run-limits): 2/5 — TurnMonitor enforces max_turns on main-thread…
uipreliga Sep 17, 2026
41b735f
feat(run-limits): 3/5 — expected_turns soft target on the persisted m…
uipreliga Sep 17, 2026
778fb30
docs(run-limits): 4/5 — max_turns / expected_turns docs, notes and th…
uipreliga Sep 17, 2026
baa9a31
feat(plugins): 5/5 — Claude Code loads each agent.plugins entry as a …
uipreliga Sep 17, 2026
636c5dd
test(harness): a fixture with a model-turn limit resolves where model…
uipreliga Sep 17, 2026
98d4c53
fix: code review fixes for max-turns
uipreliga Sep 17, 2026
73270a5
fix(sandbox): the sandbox owns the criterion PATH, so run_command see…
uipreliga Sep 17, 2026
618b421
fix(harness): audit fixes for the transport, emitter, monitor, Codex …
uipreliga Sep 17, 2026
2f52352
fix(harness): no crash retry after tool calls, empty-turn crash, enfo…
uipreliga Sep 17, 2026
6a86e76
chore(spi): reset SPI_VERSION to 1 before the first release
uipreliga Sep 17, 2026
a07fd67
feat(harness): count model turns on Codex and Antigravity
uipreliga Sep 17, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
feat(lint): 9/10 — CE072 keeps adapters on the emitter; retire CE059-…
…064; SPI drops the events; live tests type-checked

- Delete CE059, CE060, CE061, CE063 and CE064: they guarded the per-adapter
  turn accumulators that TurnEmitter replaced. Their ids are retired
  forever (runner note).
- CE072 EmitterSoleWriter: an adapter under src/coder_eval/agents/ may not
  construct an event, an AssistantMessage or an EventCollector, under any
  import spelling from coder_eval.streaming or coder_eval.models.
- coder_eval.spi no longer exports EventCollector, the seven event classes
  or CompositeStreamCallback; tests/test_spi.py pins the final list.
- The second pyright pass type-checks tests/*_live.py and the BYOA demo
  plugin, in make typecheck, make verify and CI; the live-test type
  errors are fixed.
- tests/test_harness_live.py: one tiny turn per installed harness through
  communicate, checked with assert_stream_balanced and the bucket sum.
- Docs: CLAUDE.md, README and docs/index promise, EXTENDING
  (SubprocessJsonlAgent example, coder_eval.testing sensors),
  HARNESS_PARITY and timing notes point at the emitter; two harness
  candidates closed.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
  • Loading branch information
uipreliga and claude committed Sep 17, 2026
commit 726a961bd837205247fe220079559996b810b231
13 changes: 8 additions & 5 deletions .claude/harness-candidates.md
Original file line number Diff line number Diff line change
Expand Up @@ -647,6 +647,8 @@ divergences, so the deferred-work record is one place. Measurements in
kwarg rule ever lands, extract `tests/lint/rules/_message_calls.py` at that
point rather than sooner.
Caught in: the CE060 / antigravity `message_id` run.
UPDATE: CE059 and CE060 are retired. CE072 bans an `AssistantMessage` call in
`agents/`, alias included, so only CE058's name list is still open.

- [ ] **Nothing pins that `message_id` is only ever a WITHIN-TURN identity.** Ids
repeat across retry attempts of one turn on every synthetic-id harness —
Expand Down Expand Up @@ -808,8 +810,8 @@ re-derive from scratch.
PRESENT at the call site (`ce059_generation_window_is_two_reads.py:68`), so
removing the kwarg makes `claims_a_window` true at the three legitimate
placeholder sites and forces a rule REWRITE rather than a retirement. Net
cost: five reducers, a regeneration of every golden, and a CE059 rework; net
benefit: SSOT alone. **Deferring it is safe because the seam assertion in
cost: five reducers, a regeneration of every golden, and a CE059 rework (CE059
is now retired, so that part of the cost is gone); net benefit: SSOT alone. **Deferring it is safe because the seam assertion in
`timing.subtract_tool_time` now checks the property at runtime** — a group's
raw total must equal the span its own bounds describe — which also covers a
third-party agent registered through the `coder_eval.plugins` SPI, where no
Expand All @@ -826,7 +828,8 @@ re-derive from scratch.
case was never about the clamp but about the head being measured against the
wrong instant. REVISIT IF: an inversion is observed on a live run after
CE064, which would mean a basis is still mixed somewhere the rule cannot see
(the plugin SPI, or a harness whose spans come from a CLI).
(the plugin SPI, or a harness whose spans come from a CLI). CE064 is now
retired: `TurnEmitter` stamps the bracket from the turn's one clock.

- [ ] **`test_codex_golden[a_agent_message_only]` is FLAKY, ~5% — measured, and
pre-existing.** Forty consecutive runs on an unmodified tree (`-n 0`): 2
Expand Down Expand Up @@ -967,6 +970,6 @@ re-derive from scratch.
- [ ] OpenCode: warn when an inherited `OPENCODE_CONFIG_CONTENT` `permission` / `instructions` value is not a dict / list and is replaced — today it is dropped silently; small, but needs a decision on warn vs. keep — caught in the harness-contract Phase 3 review.
- [ ] CE070 blind spot: an adapter that counts `ToolEndEvent`s (or tokens) under a new name to cap or stop a run itself — the rule matches identifiers only; needs a data-flow check that a counter in `agents/` feeds a break or an end status — caught in the central-enforcement plan (Phase 5).
- [ ] CE070 blind spot: an adapter that re-grows a skill scanner through `glob("*.md")`, `rglob`, or a file name built from parts — the rule matches the literal `"SKILL.md"` only; needs a filesystem-walk classifier scoped to `agents/` — caught in the central-enforcement plan (Phase 5).
- [ ] Every harness's `TurnEndEvent.tokens` must be a per-report DELTA: over a turn, their sum per bucket must not exceed `AgentEndEvent.usage` (the TurnMonitor latches budgets on the sum) — nothing checks it; the golden-stream runners return only the TurnRecord, so each of the five `run_*_scenario` helpers needs an event sink first — caught in the central-enforcement final review (Claude re-reported an interleaved message id's tokens).
- [ ] Live tests (`-m live`) are neither run nor type-checked in `make verify`, so an SPI signature change (`communicate(max_turns=)`, bool `should_stop`) leaves them broken until someone runs them with credentials — needs pyright over `tests/*_live.py` or an import-time signature smoke test — caught in the central-enforcement live verification (Phase 6).
- [x] ~~Every harness's `TurnEndEvent.tokens` must be a per-report DELTA: over a turn, their sum per bucket must not exceed `AgentEndEvent.usage` (the TurnMonitor latches budgets on the sum) — nothing checks it; the golden-stream runners return only the TurnRecord, so each of the five `run_*_scenario` helpers needs an event sink first — caught in the central-enforcement final review (Claude re-reported an interleaved message id's tokens).~~ **DONE.** Closed by the emitter plus `assert_stream_balanced`: `TurnEmitter` is the one per-turn accumulator on every harness and logs a WARNING when a bucket of the summed `TurnEndEvent.tokens` exceeds the published usage, and `coder_eval.testing.assert_stream_balanced` fails on the same condition over a replayed or live event stream.
- [x] ~~Live tests (`-m live`) are neither run nor type-checked in `make verify`, so an SPI signature change (`communicate(max_turns=)`, bool `should_stop`) leaves them broken until someone runs them with credentials — needs pyright over `tests/*_live.py` or an import-time signature smoke test — caught in the central-enforcement live verification (Phase 6).~~ **DONE.** Closed by pyright over live tests in `make verify` (a second pass whose generated config includes `tests/*_live.py` and the byoa demo fixture) plus `tests/test_harness_live.py`, which runs one tiny turn per installed harness through `communicate` and checks it with `assert_stream_balanced` and the bucket sums.

57 changes: 30 additions & 27 deletions .claude/notes/timing.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,14 +28,15 @@ each turn re-anchors. That is intended — do not "fix" it by re-reading the wal
which is the property being removed.

One per turn, never module-level and never reused: a long run would accumulate drift
between the pair and real wall time. The turn-state constructors take it as an argument
so the lifetime is visible in the signature, and so a unit test can pass a fake straight
in. An end-to-end test driving `communicate()` cannot — the state is built inside it, out
of the caller's reach — so those replace the class through the agent module instead
(`tests/_bracket_clock.py`). Both reach the same object.

Which harnesses use it — for their window bounds and, since **CE064**, for their turn
bracket — is stated in `docs/agents/HARNESS_PARITY.md` (the `clock basis for recorded
between the pair and real wall time. `Agent._open_emitter` is the one place an agent
constructs it: it gives a fresh `TurnClock` to the turn's `TurnEmitter` under
`TimingBasis.TURN_CLOCK`, and the host wall clock under `CLI_EPOCH_MS`. The emitter stamps
every event from that clock, the `AgentStartEvent` / `AgentEndEvent` bracket included, so
an adapter cannot put the bracket on a different basis from its window bounds. A
decoder-level test passes a `ScriptedClock` through `coder_eval.testing.replay`; a test
driving `communicate()` replaces the class through `coder_eval.agent.TurnClock`.

Which harnesses use it — for their window bounds and their turn bracket — is stated in `docs/agents/HARNESS_PARITY.md` (the `clock basis for recorded
stamps` and `turn bracket` rows), the designated SSOT for per-harness composition.
Asserting it anywhere else is the drift that put a wrong OpenCode row in that table for
months.
Expand Down Expand Up @@ -109,9 +110,12 @@ never push the window start past the first item and invert the span. claude-code
none — its stream carries no per-emission item start — so its window opens exactly at the
mark.

It deliberately does not return `completed`. The window always ends at `now`, which the
caller passed in, so handing it back would be an argument returned unchanged —
redundancy dressed as symmetry.
It returns a `Window`: a frozen dataclass that holds the two bounds and nothing else.
`duration_ms` is a property computed from them, `completed_at - started_at` clamped at
`0.0`, so an inverted window keeps its real bounds and reads as a measured zero. A caller
cannot set a duration apart from the bounds, because there is no field to set.
`TurnEmitter.add_generation` accepts only a `Window`, and in-tree only `close_window` returns
one, so every measured `generation_duration_ms` comes from this shape.

## decompose_turn

Expand Down Expand Up @@ -183,7 +187,7 @@ itself: when to reset a span list, when to clear a start stamp, when to advance
reducer now publishes the RAW window and keeps only the genuinely harness-shaped decision,
which is where its window opens.

Non-mutating for aliasing reasons rather than repeated calls. Every agent builds its
Non-mutating for aliasing reasons rather than repeated calls. `TurnEmitter` builds the
terminal event as `AgentEndEvent(messages=list(...))` — that copies the LIST, not the
message objects — so writing in place would reach back into the agent's own live state
from the collector, which is exactly the layering "the collector is the sole capture seam"
Expand All @@ -202,29 +206,28 @@ keying on the id would silently collapse every id-less message of a turn into on

That equality is what lets `generation_duration_ms` stay a PUBLISHED field rather than one
the collector derives from the bounds. Deriving it instead was considered and cut — it
would cost five reducers, a regeneration of every golden and a rewrite of CE059, whose
exemption keys on the kwarg being present at the call site — and the assertion is the
would cost five reducers and a regeneration of every golden — and the assertion is the
sensor that makes deferring that safe. A mismatch means a reducer narrowed or widened a
window without moving its bounds, which is the drift
`tests/_fixtures/golden_streams/_scrub.py::assert_timing_captured`'s "bounds that span it"
check catches one replay at a time.

It OVERLAPS with CE061 and is kept anyway. All five reducers build the window with
`close_window(mark=…, now=…)` and write `started_at=started, completed_at=now`, and CE061
— now exemption-free — forces that shape statically, so the equality is largely true by
construction. What the runtime check adds is the half an import-level check cannot see: a
reducer that bypasses `close_window`, and a third-party agent registered through the
`coder_eval.plugins` SPI, which lives outside `src/coder_eval/agents/` where no lint rule
reaches it. It is not load-bearing on its own.
It OVERLAPS with `TurnEmitter` and is kept anyway. All five reducers pass
`TurnEmitter.add_generation` a `Window` from `close_window(mark=…, now=…)`, and the emitter
writes the bounds and the duration from that one object, so the equality is largely true by
construction. What the runtime check adds is the half the emitter cannot see: a
third-party agent registered through the `coder_eval.plugins` SPI that builds an
`AssistantMessage` itself, outside `src/coder_eval/agents/` where CE072 does not reach. It
is not load-bearing on its own.

Raising kills the turn, and that is accepted — the same trade `_require_same_awareness`
makes at this seam. The condition is unreachable without a reducer bug; all five are
exercised by the golden corpus and by the ms-exact identity contract.

### Why the zero-total skip runs before the equality check

`close_window` clamps an inverted window — `now` before `mark`, two clocks disagreeing —
to `0.0` while the bounds it writes still say `completed_at < started_at`, so the bounds
`Window.duration_ms` clamps an inverted window — `now` before `mark`, two clocks disagreeing —
to `0.0` while its bounds still say `completed_at < started_at`, so the bounds
span is NEGATIVE and the equality fails. That is a measured inversion, the case
`decompose_turn` deliberately clamps because both ends were observed; raising on it would
kill turns on exactly the shape the clamp exists to tolerate. The cost is that a `0.0`
Expand All @@ -247,13 +250,13 @@ carries no timestamps at all, so indexing the raw list would measure the wrong t
raise.

A message whose `generation_duration_ms` is `None` is skipped. That field is the
codebase's own marker for "no window was measurable here", and every producer of one
stamps `started_at == completed_at == datetime.now()` at *append* time as an admitted
placeholder — Codex's rollout rebuild (`_messages_from_items`), both Codex sub-agent
codebase's own marker for "no window was measurable here", and its only in-tree writer is
`TurnEmitter.add_unmeasured_generation`, which stamps `started_at == completed_at ==
clock.now()` at *append* time as an admitted placeholder. Its callers are Codex's rollout rebuild (`_messages_from_items`), both Codex sub-agent
recovery builders, and Claude's synthesized sub-agent terminal (`_subagent_terminal_part`). Reading those
stamps as window bounds turns a placeholder into a measurement: a Codex turn rebuilt from
its rollout stamps every message at turn END, which would book the entire turn as harness
startup. It is the same exemption CE059 makes for the same reason.
startup.

`min` / `max` rather than the first and last list entries, because the list is not ordered
by time — Codex appends recovered sub-agent messages after the parent's last flush.
Expand Down
6 changes: 3 additions & 3 deletions .github/workflows/pr-checks.yml
Original file line number Diff line number Diff line change
Expand Up @@ -101,11 +101,11 @@ jobs:
- name: Type check with pyright
run: .venv/bin/pyright

# The CE036 contract engine lives under tests/, which [tool.pyright] excludes
# The CE036 contract engine and the live tests live under tests/, which [tool.pyright] excludes
# -- and `exclude` beats both a CLI file arg and an `include` entry, so it can
# only be reached through a config of its own, derived from [tool.pyright] so
# the two passes cannot drift. Mirrors `make typecheck`.
- name: Type check the CE036 contract engine
- name: Type check the CE036 contract engine and the live tests
run: |
.venv/bin/python -m tests.lint.pyright_config .pyright-tests.json
.venv/bin/pyright -p .pyright-tests.json
Expand Down Expand Up @@ -405,7 +405,7 @@ jobs:
- name: Type check with pyright
run: .venv/Scripts/pyright

- name: Type check the CE036 contract engine
- name: Type check the CE036 contract engine and the live tests
run: |
.venv/Scripts/python -m tests.lint.pyright_config .pyright-tests.json
.venv/Scripts/pyright -p .pyright-tests.json
Expand Down
29 changes: 22 additions & 7 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,8 @@ data-driven analysis.
import from `coder_eval.models`, never from its submodules.
- **`criteria/`** auto-discovers one checker per type via `pkgutil`.
- **`cli/`** holds Typer commands; each has a plain-Python twin (CE048).
- **`timing.py`** owns the single subtraction seam (CE063).
- **`timing.py`** owns the single subtraction seam and `Window`, the bounds of one
generation window.
- **`argv_match.py`** is a STDLIB-ONLY sidecar copied beside the recorder (CE057).
- **`fs_permissions.py`** is `set_permissions`, the stacked chmod window.
- **`path_utils.py`** owns run ids, atomic writes and tree digests — and every run-record
Expand All @@ -50,7 +51,11 @@ data-driven analysis.
- **`durations.py`** is `format_ms`, split from `formatting.py` so the reports layer
does not reach through an SDK-shaped module for it.
- **`isolation/`** is `driver: docker`, one container per task.
- **`streaming/`** is the event protocol and `EventCollector`.
- **`streaming/`** is the event protocol and `EventCollector`; **`streaming/emitter.py`**
is `TurnEmitter`, the per-turn kernel every agent writes its turn through.
- **`testing.py`** is `coder_eval.testing`, the adapter test sensors (`replay`,
`assert_identity_closes`, `assert_stream_balanced`, `conformance`). In-tree suites and
plugins use the same module.

Outside the package: `tasks/`, `experiments/`, `templates/`, `tests/`, `docs/`,
`evalboard/`, `plugins/coder-eval/` (the published plugin), `action.yml` (the published
Expand All @@ -65,8 +70,8 @@ Each entry is a pointer. Full rationale: `.claude/notes/` (index: `.claude/notes
- **Strategy pattern**: `Agent` ABC, implementations in `agents/`.
- **Separation of concerns**: `models/` is pure Pydantic; logic lives in `criteria/`,
`evaluation/`, `orchestration/`.
- **Callback streaming**: the agent is the sole emitter of the event protocol;
`EventCollector` reduces the stream into a `TurnRecord`. Never hand-assemble one.
- **Callback streaming**: `TurnEmitter` is the sole writer of the event protocol (CE072);
its `EventCollector` reduces the stream into a `TurnRecord`. Never hand-assemble one.
- **All core models import from `coder_eval.models`** — never from submodules.
- **Single declarative merge resolver**: all five config layers (default → experiment
defaults → task → variant → CLI) merge through `orchestration/config_merge.py`.
Expand All @@ -80,8 +85,8 @@ Each entry is a pointer. Full rationale: `.claude/notes/` (index: `.claude/notes
`aggregate()`; classification criteria layer accuracy / P/R/F1.
- **Reconciliation message**: summing token buckets across `TurnRecord.messages`
equals `token_usage` exactly, on every backend. `EventCollector` is the single writer.
- **Timing has one subtraction seam** (`timing.py`); agents must not do their own
(CE063). An unmeasured duration is `None`, never `0.0` (CE058).
- **Timing has one subtraction seam** (`timing.py`); agents must not do their own.
`TurnEmitter` stamps the turn from one clock. An unmeasured duration is `None`, never `0.0` (CE058).
- **Reference solutions are directory-only** and chmod-shielded during `communicate`.
Defense-in-depth, not a boundary — the known gaps are documented in the notes.
Authoring reference: [Reference Solutions](docs/TASK_DEFINITION_GUIDE.md#reference-solutions).
Expand Down Expand Up @@ -228,6 +233,13 @@ A few rules constrain routine edits, so they are worth knowing before you start:
- **CE070** keeps agent adapters from counting caps (`max_tool_calls`, `RunLimits`,
`tool_calls_exhausted`, …) or scanning for `SKILL.md`: the `TurnMonitor` owns caps and
`orchestration/plugin_staging.py` owns skill discovery.
- **CE071** keeps `calculate_cost` out of `agents/` and `orchestration/turn_monitor.py`.
Call `pricing.price_turn`: one rule for the cost of a turn.
- **CE072** keeps agent adapters from constructing the events (`AgentStartEvent`,
`ToolEndEvent`, …), an `AssistantMessage` (an import alias too) or an `EventCollector`.
Call `TurnEmitter` instead.
- **CE073** requires every asyncio subprocess spawn in `src/` to pass `stdin=`. An
inherited stdin stopped a CLI turn with no events until its timeout.

**Docs index SSOT.** `nav:` plus `extra.docs_index` in `mkdocs.yml` are the single
source of truth for `README.md`'s Documentation table, `docs/index.md`'s "Where to go
Expand Down Expand Up @@ -266,7 +278,10 @@ A live criterion also needs `ContractCase`s (CE036) and `make plugin-reference`.
**A new agent**: agents register through the plugin SPI (entry-point group
`coder_eval.plugins`) — there is no closed enum or dispatch to edit, and in-tree and
third-party agents take the same path. It declares a `HarnessContract` (registration
fails without one) and imports from `coder_eval.spi`. A new agent must be named on every
fails without one) and imports from `coder_eval.spi`. It writes each turn through one
`TurnEmitter` from `Agent._open_emitter` and returns its `TurnOutcome`; a JSONL CLI
subclasses `SubprocessJsonlAgent`. Test it with the `coder_eval.testing` sensors. A new
agent must be named on every
onboarding surface CE047 tracks, and its run-limit behaviour recorded in
[Run-Limit Parity](docs/agents/HARNESS_PARITY.md).

Expand Down
Loading