Skip to content
Open
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
feat(orchestration): 2/6 — one TurnMonitor answers should_stop; run_l…
…imits.max_tool_calls replaces max_turns

StopReason and end_status_for join the event protocol; communicate() loses
max_turns and its should_stop returns a reason. TurnMonitor (from
EarlyStopWatcher) is attached on every run and latches the armed early stop or
the cumulative tool-call cap; every adapter drops its own cap and finalizes
with end_status_for. sdk_options.max_turns is allowed on Claude Code and used
by the simulator and the judge. SPI_VERSION 2.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
  • Loading branch information
uipreliga and claude committed Sep 16, 2026
commit dda40e4d79221eee892c14403e18bee0c1198ddc
29 changes: 11 additions & 18 deletions .claude/notes/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,17 +61,9 @@ intentionally brief and out of scope; trimming for DISPLAY belongs in the render

- **Harness run-limit parity**: a shared `BaseAgentConfig` field must mean the same
thing on every backend, so a divergence is either fixed or documented — never silent.
**`run_limits.max_turns` on Codex/Antigravity counts VISIBLE turns** (resolved tool
calls, read live off the shared `EventCollector.visible_turn_count`, the same list
`TurnRecord.commands` holds) because one `communicate()` is a single SDK turn on both,
so a native counter would clamp at 1; claude-code keeps its native SDK cap, whose unit
(an agent-loop turn) absorbs arbitrarily many parallel calls — the same number is NOT
the same budget across harnesses. OpenCode and Pi each keep a native unit too, because
their CLIs stream a real multi-step loop per `communicate()`
(`step_start`/`step_finish`, `turn_start`/`turn_end`). The cap is enforced on the same
loop boundary as the cooperative early stop and finalizes cleanly as
`tool_calls_exhausted` (no crash, no retry); on Antigravity that boundary lives in
`_drain()`, so the background-work poll loop honors it too.
**`run_limits.max_tool_calls` is the `TurnMonitor`'s cap, in resolved tool calls, on
every harness**: no adapter counts it; each stops at its next `should_stop` poll and
finalizes cleanly as `tool_calls_exhausted` (no crash, no retry).

The **known unfixed divergences** — which config fields each harness does and does not
enforce, and the per-harness `agent.plugins[].path` depth (claude-code REQUIRES a
Expand Down Expand Up @@ -113,15 +105,16 @@ bucketing to COMPLETED.

Status precedence is the same everywhere: timeout > stopped_early > tool_calls_exhausted >
completed. `stopped_early` outranks the cap because an armed criterion deciding the
outcome is the more specific reason to have cut the run, and every loop checks it first.
outcome is the more specific reason to have cut the run; the `TurnMonitor` evaluates the
armed criteria before the cap, so an armed stop wins a tie and the first latched reason is
final.

## Why a post-stop exception is not a crash

Once the loop has broken on purpose — a cooperative stop or the turn cap — an exception
raised while tearing the stream down must NOT be escalated. Escalating triggers the
orchestrator's retry with the watcher's decision still latched, so the retry stops at turn
0 having spent nothing useful; a cap-break is the same shape, where the retry burns the
budget again and re-hits the cap. `ended_cleanly` is the guard.
Once the loop has broken on purpose — any `should_stop` reason, including the tool-call
cap — an exception raised while tearing the stream down must NOT be escalated. Escalating
triggers the orchestrator's retry with the monitor's decision still latched, so the retry
stops at its first poll having spent nothing useful. `ended_cleanly` is the guard.

## Why the constructors declare every kwarg

Expand Down Expand Up @@ -264,7 +257,7 @@ matter how much the run actually billed. So the CLI harnesses crash rather than
Every arm is gated on `stopped_early` / `tool_calls_exhausted`, because an intentional cut
can land before the clearing event arrives. Pi's error case shows why: `error_message` is
set at an error `turn_end` and cleared only by a LATER non-error `turn_end`, but a
`max_turns` / `should_stop` cut can fire at the next `turn_start`, leaving a stale error
`should_stop` cut (an early stop or the tool-call cap) can fire at the next `turn_start`, leaving a stale error
from a turn Pi was still retrying. Without the guard that clean, budget-exhausted cut
would crash and burn retries, contradicting the documented "finalizes cleanly as
`tool_calls_exhausted`, no crash" contract.
Expand Down
4 changes: 2 additions & 2 deletions .claude/notes/contracts.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ answers pass or fail, it answers the same for every longer prefix); `undecided`
verdict allowed to change on a later call. Both properties are stated at the definition
site in `criteria/base.py`, because they are the contract an author has to satisfy.

`EarlyStopWatcher`'s deferred fail-stop and its pass/fail flip-attribution are correct ONLY
`TurnMonitor`'s deferred fail-stop and its pass/fail flip-attribution are correct ONLY
because the two shipped implementations honor them. A non-monotonic or non-deterministic
override compiles, passes CE025, and silently corrupts the stop logic.

Expand Down Expand Up @@ -92,7 +92,7 @@ matches, but the same haystacks feed `exclude_pattern` and the `max_count` gate,
normalized form can newly satisfy an exclusion or trip a cap — a command that counted on
the raw text alone can stop counting.

It is memoized because the early-stop watcher re-scans the whole accumulated trajectory on
It is memoized because the `TurnMonitor` re-scans the whole accumulated trajectory on
every tool-call event, normalizing the same command many times per run.

The regex search window is capped to bound ReDoS on a large command string, and
Expand Down
4 changes: 2 additions & 2 deletions .claude/notes/isolation.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,11 +29,11 @@
**`TOOL_CALLS_EXHAUSTED` is deliberately NOT one of them — anywhere**.
`_EXECUTION_FACT_STATUSES` maps it to `False`, and the table and the chain that reads
it must agree: it shipped as `True` while `_terminal_status`'s own docstring argued
the opposite, and the disagreement pinned a re-graded max-turns row at
the opposite, and the disagreement pinned a re-graded capped row at
TOOL_CALLS_EXHAUSTED *while holding `weighted_score` 1.000* and exit 1 — a combination
`run` can never produce for the same trajectory. Under `execute`: `_terminal_status`
puts the `grade=False` arm ABOVE it, because on the graded path it is subordinate to
the verdict — `run` returns SUCCESS for a max-turns trajectory whose criteria pass —
the verdict — `run` returns SUCCESS for a capped trajectory whose criteria pass —
so it is not knowable without grading. Consuming it first made it terminal AND
permanent (the `is_execution_fact` arm then pinned it), so identical agent output
scored SUCCESS/1.0 under `run` and TOOL_CALLS_EXHAUSTED under `execute` → `evaluate`.
Expand Down
35 changes: 20 additions & 15 deletions .claude/notes/orchestration.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@

- **Generic CLI overrides (`-D`/`--set`)**: Layer 5 is a thin wrapper
(`orchestration/overrides.py`) over the resolver above. `coder-eval run -D
agent.model=opus -D run_limits.max_turns=30` overrides any field on the resolved
agent.model=opus -D run_limits.max_tool_calls=30` overrides any field on the resolved
`TaskDefinition` (`agent`/`run_limits`/`sandbox` roots), schema-validated with
did-you-mean. Only `--model` (→ `agent.model`) and `--driver` (→ `sandbox.driver`)
survive as active thin aliases that emit the equivalent `-D` entry; an alias and `-D`
Expand Down Expand Up @@ -91,7 +91,7 @@ workspace reports SUCCESS — with the original `error_message` still attached.
**The NOT_GRADED arm sits ABOVE `tool_calls_exhausted`, and that order is what makes
`execute` + `evaluate` equal a single `run`.** TOOL_CALLS_EXHAUSTED reads like an execution
fact but is not one: on the graded path it is subordinate to the verdict — `run` returns
SUCCESS for a max-turns trajectory whose criteria pass, and only falls through to
SUCCESS for a capped trajectory whose criteria pass, and only falls through to
TOOL_CALLS_EXHAUSTED when they do not — so it is not knowable under `grade=False`.
Consuming it first made it terminal AND permanent, so the same agent output scored
SUCCESS/1.0 under `run` and TOOL_CALLS_EXHAUSTED under `execute` → `evaluate`; being
Expand Down Expand Up @@ -279,8 +279,8 @@ silent. A missing stamp (a run predating the feature) is tolerated.

- **Early stop on criterion (opt-in, per-criterion arming)**: a `stop_early:` block
(`StopEarlyPolicy`) on a criterion ends a single-shot run early once the run's
**armed** criteria decide the outcome, so a raised `max_turns` isn't wasted on the
smoke flavor. The block's PRESENCE is the arming and alone activates the watcher —
**armed** criteria decide the outcome, so a raised `max_tool_calls` isn't wasted on the
smoke flavor. The block's PRESENCE is the arming and alone arms the monitor —
there is **no run-level master switch**: `run_limits.stop_early: false` is the
run-level KILL SWITCH that force-disarms every block (the one-line
experiment-variant/`-D` override for an authoritative full run), and
Expand Down Expand Up @@ -313,8 +313,9 @@ silent. A missing stamp (a run predating the feature) is tolerated.
never freezes a sibling `on_pass: continue` criterion's signal out of the trajectory).
A fail-stop is therefore verdict-preserving; a pass-stop can miss a *later* distractor
misfire, so authoritative P/R/F1 comes from a kill-switched (`stop_early: false`) run.
Driven by `orchestration/early_stop.py::EarlyStopWatcher` (built when
`early_stop_active(task)`: ≥1 armed criterion, kill switch not thrown) through the
Driven by `orchestration/turn_monitor.py::TurnMonitor` (built by `_build_monitor` in
`_setup` on every run; its criteria are armed when `early_stop_active(task)`: ≥1 armed
criterion, kill switch not thrown, and grading on) through the
agent's cooperative `should_stop` seam (tool-call granularity, no SIGKILL); live
verdicts only *trigger* the stop — the standard `check_all_async` on the frozen
trajectory is authoritative. Gating is **FIRED-ONLY**: a run the watcher actually cut
Expand Down Expand Up @@ -359,11 +360,14 @@ is never assigned there, so an armed simulation task gates strict-AND on a possi
truncated trajectory. Wiring the dialog path through it means also setting `early_stop`
there; until then the limit is stated rather than implied.

The watcher is built ONCE, in `_setup`, so its turn/tool counters and wall-clock origin
accumulate across retry attempts. It is built before the evaluate-only early return, so an
armed evaluate-only re-grade builds an inert, never-fed watcher — harmless, and one
creation point. Under `execute` it is armed but stays disabled: there is no outcome to
decide and the trajectory is the deliverable, so an armed criterion must not truncate it.
The `TurnMonitor` is built ONCE, by `_build_monitor` in `_setup`, on every run, so its
tool-call counters and wall-clock origin accumulate across retry attempts and dialog turns
(that is what makes `run_limits.max_tool_calls` cumulative per task). It is built before
the evaluate-only early return, so an evaluate-only re-grade builds an inert, never-fed
monitor — harmless, and one creation point. Under `execute` its criteria are not armed
(`arm=self.grade`): there is no outcome to decide and the trajectory is the deliverable, so
an armed criterion must not truncate it. The tool-call cap still applies there, because it
is a run limit, not a verdict.

### Verdicts latch, and the decision happens on the CALL

Expand Down Expand Up @@ -462,14 +466,15 @@ distractor rows (fail live, pass and timeout inert) without per-row conditionals
why the validator carries NO per-instance polarity guards. Arming an unobservable criterion
is structurally impossible, since the block exists only on `LiveSuccessCriterion`, so a
`file_exists` criterion carrying one is an `extra='forbid'` error at load. An armed-but-
empty set needs no guard either: with no blocks present there is simply no watcher.
empty set needs no guard either: with no blocks present the monitor has nothing armed and
only its run-limit cap can stop the run.

The watcher keeps its OWN `EventCollector`, independent of the one the agent builds its
The monitor keeps its OWN `EventCollector`, independent of the one the agent builds its
returned `TurnRecord` from, so each `live_verdict` sees a fresh single-element partial
trajectory.

**Fail-open:** a `live_verdict` that raises disarms the watcher, logs loudly, and degrades
to a full run. Because live verdicts are triggers and not truth, this can never produce a
**Fail-open:** a `live_verdict` that raises disarms the armed criteria, logs loudly, and
degrades to a full run. The tool-call cap reads counters, so it keeps running. Because live verdicts are triggers and not truth, this can never produce a
FALSE early stop — it only ever errs toward running more.

### Why the guardrails are not model validators
Expand Down
12 changes: 7 additions & 5 deletions .claude/notes/reporting.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,16 +15,16 @@

- **Run-time caps (non-criterion enforcement)**: `TaskDefinition.run_limits`
(`RunLimits` model) is the single namespace for all *task-level* run-time caps —
`max_turns` / `task_timeout` / `turn_timeout` (structural) and `max_input_tokens` /
`max_tool_calls` / `task_timeout` / `turn_timeout` (structural) and `max_input_tokens` /
`max_output_tokens` / `max_total_tokens` / `max_usd` (cumulative budget). Token/USD
breaches abort with `FinalStatus.TOKEN_BUDGET_EXCEEDED` or `COST_BUDGET_EXCEEDED`
(both `category == "failed"`). Structural caps are set from the CLI via `-D
run_limits.max_turns=…` / `-D run_limits.task_timeout=…` / `-D
run_limits.max_tool_calls=…` / `-D run_limits.task_timeout=…` / `-D
run_limits.turn_timeout=…` (field-merged into `run_limits`); budget caps via `-D
run_limits.max_usd=…` etc. or YAML. Layered config uses field-merge — a variant block
overrides individual keys without replacing the task's block. The one *per-criterion*
cap, `stop_early.decide_within`, deliberately lives on `LiveSuccessCriterion` instead
(see [orchestration.md](orchestration.md) § Early stop on criterion) — the watcher
(see [orchestration.md](orchestration.md) § Early stop on criterion) — the monitor
must attribute a decision-step timeout to a specific criterion, which `RunLimits`
(task-scoped, criterion-agnostic) cannot express.

Expand Down Expand Up @@ -332,10 +332,12 @@ means the whole bill.

### The claims the reports do NOT make

An early-stopped row does not advertise "N turns avoided". That derived from
An early-stopped row does not advertise "N turns avoided". That claim once derived from
`max_turns - sdk_turn_index`, and on harnesses where one `communicate()` is a single SDK
turn it advertised dozens of avoided turns when all that was cut was a tool-call tail. The
upper bound is still persisted, labelled as the bound it is.
upper bound is still persisted as `tool_calls_remaining_at_stop`
(`max_tool_calls - tool_call_index`, null when the cap is unset), labelled as the bound it
is.

Missing spend is worded cause-agnostically, because an unpriced turn and a hard kill reach
the same conclusion and the report cannot always tell which applied.
Expand Down
2 changes: 1 addition & 1 deletion .claude/notes/timing.md
Original file line number Diff line number Diff line change
Expand Up @@ -188,7 +188,7 @@ terminal event as `AgentEndEvent(messages=list(...))` — that copies the LIST,
message objects — so writing in place would reach back into the agent's own live state
from the collector, which is exactly the layering "the collector is the sole capture seam"
exists to prevent. It is also unconditionally safe for a caller that builds a record
twice: `EarlyStopWatcher` holds one collector across a turn's tool-call rounds and calls
twice: `TurnMonitor` holds one collector across a task's tool-call rounds and calls
`build_turn_record` on every one.

Grouping by identical bounds rather than `message_id`: Codex splits one window across two
Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/verify-published-action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -344,10 +344,10 @@ jobs:
# caps are ~10x what the one-file task needs; tripping the spend or wall-clock
# one produces a COST_BUDGET_EXCEEDED / TIMEOUT row that the gate's run-limit
# branch prints and fails on (both statuses report as "failed", so nothing else
# in the gate would). MAX_TURNS_EXHAUSTED is the one deliberate exception: it is
# in the gate would). TOOL_CALLS_EXHAUSTED is the one deliberate exception: it is
# the classic model-quality outcome, tolerated like an unmet criterion.
run_limits:
max_turns: 5
max_tool_calls: 5
task_timeout: 300
max_usd: 0.25

Expand Down
Loading