Skip to content
Open
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
feat(lint): 5/6 — generate the run-limits table from RunLimits and th…
…e contract; CE070 keeps caps and skill scans out of adapters

HARNESS_PARITY's hand-written run-limit table becomes a third CE069-checked
block; a RunLimits field without a cell rule fails the render. CE070 flags any
cap identifier or SKILL.md literal under agents/. CLAUDE.md and the notes
describe the TurnMonitor and plugin staging.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
  • Loading branch information
uipreliga and claude committed Sep 16, 2026
commit 4c7f934521698267fb76cc706d2adca6d3e86f44
2 changes: 2 additions & 0 deletions .claude/harness-candidates.md
Original file line number Diff line number Diff line change
Expand Up @@ -965,3 +965,5 @@ re-derive from scratch.
- [ ] A `TaskDefinition` serialized for a later reload (docker `_stage_inputs`, Harbor `environment/task.yaml`) must dump `agent` with `exclude_unset=True`, or the reload marks model defaults as set and the harness contract check rejects the task — two round-trip tests guard today's two sites, but nothing flags a third `task.model_dump(` written for reload; needs a call-site classifier, not a name match — caught in the harness-contract final review.
- [ ] A real-SDK Antigravity policy test: run `policy.enforce(agent._policies(real_policy))` to prove deny-beats-allow and `finish` approval against the installed SDK instead of a SimpleNamespace fake — nothing exercises the SDK's own bucket precedence; needs study of the hook-evaluation API — caught in the harness-contract Phase 3 review.
- [ ] OpenCode: warn when an inherited `OPENCODE_CONFIG_CONTENT` `permission` / `instructions` value is not a dict / list and is replaced — today it is dropped silently; small, but needs a decision on warn vs. keep — caught in the harness-contract Phase 3 review.
- [ ] CE070 blind spot: an adapter that counts `ToolEndEvent`s (or tokens) under a new name to cap or stop a run itself — the rule matches identifiers only; needs a data-flow check that a counter in `agents/` feeds a break or an end status — caught in the central-enforcement plan (Phase 5).
- [ ] CE070 blind spot: an adapter that re-grows a skill scanner through `glob("*.md")`, `rglob`, or a file name built from parts — the rule matches the literal `"SKILL.md"` only; needs a filesystem-walk classifier scoped to `agents/` — caught in the central-enforcement plan (Phase 5).
21 changes: 5 additions & 16 deletions .claude/notes/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,22 +59,11 @@ intentionally brief and out of scope; trimming for DISPLAY belongs in the render

## Harness run-limit parity

- **Harness run-limit parity**: a shared `BaseAgentConfig` field must mean the same
thing on every backend, so a divergence is either fixed or documented — never silent.
**`run_limits.max_tool_calls` is the `TurnMonitor`'s cap, in resolved tool calls, on
every harness**: no adapter counts it; each stops at its next `should_stop` poll and
finalizes cleanly as `tool_calls_exhausted` (no crash, no retry).

The **known unfixed divergences** — which config fields each harness does and does not
enforce — are the table's to state, not this file's. Full table + rationale:
docs/agents/HARNESS_PARITY.md. The per-harness `agent.plugins[].path` depth is no longer
a divergence: staging hands every harness one layout (§ Skills, per harness).

The agent-field half of parity is now the `HarnessContract` each agent class declares:
a field, `permission_mode` value or tool name a harness cannot honor is a resolution
error, and `make parity-table` renders the contract (CE069 checks it), so the page can no
longer drift from the adapters. The run-limit half is still the hand-written table above;
Plan 2 moves it onto the contract.
A shared field must mean the same thing on every backend. Both halves are generated
tables in docs/agents/HARNESS_PARITY.md (`make parity-table`, CE069): `run_limits` from
`RunLimits` and each agent's `HarnessContract`, the agent fields from the contract. Every
cap and budget is the `TurnMonitor`'s (orchestration.md § The watcher became the
TurnMonitor); CE070 keeps adapters from counting one again.

## Shared turn lifecycle

Expand Down
10 changes: 10 additions & 0 deletions .claude/notes/isolation.md
Original file line number Diff line number Diff line change
Expand Up @@ -676,6 +676,16 @@ never `source_yaml`, because the raw on-disk text predates `--model` and `-D` mu
container must see. `source_yaml` is forwarded separately so `task.json`'s audit trail
matches the in-process driver's.

### Plugin staging under docker

The in-container orchestrator stages `agent.plugins` itself, under `/work/output`, which is
the host run dir bind-mounted. Each `plugins[].path` is dumped into the container's
`task.yaml` as the absolute host path (best-effort: a path that does not resolve on the
dumping host is left as authored, because a detached grade never uses it), and that
path is auto-mounted read-only at the same absolute path, together with any staged
skill whose resolved source sits outside every plugin root. So the stage's symlinks
resolve inside the container and, afterwards, on the host.

## Trusting what the container sends back

### The stdout line limit
Expand Down
17 changes: 17 additions & 0 deletions .claude/notes/orchestration.md
Original file line number Diff line number Diff line change
Expand Up @@ -458,6 +458,23 @@ the pass-stop each round — and if none ever decides, the run simply continues
A row with zero pass-capable armed criteria (a negative row stacking only distractors) has
nothing to defer for and fail-stops on the first misfire.

### The watcher became the TurnMonitor

`EarlyStopWatcher` answered one question on the `should_stop` channel. Every harness
also counted its own turn cap in its own unit, and the budgets were checked by the
orchestrator after a turn had already spent the money. `TurnMonitor` answers all four
reasons (`EARLY_CRITERION`, `TOOL_CALL_CAP`, `TOKEN_BUDGET`, `USD_BUDGET`) from ONE
collector, so a cap means the same number of resolved tool calls on every harness and a
budget stops the agent at its next poll. It is cumulative because one instance serves
every retry attempt and every dialog turn of a task: the cap, the budgets and
`expected_tool_calls` all measure the task, not an attempt. On one round the armed stop
wins, then the cap, then the token budgets, then USD, and the first latched reason is
final, so the status an adapter finalizes with cannot flip after the fact. Fail-open
covers only the armed criteria: a raising `live_verdict` is agent-output-dependent code,
while the cap and budgets read counters and must keep running on a run that has lost its
criteria. `result.tool_calls_exhausted` still comes from the turn's end status, not the
latch, because a cap latched after the agent's last poll stopped nothing.

### Inert triggers are by design, and the watcher fails open

A trigger whose polarity an instance can never decide is INERT, not an error — one
Expand Down
28 changes: 17 additions & 11 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,12 +85,14 @@ Each entry is a pointer. Full rationale: `.claude/notes/` (index: `.claude/notes
- **Reference solutions are directory-only** and chmod-shielded during `communicate`.
Defense-in-depth, not a boundary — the known gaps are documented in the notes.
Authoring reference: [Reference Solutions](docs/TASK_DEFINITION_GUIDE.md#reference-solutions).
- **Harness run-limit parity**: a shared config field must mean the same thing on every
backend, or the divergence is documented. Every agent declares a `HarnessContract`; a
base field the harness marks unsupported is rejected at resolution. Table:
[Run-Limit Parity](docs/agents/HARNESS_PARITY.md) § Agent-field contract
(generated). Caps are authored under
[Run Limits](docs/TASK_DEFINITION_GUIDE.md#run-limits).
- **Harness run-limit parity**: every structural cap and budget is one `TurnMonitor`
answer on the `should_stop` channel, in tool calls or tokens, on every harness;
`run_limits` and agent-field meanings are both generated tables in
[Run-Limit Parity](docs/agents/HARNESS_PARITY.md). Every agent declares a
`HarnessContract`; a base field the harness marks unsupported is rejected at
resolution. Caps are authored under [Run Limits](docs/TASK_DEFINITION_GUIDE.md#run-limits).
- **Plugin staging**: `stage_plugins` hands every harness one canonical plugin root;
`skills_offered` is the positive control `skill_triggered` checks.
- **Execute vs. run**: `execute` is `run` with grading off — rows finalize as
`NOT_GRADED` and leave both sides of every rate. Per-command behaviour:
[CLI Commands](docs/USER_GUIDE.md#cli-commands).
Expand Down Expand Up @@ -134,8 +136,9 @@ CLI → ExperimentRunner (task × variant, 5-layer merge) → run_batch → Orch

Per-task (single iteration; simulation mode runs a multi-turn dialog):
1. Orchestrator._communicate_with_retry(prompt, iteration) → TurnRecord
(wraps agent.communicate with retry, per-attempt turn_timeout, and
on_attempt_error → preserves crashed=True partial TurnRecords)
(wraps agent.communicate with retry, per-attempt turn_timeout, the task's
TurnMonitor as the should_stop poll, and on_attempt_error → preserves
crashed=True partial TurnRecords)
2. SuccessChecker.check_all_async() → List[CriterionResult]

Cleanup: stop agent, save EvaluationResult, generate reports.
Expand Down Expand Up @@ -164,7 +167,7 @@ make evalboard-verify # the JS half: tsc --noEmit + vitest + next build
make docs-indexes # README/docs index tables from the mkdocs nav (CE028)
make plugin-reference # the plugin's criteria reference from the models (CE033)
make pricing-mirror # the evalboard's rate table from pricing.py (CE065)
make parity-table # the agent-field contract tables from the agent classes (CE069)
make parity-table # the run-limit and agent-field tables from RunLimits and the agent classes (CE069)

make docs-budget # per-file comment budget + docstring essay check (fails `make verify`)
```
Expand Down Expand Up @@ -220,8 +223,11 @@ A few rules constrain routine edits, so they are worth knowing before you start:
`stats.py` and `run_record.py`.
- **CE068** keeps `orchestration/`, `streaming/` and `timing.py` free of concrete agent
config classes and `AgentKind` members (except `UNKNOWN`); ask the registry instead.
- **CE069** diffs the generated contract tables in `docs/agents/HARNESS_PARITY.md` against
the agent classes. Regenerate with `make parity-table`.
- **CE069** diffs the generated run-limit and contract tables in `docs/agents/HARNESS_PARITY.md`
against `RunLimits` and the agent classes. Regenerate with `make parity-table`.
- **CE070** keeps agent adapters from counting caps (`max_tool_calls`, `RunLimits`,
`tool_calls_exhausted`, …) or scanning for `SKILL.md`: the `TurnMonitor` owns caps and
`orchestration/plugin_staging.py` owns skill discovery.

**Docs index SSOT.** `nav:` plus `extra.docs_index` in `mkdocs.yml` are the single
source of truth for `README.md`'s Documentation table, `docs/index.md`'s "Where to go
Expand Down
2 changes: 1 addition & 1 deletion Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ plugin-reference: ## Regenerate the plugin's bundled criteria reference from th
pricing-mirror: ## Regenerate the evalboard's rate table from pricing.py (SSOT)
uv run python -m tests.lint.pricing_mirror

parity-table: ## Regenerate the agent-field contract table from the agent classes (SSOT)
parity-table: ## Regenerate the agent-field, tool-name and run-limit tables (SSOT)
uv run python -m tests.lint.harness_parity

docs-budget: ## Report the docstring/comment prose budget and check it against the baseline
Expand Down
Loading