Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .claude/notes/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ rule docstrings in `tests/lint/rules/`, then the guides under `docs/`.
## Contents

- [agents.md](agents.md) — agent adapters, the turn lifecycle, token reconciliation, harness parity
- [context-window.md](context-window.md) — the `agent.context_window` cap, per harness
- [contracts.md](contracts.md) — criteria, datasets, aggregation, judging
- [isolation.md](isolation.md) — the docker driver, the sandbox, detached grading
- [lint-rules.md](lint-rules.md) — why each CE lint rule exists
Expand Down
193 changes: 193 additions & 0 deletions .claude/notes/context-window.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,193 @@
# Capping the agent's context window

Analysis and design for `agent.context_window`: one harness-neutral field that caps the
context window an agent works within, set per agent config or per experiment variant.

## Problem and evidence

A benchmark of Claude Code on Sonnet 5 (Bedrock, 1M-token window) let runs grow to
250k-480k tokens of context. Above about 150k tokens each extended-thinking step got
about 10x slower, so most runs hit the 2-hour task limit. The operator wants to measure
pass rate and cost at several context caps, for example one variant at 200k and one at
400k, on the same tasks.

Before this change no harness-neutral way existed:

- **claude-code**: `claude_settings` (`--settings`) could carry `autoCompactWindow`. That
works for one harness only, and when `claude_settings` is a file path a variant cannot
merge one key into it.
- **codex**: `CodexAgentConfig` has no pass-through. `CodexAgent._build_thread_options`
builds the thread `config` dict itself, so no value reached Codex.
- **pi**: a window can be set per model in `models.json`, but coder_eval wrote none.
- **antigravity, opencode, delegate**: no knob is wired.

## How each harness controls its window

### claude-code

Verified in the CLI the SDK spawns. `claude-agent-sdk` runs its **bundled** CLI before it
looks for `claude` on `PATH` (`SubprocessCLITransport._find_cli`), so the CLI that runs is
the one the SDK wheel bundles, not the one the agent image installs with npm:

| coder_eval | claude-agent-sdk (uv.lock) | bundled CLI | image npm CLI |
|---|---|---|---|
| 0.12.1 | 0.2.124 | 2.1.216 | 2.1.177 |
| 0.12.12 | 0.2.159 | 2.1.281 | 2.1.281 |

`environment_info.claude_code_cli` records `claude -v` from `PATH`, so a 0.12.1 run
records 2.1.177 while 2.1.216 ran. That is a separate defect and is not fixed here.

Three inputs set the auto-compact window. Both bundled CLIs (2.1.216 and 2.1.281) contain
the same resolution function. The highest precedence comes first:

1. `CLAUDE_CODE_AUTO_COMPACT_WINDOW` environment variable: a plain token count. It
outranks the flag and every settings scope.
2. `--autocompact <auto|tokens>`: added in 2.1.221 (CLI reference). It is absent from
2.1.216 as a parsed option.
3. `autoCompactWindow` settings key. The settings schema is
`int().min(100000).max(1000000).optional().catch(undefined)`, so an out-of-range
value is **silently dropped**.

The effective window is `min(model context window, configured value)`. Auto-compaction
starts when the context approaches that window.

### codex

Verified in the `openai-codex-cli-bin` 0.156.1 binary (the pin) and its source at tag
`rust-v0.156.1`:

- `model_context_window` ("Size of the context window for the model, in tokens"):
`with_config_overrides` sets the model's `context_window` to
`min(value, max_context_window)`. Three things follow from it. The default
auto-compact limit becomes 90% of it. The hard cap, which forces compaction, becomes
`context_window * effective_context_window_percent / 100`. The remaining-window figure
the model is shown also changes.
- `model_auto_compact_token_limit` ("Token usage threshold triggering auto-compaction"):
the trigger only, clamped to 90% of the resolved window. It does not change the window.
- `model_auto_compact_token_limit_scope` (`total` by default, or `body_after_prefix`) and
`model_post_turn_compact_threshold_percent` (turn-end compaction, off by default) tune
when compaction starts. They are not caps.

The thread `config` dict `thread_start` takes is the same override surface that
coder_eval already uses for `enabled_tools` and `model_providers`.

### pi

Pi 0.87.1 reads `models.json` from its agent dir (`PI_CODING_AGENT_DIR`, default
`~/.pi/agent`). `providers.<provider>.modelOverrides.<model>.contextWindow` replaces a
built-in model's window as the topmost config layer (`dist/core/provider-composer.js`,
`applyModelOverride`), and Pi compacts by summarising once the context passes
`contextWindow - compaction.reserveTokens` (16384 by default;
`dist/core/compaction/compaction.js`). That is the same effect as Codex's
`model_context_window`.

Pi has no separate path for `models.json`, so a capped agent runs from a per-agent
temp dir that links every entry of the host's agent dir (auth, settings, extensions,
prompts) and holds the host's `models.json` plus the override. Links rather than
copies keep credentials in one place and keep the capped variant's setup identical to
the uncapped one's; where the OS refuses a link the entry is copied. The temp dir is
removed in `stop()`.

Pi silently ignores an override for a model id it does not know, so `start()` runs
`pi --list-models <model>` against the mirror and fails unless that exact
provider/model row shows the capped window. A host `models.json` with comments (Pi
strips them) cannot be merged as JSON and fails `start()` the same way.

### antigravity, opencode, delegate

None of them has a context-window knob wired in coder_eval, and they reject the field.

- **opencode**: `provider.<id>.models.<model>.limit.context` (and `limit.input`, which
models that declare it use instead) would work, but the schema also requires
`limit.output`, which coder_eval does not know per model, and the OpenCode CLI is not
version-pinned.
- **antigravity**: `CompactionConfig(token_threshold=N)` in `LocalAgentConfig`
(google-antigravity 0.1.20) moves the compaction trigger without lowering the
window; what the closed binary does at the threshold has not been observed.
- **delegate**: the conversation lives and is summarised in the UiPath backend, and
no client option controls it.

## Options considered

**Per-harness pass-through** (a Codex `config` pass-through next to `claude_settings`).
It is flexible, but a variant would have to name a different key per harness, and the
recorded value would mean something different on each. It also opens a free-form
pass-through on Codex with no denylist. Rejected.

**A field under `run_limits`.** Run limits are caps that the orchestrator or the agent
enforces on a run (turns, time, tokens, USD), and the parity table lists them. A context
window is not a spend cap. It is a harness setting that changes what the agent does, and
it depends on the agent type, which `run_limits` does not know. The guide also says that
`agent:` holds no run-time caps. Rejected.

**A field on the agent config.** One name, accepted on a type-less config so that
experiment defaults and variants can set it, and validated against the concrete type
when the config resolves. **Chosen.**

Naming: `context_window`, in tokens. It names the quantity the operator caps, matches
Codex's `model_context_window`, and avoids "auto-compact", which describes a mechanism of
one harness. `max_context_tokens` was considered, but it reads as a per-request input
limit, which neither harness enforces.

## Recommended design

`BaseAgentConfig.context_window: int | None` (default `None`, `gt=0`).

- Each config class declares `_context_window_range: ClassVar[tuple[int, int | None] |
None]`. `None` (the base default) means that the harness cannot apply the field.
`ClaudeCodeAgentConfig` declares `(100_000, 1_000_000)` and `CodexAgentConfig` declares
`(1, None)`.
- `check_context_window_supported` (model validator) rejects an unsupported type or an
out-of-range value as soon as the type is known. A plugin agent inherits `None`, so it
also rejects the field until it opts in.
- The field uses the default `replace` merge strategy. It is set at any of the five
layers, and with `-D agent.context_window=N`.

Per-harness mapping:

- **claude-code**: `CLAUDE_CODE_AUTO_COMPACT_WINDOW=<N>` in the SDK `env`. The
environment variable was chosen over the settings key and the flag for four reasons.
It has the highest precedence, so host, project and `claude_settings` values cannot
override the cap. It works whether `claude_settings` is a dict or a file path. Both
bundled CLIs read it, which the flag does not. `extra_args` stays framework-owned.
`claude_settings.autoCompactWindow` together with `context_window` is a load error,
because the environment variable would silently override it.
- **codex**: `model_context_window: <N>` in the thread `config`. No trigger key is set,
so Codex compacts at its default 90% of the window.

## Validation

- `gt=0` on every config, so a type-less config rejects nonsense as well.
- claude-code: 100000-1000000, the CLI's documented range. Outside that range the CLI
drops a settings value silently, so coder_eval enforces the range itself.
- codex: any positive integer. Codex clamps the value to the model's catalog maximum.
- pi: at least 32768, twice Pi's default compaction reserve, and `model` must be in
`provider/model` form, since the override is per model.
- every other type: rejected at load, naming the supported types. A variant that
switches `type` to an unsupported harness fails when the experiment resolves.

## Recording

The resolved agent config is persisted as `agent_config` on every task row (`run.json`
and `task.json`). `agent.context_window` therefore appears per task, and config lineage
records which layer set it (for example, the variant). The experiment report groups by
variant, so a cap per variant groups without new report code. On claude-code the
`sdk_options` dump also shows the environment variable in `env`.

## Open questions

- **A silent clamp to the model's window.** Neither harness reports that the model's
window lowered the cap. A cap above the model's window runs at the model's window.
Recording the effective window would need a per-model catalog in coder_eval.
- **Trigger points differ.** Codex compacts at 90% of the window. Claude Code compacts
when the context approaches the window, minus an internal buffer. Equal caps therefore
do not compact at exactly the same token count. A separate harness-neutral trigger
field is possible, but it is not added until a benchmark needs it.
- **A host `CLAUDE_CODE_AUTO_COMPACT_WINDOW` leaks into uncapped runs.** The SDK starts
the CLI with `os.environ` underneath `options.env`. When the field is unset, a host
value still applies. Clearing it would change behaviour for current users. It is
recorded in this note but not changed.
- **The recorded CLI version is wrong.** See the version table above:
`environment_info.claude_code_cli` should record the bundled CLI version.
- **Other harnesses.** OpenCode can opt in once its CLI is pinned and coder_eval can
supply `limit.output`; Antigravity once one run shows what `token_threshold` does.
24 changes: 24 additions & 0 deletions docs/AB_EXPERIMENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@ to the orchestrator are involved.
- [What a Variant Can Override](#what-a-variant-can-override)
- [Recipe: A/B a Skill](#recipe-ab-a-skill)
- [Recipe: A/B a Model](#recipe-ab-a-model)
- [Recipe: A/B a Context Window](#recipe-ab-a-context-window)
- [Recipe: A/B a Prompt](#recipe-ab-a-prompt)
- [Recipe: Smoke vs. e2e Flavors (Early Stop)](#recipe-smoke-vs-e2e-flavors-early-stop)
- [Replicates (Statistical Power)](#replicates-statistical-power)
Expand Down Expand Up @@ -222,6 +223,29 @@ variants:
agent: { model: claude-opus-5 }
```

## Recipe: A/B a Context Window

`agent.context_window` caps the context the agent works within, in tokens: the
harness compacts the conversation before it outgrows the cap. Use it to measure
pass rate and cost against the cap on a model with a large window. Each task row
records the cap in `agent_config.context_window`.

```yaml
experiment_id: context-window
description: "The same model with a 200k and a 400k context cap"

variants:
- variant_id: cap-200k
agent: { context_window: 200000 }
- variant_id: cap-400k
agent: { context_window: 400000 }
```

Only `claude-code` and `codex` can apply the cap. On every other agent type the
experiment fails at load. See
[Run-Limit Parity](agents/HARNESS_PARITY.md#the-context-window-cap-per-harness) for what
the cap does on each harness.

## Recipe: A/B a Prompt

Use `prompt_mutations` (transform the task prompt) or `initial_prompt` /
Expand Down
8 changes: 8 additions & 0 deletions docs/TASK_DEFINITION_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -170,6 +170,7 @@ agent:
- "Write"
- "Bash"
model: "claude-sonnet-5" # Optional: specific model
context_window: 200000 # Optional: cap the context window, in tokens (claude-code, codex, pi)
sdk_options: # Optional: agent SDK pass-through (keys depend on `type`)
effort: high # claude-code: any non-framework-managed ClaudeAgentOptions field
```
Expand All @@ -188,6 +189,13 @@ validates its own keys at YAML load:
Deep-merged across the 5-layer config chain. Override via CLI with the
repeatable `-D agent.sdk_options.KEY=VALUE`.

**`context_window`** caps the context window the agent works within, in tokens: the
harness compacts the conversation before it outgrows the cap. Unset, the harness uses
the model's full window. `claude-code` accepts 100000-1000000 and `codex` any positive
value; every other agent type rejects the field at load. The cap never raises the
model's own window. Override via CLI with `-D agent.context_window=400000`. What each
harness does with it: [Run-Limit Parity](agents/HARNESS_PARITY.md#the-context-window-cap-per-harness).

**Permission Modes:**
- `default` — Default permission handling
- `acceptEdits` — Auto-accept file edits (recommended for evaluations)
Expand Down
1 change: 1 addition & 0 deletions docs/agents/CLAUDE_CODE.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,6 +104,7 @@ agent:
| `system_prompt_file` | `str \| null` | Path (relative to the task YAML) loaded into `system_prompt` at resolution. Works with either `system_prompt_mode`. |
| `setting_sources` | `list["user"\|"project"\|"local"] \| null` | Which host setting sources the SDK reads. Default resolves to `["project"]`. See [Sandbox isolation](#sandbox-isolation). |
| `claude_settings` | `str \| dict \| null` | Passed to the SDK `--settings`. A dict is JSON-serialized; a str is a settings file path. Use `permissions.deny` to block tools/paths. |
| `context_window` | `int \| null` (100000-1000000) | Auto-compact window, in tokens, sent as `CLAUDE_CODE_AUTO_COMPACT_WINDOW`. It outranks `--autocompact` and every settings scope, and is capped to the model's window. Setting `claude_settings.autoCompactWindow` too is a load error. See [Harness parity](HARNESS_PARITY.md#the-context-window-cap-per-harness). |
| `sdk_options` | `dict` (default `{}`) | Pass-through for `ClaudeAgentOptions` fields Coder Eval doesn't own (e.g. `effort`). Validated at load — an unknown or framework-owned key is a hard error. |
| `ignore_patterns` | `list[str] \| null` | Gitignore-style overrides for the workspace copy used by judge sub-agents (supports `!` negation). |

Expand Down
7 changes: 7 additions & 0 deletions docs/agents/CODEX.md
Original file line number Diff line number Diff line change
Expand Up @@ -192,6 +192,13 @@ The agent maps `permission_mode` to the Codex SDK's `Sandbox`. The approval mode

`allowed_tools` / `disallowed_tools` are normalized (`Bash` → `shell`, `Write`/`Edit` → `apply_patch`, etc.) and passed as `enabled_tools` / `disabled_tools` in the thread `config`. **Note:** the Codex SDK does not currently enforce `disabled_tools`; do not rely on it as a security boundary (the agent logs a warning when it is set).

### Context Window

`agent.context_window` is passed as `model_context_window` in the thread `config`. Codex
uses it as the model's context window, capped to the model's catalog maximum, and
auto-compacts at 90% of it. See
[Harness parity](HARNESS_PARITY.md#the-context-window-cap-per-harness).

### Skills Discovery

The agent sets up SKILL.md files (Agent Skills open standard) in `.agents/skills/` directory:
Expand Down
28 changes: 28 additions & 0 deletions docs/agents/HARNESS_PARITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -637,6 +637,34 @@ A timeout is a *failure* (partial turn captured, error status); the turn cap is
*clean stop*. Conflating them is the mistake this page exists to prevent: a task
whose cap fires should not look like a task whose harness hung.

## The context window cap per harness

`agent.context_window` is not a run limit, but the same promise applies: one value
must mean the same cap on every harness that accepts it. It is the context window, in
tokens, that the harness compacts the conversation within. A harness that cannot apply
it rejects it when the config loads. It is never ignored, because a variant that runs
uncapped under a "capped" label is a wrong measurement.

| | claude-code | codex | antigravity | opencode | pi | delegate |
|---|---|---|---|---|---|---|
| mechanism | `CLAUDE_CODE_AUTO_COMPACT_WINDOW` in the CLI environment | thread config `model_context_window` | rejected | rejected | `modelOverrides.<model>.contextWindow` in a per-agent `models.json` | rejected |
| accepted range | 100000-1000000 | any positive integer | | | >= 32768, `model` as `provider/model` | |
| when compaction starts | when the context approaches the window | at 90% of the window (Codex's default auto-compact limit) | | | past the window minus `compaction.reserveTokens` (16384 by default) | |
| above the model's own window | capped to the model's window by the CLI | capped to the model's catalog maximum by Codex | | | replaces the model's window | |

- **claude-code**: the environment variable outranks `--autocompact` and every settings
scope, so a host or project `autoCompactWindow` cannot change the cap. Setting
`claude_settings.autoCompactWindow` together with `context_window` fails at load.
- **codex**: the value replaces the model's context window, so the hard-cap
compaction at the usable window (a per-model share of it, 95% for Codex's fallback
model metadata) also moves.
- **pi**: Pi reads `models.json` from its agent dir only, so a capped agent runs from a
temp dir that links the host's agent dir (auth, settings, extensions) and adds the
override. `start()` fails unless `pi --list-models` shows the capped window for that
exact model, because Pi silently ignores an override for an unknown model id.
- Neither claude-code nor codex reports when the model's own window has lowered the cap. A cap
above the model's window runs at the model's window.

## `agent.plugins[].path` accepts different depths per harness

Not a run limit, but the same promise: one task file, three harnesses, same meaning.
Expand Down
14 changes: 14 additions & 0 deletions docs/agents/PI.md
Original file line number Diff line number Diff line change
Expand Up @@ -112,6 +112,20 @@ Pi's reasoning effort, forwarded as `--thinking`. Accepts the seven-value set
`off` / `minimal` / `low` / `medium` / `high` / `xhigh` / `max` (a strict
superset of the Antigravity `thinking_level`), defaulting to `medium`.

### `context_window`

`agent.context_window` caps the model's context window, in tokens (at least 32768).
It needs `model` in `provider/model` form, because Pi applies it as
`providers.<provider>.modelOverrides.<model>.contextWindow` in `models.json`. Pi then
compacts once the context passes the window minus `compaction.reserveTokens`.

Pi reads `models.json` only from its agent dir, so a capped agent runs from a temp dir
that links the host's agent dir (auth, settings, extensions) and holds the host's
`models.json` plus the override. `start()` fails unless `pi --list-models` shows the
capped window for that exact model, since Pi silently ignores an override for a model
id it does not know. See
[Harness parity](HARNESS_PARITY.md#the-context-window-cap-per-harness).

### Enforced config fields

Pi forwards these config knobs to real CLI flags:
Expand Down
Loading
Loading