Skip to content

feat(agent): cap the agent's context window with agent.context_window - #219

Draft
migheorghe wants to merge 2 commits into
mainfrom
feat/agent-context-window
Draft

migheorghe wants to merge 2 commits into
mainfrom
feat/agent-context-window

Conversation

@migheorghe

@migheorghe migheorghe commented Oct 7, 2026 •

Copy link
Copy Markdown

feat(agent): cap the agent's context window with agent.context_window

Problem

In a benchmark of Claude Code on Sonnet 5 (Bedrock, 1M-token window), runs grew to 250k-480k tokens of context. Above about 150k tokens each extended-thinking step got about 10x slower, so most runs hit the 2-hour task limit. We want to measure pass rate and cost at several context caps (for example 200k and 400k) on the same tasks, as experiment variants.

Before this PR there was no harness-neutral way to do that. On claude-code the only route was claude_settings.autoCompactWindow, which a variant cannot merge into a claude_settings file path. On codex there was no route at all, because _build_thread_options builds the thread config itself and has no pass-through.

Design

A new field, BaseAgentConfig.context_window: int | None (tokens, gt=0, default None = harness default). You can set it at any of the 5 config layers, and with -D agent.context_window=N.

  • Each config class declares _context_window_range (a ClassVar). None, the base default, means the harness cannot apply the field.
  • A model validator rejects an unsupported type, or a value out of range, as soon as the concrete type is known. A type-less config defers the check until the type resolves. A plugin agent inherits None, so it rejects the field until it opts in.
  • The value is never silently ignored. A variant that runs uncapped under a "capped" label would be a wrong measurement.

Recording: the field is part of agent_config on every task row (run.json / task.json), and config lineage records the layer that set it. Variant grouping in the experiment report needs no new code.

Rationale and the options considered (a per-harness pass-through, a field under run_limits, naming) are in .claude/notes/context-window.md.

Per-harness mapping

harness mechanism range compaction trigger
claude-code CLAUDE_CODE_AUTO_COMPACT_WINDOW=<N> in the SDK env 100000-1000000 when the context approaches the window, which is capped to the model's window
codex thread config.model_context_window = N any positive integer at 90% of the window, which Codex caps to the model's catalog maximum
pi providers.<provider>.modelOverrides.<model>.contextWindow = N in a per-agent models.json (PI_CODING_AGENT_DIR) >= 32768, model as provider/model past the window minus compaction.reserveTokens (16384 by default)
antigravity, opencode, delegate, none rejected at load

Pi reads models.json only from its agent dir, so a capped Pi agent runs from a temp dir that links every entry of the host's agent dir (auth, settings, extensions) and holds the host's models.json plus the override; the dir is removed in stop(). Pi silently ignores an override for an unknown model id, so start() runs pi --list-models <model> against that dir and fails unless that exact provider/model row shows the capped window.

Why the other three stay rejected (details in the note): OpenCode's limit.context would work, but its schema also requires limit.output, which coder_eval does not know per model, and its CLI is not pinned. Antigravity's CompactionConfig(token_threshold=N) only moves the compaction trigger, and nobody has seen what the binary does at it. Delegate keeps and summarises the history in the UiPath backend, and no client option controls it.

Why the env var rather than the autoCompactWindow settings key or --autocompact on claude-code:

  • The env var has the highest precedence (env > --autocompact > settings), so host, project and claude_settings values cannot override the cap.
  • It works whether claude_settings is a dict or a file path.
  • Both bundled CLIs read it (2.1.216 and 2.1.281). --autocompact only exists from 2.1.221, and extra_args stays framework-owned.
  • claude_settings.autoCompactWindow set together with context_window is a load error, because the env var would silently override it.
  • An uncapped claude-code variant blanks a host-set CLAUDE_CODE_AUTO_COMPACT_WINDOW (empty is falsy to the CLI), so a run recorded as uncapped is never capped by the host.

Verification

  • New tests/test_agent_context_window.py (26 scenarios): what reaches the claude-code SDK env and the codex thread config; Pi's mirrored agent dir, merged models.json (other overrides kept, host files untouched, dir removed on stop()), the failed --list-models check, the uncapped default, and Pi's config errors; range errors, the conflicting-source error, rejection by every unsupported harness, deferral on a type-less config, the variant layer plus lineage, a variant switching to an unsupported type, and a -D override (both valid and out of range).
  • uv run pytest tests/test_agent_context_window.py: 26 passed. tests/test_pi_agent.py and tests/test_pi_agent_config.py still pass.
  • Live, on claude-code and codex: the devtools-coding-agent-bench pipeline carried this change as a patch on coder_eval 0.12.1 and ran Sonnet 5 and gpt-5.5 at context_window: 200000 on one task (Azure DevOps build 13703588). Both recorded context_window: 200000 in agent_config. Claude Code compacted at 161,677 → 59,806 and 155,083 → 58,987 tokens (it had peaked at 319k on the same task uncapped), and Codex at 175,917 → 15,274 (88% of the cap; 225k uncapped). Pi has not been run live.
  • uv run ruff format --check / uv run ruff check on src/ tests/ .github/scripts/: clean.
  • uv run python -m tests.lint.prose_budget: clean.
  • uv run pytest tests/test_custom_lint.py: 734 passed, 1 failed. The failure is TestCE033PluginReferenceParity::test_drift_is_detected: the test calls write_text with no encoding, so it fails on Windows. It also fails on origin/main.
  • uv run pytest -n auto -m "not live and not lint" tests/: 5960 passed, 118 failed, 5 errors. That failure set is identical, test for test, to the same files run on origin/main on this Windows host (shell/action/litellm-extra tests). It is not caused by this change.
  • uv run pyright: 6 errors, all missing optional extras (litellm, harbor) in the local env. They are not in the touched files.
  • No live or paid-API test was run from this repo; the live run above was a devtools-coding-agent-bench pipeline build.

Facts verified for the design:

  • Claude Code: in the bundled CLI binaries of claude-agent-sdk 0.2.124 (CLI 2.1.216, used by coder_eval 0.12.1) and 0.2.159 (CLI 2.1.281), the settings schema is autoCompactWindow: int().min(1e5).max(1e6).optional().catch(undefined). The window resolver gives the env var precedence over settings and takes min(model window, value). Docs: model-config "Set the auto-compact window" and cli-reference --autocompact ("Requires Claude Code v2.1.221 or later").
  • Codex: in openai-codex-cli-bin 0.156.1 and openai/[email protected], model_context_window replaces the model's context_window (capped to max_context_window) and drives both the default auto-compact limit (90%) and the hard cap. model_auto_compact_token_limit only moves the trigger, and is capped to 90% of the window. The key is also in 0.144.4, the CLI coder_eval 0.12.1 pins.
  • Pi 0.87.1 (the Dockerfile's PI_VERSION): modelOverrides.<id>.contextWindow is the topmost layer over built-in models (dist/core/provider-composer.js, applyModelOverride); compaction fires at contextTokens > contextWindow - reserveTokens (dist/core/compaction/compaction.js); models.json is read only from getAgentDir() = PI_CODING_AGENT_DIR or ~/.pi/agent (dist/config.js); --list-models prints the context column as 200K / 1M (dist/cli/list-models.js), which is what start() compares.

Unverified / risks

  • Pi has not been run live with a cap; its behaviour is read from the 0.87.1 package source. --list-models lists only models whose provider has credentials, so a capped Pi agent without them fails at start() instead of running.
  • Neither claude-code nor codex reports when the model's own window has lowered the cap.
  • Equal caps do not compact at the same token count on both harnesses: Codex compacts at 90% of the window, and Claude Code at its internal buffer below the window.
  • A finding outside this PR: claude-agent-sdk runs its bundled CLI ahead of claude on PATH, so environment_info.claude_code_cli (claude -v) can record the wrong version. For example, 0.12.1 records 2.1.177 while 2.1.216 ran.
  • I could not run mkdocs build locally (mkdocs is not installed), so the new anchor #the-context-window-cap-per-harness is not checked in the built HTML.

🤖 Generated with Claude Code

https://claude.ai/code/session_01RarvxSsr4K2KkBisTbXxQY

migheorghe and others added 2 commits October 7, 2026 17:26
Add a harness-neutral `agent.context_window` field (tokens) so a benchmark
operator can run the same tasks at different context caps, per agent config
or per experiment variant.

- claude-code: sent as CLAUDE_CODE_AUTO_COMPACT_WINDOW in the CLI env, which
  outranks --autocompact and every settings scope; range 100000-1000000
  (the CLI silently drops out-of-range settings values, so the range is
  enforced at load). Setting claude_settings.autoCompactWindow too is a
  load error.
- codex: sent as model_context_window in the thread config; Codex caps it to
  the model's catalog maximum and auto-compacts at 90% of it.
- every other agent type rejects the field at load instead of running
  uncapped under a capped label.

The resolved value is recorded in agent_config on every task row and in the
config lineage. Rationale: .claude/notes/context-window.md

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01RarvxSsr4K2KkBisTbXxQY
Pi 0.87.1 can lower a model's window with
providers.<provider>.modelOverrides.<model>.contextWindow in models.json, and
then compacts once the context passes the window minus
compaction.reserveTokens, the same effect as Codex's model_context_window.

Pi reads models.json only from its agent dir, so a capped agent runs from a
temp dir that links every entry of the host's agent dir (auth, settings,
extensions) and holds the host's models.json plus the override, pointed at by
PI_CODING_AGENT_DIR and removed in stop(). Links keep credentials in one place
and the capped variant's setup identical to the uncapped one's.

Pi silently ignores an override for an unknown model id, so start() runs
`pi --list-models <model>` against the mirror and fails unless that exact
provider/model row shows the capped window. The config needs model as
provider/model and a window of at least 32768 (twice the default reserve).

OpenCode, Antigravity and delegate still reject the field; the note records why.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01RarvxSsr4K2KkBisTbXxQY

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant