Repository navigation
feat(agent): cap the agent's context window with agent.context_window - #219
Draft
migheorghe wants to merge 2 commits into
Draft
migheorghe wants to merge 2 commits into
migheorghe wants to merge 2 commits into
Conversation
Add a harness-neutral `agent.context_window` field (tokens) so a benchmark operator can run the same tasks at different context caps, per agent config or per experiment variant. - claude-code: sent as CLAUDE_CODE_AUTO_COMPACT_WINDOW in the CLI env, which outranks --autocompact and every settings scope; range 100000-1000000 (the CLI silently drops out-of-range settings values, so the range is enforced at load). Setting claude_settings.autoCompactWindow too is a load error. - codex: sent as model_context_window in the thread config; Codex caps it to the model's catalog maximum and auto-compacts at 90% of it. - every other agent type rejects the field at load instead of running uncapped under a capped label. The resolved value is recorded in agent_config on every task row and in the config lineage. Rationale: .claude/notes/context-window.md Co-Authored-By: Claude Opus 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01RarvxSsr4K2KkBisTbXxQY
Pi 0.87.1 can lower a model's window with providers.<provider>.modelOverrides.<model>.contextWindow in models.json, and then compacts once the context passes the window minus compaction.reserveTokens, the same effect as Codex's model_context_window. Pi reads models.json only from its agent dir, so a capped agent runs from a temp dir that links every entry of the host's agent dir (auth, settings, extensions) and holds the host's models.json plus the override, pointed at by PI_CODING_AGENT_DIR and removed in stop(). Links keep credentials in one place and the capped variant's setup identical to the uncapped one's. Pi silently ignores an override for an unknown model id, so start() runs `pi --list-models <model>` against the mirror and fails unless that exact provider/model row shows the capped window. The config needs model as provider/model and a window of at least 32768 (twice the default reserve). OpenCode, Antigravity and delegate still reject the field; the note records why. Co-Authored-By: Claude Opus 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01RarvxSsr4K2KkBisTbXxQY
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
feat(agent): cap the agent's context window with agent.context_window
Problem
In a benchmark of Claude Code on Sonnet 5 (Bedrock, 1M-token window), runs grew to 250k-480k tokens of context. Above about 150k tokens each extended-thinking step got about 10x slower, so most runs hit the 2-hour task limit. We want to measure pass rate and cost at several context caps (for example 200k and 400k) on the same tasks, as experiment variants.
Before this PR there was no harness-neutral way to do that. On claude-code the only route was
claude_settings.autoCompactWindow, which a variant cannot merge into aclaude_settingsfile path. On codex there was no route at all, because_build_thread_optionsbuilds the threadconfigitself and has no pass-through.Design
A new field,
BaseAgentConfig.context_window: int | None(tokens,gt=0, defaultNone= harness default). You can set it at any of the 5 config layers, and with-D agent.context_window=N._context_window_range(a ClassVar).None, the base default, means the harness cannot apply the field.None, so it rejects the field until it opts in.Recording: the field is part of
agent_configon every task row (run.json/task.json), and config lineage records the layer that set it. Variant grouping in the experiment report needs no new code.Rationale and the options considered (a per-harness pass-through, a field under
run_limits, naming) are in.claude/notes/context-window.md.Per-harness mapping
CLAUDE_CODE_AUTO_COMPACT_WINDOW=<N>in the SDKenvconfig.model_context_window = Nproviders.<provider>.modelOverrides.<model>.contextWindow = Nin a per-agentmodels.json(PI_CODING_AGENT_DIR)modelasprovider/modelcompaction.reserveTokens(16384 by default)Pi reads
models.jsononly from its agent dir, so a capped Pi agent runs from a temp dir that links every entry of the host's agent dir (auth, settings, extensions) and holds the host'smodels.jsonplus the override; the dir is removed instop(). Pi silently ignores an override for an unknown model id, sostart()runspi --list-models <model>against that dir and fails unless that exact provider/model row shows the capped window.Why the other three stay rejected (details in the note): OpenCode's
limit.contextwould work, but its schema also requireslimit.output, which coder_eval does not know per model, and its CLI is not pinned. Antigravity'sCompactionConfig(token_threshold=N)only moves the compaction trigger, and nobody has seen what the binary does at it. Delegate keeps and summarises the history in the UiPath backend, and no client option controls it.Why the env var rather than the
autoCompactWindowsettings key or--autocompacton claude-code:--autocompact> settings), so host, project andclaude_settingsvalues cannot override the cap.claude_settingsis a dict or a file path.--autocompactonly exists from 2.1.221, andextra_argsstays framework-owned.claude_settings.autoCompactWindowset together withcontext_windowis a load error, because the env var would silently override it.CLAUDE_CODE_AUTO_COMPACT_WINDOW(empty is falsy to the CLI), so a run recorded as uncapped is never capped by the host.Verification
tests/test_agent_context_window.py(26 scenarios): what reaches the claude-code SDK env and the codex thread config; Pi's mirrored agent dir, mergedmodels.json(other overrides kept, host files untouched, dir removed onstop()), the failed--list-modelscheck, the uncapped default, and Pi's config errors; range errors, the conflicting-source error, rejection by every unsupported harness, deferral on a type-less config, the variant layer plus lineage, a variant switching to an unsupported type, and a-Doverride (both valid and out of range).uv run pytest tests/test_agent_context_window.py: 26 passed.tests/test_pi_agent.pyandtests/test_pi_agent_config.pystill pass.context_window: 200000on one task (Azure DevOps build 13703588). Both recordedcontext_window: 200000inagent_config. Claude Code compacted at 161,677 → 59,806 and 155,083 → 58,987 tokens (it had peaked at 319k on the same task uncapped), and Codex at 175,917 → 15,274 (88% of the cap; 225k uncapped). Pi has not been run live.uv run ruff format --check/uv run ruff checkonsrc/ tests/ .github/scripts/: clean.uv run python -m tests.lint.prose_budget: clean.uv run pytest tests/test_custom_lint.py: 734 passed, 1 failed. The failure isTestCE033PluginReferenceParity::test_drift_is_detected: the test callswrite_textwith no encoding, so it fails on Windows. It also fails onorigin/main.uv run pytest -n auto -m "not live and not lint" tests/: 5960 passed, 118 failed, 5 errors. That failure set is identical, test for test, to the same files run onorigin/mainon this Windows host (shell/action/litellm-extra tests). It is not caused by this change.uv run pyright: 6 errors, all missing optional extras (litellm,harbor) in the local env. They are not in the touched files.Facts verified for the design:
autoCompactWindow: int().min(1e5).max(1e6).optional().catch(undefined). The window resolver gives the env var precedence over settings and takesmin(model window, value). Docs: model-config "Set the auto-compact window" and cli-reference--autocompact("Requires Claude Code v2.1.221 or later").openai-codex-cli-bin0.156.1 andopenai/[email protected],model_context_windowreplaces the model'scontext_window(capped tomax_context_window) and drives both the default auto-compact limit (90%) and the hard cap.model_auto_compact_token_limitonly moves the trigger, and is capped to 90% of the window. The key is also in 0.144.4, the CLI coder_eval 0.12.1 pins.PI_VERSION):modelOverrides.<id>.contextWindowis the topmost layer over built-in models (dist/core/provider-composer.js,applyModelOverride); compaction fires atcontextTokens > contextWindow - reserveTokens(dist/core/compaction/compaction.js);models.jsonis read only fromgetAgentDir()=PI_CODING_AGENT_DIRor~/.pi/agent(dist/config.js);--list-modelsprints the context column as200K/1M(dist/cli/list-models.js), which is whatstart()compares.Unverified / risks
--list-modelslists only models whose provider has credentials, so a capped Pi agent without them fails atstart()instead of running.claude-agent-sdkruns its bundled CLI ahead ofclaudeonPATH, soenvironment_info.claude_code_cli(claude -v) can record the wrong version. For example, 0.12.1 records 2.1.177 while 2.1.216 ran.mkdocs buildlocally (mkdocs is not installed), so the new anchor#the-context-window-cap-per-harnessis not checked in the built HTML.🤖 Generated with Claude Code
https://claude.ai/code/session_01RarvxSsr4K2KkBisTbXxQY