You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
bug(core): chat compression can exceed its target model context window #9455
Both automatic reactive compression and manual /compress can submit a full-history summarization request that exceeds the compression model's own context window. If that request fails, the dedicated cold/slim fallback can also exceed the same window, leaving compression unable to recover an already oversized session.
The compressor currently performs a size guard when a separately configured compaction model differs from the main model. When the active/main model is also the compression model, that admission check is skipped. Cache-sharing admission can additionally use a retained provider token count rather than the actual current-route payload.
A dogfood session switching to an Anthropic model with a 1M prompt maximum produced the following sanitized sequence:
a normal request exceeded the limit by about 101k;
reactive compression's cold request still exceeded it by about 18k;
manual /compress later attempted a cache-sharing request that exceeded the limit by about 102k, then a cold fallback that exceeded it by about 18k;
a later reactive cold request exceeded the limit by about 76k.
These numbers are endpoint-reported request overages. They are not evidence that a local estimator undercounted by those exact amounts, because no same-main-model cold admission estimate ran. Separately, source inspection found that the local content estimator returns a thought part's text length before considering a co-located thoughtSignature; any future admission guard must count both fields rather than inheriting that omission.
What did you expect to happen?
Compression should not send a summarization request that is already known to exceed the selected model's context window.
Before either automatic compression or /compress sends a request, Qwen should account for the complete request and enough room for the summary, even when the compression model is the active chat model. If the full conversation still cannot fit, Qwen should reduce it in bounded stages while preserving complete tool interactions, the original user intent, and both the newest and oldest relevant context. If it cannot make safe progress, it should stop locally with a clear input-too-large error instead of sending a request expected to fail.
Conversations that already fit should continue to compress in one pass.
Client information
Client Information
Qwen Code development/dogfood build on macOS, based on commit 78f7926962b42d309f5a35edb80f775cf0a6ed9d.
Login information
Observed with an Anthropic-compatible target model. The issue is in shared compression admission and recovery behavior.
#2565 proposes pre-flight trimming for large tool results before normal Anthropic/OpenAI requests. This issue is related but not a duplicate: it covers the compression side-query itself exceeding the target window, including 1M-token models that #2565 explicitly excludes, and defines recovery when the full-history summarizer cannot fit.
这些数字是端点报告的请求超出量,不能解释为本地估算器恰好低估了这些数值,因为同主模型的冷路径根本没有运行准入估算。另一个由源代码确认的问题是:当同一个 thought Part 同时包含文本和 thoughtSignature 时,本地内容估算器会先返回文本长度而忽略签名;未来的准入检查必须同时计算两者,不能继承该遗漏。
What happened?
Both automatic reactive compression and manual
/compresscan submit a full-history summarization request that exceeds the compression model's own context window. If that request fails, the dedicated cold/slim fallback can also exceed the same window, leaving compression unable to recover an already oversized session.The compressor currently performs a size guard when a separately configured compaction model differs from the main model. When the active/main model is also the compression model, that admission check is skipped. Cache-sharing admission can additionally use a retained provider token count rather than the actual current-route payload.
A dogfood session switching to an Anthropic model with a 1M prompt maximum produced the following sanitized sequence:
/compresslater attempted a cache-sharing request that exceeded the limit by about 102k, then a cold fallback that exceeded it by about 18k;These numbers are endpoint-reported request overages. They are not evidence that a local estimator undercounted by those exact amounts, because no same-main-model cold admission estimate ran. Separately, source inspection found that the local content estimator returns a thought part's text length before considering a co-located
thoughtSignature; any future admission guard must count both fields rather than inheriting that omission.What did you expect to happen?
Compression should not send a summarization request that is already known to exceed the selected model's context window.
Before either automatic compression or
/compresssends a request, Qwen should account for the complete request and enough room for the summary, even when the compression model is the active chat model. If the full conversation still cannot fit, Qwen should reduce it in bounded stages while preserving complete tool interactions, the original user intent, and both the newest and oldest relevant context. If it cannot make safe progress, it should stop locally with a clear input-too-large error instead of sending a request expected to fail.Conversations that already fit should continue to compress in one pass.
Client information
Client Information
Qwen Code development/dogfood build on macOS, based on commit
78f7926962b42d309f5a35edb80f775cf0a6ed9d.Login information
Observed with an Anthropic-compatible target model. The issue is in shared compression admission and recovery behavior.
Anything else we need to know?
Relationship to existing issue #2565
#2565 proposes pre-flight trimming for large tool results before normal Anthropic/OpenAI requests. This issue is related but not a duplicate: it covers the compression side-query itself exceeding the target window, including 1M-token models that #2565 explicitly excludes, and defines recovery when the full-history summarizer cannot fit.
Related work
中文
发生了什么?
自动响应式压缩和手动
/compress都可能发送一个超过压缩模型自身上下文窗口的完整历史摘要请求。如果该请求失败,专用冷/精简回退请求也可能超过同一窗口,导致已经过大的会话无法通过压缩恢复。当单独配置的压缩模型不同于主模型时,压缩器会进行大小检查;但当当前/主模型同时也是压缩模型时,该准入检查会被跳过。缓存共享准入还可能使用之前提供商返回的 token 计数,而不是当前路由的实际请求体。
一个切换到 1M 提示上限 Anthropic 模型的试用会话出现了以下经脱敏的序列:
/compress随后发送的缓存共享请求超过约 102k,冷回退仍超过约 18k;这些数字是端点报告的请求超出量,不能解释为本地估算器恰好低估了这些数值,因为同主模型的冷路径根本没有运行准入估算。另一个由源代码确认的问题是:当同一个 thought Part 同时包含文本和
thoughtSignature时,本地内容估算器会先返回文本长度而忽略签名;未来的准入检查必须同时计算两者,不能继承该遗漏。期望行为
压缩不应发送一个已知会超过所选模型上下文窗口的摘要请求。
无论是自动压缩还是
/compress,在发送请求前都应计算完整请求并为摘要保留足够空间,即使压缩模型就是当前聊天模型。如果完整会话仍无法放入窗口,Qwen 应分阶段且有界地缩减内容,同时保留完整的工具交互、最初的用户意图以及最新和最早的相关上下文。如果无法安全推进,应在本地停止并返回清晰的 input-too-large 错误,而不是发送预期会失败的请求。本来可以放入窗口的会话应继续只压缩一次。
客户端信息
macOS 上的 Qwen Code 开发/试用版本,基于提交
78f7926962b42d309f5a35edb80f775cf0a6ed9d。登录信息
在 Anthropic 兼容目标模型上观察到。问题位于共享压缩准入和恢复行为。
其他信息
与 #2565 的关系:#2565 提议在普通 Anthropic/OpenAI 请求前修剪大型工具结果;本 issue 与其相关但不重复。本 issue 关注压缩侧查询本身超过目标窗口,包括 #2565 明确排除的 1M-token 模型,并定义完整历史摘要无法放入窗口时的恢复行为。
相关工作:该问题在验证 PR #8169 并切换保留会话线路时发现。另有独立 issue 跟踪路由作用域提示/输出 token 所有权和外来推理签名兼容性。本文不声称外来签名导致了已测得的压缩超限;该结论需要相同历史的线路/tokenizer A/B 捕获。