Repository navigation
fix(cli): keep one-shot system-reminder prefixes out of shell mode - #12605
Conversation
Fixes #11626. Two gaps let `<system-reminder>` text reach bash or misroute queued input around shell mode: - handleFinalSubmit prepended the recovered-agents and the worktree-restore one-shot notices even in shell mode, where the leading `<system-reminder>` is a bash syntax error — and consuming the one-shot latch there dropped the notice before any model turn could see it. Both notices are now gated on `!shellModeActive` and stay armed for the next model-bound prompt. (The workflow-keyword reminder was already shell-gated.) - The queue drain routed submissions on the live shell-mode flag at drain time, so a flip while an entry waited in the queue could send a shell command to the model or a model prompt to bash. Shell intent is now recorded at submit time (addMessage/restoreMessages), travels with the queued entry, and the drain/submitQuery route on `shellModeIntent ?? shellModeActive`; batches stay intent-homogeneous so one routing decision never misroutes another entry. Producers that record no intent keep routing on the live flag, as before. Tests: packages/cli src/ui/AppContainer.test.tsx, hooks/use-llm-stream.test.tsx, hooks/useMessageQueue.test.ts — the new cases are red on the unfixed source (prefix reached shell-mode submissions; recorded intent ignored at drain) and green with the fix (544 passed, 0 failed). Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-issue-patrol/jmuf0dfzl0a
…-mode-reminder-prefix
Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmufqt39629
popNextSubmission selected every queue entry matching the first plain entry's recorded shell intent, so a later same-intent entry overtook an earlier different-intent one. Queueing a model prompt, a shell command and a second model prompt while a turn was responding drained both prompts as one batch and left the command queued behind them: the model answered the second prompt against the pre-command tree, and the command then ran with no prompt consuming its output. Take the contiguous same-intent run from the head instead. That keeps per-entry routing and submission order together; the only cost is one extra model turn when a shell command sits between two prompts, which no longer merge. batch[0] is still plainMessages[0], so the batch turnKey and the key peekNextUserBatchKey reserves for the Goal runtime are unchanged. The case that pinned the old reorder now asserts order preservation, and a new case pins that adjacent same-intent entries still aggregate. Review finding: R1-5 on PR 12605. Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmufqt39629
…ents The submitQuery metadata spread in the queued-submission drain is the only production hop that carries a queue entry's recorded shell intent to the router, and nothing pinned it: deleting those three lines left all three suites green, because the use-llm-stream cases inject the intent into submitQuery by hand and the existing drain assertions use objectContaining on a different key, which cannot fail on a missing one. Extend the existing persistent-admission-failure drain case instead of adding a second harness: its submission now records shellMode, the new assertion pins the metadata key itself, and the restore leg is asserted on intent rather than arity (third argument stays true, as that leg defers). Also correct two comments in this file that this change falsified: they said the other notices handleFinalSubmit prepends do not check shell mode and that the gate is load-bearing only through the workflow reminder's call site. Every notice the handler prepends is now gated at its own call site, and the two #11626 cases below depend on the recovered-agents and worktree gates rather than on this one. Review findings: R1-2 and R1-1 on PR 12605. Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmufqt39629
Bring the branch up to date with main (ffea2d0) so the CI lint/freshness gate and the review lane evaluate the diff against the current base. Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmufv3et42h
Ctrl+Q is not shell-gated (InputPrompt.tsx binds QUEUE_MESSAGE after the `if (!shellModeActive)` block closes), so a shell-mode submission can take the deferUntilIdle branch of handleFinalSubmit. That leg's addMessage call was only asserted with the 4th argument false, which is indistinguishable from the parameter's default: replacing shellModeActive with false at AppContainer.tsx:3253 kept the whole suite green. Add the shell-mode arm so both directions of the argument are pinned. Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmufv3et42h
Local real-environment verification: PR head
|
| # | Input | BASE ffea2d0 |
PR d46a3b4 |
|---|---|---|---|
| s1 | qwen --worktree probe-wt (arms the startup worktree notice, which uses the same latch path) → ! echo PR12605_SHELL_RAN; basename "$PWD" → ! off → a model prompt |
❌ bash: syntax error near unexpected token 'newline' / `{ <system-reminder>'. The notice reaches the model only as the text of the failed command (I ran the following shell command: ```sh <system-reminder>…); the model prompt that follows carries no reminder |
✅ the command runs (PR12605_SHELL_RAN, probe-wt). The next model request ends with <system-reminder>[Startup] Active worktree …</system-reminder>\n\nfirst model prompt after shell |
| s2q | Ctrl+Q in shell mode while the model responds, shell mode off before the drain | ❌ sent to the model as a prompt | ✅ runs in bash |
| s3q | Ctrl+Q a prompt, shell mode on before the drain | ❌ bash: summarize: command not found (exit 127) |
✅ sent to the model |
| s2bq | Ctrl+Q in shell mode, shell mode stays on | bash | bash (no change) |
| s2 / s2b | Enter in shell mode while responding (s2: shell mode off before the drain; s2b: stays on) | ❌ sent to the model as a steer | ❌ same: sent to the model as a steer |
| s3 | Enter a prompt, shell mode on before the drain | model | model (no change) |
| s4q / s5q | Ctrl+Q sleep 8; echo … in shell mode, then Ctrl+Q a prompt with shell mode off |
both merged into one blob sent to the model (sleep 8; echo …\n\nsummarize the diff); Esc cancels it normally |
shell runs and the prompt starts as a concurrent model turn during sleep 8. When the command finishes the UI shows idle, Esc is a no-op, and the reply lands about 20 s later (R1-4) |
Finding 1: the Enter / steer path still sends shell-mode commands to the model (pre-existing, not addressed)
When a model stream ends with no tool calls, client.ts runs takeSteerInput() → getSteerInput: drainSteerAtBoundary → midTurnDrainRef → useMessageQueue.drainQueue. drainQueue returns plain text for every entry that is not deferred-until-idle, and ignores shellMode. So an entry queued with Enter never reaches popNextSubmission, and the new intent field is never read. In s2b the footer reads shell mode enabled · ⏳1 queued and shell mode is never turned off. The command still goes to the model (ledger request #1: user content echo PR12605_QUEUED_SHELL_RAN).
Candidate fix: let the steer drain leave shell-intent entries for the queue drain, which already routes them correctly:
@@ packages/cli/src/ui/hooks/useMessageQueue.ts (drainQueue)
- (includeDeferred || !message.deferUntilIdle);
+ (includeDeferred || !message.deferUntilIdle) &&
+ message.shellMode !== true;Measured in the real TUI (rebuilt bundle, same driver):
- s2 and s2b now run in bash (
✓ Shell Command … PR12605_QUEUED_SHELL_RAN, and the command never appears in a model request). - s3 still goes to the model.
- The three suites this PR names stay green: 546/546.
Nothing in the suites tests this, so it would need its own test. There is also one trade-off: at a steer boundary, a later model-intent entry can overtake an earlier shell entry, which then waits for the turn to end. This can land here or as a follow-up. If it becomes a follow-up, please open the issue the bot suggested so the gap is tracked.
Finding 2: R1-4 / #12664 confirmed end to end
s5q on the PR arm, probing the screen at fixed times:
- During
sleep 8, the UI shows the promptHOLD_TURN summarize the diffrunning as a model turn (spinner,esc to cancel) next to the runningShell Command. - About 8 s later the command finishes. The spinner is gone, the prompt reads idle, and the fake server still has the request open.
- Esc changes nothing. The reply
MODEL_REPLY_HOLD_DONEarrives about 20 s after the request started.
The merge base took the same input as one blob (itself mis-routed: a shell command sent to the model), and Esc cancelled it. This agrees with the attribution in #12664: the idle-state overwrite is pre-existing, and this PR makes a mixed Ctrl+Q queue walk into it routinely. It is already tracked as P1 with need-discussion, so I don't see it as blocking here.
Unit-level red/green check
In packages/cli: vitest run src/ui/AppContainer.test.tsx src/ui/hooks/use-llm-stream.test.tsx src/ui/hooks/useMessageQueue.test.ts
- PR head: 546/546 passed.
- The three production files reverted to
ffea2d0, tests kept: 27 failed. These are all the new fix(cli): system-reminder prefixes reach bash in shell mode #11626 cases (both shell-mode reminder cases, both intent-routing directions, homogeneous batching, restore, the Ctrl+Q shell pin, the drain metadata pin), plus existing cases whose assertions now expect the fourthaddMessageargument. - The tree was restored afterwards.
Not covered
- The recovered-agents notice was not triggered live (it needs a paused background agent from an earlier session). It uses the same gating pattern as the worktree notice verified in s1, and the unit suite pins it.
- Linux and Windows were not run locally; CI on this head is green.
Recommendation
OK to merge as the fix for the reminder half of #11626 and for Ctrl+Q queue routing. Neither finding is a regression in what the PR touches:
- Finding 1 is pre-existing and was scoped out.
- Finding 2 is tracked in Shell-mode commands never hold the session busy: streamingState reads Idle while the command runs, so the queue drain admits a concurrent model turn #12664.
Before or right after merging:
- Narrow the description's "queued submissions now carry the shell intent" to the Ctrl+Q /
popNextSubmissionpath. - Either take the one-line
drainQueuecandidate with a test, or file it as a follow-up. Without it, the everyday way of queuing a shell command while a turn runs still sends the command to the model.
中文说明
本地真实环境验证:PR head d46a3b4 对照合并基线 ffea2d0
结论: 修复的两部分在真实 TUI 里都生效了:
- shell 模式提交不再带一次性提醒,且提醒保留到了下一轮模型请求;
- Ctrl+Q 排队的条目按提交时记录的意图路由。
真实环境里有两点影响合并判断,都不必阻塞:
- 模型回复期间按 Enter(页脚写的是 "Enter to steer")完全绕开了本修复。这类条目走 steer drain(
drainQueue),所以两臂上 shell 模式下输入的命令都照样被当成提示词发给模型,即使 shell 模式全程开着。PR 描述确实把 steer drain 列为范围外,但这是轮次进行中最常用的排队方式,而且目前还没有后续 issue。下面附一行候选修复,已在真实 TUI 实测。 - R1-4(Shell-mode commands never hold the session busy: streamingState reads Idle while the command runs, so the queue drain admits a concurrent model turn #12664)在真实 TUI 里复现:Ctrl+Q 混合队列时,shell 命令运行期间,排队的提示词作为并发模型轮次被放行;命令结束后 UI 显示空闲,但模型请求还没结束,Esc 无效,约 20 秒后回复才到。
装置
- 两个 worktree:
d46a3b4(PR head)与ffea2d0(它合入的 merge base)。 - 每臂依次
pnpm install --frozen-lockfile、npm run build、npm run bundle,然后在 PTY 里跑真实的dist/cli.js。 - 用仓库自带的
integration-tests/terminal-capture(node-pty + xterm.js)驱动。 integration-tests/fake-openai-server充当模型:含HOLD_TURN的提示词挂起 20 秒,以便在轮次进行中测队列;记录每个主循环请求体,路由按这份台账加屏幕判定。- 独立
HOME,全新 git 仓库;macOS arm64,Node 22。
场景矩阵(摘要)
- s1(
--worktree预置启动 worktree 通知,与 resume 的通知走同一开关路径):- BASE:bash 报
syntax error near unexpected token 'newline';通知只以失败命令正文的形式进入模型上下文,随后那条模型提示词不带提醒。 - PR:命令正常执行,下一条模型请求末尾带
<system-reminder>[Startup] Active worktree …</system-reminder>。
- BASE:bash 报
- s2q(Shell 模式下 Ctrl+Q,出队前关掉 shell 模式):BASE 发给模型 ❌,PR 在 bash 执行 ✅。
- s3q(Ctrl+Q 普通提示词,出队前打开 shell 模式):BASE 进 bash,报
summarize: command not found❌;PR 发给模型 ✅。 - s2bq(Shell 模式下 Ctrl+Q,shell 模式一直开着):两臂都在 bash 执行,无变化。
- s2 / s2b(Shell 模式下按 Enter 排队;s2 出队前关掉 shell 模式,s2b 一直开着):两臂都作为 steer 发给模型 ❌,本 PR 未覆盖。
- s3(Enter 排队普通提示词,出队前打开 shell 模式):两臂都发给模型,无变化。
- s4q / s5q(Ctrl+Q 排
sleep 8; echo …,再关掉 shell 模式后 Ctrl+Q 一条提示词):- BASE:两条合成一坨发给模型,Esc 可以正常取消。
- PR:shell 执行的同时,提示词作为并发模型轮次启动;命令结束后 UI 显示空闲、Esc 无效,约 20 秒后回复才到(R1-4)。
发现 1:Enter / steer 路径仍把 shell 模式命令发给模型(既有问题,未覆盖)
模型流结束且没有工具调用时,client.ts 的 takeSteerInput() 经 drainSteerAtBoundary 调到 useMessageQueue.drainQueue。后者把所有非 deferUntilIdle 的条目以纯文本取走,不看 shellMode,所以用 Enter 排队的条目根本到不了 popNextSubmission,新加的意图字段不会被读取。s2b 中页脚显示 shell mode enabled · ⏳1 queued,shell 模式始终没关,命令仍被发给模型。
候选修复:drainQueue 的 shouldDrain 加上 && message.shellMode !== true,让 shell 意图条目留给已经能正确路由的队列出队路径。真实 TUI 实测:
- s2、s2b 改为在 bash 执行;
- s3 仍发给模型;
- 三个测试套件 546/546 通过。
现有套件不覆盖这一点,需要补测试。另有一个取舍:在 steer 边界,后排的模型条目可能越过前排的 shell 条目,shell 条目要等轮次结束才执行。可以在本 PR 做,也可以作为后续;如果走后续,请按 bot 的建议开 issue 跟踪。
发现 2:R1-4 / #12664 端到端确认
PR 臂 s5q 按固定时间点探测屏幕:
sleep 8期间,提示词HOLD_TURN summarize the diff作为模型轮次运行(有 spinner 和esc to cancel),旁边是正在运行的Shell Command;- 约 8 秒后命令结束,spinner 消失、输入框显示空闲,但假服务端的请求仍未结束;
- 按 Esc 无变化,约 20 秒后回复
MODEL_REPLY_HOLD_DONE才到。
合并基线把同样的输入合成一坨(本身也路由错了:shell 命令被发给模型),但 Esc 能正常取消。这与 #12664 的归因一致:状态覆盖是既有问题,本 PR 让 Ctrl+Q 混合队列必然走进这个窗口。该 issue 已是 P1 + need-discussion,我认为不阻塞本 PR。
单测红绿验证
在 packages/cli 运行三个套件:
- PR head:546/546 通过;
- 把三个生产文件还原到
ffea2d0、测试保留:27 个失败。包括全部新增的 fix(cli): system-reminder prefixes reach bash in shell mode #11626 用例,以及因断言改为期望addMessage第四个参数而失败的既有用例; - 验证后工作树已还原。
未覆盖
- 恢复 agents 通知没有在真机触发(需要上一会话留下的暂停后台 agent)。它与 s1 验证的 worktree 通知同一门控写法,单测已覆盖。
- 未在本地跑 Linux/Windows;本 head 的 CI 全绿。
建议
可以合并,作为 #11626 提醒部分与 Ctrl+Q 队列路由的修复。两个发现都不是本 PR 所改部分的回归:
- 发现 1 是既有问题,且已声明在范围外;
- 发现 2 已由 Shell-mode commands never hold the session busy: streamingState reads Idle while the command runs, so the queue drain admits a concurrent model turn #12664 跟踪。
合并前或合并后尽快:
- 把描述中「排队提交携带 shell 意图」收窄到 Ctrl+Q /
popNextSubmission路径; - 要么采用这一行
drainQueue候选修复并补测试,要么开后续 issue。否则,轮次进行中最常用的 shell 命令排队方式仍会把命令发给模型。
|
@qwen-code /triage |
main's #12605 recorded shell intent on the same message-queue and submitQuery APIs this branch extended with the producer's system-reminder decomposition. Keep both features: shellMode stays in trunk's positional slot and reminders appends after it, so the drain routes on recorded intent while the restored aggregate still carries its per-member envelope run.




What this PR does
In shell mode,
handleFinalSubmitno longer prepends the two one-shot<system-reminder>prefixes (the recovered-agents notice and the worktree-restore notice) to the submitted text, and no longer consumes their one-shot latches; both stay armed for the next model-bound prompt. Separately, queued submissions now carry the shell intent recorded at submit time:addMessage/restoreMessagesstore it on the queue entry, the drain forwards it assubmitQuerymetadata, and shell routing usesshellModeIntent ?? shellModeActiveinstead of the live flag alone. Queue batches are kept intent-homogeneous — each batch is the contiguous same-intent run from the head of the queue — so a shell command and a model prompt never merge into one blob that can only be routed a single way, and no entry overtakes an earlier one queued with a different intent. Producers that record no intent (e.g. remote input) keep routing on the live flag, as before.Why it's needed
Fixes #11626. In shell mode the input goes to bash, where a leading
<system-reminder>line is a syntax error — and because the one-shot latch was consumed by that submission, the notice never reached any model turn at all. The second gap: the queue drain routed on the liveshellModeActiveflag at drain time, so if the user toggled shell mode while an entry waited in the queue, a shell command could be sent to the model or a model prompt to bash. (The workflow-keyword reminder was already shell-gated viabuildWorkflowKeywordPrefix; this PR covers the two remaining reminders and the drain routing.)Reviewer Test Plan
How to verify
Red-green at the test level, in
packages/cli:With the fix: 3 files, 545 tests passed, 0 failed. With the three source files stashed (tests kept), the new cases fail for exactly the reported reasons: the
<system-reminder>prefix is prepended to shell-mode submissions (gh workflow list,git status) and the recorded shell intent is ignored at drain.tsc --noEmitforpackages/clipasses.New coverage: shell-mode submissions skip both one-shot reminders and leave them armed for the next model-bound prompt; submit-time shell intent overrides the live flag in both directions at drain; producers without recorded intent still route on the live flag; queue batches stay intent-homogeneous and preserve submission order; intent survives an admission-failure restore.
Evidence (Before & After)
N/A — no visual/rendering change. The change alters what text reaches bash in shell mode and how queued entries route; behavior is pinned by the unit tests above.
Tested on
Environment (optional)
N/A — unit tests only.
Risk & Scope
drainQueue) still returns raw text without intent — pre-existing behavior, unchanged by this fix. The direct (non-queued)submitQuerypath passes no intent because it is synchronous, so the live flag is the submit-time intent.<system-reminder>envelope inside the user's own message #11553); whichever lands second may need a small rebase.Linked Issues
Fixes #11626
中文说明
这个 PR 做了什么
在 shell 模式下,
handleFinalSubmit不再给提交文本前置两条一次性<system-reminder>前缀(会话恢复 agents 提示和 worktree 恢复提示),也不再消耗它们的一次性开关——两条提示都会保留到下一个发往模型的输入。另一方面,排队提交现在携带提交时记录的 shell 意图:addMessage/restoreMessages把它存到队列条目上,drain 把它作为submitQuery元数据传递,shell 路由使用shellModeIntent ?? shellModeActive而不是只看实时开关。队列批次保持意图同质(每批取队首起连续同意图的一段),避免 shell 命令和模型输入合并成只能单向路由的一坨,也不会有后排条目越过前排的异意图条目。没有记录意图的生产者(如远程输入)仍按实时开关路由,行为不变。为什么需要
修复 #11626。shell 模式下输入会交给 bash,前导的
<system-reminder>在 bash 里是语法错误;而且一次性开关被这次提交消耗掉后,该提示永远不会到达任何模型轮次。第二个问题:队列 drain 在取出时按实时shellModeActive路由,如果用户在条目排队期间切换了 shell 模式,shell 命令可能被发给模型、模型输入可能被发给 bash。(workflow 关键词提示此前已通过buildWorkflowKeywordPrefix做了 shell 门控;本 PR 覆盖剩余两条提示和 drain 路由。)验证方式
在
packages/cli下做红绿验证:npx vitest run src/ui/AppContainer.test.tsx src/ui/hooks/use-llm-stream.test.tsx src/ui/hooks/useMessageQueue.test.ts。带修复:3 个文件 545 个测试全部通过;stash 掉三个源文件(保留测试)后,新增用例按 issue 报告的确切原因失败——<system-reminder>前缀被加到 shell 模式提交上、记录的 shell 意图在 drain 时被忽略。packages/cli的tsc --noEmit通过。风险与范围
主要权衡:意图同质分批按「队首起连续同意图段」切批,交错队列会多一个模型轮次(被 shell 命令隔开的两条模型输入不再合并),跨意图的提交顺序保持不变。未覆盖:mid-turn steer 的
drainQueue仍返回不带意图的原始文本——属既有行为,本 PR 不变。直接(非排队)submitQuery路径不传意图,因为它是同步的,实时开关即提交时意图。无破坏性变更。与进行中的 PR #11562 触及相同文件,但它是展示层剥离、范围不同(关闭 #11553),后合并的一方可能需要小幅 rebase。关联 issue:修复 #11626。