You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 3e63078
Browse filesBrowse the repository at this point in the historyBrowse files
fix(cli): merge main's Monitor runtime into the takeover close path
Main's H3 background Shell and Monitor runtime (#13265, b2c95e0) and
this branch's owed file-history retirement (ee3183b) both inserted
cleanup steps at the same point in the Session close sequence, right
after the lease release. Keep both: the retirement runs first because it
is guarded by its own try/catch, while the observation-loop stop is not,
so a throw there must not strand the retirement and leave the durable
marker outliving the Session.
|`## Using Your Tools` bullets | A bullet is kept only when **every** tool it names is declared — a bullet that still named a missing tool would send the model after something it cannot call, which is the defect being fixed. Sub-bullets are keyed on the single tool they name, and the "prefer dedicated tools" bullet goes when none of its sub-bullets survive. |
85
+
|`## Using Your Tools` bullets | A bullet is kept only when **every** tool it names is declared — a bullet that still named a missing tool would send the model after something it cannot call, which is the defect being fixed. Sub-bullets are keyed on the single tool they name, and the "prefer dedicated tools" bullet goes when none of its sub-bullets survive. One later exception: bridge reachability satisfies the Agent prerequisite in the two Agent bullets (§4.6); Codebase Search still requires declared `grep_search` and `glob`.|
83
86
|`# Examples` transcripts | A block is kept only when every tool it calls is declared. Blocks are matched as `<example>`/`</example>` pairs, not split on blank lines: an example can contain blank lines of its own, and splitting on them orphans the tool calls in its later paragraphs from the tag that gates them (caught by the test suite while implementing). Every surviving block calls at least one tool, so any of them can be gated, and when none survive the filter the `# Examples` heading is dropped with the section; the model-specific XML and JSON formats use no `[tool_call: …]` notation so they are left ungated. |
|`getActionsSection`, security and safety rules, Core Mandates | Unconditional. A dangerous-action or denied-call clause must never depend on which tools are declared. |
@@ -94,6 +97,21 @@ The prompt depends only on the session-start snapshot, so `setStaticSystemPrefix
94
97
95
98
`collectContextData` builds the prompt through `getMainSessionBaseSystemPrompt` and reads declarations separately, and it does not warm the registry. It therefore reads the same snapshot, which keeps its system-prompt row consistent with the request. This lands on top of the breakdown rework in [#12119](https://github.com/QwenLM/qwen-code/pull/12119) (#12033), whose numbers are the measurement instrument for §7.
96
99
100
+
### 4.6 Agent bridge reachability (#13033)
101
+
102
+
#13033 defers `agent` by default, which made the absolute form of the §4.3
103
+
rule self-defeating: the Subagent Delegation and Codebase Search bullets are
104
+
the policy that sends the model to the deferred-tool bridge to discover Agent,
105
+
so dropping them whenever `agent` is undeclared would turn deferral into
106
+
silent removal. `PromptToolSurface` therefore carries one more input,
107
+
`agentReachable`, set in `startChat` when `agent` is declared or is registered
108
+
behind both bridge halves and listed in the deferred summary. That input
109
+
satisfies only their Agent prerequisite: Codebase Search still requires
110
+
declared `grep_search` and `glob`. No other gated line gains an exception —
111
+
`monitor` stays gated on declaration — and an Agent withheld from the eager
112
+
reveal in an incomplete-bridge session is neither declared nor
113
+
bridge-reachable, so both lines still drop there.
114
+
97
115
## 5. Design decisions
98
116
99
117
| Decision | Rationale | Alternative rejected |
@@ -120,15 +138,15 @@ The prompt depends only on the session-start snapshot, so `setStaticSystemPrefix
120
138
Items 1-4 are automated in this PR's `prompts.test.ts`, so every push re-checks them; item 5 needs a real session and is handed off in [`docs/verification/resident-tool-prompt-assembly/README.md`](../verification/resident-tool-prompt-assembly/README.md).
121
139
122
140
1.**Default-session regression (in CI).** The 17 existing full-prompt snapshots cover the no-snapshot path, and `renders identically when every tool is declared` covers the all-declared path. Together they are the guard that makes the change safe for the common case.
123
-
2.**Effect, and no drift outside it (in CI).** Two tests bracket the saving: a file-work allowlist must drop 900-1,400 characters (measured 1,104, ~276 tokens — policy bullets only, since that allowlist keeps every example), and a narrower allowlist must drop 3,800-5,000 (measured 4,327, ~1,082 tokens, three example blocks included). A third asserts all four model-specific example notations are gated, not just the bracket form. Together they fail on a lost saving and on newly added ungated tool text. `changes nothing outside the two gated sections` strips `## Using Your Tools` and `# Examples` from both renders and asserts the remainder is identical.
124
-
3.**Invariant, both directions (in CI).**`never names an undeclared tool inside the gated sections` sweeps every `ToolNames` value against the gated text with a word-boundary match, and `gates every tool name the gated sections can mention, on every example set` makes that config-independent by withholding each of the 66 names in turn against all four example sets — the check that would have caught the model-specific notations going ungated. `keeps the policy text of every tool that is declared` pins the opposite direction so gating cannot over-reach. Scoped to those sections because of the residues in §6.
125
-
4.**Reverse checks and plumbing (in CI).**`leaves CodeModeOnly guidance untouched by the declared set` asserts code mode renders identically with and without a snapshot, and `takes the declared set from the Config snapshot` asserts `getMainSessionBaseSystemPrompt` reads `Config.getPromptToolSnapshot()` — the property that keeps `/context` and the request on one source.
141
+
2. **Effect, and no drift outside it (in CI).** The file-work bracket is pinned in the two states the reachability exception (§4.6) distinguishes, and the pinned bounds below are exactly the ones `prompts.test.ts` asserts: with `agentReachable` unset the drop must stay within 900-1,500 characters (measured 1,135 — delegation 448 + codebase search 368 + monitor 317 of bullet text, plus dropped line breaks); with `agentReachable: true` — the default in practice, since the bridge pair is exempt from `tools.eager` — both Agent bullets survive and only the monitor policy drops (must stay within 250-450; measured 317). A narrower allowlist drops 3,158 characters (Agent unreachable) or 2,709 (reachable — the default this PR creates), two example blocks included, measured at this commit; no CI bracket pins the narrow case, so those figures are descriptive (the pre-exception measurement was 4,327 with three example blocks). A third test asserts all four model-specific example notations are gated, not just the bracket form. Together they fail on a lost saving and on newly added ungated tool text. `changes nothing outside the two gated sections` strips `## Using Your Tools` and `# Examples` from both renders and asserts the remainder is identical.
142
+
3.**Invariant, both directions (in CI).**Undeclared Agent is permitted only under the reachability exception (§4.6); other named tools must remain declared. `never names an undeclared tool inside the gated sections` sweeps every `ToolNames` value against the gated text with a word-boundary match, and `gates every tool name the gated sections can mention, on every example set` makes that config-independent by withholding each of the 66 names in turn against all four example sets — the check that would have caught the model-specific notations going ungated. `keeps the policy text of every tool that is declared` pins the opposite direction so gating cannot over-reach. Scoped to those sections because of the residues in §6.
143
+
4.**Reverse checks and plumbing (in CI).**`leaves CodeModeOnly guidance untouched by the declared set` asserts code mode renders identically with and without a snapshot, and `keeps Agent guidance when Agent is bridge-reachable` asserts `getMainSessionBaseSystemPrompt` reads the `Config` session snapshots (`getPromptToolSnapshot()`, plus `getPromptAgentReachable()` since #13033) — the property that keeps `/context` and the request on one source.
126
144
5.**Token measurement (handed off).** On a session with a trimmed `tools.eager` allowlist, compare the system-prompt row before and after, anchored on the provider's `input_token_count` (the category ruler itself is being fixed in #12119). The brief also carries the three-way run that separates this change's saving from `tools.eager`'s own, and the weakened recall check that is all the repo's missing eval harness allows.
127
145
128
146
## 8. Acceptance criteria
129
147
130
-
-A default session's base prompt is unchanged, byte for byte.
131
-
- In a trimmed session, no bullet or example names a tool that is not declared, and every declared tool's policy text is still present.
148
+
-No-snapshot and all-declared base prompts are unchanged, byte for byte.
149
+
- In a trimmed session, no bullet or example names a tool the session cannot call — declared, or bridge-reachable for the Agent bullets (§4.6) — and every declared tool's policy text is still present.
132
150
- Safety, permission, and dangerous-action text is present in every configuration.
133
151
-`setStaticSystemPrefix` is written no more often than before this change.
134
152
-`/context`'s system-prompt row and the request's system instruction come from the same snapshot.
0 commit comments