diff --git a/docs/design/2026-09-12-agent-container-execution.md b/docs/design/2026-09-12-agent-container-execution.md index 14d2f51ed77..6222258bb35 100644 --- a/docs/design/2026-09-12-agent-container-execution.md +++ b/docs/design/2026-09-12-agent-container-execution.md @@ -79,8 +79,9 @@ values. The policy composes with `isolation: "worktree"` and `working_dir`. It does not extend the model-visible isolation enum or change `isolation: "remote"`. -Combining it with `tools.codeModeOnly` is rejected before container startup; -the first container registry supports direct tool calls only. +Combining it with `tools.mode: "code_mode_only"` is rejected before container startup; +`tools.mode: "code_mode"` warns and continues with direct tools only, registering +no `exec`. The first container registry supports direct tool calls only. ## Backend policy and definition boundaries diff --git a/docs/design/2026-09-12-agent-container-execution.zh-CN.md b/docs/design/2026-09-12-agent-container-execution.zh-CN.md index 0ca5f028b85..382b7796729 100644 --- a/docs/design/2026-09-12-agent-container-execution.zh-CN.md +++ b/docs/design/2026-09-12-agent-container-execution.zh-CN.md @@ -58,7 +58,8 @@ CLI 重启时保留环境变量值,供 Node TLS 证书、设置插值等启动 该策略可与 `isolation: "worktree"` 和 `working_dir` 组合。 它不扩展面向模型的 isolation 枚举,也不改变 `isolation: "remote"`。 -与 `tools.codeModeOnly` 的组合会在容器启动前被拒绝;首版容器注册表只支持直接工具调用。 +与 `tools.mode: "code_mode_only"` 的组合会在容器启动前被拒绝;`tools.mode: "code_mode"` +会告警并仅以直接工具继续,不注册 `exec`。首版容器注册表只支持直接工具调用。 ## 后端策略与定义边界 diff --git a/docs/design/code-mode-only.md b/docs/design/code-mode-only.md index 05e13337d42..51527d013eb 100644 --- a/docs/design/code-mode-only.md +++ b/docs/design/code-mode-only.md @@ -1,5 +1,9 @@ # CodeModeOnly MVP +> Current behavior: Only discovers schemas through top-level tool_search and invokes tools through exec. Full signatures are included when search is unavailable in the current scope; tools.eager can reduce the initial declaration. Hybrid keeps direct tools plus exec with its existing bridge and eager permission boundaries. The original MVP statements below about hiding tool_search or always including full schemas are superseded by the [lazy-loading design](lazy-code-mode.md). + +[English](code-mode-only.md) | [简体中文](code-mode-only.zh-CN.md) + ## Status Implemented for [#10377](https://github.com/QwenLM/qwen-code/issues/10377). @@ -12,20 +16,20 @@ describe this MVP; `tool_call` stays hidden. ## Goal -Add a `tools.codeModeOnly` setting that replaces the ordinary model-facing -tool surface with one `exec` JavaScript tool plus the small set of tools that -must remain direct control-plane calls. `exec` code can call ordinary tools +Add a `tools.mode: "code_mode_only"` setting that replaces the ordinary +model-facing tool surface with one `exec` JavaScript tool plus the small set of +tools that must remain direct control-plane calls. `exec` code can call ordinary tools through `tools.(args)` without bypassing Qwen Code's validation, permissions, approvals, hooks, telemetry, cancellation, concurrency, or output budgets. -Direct mode is a compatibility boundary: when the setting is false, tool +Direct mode is a compatibility boundary: when `tools.mode` is `direct`, tool registration, deferred-tool behavior, provider requests, and execution remain unchanged. ## Non-goals -- Hybrid direct/code exposure. +- Defining hybrid direct/code exposure; see [Code Mode](code-mode.md). - Persistent cells, globals, or values between `exec` calls. - Background jobs, `wait`, `yield`, `store`, or `load`. - Raw/freeform provider calls. @@ -37,27 +41,27 @@ unchanged. ```json { "tools": { - "codeModeOnly": true + "mode": "code_mode_only" } } ``` -The setting resolves once to the effective `ToolMode` value `direct` or -`code_mode_only`. `ToolRegistry` and the execution surfaces consume that mode. -`exec` is only registered when the setting is enabled, so disabling the setting -also removes it from diagnostics and registry listings. +The setting resolves once to the effective `ToolMode` value. `ToolRegistry` +and the execution surfaces consume that mode. `exec` is only registered when a +code mode is enabled, so selecting `direct` also removes it from diagnostics +and registry listings. ## Exposure policy The registry remains the source of truth. Exposure is a view over registered tools, never a second registry. -| Category | Model top level | `tools.*` inside `exec` | -| ------------------------------------------ | ----------------- | ----------------------- | -| `exec` | CodeModeOnly only | No | -| Direct control | Yes | No | -| Ordinary registered tool | No | Yes | -| Hidden bridge (`tool_search`, `tool_call`) | No | No | +| Category | Model top level | `tools.*` inside `exec` | +| ------------------------------------------ | ------------------------------------- | ----------------------- | +| `exec` | CodeMode and CodeModeOnly | No | +| Direct control | Yes | No | +| Ordinary registered tool | CodeMode only | Yes | +| Hidden bridge (`tool_search`, `tool_call`) | Existing behavior outside strict mode | No | The direct-control allowlist is centralized and deliberately small. It covers user interaction (`ask_user_question`), delegation (`agent`), terminal output @@ -80,7 +84,8 @@ Before each provider tool sync, the `exec` description is generated from the current registry. Tools are sorted by canonical name. A name is normalized to a JavaScript property by replacing invalid identifier characters and prefixing names that begin with a digit. If two canonical names normalize to the same -property, the lexicographically first name wins and one warning names the +property, an exact canonical match wins over rewritten names. If neither is an +exact match, the lexicographically first name wins. The description names the omitted collision. The description defines: @@ -88,9 +93,11 @@ The description defines: - a fresh async JavaScript execution environment; - `tools.(args)` for nested calls; - `ALL_TOOLS`, including canonical and JavaScript names; -- `text(value)`, `image(value)`, `audio(value)`, and `exit()`; +- `text(value)`, `image(value)`, `audio(value)`, `generatedImage(value)`, + `setTimeout(callback, delayMs)`, `clearTimeout(timeoutId)`, and `exit()`; - TypeScript-like signatures generated deterministically from JSON Schema; -- the absence of Node.js, imports, network APIs, timers, and persistent state. +- the absence of Node.js, `process`, `require`, filesystem, network, imports, + `console`, `WebAssembly`, `Atomics`, and persistent state. Pending timers do not keep `exec` alive by themselves. The nested call returns a JSON-safe object containing the real call id, tool name, status, output, and structured content. Failed and cancelled calls reject @@ -105,8 +112,8 @@ configuration. The parent maps JavaScript names back to canonical registry names and dispatches each call. The guest has no Node globals, `require`, `process`, filesystem, sockets, -module loader, `console`, timers, `Atomics`, `SharedArrayBuffer`, or -`WebAssembly`. Dynamic and static imports fail because no module loader is +module loader, `console`, `Atomics`, `SharedArrayBuffer`, or `WebAssembly`. +Dynamic and static imports fail because no module loader is installed. Runtime memory and stack limits are fixed. QuickJS's interrupt hook enforces a guest CPU budget. That budget and the parent's fallback watchdog pause while the guest is suspended on registered host tools, whose own @@ -162,14 +169,13 @@ OpenAI-compatible, and Anthropic adapters all receive the structured `exec` declaration without provider-specific prompting. Filtered subagent declarations apply the same policy. For a read-only teammate -or a fork with an execution allowlist, `exec` is the audited gateway while the -exact allowed nested names are carried in its invocation context. The same set -generates the description and is checked again before Core dispatch, so an -explicit allowlist can narrow code-mode-callable nested tools without becoming -prompt-only policy, exposing a hidden bridge, or making `exec` recursive. -For cache-compatible forks, an inherited `exec` declaration represents its -ordinary bindings: an omitted `fork_tools` inherits them, while an explicit -list replaces them with the requested subset. +or a fork with an execution allowlist, `exec` is the audited gateway and the +nested names carried in its invocation context are checked again before Core +dispatch. Explicit ordinary-tool entries narrow that nested set. An inherited +or explicitly allowed `exec` instead represents all surviving ordinary +code-mode-callable bindings, while the agent's own `tools` list still narrows +its direct surface. Hidden bridges remain unavailable and `exec` cannot call +itself. ## Failure and rollback @@ -178,9 +184,9 @@ closed before scheduling. Invalid arguments continue to fail in the normal execution chain. A sandbox startup, protocol, timeout, memory, or teardown failure becomes an `exec` tool error. -Rollback is setting `tools.codeModeOnly` to false. No session migration or -registry cleanup is required because code mode has no persistent state and the -ordinary registry was never replaced. +Rollback is setting `tools.mode` to `direct`. No session migration or registry +cleanup is required because code mode has no persistent state and the ordinary +registry was never replaced. ## Verification diff --git a/docs/design/code-mode-only.zh-CN.md b/docs/design/code-mode-only.zh-CN.md new file mode 100644 index 00000000000..7238c2f8047 --- /dev/null +++ b/docs/design/code-mode-only.zh-CN.md @@ -0,0 +1,172 @@ +# CodeModeOnly MVP + +> 当前行为:Only 模式通过顶层 tool_search 按需加载 schema,并通过 exec 调用。搜索在当前范围不可用时才提供完整签名;tools.eager 可缩小初始声明。Hybrid 继续使用直接工具加 exec,保留其 bridge 与 eager 权限边界。本文以下 MVP 中“隐藏 tool_search / 始终完整 schema”的旧约定已由 [延迟加载设计](lazy-code-mode.zh-CN.md) 取代。 + +[English](code-mode-only.md) | [简体中文](code-mode-only.zh-CN.md) + +## 状态 + +已为 [#10377](https://github.com/QwenLM/qwen-code/issues/10377) 实现。 +该功能为可选功能,默认关闭。 + +## 目标 + +新增 `tools.mode: "code_mode_only"` 设置,用一个 `exec` JavaScript 工具和少量 +必须保留为直接调用的控制面工具,取代面向模型的普通工具面。`exec` 中的代码可通过 +`tools.(args)` 调用普通工具,同时不会绕过 Qwen Code 的校验、权限、审批、 +hook、遥测、取消、并发或输出预算。 + +直接模式是兼容性边界:当 `tools.mode` 为 `direct` 时,工具注册、延迟工具行为、 +provider 请求和执行均保持不变。 + +## 非目标 + +- 定义混合的直接调用和代码调用;详见 [Code Mode](code-mode.zh-CN.md)。 +- 在多次 `exec` 调用间持久化 cell、全局变量或值。 +- 后台任务、`wait`、`yield`、`store` 或 `load`。 +- 原始/freeform provider 调用。 +- 在 code mode 中提供 `tool_search` 或 `tool_call` bridge。 +- 提供兼容 Node.js 的 sandbox。 + +## 配置 + +```json +{ + "tools": { + "mode": "code_mode_only" + } +} +``` + +该设置会解析一次,得到有效的 `ToolMode` 值;`ToolRegistry` 和各执行面使用这一 +模式。只有启用某个 code mode 时才注册 `exec`,因此选择 `direct` 也会将它从 +诊断信息和 registry 列表中移除。 + +## 暴露策略 + +registry 仍是事实来源。暴露只是注册工具之上的视图,而不是第二套 registry。 + +| 类别 | 模型顶层调用 | `exec` 内的 `tools.*` | +| ----------------------------------------- | ------------------------ | --------------------- | +| `exec` | CodeMode 和 CodeModeOnly | 否 | +| 直接控制工具 | 是 | 否 | +| 已注册的普通工具 | 仅 CodeMode | 是 | +| 隐藏 bridge(`tool_search`、`tool_call`) | 严格模式之外沿用既有行为 | 否 | + +直接控制 allowlist 集中维护且刻意保持精简。它覆盖用户交互 +(`ask_user_question`)、委派(`agent`)、终止输出约定、 +plan/goal/task/team/worktree/session 控制,以及生命周期无法安全隐藏在解释执行程序 +后的 ACP host 控制。新工具默认可在 code mode 中调用;增加仅直接调用或隐藏工具时, +必须显式修改策略。 + +延迟工具保持注册状态和延迟 registry 状态,并可从 `exec` 调用。生成的 `exec` +描述仍会携带这些工具的完整 schema,以及它们在 `ALL_TOOLS` 中的名称和描述: +CodeModeOnly 会隐藏 `tool_search`,嵌套调用也不会以 `functionCall` 出现在历史记录 +中,因此后续 reveal 无法补充描述里遗漏的 schema。CodeModeOnly 会跳过延迟预加载 +和 ToolSearch 提醒,因为二者都不属于它面向模型的协议。 + +## 确定性的 JavaScript 接口 + +每次向 provider 同步工具前,都会从当前 registry 生成 `exec` 描述。工具按规范 +名称排序。名称会通过替换无效标识符字符来规范化为 JavaScript 属性;如果名称以 +数字开头,还会添加前缀。如果两个规范名称映射到同一个属性,优先保留与属性精确 +一致的规范名称;若都不是精确匹配,则字典序靠前的名称胜出。描述会指出被省略的 +冲突项。 + +描述会定义: + +- 全新的异步 JavaScript 执行环境; +- 用于嵌套调用的 `tools.(args)`; +- 包含规范名称和 JavaScript 名称的 `ALL_TOOLS`; +- `text(value)`、`image(value)`、`audio(value)`、`generatedImage(value)`、 + `setTimeout(callback, delayMs)`、`clearTimeout(timeoutId)` 和 `exit()`; +- 从 JSON Schema 确定性生成的类 TypeScript 签名; +- 不提供 Node.js、`process`、`require`、文件系统、网络、import、`console`、 + `WebAssembly`、`Atomics` 和持久状态。待处理的 timer + 本身不会让 `exec` 保持运行。 + +嵌套调用返回一个 JSON-safe 对象,其中包含真实 call id、工具名、状态、输出和 +structured content。调用失败或取消时,guest promise 会使用 scheduler/ACP 错误 +reject。 + +## Sandbox 与传输 + +`exec` 在独立子进程中运行编译为 WebAssembly 的 QuickJS。每次调用都会创建全新 +的 QuickJS runtime 和 context。子进程通过 stdio 接收精简的 framed JSON 协议; +其中不包含工具实现或 Qwen 配置。父进程把 JavaScript 名称映射回规范 registry +名称,并分派每次调用。 + +guest 不提供 Node 全局变量、`require`、`process`、文件系统、socket、模块加载器、 +`console`、`Atomics`、`SharedArrayBuffer` 或 `WebAssembly`。由于没有安装 +模块加载器,动态和静态 import 都会失败。runtime 的内存和 stack 限制固定。 +QuickJS 的 interrupt hook 会限制 guest CPU 预算。当 guest 挂起等待已注册的 host +工具时,该预算和父进程的兜底 watchdog 会暂停;在 guest job 再次运行前恢复。 +因此,长时间 build 可以继续使用工具声明的 timeout,同时 guest CPU 死循环无法 +逃逸固定预算。源码、协议 frame、helper 输出和最终结果都有上限。 + +取消操作会中止所有嵌套调用并终止子进程。顶层 promise settled 后,也会在 teardown +前取消未 await 的嵌套调用。子进程、timer、promise handle 或 guest 全局变量都不 +会在调用结束后存活。 + +子进程只接收最小化且经过清理的环境,guest 无法检查该环境。独立进程为解释器故障 +提供纵深防御;QuickJS/WASM 是 guest 的 capability 边界。 + +## 重入式分派 + +`CoreToolScheduler.schedule()` 无法递归调用:子调用会排在仍在运行的父调用之后, +从而形成死锁。因此,scheduler 会在 invocation 边界绑定 async-local +`ToolCallRuntime` context,`exec` 只与该 context 通信。 + +scheduler runtime 会将同一 event-loop turn 中收到的嵌套调用合并为一批,并通过 +使用相同 `Config` 和 observer 配置的 sibling scheduler 执行。这样无需直接调用 +`tool.execute()`,仍能沿用现有的构建/校验、权限、确认、hook、执行、截断、遥测和 +并发链路。guest 的连续 await 会生成连续 batch;`Promise.all` 调用会进入同一个 +batch,由现有的只读并发分类器处理。嵌套 request id 包含父 id,并携带 +`source: code_mode` 和 `parentCallId`。 + +嵌套 scheduler 更新会合并到所属 scheduler 的可见调用中,嵌套 id 的确认响应也会 +委派给它。外层模型仍只接收完整的 `exec` 响应。 + +ACP 沿用其独立审计过的执行链。它的 invocation 边界会绑定相同的 runtime 接口, +嵌套分派通过 `Session.runTool` 重入。因此,ACP 顶层和嵌套调用使用相同的 ACP 权限、 +审批、hook、遥测、持久化和取消机制,而不是借用 CLI scheduler 状态或直接执行工具。 +ACP 会串行执行普通嵌套调用,与现有直接工具顺序一致;Core scheduler 则保留现有的 +安全只读并行 batch。 + +## Provider 行为 + +所有 provider 继续使用 `ToolRegistry` 提供的 `FunctionDeclaration[]`。在 Direct +模式下,该数组与现有视图逐字节一致。在 CodeModeOnly 中,它是暴露策略生成的视图, +因此 Gemini/Qwen、OpenAI-compatible 和 Anthropic adapter 都会收到结构化的 +`exec` declaration,不需要 provider 专用 prompt。 + +经过过滤的子智能体 declaration 使用相同策略。对于只读 teammate 或带执行 +allowlist 的 fork,`exec` 是经过审计的 gateway,其 invocation context 中携带的嵌套 +名称会在 Core 分派前再次校验。显式列出的普通工具会收窄该嵌套集合;继承或显式允许 +的 `exec` 则代表所有仍可用的普通 code-mode-callable binding,而智能体自身的 +`tools` 列表仍会收窄直接调用面。隐藏 bridge 始终不可用,`exec` 也不能调用自身。 + +## 失败与回滚 + +未知、冲突、隐藏、仅直接调用和递归请求的工具都会在调度前 fail closed。无效参数 +继续在正常执行链中失败。sandbox 启动、协议、timeout、内存或 teardown 失败会转为 +`exec` 工具错误。 + +回滚方式是将 `tools.mode` 设为 `direct`。无需迁移 session 或清理 registry, +因为 code mode 没有持久状态,且普通 registry 从未被替换。 + +## 验证 + +单元和集成测试必须覆盖: + +- Direct 和 CodeModeOnly 暴露、延迟保留、冲突处理、确定性描述和显式子智能体过滤。 +- 有效/无效 JavaScript、异步顺序、`Promise.all`、helper 输出、throw error、CPU + 死循环、内存限制、import、不可用全局变量、隔离、输出限制、取消、未 await 工作和 + 递归调用。 +- 嵌套权限、确认、Pre/Post hook、失败、输出预算、真实名称遥测/UI 更新、MCP 工具 + 和 scheduler 死锁回归。 +- Gemini/Qwen、OpenAI-compatible、Anthropic、headless、interactive、subagent + 和 ACP 调用面。 + +只有通过相关 package 测试、build、typecheck、E2E probe 和两轮干净的完整 diff +自审,才算实现完成。 diff --git a/docs/design/code-mode.md b/docs/design/code-mode.md new file mode 100644 index 00000000000..eda51ddaceb --- /dev/null +++ b/docs/design/code-mode.md @@ -0,0 +1,121 @@ +# Code Mode + +> Current behavior: Only discovers schemas through top-level tool_search and invokes tools through exec. Full signatures are included when search is unavailable in the current scope; tools.eager can reduce the initial declaration. Hybrid keeps direct tools plus exec with its existing bridge and eager permission boundaries. The original MVP statements below about hiding tool_search or always including full schemas are superseded by the [lazy-loading design](lazy-code-mode.md). + +[English](code-mode.md) | [简体中文](code-mode.zh-CN.md) + +## Status + +Implemented. The feature is experimental and opt-in. + +## Problem + +Qwen Code currently supports direct tool calls and `CodeModeOnly`. The latter +replaces ordinary top-level tools with `exec`, which is useful for orchestration +but prevents a model from choosing direct calls when they are clearer or more +efficient. Codex models this as three tool modes: direct, hybrid Code Mode, and +CodeModeOnly. + +## Goal + +Add a hybrid `CodeMode` option while preserving the direct default and the +strict `CodeModeOnly` exposure policy. Both code modes share the corrected +`exec` failure guidance and exact-name collision precedence described below. + +## Configuration + +```json +{ + "tools": { + "mode": "code_mode" + } +} +``` + +The effective modes are: + +| `tools.mode` value | Effective mode | +| ------------------ | ---------------- | +| omitted / `direct` | `direct` | +| `code_mode` | `code_mode` | +| `code_mode_only` | `code_mode_only` | + +This follows Codex's `ToolMode` serialized values. Safe mode and bare mode +force `direct`. Container execution exposes direct tools only: selecting +`code_mode` emits a warning and provides no `exec`; `code_mode_only` is rejected. + +## Exposure policy + +| Surface | `direct` | `code_mode` | `code_mode_only` | +| -------------------- | --------------------------- | ------------------------------- | ---------------------------------- | +| Ordinary eager tools | Direct | Direct and nested | Nested only | +| Deferred tools | `tool_search` + `tool_call` | Bridge and nested | Nested only | +| Direct-control tools | Direct | Direct only | Direct only | +| `exec` | Not registered | Direct | Direct | +| Discovery bridge | Existing direct behavior | Direct where already applicable | Top-level search; tool_call hidden | + +In `code_mode`, ordinary visible tool descriptions gain an `exec` declaration +for that tool. The `exec` description keeps every `ALL_TOOLS` name, JavaScript name and +deferral flag. On the session surface with both bridge tools available, deferred descriptions are +omitted from this metadata, and `tool_search` +returns deferred parameter schemas without changing the top-level declarations. +Match each returned schema name exactly to `ALL_TOOLS.name`, then invoke +`tools[entry.jsName]` with arguments shaped by that schema. If there is no entry, +do not normalize or guess a binding; use `tool_call` outside `exec`, or an +available direct tool, subject to normal validation and approval. + +When either bridge half is absent or the agent surface is filtered, `exec` +includes signatures for nested tools whose schemas are otherwise unavailable. +In `code_mode_only`, top-level `tool_search` discovers deferred schemas and +`exec` calls the tools. Only `tool_call` is hidden. Deferred reminders, budget +preload, and the incomplete-bridge warning are skipped. `tools.eager` and +`tools.visible` control which signatures appear initially; if search is absent +from the current scope, all allowed signatures are included. Deferred tools +remain callable through `exec`. Hybrid AgentCore surfaces exclude tools hidden +by `tools.eager` from nested bindings; Only retains these deferred targets. +Both modes apply the agent allowlist rules below. + +Nested bindings prefer an exact canonical JavaScript name over names rewritten +to that property. Other collisions retain canonical-name ordering; omitted +bindings remain absent from nested signatures and are never advertised as +reachable through `exec`. + +Filtered subagent declarations preserve the same mode. The agent's `tools` +list narrows its direct surface. Agent allowlists that do not grant `exec` +narrow nested bindings. Inheriting or explicitly granting `exec` keeps all +otherwise admitted ordinary code-mode-callable bindings. An execution +allowlist that mentions any MCP tool additionally restricts MCP bindings to +matching exact names or server patterns. Forks inherit the parent's direct-call +bound separately from its nested binding set; nested access never widens +the child's direct-call grant. Both bounds persist through background resume. + +## Constraints and risks + +- Normal direct declarations must remain byte-for-byte unchanged in `direct`. +- Nested calls continue through the existing scheduler or ACP execution path; + the mode must not bypass validation, permissions, hooks, cancellation, or + telemetry. +- Duplicating every schema in hybrid mode would increase prompt size, so only + the per-tool nested declaration is appended there. +- A denied or failed nested call rejects its promise. An uncaught rejection + aborts `exec`; catching an expected rejection lets the program continue. Keep + calls that may be refused out of a batch that would need to be repeated. + +## Validation + +- Verify mode resolution and safe/bare-mode fallback. +- Verify direct, hybrid, and CodeModeOnly declaration surfaces. +- Verify filtered subagent declarations and nested allowlists. +- Run focused Core and CLI tests, then build and typecheck. + +## Acceptance criteria + +- `tools.mode: "code_mode"` registers `exec` while retaining ordinary direct + tools and deferred discovery. +- Ordinary visible tools advertise their nested JavaScript signature. +- `tools.mode: "code_mode_only"` selects the stricter behavior. +- The default, safe-mode, and bare-mode surfaces remain direct-only. +- Both code modes dispatch an exact canonical name to that tool under a + normalization collision and describe caught versus uncaught failures accurately. +- Context accounting charges only emitted declarations; container fallback and + withheld-tool warnings describe the actual available bindings. diff --git a/docs/design/code-mode.zh-CN.md b/docs/design/code-mode.zh-CN.md new file mode 100644 index 00000000000..12f180cca2c --- /dev/null +++ b/docs/design/code-mode.zh-CN.md @@ -0,0 +1,108 @@ +# Code Mode + +> 当前行为:Only 模式通过顶层 tool_search 按需加载 schema,并通过 exec 调用。搜索在当前范围不可用时才提供完整签名;tools.eager 可缩小初始声明。Hybrid 继续使用直接工具加 exec,保留其 bridge 与 eager 权限边界。本文以下 MVP 中“隐藏 tool_search / 始终完整 schema”的旧约定已由 [延迟加载设计](lazy-code-mode.zh-CN.md) 取代。 + +[English](code-mode.md) | [简体中文](code-mode.zh-CN.md) + +## 状态 + +已实现。该功能为实验性功能,默认关闭。 + +## 问题 + +Qwen Code 目前支持直接工具调用和 `CodeModeOnly`。后者会用 `exec` 替换普通 +顶层工具,适合编排调用,但模型无法在直接调用更清晰或更高效时选择直接调用。 +Codex 将这套能力建模为三种工具模式:直接模式、混合 Code Mode 和 +CodeModeOnly。 + +## 目标 + +新增混合 `CodeMode`,同时保持直接调用默认模式和 `CodeModeOnly` 的严格暴露策略。 +两种代码模式共享下述修正后的 `exec` 失败说明和精确名称碰撞优先级。 + +## 配置 + +```json +{ + "tools": { + "mode": "code_mode" + } +} +``` + +有效模式如下: + +| `tools.mode` 值 | 有效模式 | +| ---------------- | ---------------- | +| 省略 / `direct` | `direct` | +| `code_mode` | `code_mode` | +| `code_mode_only` | `code_mode_only` | + +这些值与 Codex 的 `ToolMode` 序列化值一致。安全模式和 bare 模式会强制使用 +`direct`。容器执行仅支持直接工具:选择 `code_mode` 会显示警告且不提供 `exec`; +选择 `code_mode_only` 则会报错。 + +## 暴露策略 + +| 调用面 | `direct` | `code_mode` | `code_mode_only` | +| ------------ | --------------------------- | -------------------- | ------------------------ | +| 普通即时工具 | 直接调用 | 直接调用和嵌套调用 | 仅嵌套调用 | +| 延迟工具 | `tool_search` + `tool_call` | 桥接和嵌套调用 | 仅嵌套调用 | +| 直接控制工具 | 直接调用 | 仅直接调用 | 仅直接调用 | +| `exec` | 未注册 | 直接调用 | 直接调用 | +| 发现桥接工具 | 保持现有直接模式行为 | 在原本适用时直接调用 | 顶层搜索;隐藏 tool_call | + +在 `code_mode` 中,普通可见工具的描述会附加该工具的 `exec` 调用声明。 +`exec` 描述保留每个 `ALL_TOOLS` 条目的工具名、JavaScript 名称和延迟标记。 +在会话调用面,两个桥接工具都可用时,元数据不包含延迟工具的完整描述;`tool_search` 按需返回 +其参数 schema,不改变顶层声明列表。 +用返回的 schema 名称精确匹配 `ALL_TOOLS.name`,再按该 schema 构造参数调用 +`tools[entry.jsName]`。没有匹配项时不能自行规范化或猜测绑定名称;应在 `exec` +外使用 `tool_call`,或调用可用的直接工具,沿用正常的参数校验和审批。 + +任一桥接工具缺失或智能体调用面经过过滤时,`exec` 会为无法从其他途径取得 +schema 的嵌套工具附上签名。在 `code_mode_only` 中,顶层 `tool_search` 按需发现 +延迟 schema,`exec` 负责调用;只有 `tool_call` 被隐藏。延迟提醒、预算预载与 +桥接不完整的警告被跳过。`tools.eager` 和 `tools.visible` 控制初始签名,当前范围 +不能搜索时则提供全部允许工具的签名。延迟工具仍可通过 `exec` 调用。 +Hybrid 的 AgentCore 调用面从嵌套绑定中排除被 `tools.eager` 隐藏的工具;Only +保留这些延迟目标。两种模式均应用下文的智能体 allowlist 规则。 + +嵌套绑定优先保留与 JavaScript 属性精确一致的规范名称,再考虑改写为该属性的 +名称。其他碰撞沿用规范名称字典序;被省略的绑定不出现在嵌套签名中, +也不会被警告描述为可通过 `exec` 调用。 + +经过过滤的子智能体声明沿用相同模式。智能体的 `tools` 列表会收窄直接调用面。 +未授权 `exec` 的智能体 allowlist 会收窄嵌套绑定;继承或显式允许 `exec` 时, +会保留所有通过其他准入检查的普通 code-mode-callable binding。若执行 allowlist +包含任一 MCP 工具项,MCP 绑定还必须匹配其中的精确工具名或服务器模式。 +Fork 分别继承父级的直接调用边界与嵌套绑定集合;嵌套访问权限不会扩大子级 +的直接调用授权。后台恢复会保留这两个边界。 + +## 约束与风险 + +- `direct` 模式下的普通工具声明必须保持完全不变。 +- 嵌套调用继续经过现有 scheduler 或 ACP 执行链,不能绕过校验、权限、hook、 + 取消和遥测。 +- 混合模式若重复全部 schema 会增大提示词,因此只在各顶层工具描述中附加对应的 + 嵌套声明。 +- 被拒绝或失败的嵌套调用会使 promise reject。未捕获的 rejection 会中止 `exec`; + 捕获预期 rejection 后可以继续执行。可能被拒绝的调用不应加入随后必须整体重试 + 的批次。 + +## 验证 + +- 验证模式解析以及安全模式/bare 模式回退。 +- 验证直接、混合和 CodeModeOnly 三种声明面。 +- 验证子智能体过滤后的声明和嵌套 allowlist。 +- 运行 Core 和 CLI 的相关测试,然后执行构建和类型检查。 + +## 验收标准 + +- `tools.mode: "code_mode"` 注册 `exec`,同时保留普通直接工具和延迟发现。 +- 普通可见工具会声明其嵌套 JavaScript 调用签名。 +- `tools.mode: "code_mode_only"` 选择严格模式。 +- 默认模式、安全模式和 bare 模式仍然只使用直接调用。 +- 两种代码模式在规范化碰撞时都将精确规范名称分派给该工具,并准确区分已捕获与 + 未捕获失败的行为。 +- 上下文计数仅计入实际发送的声明;容器回退和被保留工具(withheld)的警告与实际可用绑定一致。 diff --git a/docs/design/freeform-exec-experiment.md b/docs/design/freeform-exec-experiment.md index 32d740fca8d..de48e5207b5 100644 --- a/docs/design/freeform-exec-experiment.md +++ b/docs/design/freeform-exec-experiment.md @@ -13,7 +13,7 @@ evaluation inside exec remains. ## Design `tools.freeform` opts into a Responses custom tool with text input for exec. -The setting is effective only when `tools.codeModeOnly` is true and the current +The setting is effective only when `tools.mode` is `code_mode_only` and the current model uses `wireApi: "responses"`. Other tools and the default behavior remain function tools. The Responses converter maps completed custom input to internal `{ source }`, leaving scheduler validation, permissions, runtime execution, @@ -35,7 +35,7 @@ types, converter, pipeline, and their tests. There is no automatic capability detection, fallback, grammar constraint, or change to other provider adapters. The caller enables the setting only on an endpoint that supports [Responses Custom Tools](https://developers.openai.com/api/docs/guides/function-calling#custom-tools). -When `tools.codeModeOnly` is false, `tools.freeform` is stored but has no effect. +When `tools.mode` is `direct` or `code_mode`, `tools.freeform` is stored but has no effect. ## Validation and acceptance diff --git a/docs/design/freeform-exec-experiment.zh-CN.md b/docs/design/freeform-exec-experiment.zh-CN.md index a21c6cdfbdd..ecc21e91378 100644 --- a/docs/design/freeform-exec-experiment.zh-CN.md +++ b/docs/design/freeform-exec-experiment.zh-CN.md @@ -12,7 +12,7 @@ Code mode 当前将 exec 暴露为接收含 `source` 字段的 JSON 对象的函 ## 设计 通过 `tools.freeform` 将 exec 切换为接收文本的 Responses custom tool。 -该配置只有在 `tools.codeModeOnly` 为 true 且当前模型的 `wireApi` 为 +该配置只有在 `tools.mode` 为 `code_mode_only` 且当前模型的 `wireApi` 为 `"responses"` 时生效。其他工具和默认行为仍使用 function tool。 Responses 转换器将完成的 custom 输入映射为内部 `{ source }`,保留调度器校验、 权限、运行时执行、工具调用显示和记录方式。启用时,适配器将 exec 历史回传为配对的 @@ -30,7 +30,7 @@ Responses 转换器将完成的 custom 输入映射为内部 `{ source }`,保 用户只应在确认接口支持 [Responses Custom Tools](https://developers.openai.com/api/docs/guides/function-calling#custom-tools) 时启用配置。 -当 `tools.codeModeOnly` 为 false 时,`tools.freeform` 会被保存但不生效。 +当 `tools.mode` 为 `direct` 或 `code_mode` 时,`tools.freeform` 会被保存但不生效。 ## 验证与验收 diff --git a/docs/design/lazy-code-mode.md b/docs/design/lazy-code-mode.md index 0490f51b850..a944ab514ca 100644 --- a/docs/design/lazy-code-mode.md +++ b/docs/design/lazy-code-mode.md @@ -4,14 +4,14 @@ ## Problem and scope -CodeModeOnly currently puts every callable tool's signature and description in +Before this change, CodeModeOnly put every callable tool's signature and description in the initial `exec` declaration, including deferred tools. This removes the prompt-size benefit of deferral. The existing `tool_call` bridge also describes target arguments as a generic object; issue #12889 reports repeated empty arguments with Responses providers. Code Mode lets the model express these arguments as JavaScript, with ordinary runtime validation still enforced. -This change makes the existing experimental `tools.codeModeOnly` mode load +This change makes the existing experimental `tools.mode: "code_mode_only"` mode load deferred descriptions and schemas through `tool_search`. Direct mode keeps its current protocol. This extends the original CodeModeOnly MVP's scope. diff --git a/docs/design/lazy-code-mode.zh-CN.md b/docs/design/lazy-code-mode.zh-CN.md index 266f7f4d2fa..b380c6e3564 100644 --- a/docs/design/lazy-code-mode.zh-CN.md +++ b/docs/design/lazy-code-mode.zh-CN.md @@ -4,9 +4,9 @@ ## 问题与范围 -CodeModeOnly 目前把所有可调用工具的签名和描述写入初始 `exec` 声明,包括 deferred 工具,因此失去了延迟加载节省提示词的效果。现有 `tool_call` 桥接层还把目标参数描述为通用对象;issue #12889 记录了 Responses provider 下反复生成空参数的问题。Code Mode 允许模型通过 JavaScript 表达参数,并保留正常的运行时校验。 +本次修改前,CodeModeOnly 把所有可调用工具的签名和描述写入初始 `exec` 声明,包括 deferred 工具,因此失去了延迟加载节省提示词的效果。现有 `tool_call` 桥接层还把目标参数描述为通用对象;issue #12889 记录了 Responses provider 下反复生成空参数的问题。Code Mode 允许模型通过 JavaScript 表达参数,并保留正常的运行时校验。 -本次修改让现有实验性 `tools.codeModeOnly` 模式通过 `tool_search` 按需加载 deferred 工具的描述和 schema。Direct 模式保持现有协议。这扩展了最初 CodeModeOnly MVP 的范围。 +本次修改让现有实验性 `tools.mode: "code_mode_only"` 模式通过 `tool_search` 按需加载 deferred 工具的描述和 schema。Direct 模式保持现有协议。这扩展了最初 CodeModeOnly MVP 的范围。 ## 设计 diff --git a/docs/developers/sdk-typescript.md b/docs/developers/sdk-typescript.md index f34bac68d45..8e421cedc63 100644 --- a/docs/developers/sdk-typescript.md +++ b/docs/developers/sdk-typescript.md @@ -53,27 +53,27 @@ Creates a new query session with the Qwen Code. #### QueryOptions -| Option | Type | Default | Description | -| ------------------------ | -------------------------------------------------------------------- | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| `cwd` | `string` | `process.cwd()` | The working directory for the query session. Determines the context in which file operations and commands are executed. | -| `model` | `string` | - | The AI model to use (e.g., `'qwen-max'`, `'qwen-plus'`, `'qwen-turbo'`). Takes precedence over `OPENAI_MODEL` and `QWEN_MODEL` environment variables. | -| `pathToQwenExecutable` | `string` | Bundled CLI | Path to the Qwen Code executable. Supports multiple formats: `'qwen'` (native binary from PATH), `'/path/to/qwen'` (explicit path), `'/path/to/cli.js'` (Node.js bundle), `'node:/path/to/cli.js'` (force Node.js runtime), `'bun:/path/to/cli.js'` (force Bun runtime). If not provided, the SDK uses the bundled CLI included with the package. | -| `permissionMode` | `'default' \| 'plan' \| 'auto-edit' \| 'auto' \| 'yolo'` | `'default'` | Permission mode controlling tool execution approval. See [Permission Modes](#permission-modes) for details. | -| `canUseTool` | `CanUseTool` | - | Custom permission handler for tool execution approval. Invoked when a tool requires confirmation. Must respond within 60 seconds or the request will be auto-denied. See [Custom Permission Handler](#custom-permission-handler). | -| `env` | `Record` | - | Environment variables to pass to the Qwen Code process. Merged with the current process environment. | -| `systemPrompt` | `string \| QuerySystemPromptPreset` | - | System prompt configuration for the main session. Use a string to fully override the built-in Qwen Code system prompt, or a preset object to keep the built-in prompt and append extra instructions. | -| `mcpServers` | `Record` | - | MCP (Model Context Protocol) servers to connect. Supports external servers (stdio/SSE/HTTP) and SDK-embedded servers. External servers are configured with transport options like `command`, `args`, `url`, `httpUrl`, etc. SDK servers use `{ type: 'sdk', name: string, instance: Server }`. | -| `abortController` | `AbortController` | - | Controller to cancel the query session. Call `abortController.abort()` to terminate the session and cleanup resources. | -| `debug` | `boolean` | `false` | Enable debug mode for verbose logging from the CLI process. | -| `maxSessionTurns` | `number` | `-1` (unlimited) | Maximum number of conversation turns before the session automatically terminates. Must be an integer. A turn consists of a user message and an assistant response. | -| `coreTools` | `string[]` | - | Uses the legacy `coreTools` / CLI `--core-tools` allowlist semantics. If specified, only matching core tools are registered for the session. This is the only allowlist-style option that restricts built-in tool registration; a whole-tool `permissions.deny` / `excludeTools` rule (and `tools.disabled` in settings.json) also removes a tool from the registry. `permissions.allow` in settings.json is pure auto-approval and never removes, demotes, or hides a tool (#10075). To keep a tool's schema out of the initial model request, use `tools.eager` in settings.json (requires restart, #9827) — `tool_search`, `tool_call`, `structured_output`, plan-mode lifecycle tools, `task_stop`, `mcp__*` and `computer_use__*` tools are exempt from that allowlist and keep their normal loading; tools demoted this way stay registered and reachable through `tool_search` + `tool_call` while both bridge tools are registered — when either is unregistered (`tools.toolSearch.enabled: false` denies both; a `tool_search` or `tool_call` deny rule, or a `tools.disabled` entry removes one) the demoted tools that remain hidden are not offered to the model and cannot be reached through the bridge for that session, and a warning is written to the CLI process's stderr (SDK forwards it only with piped stderr and effective `debug` logging; an explicit `logLevel` always wins); they stay registered, so a direct call by their own name is still evaluated and approved normally — except tools also listed in `tools.visible`, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares.; to remove a tool entirely, use a whole-tool `excludeTools` / `permissions.deny` rule — a rule with a specifier (such as `'Bash(rm *)'`) only denies matching invocations at runtime. MCP tools are exempt from deny-based removal: hide them with the per-server `excludeTools` / `tools.disabled` filters instead (deny still blocks their calls at runtime). Example: `['read_file', 'edit', 'run_shell_command']`. | -| `excludeTools` | `string[]` | - | Equivalent to `permissions.deny` in settings.json. Excluded tools return a permission error immediately. Takes highest priority over all other permission settings. Supports tool name aliases and pattern matching: tool name (`'write_file'`), shell command prefix (`'Bash(rm *)'`), or path patterns (`'Read(.env)'`, `'Edit(/src/**)'`). | -| `allowedTools` | `string[]` | - | Equivalent to `permissions.allow` in settings.json for auto-approval. Matching tools bypass `canUseTool` callback and execute automatically. Only applies when tool requires confirmation. Like `permissions.allow`, this is pure auto-approval and never affects which tools are registered or which schemas are sent (#10075). Supports same pattern matching as `excludeTools`. Example: `['Bash(git status)', 'Bash(npm test)']`. | -| `authType` | `'openai' \| 'anthropic' \| 'qwen-oauth' \| 'gemini' \| 'vertex-ai'` | - | Authentication type for the AI service. When provided, the SDK forwards it to the CLI as `--auth-type`. | -| `agents` | `SubagentConfig[]` | - | Configuration for subagents that can be invoked during the session. Subagents are specialized AI agents for specific tasks or domains. | -| `includePartialMessages` | `boolean` | `false` | When `true`, the SDK emits incomplete messages as they are being generated, allowing real-time streaming of the AI's response. | -| `resume` | `string` | - | Resume a previous session by providing its session ID. Equivalent to CLI's `--resume` flag. | -| `sessionId` | `string` | - | Specify a session ID for the new session. Ensures SDK and CLI use the same ID without resuming history. Equivalent to CLI's `--session-id` flag. | +| Option | Type | Default | Description | +| ------------------------ | -------------------------------------------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `cwd` | `string` | `process.cwd()` | The working directory for the query session. Determines the context in which file operations and commands are executed. | +| `model` | `string` | - | The AI model to use (e.g., `'qwen-max'`, `'qwen-plus'`, `'qwen-turbo'`). Takes precedence over `OPENAI_MODEL` and `QWEN_MODEL` environment variables. | +| `pathToQwenExecutable` | `string` | Bundled CLI | Path to the Qwen Code executable. Supports multiple formats: `'qwen'` (native binary from PATH), `'/path/to/qwen'` (explicit path), `'/path/to/cli.js'` (Node.js bundle), `'node:/path/to/cli.js'` (force Node.js runtime), `'bun:/path/to/cli.js'` (force Bun runtime). If not provided, the SDK uses the bundled CLI included with the package. | +| `permissionMode` | `'default' \| 'plan' \| 'auto-edit' \| 'auto' \| 'yolo'` | `'default'` | Permission mode controlling tool execution approval. See [Permission Modes](#permission-modes) for details. | +| `canUseTool` | `CanUseTool` | - | Custom permission handler for tool execution approval. Invoked when a tool requires confirmation. Must respond within 60 seconds or the request will be auto-denied. See [Custom Permission Handler](#custom-permission-handler). | +| `env` | `Record` | - | Environment variables to pass to the Qwen Code process. Merged with the current process environment. | +| `systemPrompt` | `string \| QuerySystemPromptPreset` | - | System prompt configuration for the main session. Use a string to fully override the built-in Qwen Code system prompt, or a preset object to keep the built-in prompt and append extra instructions. | +| `mcpServers` | `Record` | - | MCP (Model Context Protocol) servers to connect. Supports external servers (stdio/SSE/HTTP) and SDK-embedded servers. External servers are configured with transport options like `command`, `args`, `url`, `httpUrl`, etc. SDK servers use `{ type: 'sdk', name: string, instance: Server }`. | +| `abortController` | `AbortController` | - | Controller to cancel the query session. Call `abortController.abort()` to terminate the session and cleanup resources. | +| `debug` | `boolean` | `false` | Enable debug mode for verbose logging from the CLI process. | +| `maxSessionTurns` | `number` | `-1` (unlimited) | Maximum number of conversation turns before the session automatically terminates. Must be an integer. A turn consists of a user message and an assistant response. | +| `coreTools` | `string[]` | - | Uses the legacy `coreTools` / CLI `--core-tools` allowlist semantics. If specified, only matching core tools are registered for the session. This is the only allowlist-style option that restricts built-in tool registration; a whole-tool `permissions.deny` / `excludeTools` rule (and `tools.disabled` in settings.json) also removes a tool from the registry. `permissions.allow` in settings.json is pure auto-approval and never removes, demotes, or hides a tool (#10075). To keep a tool's schema out of the initial model request, use `tools.eager` in settings.json (requires restart, #9827) — `tool_search`, `tool_call`, `structured_output`, plan-mode lifecycle tools, `task_stop`, `mcp__*` and `computer_use__*` tools are exempt from that allowlist and keep their normal loading; tools demoted this way stay registered and reachable through `tool_search` + `tool_call` while both bridge tools are registered — when either is unregistered (`tools.toolSearch.enabled: false` denies both; a `tool_search` or `tool_call` deny rule, or a `tools.disabled` entry removes one) the demoted tools that remain hidden are absent from top-level declarations and cannot be reached through the bridge for that session, and a warning is written to the CLI process's stderr (SDK forwards it only with piped stderr and effective `debug` logging; an explicit `logLevel` always wins). These bridge and warning rules apply to direct and hybrid code modes. On the session surface in hybrid mode, while `exec` itself is registered (container and SSH execution warn and fall back to direct tools without it), exec retains callable nested bindings; their schemas are included in exec when either bridge tool is unavailable. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by tools.eager from nested bindings. In both code modes, agent allowlists that do not grant `exec` narrow nested bindings. Inheriting or explicitly granting `exec` keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns. In direct and hybrid modes they stay registered, so a direct call by their own name is still evaluated and approved normally — except tools also listed in `tools.visible`, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares.; to remove a tool entirely, use a whole-tool `excludeTools` / `permissions.deny` rule — a rule with a specifier (such as `'Bash(rm *)'`) only denies matching invocations at runtime. MCP tools are exempt from deny-based removal: hide them with the per-server `excludeTools` / `tools.disabled` filters instead (deny still blocks their calls at runtime). Example: `['read_file', 'edit', 'run_shell_command']`. | +| `excludeTools` | `string[]` | - | Equivalent to `permissions.deny` in settings.json. Excluded tools return a permission error immediately. Takes highest priority over all other permission settings. Supports tool name aliases and pattern matching: tool name (`'write_file'`), shell command prefix (`'Bash(rm *)'`), or path patterns (`'Read(.env)'`, `'Edit(/src/**)'`). | +| `allowedTools` | `string[]` | - | Equivalent to `permissions.allow` in settings.json for auto-approval. Matching tools bypass `canUseTool` callback and execute automatically. Only applies when tool requires confirmation. Like `permissions.allow`, this is pure auto-approval and never affects which tools are registered or which schemas are sent (#10075). Supports same pattern matching as `excludeTools`. Example: `['Bash(git status)', 'Bash(npm test)']`. | +| `authType` | `'openai' \| 'anthropic' \| 'qwen-oauth' \| 'gemini' \| 'vertex-ai'` | - | Authentication type for the AI service. When provided, the SDK forwards it to the CLI as `--auth-type`. | +| `agents` | `SubagentConfig[]` | - | Configuration for subagents that can be invoked during the session. Subagents are specialized AI agents for specific tasks or domains. | +| `includePartialMessages` | `boolean` | `false` | When `true`, the SDK emits incomplete messages as they are being generated, allowing real-time streaming of the AI's response. | +| `resume` | `string` | - | Resume a previous session by providing its session ID. Equivalent to CLI's `--resume` flag. | +| `sessionId` | `string` | - | Specify a session ID for the new session. Ensures SDK and CLI use the same ID without resuming history. Equivalent to CLI's `--session-id` flag. | > [!note] > For `coreTools`, aliases like `Read`, `Edit`, and `Bash` also work, but invocation specifiers such as `Bash(git *)` are stripped. `coreTools` restricts tool registration, not invocation patterns. diff --git a/docs/users/configuration/settings.md b/docs/users/configuration/settings.md index bde7a912b31..ca5b24412ec 100644 --- a/docs/users/configuration/settings.md +++ b/docs/users/configuration/settings.md @@ -375,46 +375,48 @@ If you are experiencing performance issues with file searching (e.g., with `@` c #### tools -Set `tools.codeModeOnly` to `true` to enable experimental Code Mode with lazy tool discovery. It defaults to `false` and requires a restart. +Set `tools.mode` to `code_mode` for hybrid direct and nested calls, or `code_mode_only` for exec-based calls with lazy tool discovery. It defaults to `direct` and requires a restart. Set `tools.freeform` to `true` to send Code Mode's `exec` input as raw text on -OpenAI Responses models. It takes effect only when `tools.codeModeOnly` is also -`true` and the selected model uses `wireApi: "responses"`. Enable it only for +OpenAI Responses models. It takes effect only when `tools.mode` is +`code_mode_only` and the selected model uses `wireApi: "responses"`. Enable it only for endpoints that support [Responses Custom Tools](https://developers.openai.com/api/docs/guides/function-calling#custom-tools). This setting is intentionally hidden from the Settings dialog; edit it in the settings file. It defaults to `false` and requires a restart. -The bridge-availability and missing-bridge warning rules below describe direct tool mode. CodeModeOnly discovers deferred descriptions and schemas through top-level `tool_search`, then invokes tools through `exec`. Search results do not change the tool declarations. It skips deferred preload and startup catalogs; `tools.eager` also reduces the initial `exec` description. When search is unavailable in the current scope, `exec` includes all allowed tool signatures. See [Lazy Code Mode](../../design/lazy-code-mode.md). - -| Setting | Type | Description | Default | Notes | -| ------------------------------------ | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `tools.freeform` | boolean | Use raw text input for the Code Mode `exec` tool on OpenAI Responses models. Effective only when `tools.codeModeOnly` is `true` and the selected model uses `wireApi: "responses"`. Enable only for endpoints that support Responses Custom Tools. | `false` | Requires restart: Yes | -| `tools.sandbox` | boolean or string | Sandbox execution environment (can be a boolean or a path string). | `undefined` | | -| `tools.sandboxImage` | string | Sandbox image URI used by Docker/Podman when `--sandbox-image` and `QWEN_SANDBOX_IMAGE` are not set. | `undefined` | | -| `tools.shell.enableInteractiveShell` | boolean | Use `node-pty` for an interactive shell experience. When unset, explicit one-shot prompts use `child_process`; interactive TUI, ACP, stream-json input, stdin-only, and file-input sessions use PTY. Platform and availability fallbacks to `child_process` still apply. | mode-based | | -| `tools.shell.defaultTimeoutMs` | number | Default timeout, in milliseconds, for foreground shell commands started by the agent. A per-call timeout on the shell tool overrides this. When unset, foreground commands time out after 120000 ms (2 minutes). Set to 0 to disable the timeout. | `undefined` | | -| `tools.shell.heartbeatIntervalMs` | number | Interval, in milliseconds, between liveness heartbeats emitted while a foreground shell command produces no output. Heartbeats are forwarded to ACP clients and stream-json consumers so they can tell a silent command from a dead session. When unset, heartbeats fire every 10000 ms (10 seconds). Set to 0 to disable heartbeats. | `undefined` | | -| `tools.core` | array of strings | **Deprecated.** Will be removed in next version. A non-empty list restricts the core tool set (file, shell, search and related built-ins) to an allowlist: core tools not in the list are disabled (fail-closed). Tools outside that set — dynamically discovered tools (MCP, skill) and synthetic/system built-ins such as `agent`, `list_agents`, plan-mode lifecycle tools, goal tools, `task_stop`, `send_message`, `tool_search` and `tool_call` — bypass the allowlist by design; use `permissions.deny` to block a tool's calls (for MCP tools it stays listed and is rejected at runtime), or `tools.disabled` / the per-server `excludeTools` filter to remove it from the registry outright. An empty list (`[]`) is treated as unset and disables nothing. `permissions.allow` cannot reproduce this restriction — it is pure auto-approval (#10075). Use `tools.eager` to restrict which eager-by-default tool schemas are sent initially (unlisted tools are deferred, not disabled — they stay reachable through `tool_search` + `tool_call` while both bridge tools are registered), and `permissions.deny` to block tools outright. | `undefined` | | -| `tools.exclude` | array of strings | **Deprecated.** Use `permissions.deny` instead. Tool names to exclude from discovery. Not automatically migrated; the legacy setting remains honoured at startup. | `undefined` | | -| `tools.disabled` | array of strings | Tool names hidden from the registry entirely. Unlike `permissions.deny` (which blocks calls at runtime), disabled tools are never registered, so they do not appear in `/tools` and cannot be discovered or called by the model. For example, `["enter_plan_mode"]` prevents the model from switching into plan mode on its own. Merged as a union across scopes. | `undefined` | | -| `tools.visible` | array of strings | Deferred tool names made visible at startup without requiring the ToolSearch + ToolCall bridge. Listed tools appear alongside core tools in the initial session. Merged as a union across scopes. | `undefined` | | -| `tools.eager` | array of strings | Allowlist of eager-by-default built-in tool names whose schemas remain eligible for the initial model request. Unlisted non-exempt tools are deferred instead: still registered, listed in `/tools`, and reachable through the `tool_search` + `tool_call` bridge. Tools already deferred by default stay on demand even when listed; use `tools.visible` to surface one at startup. `tool_search`, `tool_call`, `structured_output`, plan-mode lifecycle tools, `task_stop`, MCP tools, and `computer_use__*` tools are unaffected and keep their normal loading behaviour. An explicitly empty list (`[]`) is active and defers every non-exempt eager-by-default tool; omitting the setting means no restriction. Pairs with the ToolSearch + ToolCall bridge: when either half is not registered — `tools.toolSearch.enabled: false` (which denies both), a `tool_search` or `tool_call` deny rule, or a `tools.disabled` entry — the allowlist still withholds the schemas, but nothing can load them back, so the demoted tools that remain hidden are not offered to the model and cannot be reached through the bridge for that session (they stay registered and in `/tools`, a direct call by their own name is still evaluated and approved normally, and a warning is logged) — except tools also listed in `tools.visible`, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares. Use `permissions.deny` if you meant to remove them, or keep both bridge tools registered. Unusable entries (empty or malformed) are dropped with a warning and leave the rest of the list active. Later scopes replace earlier lists. Requires restart. | `undefined` | | -| `tools.allowed` | array of strings | **Deprecated.** Use `permissions.allow` instead. Tool names that bypass the confirmation dialog. Not automatically migrated; the legacy setting remains honoured at startup. | `undefined` | | -| `tools.approvalMode` | string | Sets the default approval mode for tool usage. | `auto` | Possible values: `plan` (analyze only, do not modify files or execute commands), `default` (require approval before file edits or shell commands run), `auto-edit` (automatically approve file edits), `auto` (LLM classifier auto-approves safe actions, blocks risky ones), `yolo` (automatically approve all tool calls) | -| `tools.discoveryCommand` | string | Command to run for tool discovery. When the `tools.eager` allowlist is active, a discovered tool not named in it is registered as deferred: it stays in `/tools` and is reachable through `tool_search` + `tool_call` while both bridge tools are registered, but its schema is not sent in the initial model request. | `undefined` | | -| `tools.callCommand` | string | Defines a custom shell command for calling a specific tool that was discovered using `tools.discoveryCommand`. The shell command must meet the following criteria: It must take function `name` (exactly as in [function declaration](https://ai.google.dev/gemini-api/docs/function-calling#function-declarations)) as first command line argument. It must read function arguments as JSON on `stdin`, analogous to [`functionCall.args`](https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/inference#functioncall). It must return function output as JSON on `stdout`, analogous to [`functionResponse.response.content`](https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/inference#functionresponse). | `undefined` | | -| `tools.useRipgrep` | boolean | Use ripgrep for file content search instead of the fallback implementation. Provides faster search performance. | `true` | | -| `tools.useBuiltinRipgrep` | boolean | Use the bundled ripgrep binary. When set to `false`, the system-level `rg` command will be used instead. This setting is only effective when `tools.useRipgrep` is `true`. | `true` | | -| `tools.workflowsEnabled` | boolean | Enable the Workflow tool, which lets the model author and run a script that orchestrates subagents in parallel. Off by default; a run can dispatch many subagents and spend tokens accordingly. | `false` | User, System, and SystemDefaults scopes only; workspace values are ignored. Requires restart: Yes. Env overrides: `QWEN_CODE_ENABLE_WORKFLOWS=1` forces on; `QWEN_CODE_DISABLE_WORKFLOWS=1` forces off (disable wins). | -| `tools.workflowSizeGuideline` | enum | Advisory size guideline for the dynamic workflows the model writes: `"small"` aims for fewer than 5 agents, `"medium"` fewer than 15, `"large"` fewer than 50, and `"unrestricted"` sends no guideline. It is not an enforced limit. It also sets the agent count at which a running workflow is flagged as a large workflow in the background-tasks view. | `"medium"` | Possible values: `"small"`, `"medium"`, `"large"`, `"unrestricted"`. Requires restart: No; a change is announced to the model with your next message. Env override for the warning threshold: `QWEN_CODE_WORKFLOW_SIZE_WARNING_AGENTS`. | -| `tools.workflowNameOnly` | boolean | Restrict the model to running named workflows — saved workflows and the workflows extensions ship, called as `Workflow({ name, args })`. The model cannot run an inline `script` or a `scriptPath`, and a running script cannot nest `workflow({ scriptPath })`, so every run the model starts can be matched by a `Workflow(name:...)` permission rule. It does not replace an approval policy: the model can still save a new workflow file and run it by name, which a rule scoped to specific names or script digests asks about. Runs a host starts over ACP (`run-saved`, `run-script`, retry, rerun) are not restricted. In such a session the `/review` workflow fan-out is unavailable. | `false` | Requires restart: Yes. Env: `QWEN_CODE_WORKFLOW_NAME_ONLY=1` turns it on too; a project `.env` cannot set it. A workspace may set this to `true` only; a workspace `false` is ignored. | -| `goals.modelProposed` | enum | Controls the `propose_goal` tool, which lets the model propose a session Goal for you to approve: `alwaysAsk` shows every proposal in an approval dialog and nothing is set until you accept it; `"disabled"` removes the tool. A typed `/goal` is unaffected. | `alwaysAsk` | User, System, and SystemDefaults scopes only; workspace values are ignored. Requires restart: Yes. | -| `tools.truncateToolOutputThreshold` | number | Truncate tool output if it is larger than this many characters. Applies to Shell, Grep, Glob, ReadFile and ReadManyFiles tools. | `25000` | Requires restart: Yes | -| `tools.truncateToolOutputLines` | number | Maximum lines or entries kept when truncating tool output. Applies to Shell, Grep, Glob, ReadFile and ReadManyFiles tools. | `1000` | Requires restart: Yes | -| `tools.toolSearch.enabled` | boolean | Review deferred tool schemas through ToolSearch and invoke them through the stable ToolCall bridge. Bridge review and invocation keep the tool list stable — the bridge never re-declares what it reveals — reducing prompt size without touching the prompt-cache prefix. The declaration list is not immutable, though: a session still re-declares on resume, whenever a tool-set refresh (MCP discovery, the first plan-mode entry in a session, a subagent definition change) finds a direct call to a still-hidden deferred tool in the live history, when a subagent definition change rewrites the agent tool's own description, and when an MCP server registers mid-session with `alwaysLoadTools: true`. | `true` | Requires restart: Yes | -| `tools.toolSearch.threshold` | number | Context-window percentage used as the session-start budget for preloading ordinary deferred tools (bundled built-ins and MCP alike). Defaults to `0`, which performs no threshold-based preload; ordinary deferred tools normally stay behind the stable ToolSearch + ToolCall bridge, at the cost of one `tool_search` round trip before first use. Raise it to `N` so that, when every eligible deferred schema fits within `N`% of the context window, all are declared upfront for direct calls with no bridge round trip; otherwise they stay behind the bridge while both bridge tools are registered. Tools demoted by `tools.eager` are excluded from this preload and stay reachable on demand through that bridge while it is registered. Separate paths can still declare deferred tools at `0`: `tools.visible`; the live-history compatibility scan on every tool-set refresh (including resume, MCP discovery, first plan-mode entry, and subagent definition changes); the incomplete-bridge eager fallback; and daemon ACP's late registration, which explicitly reveals and pins `create_sub_session`. | `0` | Requires restart: Yes | -| `tools.listDirectory.enabled` | boolean | Enable the built-in `list_directory` tool. Disabled by default because `glob` covers directory listing in most cases; the tool is also re-enabled automatically when explicitly listed in the `coreTools` allowlist (`--core-tools` / `tools.core`). | `false` | Requires restart: Yes | -| `tools.todoWrite.enabled` | boolean | Enable the built-in `todo_write` tool and its system-prompt guidance. Disabled by default. | `false` | Requires restart: Yes | +The bridge-availability and missing-bridge warning rules below describe direct and hybrid tool modes. CodeModeOnly discovers deferred descriptions and schemas through top-level `tool_search`, then invokes tools through `exec`. Search results do not change the tool declarations. It skips deferred preload and startup catalogs; `tools.eager` also reduces the initial `exec` description. When search is unavailable in the current scope, `exec` includes all allowed tool signatures. See [Lazy Code Mode](../../design/lazy-code-mode.md). + +| Setting | Type | Description | Default | Notes | +| ------------------------------------ | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `tools.freeform` | boolean | Use raw text input for the Code Mode `exec` tool on OpenAI Responses models. Effective only when `tools.mode` is `code_mode_only` and the selected model uses `wireApi: "responses"`. Enable only for endpoints that support Responses Custom Tools. | `false` | Requires restart: Yes | +| `tools.sandbox` | boolean or string | Sandbox execution environment (can be a boolean or a path string). | `undefined` | | +| `tools.sandboxImage` | string | Sandbox image URI used by Docker/Podman when `--sandbox-image` and `QWEN_SANDBOX_IMAGE` are not set. | `undefined` | | +| `tools.shell.enableInteractiveShell` | boolean | Use `node-pty` for an interactive shell experience. When unset, explicit one-shot prompts use `child_process`; interactive TUI, ACP, stream-json input, stdin-only, and file-input sessions use PTY. Platform and availability fallbacks to `child_process` still apply. | mode-based | | +| `tools.shell.defaultTimeoutMs` | number | Default timeout, in milliseconds, for foreground shell commands started by the agent. A per-call timeout on the shell tool overrides this. When unset, foreground commands time out after 120000 ms (2 minutes). Set to 0 to disable the timeout. | `undefined` | | +| `tools.shell.heartbeatIntervalMs` | number | Interval, in milliseconds, between liveness heartbeats emitted while a foreground shell command produces no output. Heartbeats are forwarded to ACP clients and stream-json consumers so they can tell a silent command from a dead session. When unset, heartbeats fire every 10000 ms (10 seconds). Set to 0 to disable heartbeats. | `undefined` | | +| `tools.core` | array of strings | **Deprecated.** Will be removed in next version. A non-empty list restricts the core tool set (file, shell, search and related built-ins) to an allowlist: core tools not in the list are disabled (fail-closed). Tools outside that set — dynamically discovered tools (MCP, skill) and synthetic/system built-ins such as `agent`, `list_agents`, plan-mode lifecycle tools, goal tools, `task_stop`, `send_message`, `tool_search` and `tool_call` — bypass the allowlist by design; use `permissions.deny` to block a tool's calls (for MCP tools it stays listed and is rejected at runtime), or `tools.disabled` / the per-server `excludeTools` filter to remove it from the registry outright. An empty list (`[]`) is treated as unset and disables nothing. `permissions.allow` cannot reproduce this restriction — it is pure auto-approval (#10075). Use `tools.eager` to restrict which eager-by-default tool schemas are sent initially (unlisted tools are deferred, not disabled — they stay reachable through `tool_search` + `tool_call` while both bridge tools are registered), and `permissions.deny` to block tools outright. | `undefined` | | +| `tools.exclude` | array of strings | **Deprecated.** Use `permissions.deny` instead. Tool names to exclude from discovery. Not automatically migrated; the legacy setting remains honoured at startup. | `undefined` | | +| `tools.disabled` | array of strings | Tool names hidden from the registry entirely. Unlike `permissions.deny` (which blocks calls at runtime), disabled tools are never registered, so they do not appear in `/tools` and cannot be discovered or called by the model. For example, `["enter_plan_mode"]` prevents the model from switching into plan mode on its own. Merged as a union across scopes. | `undefined` | | +| `tools.visible` | array of strings | Deferred tool names made visible at startup without requiring the ToolSearch + ToolCall bridge. Listed tools appear alongside core tools in the initial session. In Code Mode Only, listed tools have signatures in the initial `exec` description; other deferred schemas are discovered through top-level `tool_search`. Merged as a union across scopes. | `undefined` | | +| `tools.eager` | array of strings | Allowlist of eager-by-default built-in tool names whose schemas remain eligible for the initial model request. Unlisted non-exempt tools are deferred instead: still registered, listed in `/tools`, and reachable through the `tool_search` + `tool_call` bridge. Tools already deferred by default stay on demand even when listed; use `tools.visible` to surface one at startup. `tool_search`, `tool_call`, `structured_output`, plan-mode lifecycle tools, `task_stop`, MCP tools, and `computer_use__*` tools are unaffected and keep their normal loading behaviour. An explicitly empty list (`[]`) is active and defers every non-exempt eager-by-default tool; omitting the setting means no restriction. Pairs with the ToolSearch + ToolCall bridge: when either half is not registered — `tools.toolSearch.enabled: false` (which denies both), a `tool_search` or `tool_call` deny rule, or a `tools.disabled` entry — the allowlist still withholds the schemas, but nothing can load them back, so the demoted tools that remain hidden are absent from top-level declarations and cannot be reached through the bridge for that session (they stay registered and in `/tools`; in direct and hybrid modes a direct call by their own name is still evaluated and approved normally, and a warning is logged) — except tools also listed in `tools.visible`, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares. Use `permissions.deny` if you meant to remove them, or keep both bridge tools registered. Unusable entries (empty or malformed) are dropped with a warning and leave the rest of the list active. Later scopes replace earlier lists. Requires restart. In Code Mode, withheld tools with actual nested bindings also remain callable through `exec`, and the warning names that subset. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by `tools.eager` from nested bindings. In both code modes, agent allowlists that do not grant `exec` narrow nested bindings. Inheriting or explicitly granting `exec` keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns. | `undefined` | | +| `tools.allowed` | array of strings | **Deprecated.** Use `permissions.allow` instead. Tool names that bypass the confirmation dialog. Not automatically migrated; the legacy setting remains honoured at startup. | `undefined` | | +| `tools.approvalMode` | string | Sets the default approval mode for tool usage. | `auto` | Possible values: `plan` (analyze only, do not modify files or execute commands), `default` (require approval before file edits or shell commands run), `auto-edit` (automatically approve file edits), `auto` (LLM classifier auto-approves safe actions, blocks risky ones), `yolo` (automatically approve all tool calls) | +| `tools.discoveryCommand` | string | Command to run for tool discovery. When the `tools.eager` allowlist is active, a discovered tool not named in it is registered as deferred: it stays in `/tools` and is reachable through `tool_search` + `tool_call` while both bridge tools are registered, but its schema is not sent in the initial model request. | `undefined` | | +| `tools.callCommand` | string | Defines a custom shell command for calling a specific tool that was discovered using `tools.discoveryCommand`. The shell command must meet the following criteria: It must take function `name` (exactly as in [function declaration](https://ai.google.dev/gemini-api/docs/function-calling#function-declarations)) as first command line argument. It must read function arguments as JSON on `stdin`, analogous to [`functionCall.args`](https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/inference#functioncall). It must return function output as JSON on `stdout`, analogous to [`functionResponse.response.content`](https://cloud.google.com/vertex-ai/generative-ai/docs/model-reference/inference#functionresponse). | `undefined` | | +| `tools.useRipgrep` | boolean | Use ripgrep for file content search instead of the fallback implementation. Provides faster search performance. | `true` | | +| `tools.useBuiltinRipgrep` | boolean | Use the bundled ripgrep binary. When set to `false`, the system-level `rg` command will be used instead. This setting is only effective when `tools.useRipgrep` is `true`. | `true` | | +| `tools.workflowsEnabled` | boolean | Enable the Workflow tool, which lets the model author and run a script that orchestrates subagents in parallel. Off by default; a run can dispatch many subagents and spend tokens accordingly. | `false` | User, System, and SystemDefaults scopes only; workspace values are ignored. Requires restart: Yes. Env overrides: `QWEN_CODE_ENABLE_WORKFLOWS=1` forces on; `QWEN_CODE_DISABLE_WORKFLOWS=1` forces off (disable wins). | +| `tools.workflowSizeGuideline` | enum | Advisory size guideline for the dynamic workflows the model writes: `"small"` aims for fewer than 5 agents, `"medium"` fewer than 15, `"large"` fewer than 50, and `"unrestricted"` sends no guideline. It is not an enforced limit. It also sets the agent count at which a running workflow is flagged as a large workflow in the background-tasks view. | `"medium"` | Possible values: `"small"`, `"medium"`, `"large"`, `"unrestricted"`. Requires restart: No; a change is announced to the model with your next message. Env override for the warning threshold: `QWEN_CODE_WORKFLOW_SIZE_WARNING_AGENTS`. | +| `tools.workflowNameOnly` | boolean | Restrict the model to running named workflows — saved workflows and the workflows extensions ship, called as `Workflow({ name, args })`. The model cannot run an inline `script` or a `scriptPath`, and a running script cannot nest `workflow({ scriptPath })`, so every run the model starts can be matched by a `Workflow(name:...)` permission rule. It does not replace an approval policy: the model can still save a new workflow file and run it by name, which a rule scoped to specific names or script digests asks about. Runs a host starts over ACP (`run-saved`, `run-script`, retry, rerun) are not restricted. In such a session the `/review` workflow fan-out is unavailable. | `false` | Requires restart: Yes. Env: `QWEN_CODE_WORKFLOW_NAME_ONLY=1` turns it on too; a project `.env` cannot set it. A workspace may set this to `true` only; a workspace `false` is ignored. | +| `goals.modelProposed` | enum | Controls the `propose_goal` tool, which lets the model propose a session Goal for you to approve: `alwaysAsk` shows every proposal in an approval dialog and nothing is set until you accept it; `"disabled"` removes the tool. A typed `/goal` is unaffected. | `alwaysAsk` | User, System, and SystemDefaults scopes only; workspace values are ignored. Requires restart: Yes. | +| `tools.truncateToolOutputThreshold` | number | Truncate tool output if it is larger than this many characters. Applies to Shell, Grep, Glob, ReadFile and ReadManyFiles tools. | `25000` | Requires restart: Yes | +| `tools.truncateToolOutputLines` | number | Maximum lines or entries kept when truncating tool output. Applies to Shell, Grep, Glob, ReadFile and ReadManyFiles tools. | `1000` | Requires restart: Yes | +| `tools.toolSearch.enabled` | boolean | Review deferred tool schemas through ToolSearch and invoke them through the stable ToolCall bridge. Bridge review and invocation keep the tool list stable — the bridge never re-declares what it reveals — reducing prompt size without touching the prompt-cache prefix. The declaration list is not immutable, though: a session still re-declares on resume, whenever a tool-set refresh (MCP discovery, the first plan-mode entry in a session, a subagent definition change) finds a direct call to a still-hidden deferred tool in the live history, when a subagent definition change rewrites the agent tool's own description, and when an MCP server registers mid-session with `alwaysLoadTools: true`. | `true` | Requires restart: Yes | +| `tools.toolSearch.threshold` | number | Context-window percentage used as the session-start budget for preloading ordinary deferred tools (bundled built-ins and MCP alike). Defaults to `0`, which performs no threshold-based preload; ordinary deferred tools normally stay behind the stable ToolSearch + ToolCall bridge, at the cost of one `tool_search` round trip before first use. Raise it to `N` so that, when every eligible deferred schema fits within `N`% of the context window, all are declared upfront for direct calls with no bridge round trip; otherwise they stay behind the bridge while both bridge tools are registered. Tools demoted by `tools.eager` are excluded from this preload and stay reachable on demand through that bridge while it is registered. Separate paths can still declare deferred tools at `0`: `tools.visible`; the live-history compatibility scan on every tool-set refresh (including resume, MCP discovery, first plan-mode entry, and subagent definition changes); the incomplete-bridge eager fallback; and daemon ACP's late registration, which explicitly reveals and pins `create_sub_session`. | `0` | Requires restart: Yes | +| `tools.listDirectory.enabled` | boolean | Enable the built-in `list_directory` tool. Disabled by default because `glob` covers directory listing in most cases; the tool is also re-enabled automatically when explicitly listed in the `coreTools` allowlist (`--core-tools` / `tools.core`). | `false` | Requires restart: Yes | +| `tools.todoWrite.enabled` | boolean | Enable the built-in `todo_write` tool and its system-prompt guidance. Disabled by default. | `false` | Requires restart: Yes | +| `tools.mode` | string | Selects the tool exposure mode: `direct` uses ordinary tool calls, `code_mode` also exposes the isolated `exec` JavaScript tool, and `code_mode_only` exposes ordinary tools only through `exec` while keeping direct control tools available. Ignored in safe and bare modes. Container execution warns and uses direct tools for `code_mode`, and rejects `code_mode_only`. SSH workspaces warn and use Direct for either code mode. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools hidden by `tools.eager` from nested bindings. In both code modes, agent allowlists that do not grant `exec` narrow nested bindings. Inheriting or explicitly granting `exec` keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns. | `direct` | Experimental. Requires restart. | +| `tools.codeModeOnly` | boolean | **Deprecated.** Use `tools.mode: "code_mode_only"` instead. Removed from the settings schema; still honored at startup only while `tools.mode` is unset. Not shown in the settings dialog. | `undefined` | Requires restart. | > [!note] > diff --git a/docs/users/features/context-cost.md b/docs/users/features/context-cost.md index 7bff9641573..3d1d74b27a0 100644 --- a/docs/users/features/context-cost.md +++ b/docs/users/features/context-cost.md @@ -53,9 +53,9 @@ Measure the allowlist against the right baseline. Tools that are on demand by _t Four things to know before you use it: - **It is not a disable.** A demoted tool stays reachable. If you meant to remove a tool, use a whole-tool `permissions.deny` rule or `tools.disabled`. -- **Some tools are exempt** and keep their normal loading behaviour whatever the list says: `tool_search` and `tool_call` (the two halves of the bridge that reaches a withheld tool), `structured_output`, the plan-mode lifecycle tools (`enter_plan_mode`, `exit_plan_mode`, `ask_user_question`), `task_stop`, MCP tools (`mcp__*`), and Computer Use tools (`computer_use__*`). `task_stop` and the Computer Use family are on demand by default anyway, so gating them would save nothing; MCP tools are governed by `tools.toolSearch.*` and the per-server `includeTools` / `excludeTools` filters; and nothing in `tools.eager` can drop one of the rest: that takes a whole-tool `permissions.deny` rule, a `tools.disabled` entry, or `--exclude-tools` — and for the bridge pair, `tools.toolSearch.enabled: false` removes both halves at once. Removing a bridge half is not a neutral way out of the allowlist: every ordinary deferred tool (all MCP tools, the deferred Computer Use family) is then force-declared into every request, while a withheld tool is not offered to the model and cannot be loaded back through the bridge — it stays registered, so a direct call by name still goes through normal approval. +- **Some tools are exempt** and keep their normal loading behaviour whatever the list says: `tool_search` and `tool_call` (the two halves of the bridge that reaches a withheld tool), `structured_output`, the plan-mode lifecycle tools (`enter_plan_mode`, `exit_plan_mode`, `ask_user_question`), `task_stop`, MCP tools (`mcp__*`), and Computer Use tools (`computer_use__*`). `task_stop` and the Computer Use family are on demand by default anyway, so gating them would save nothing; MCP tools are governed by `tools.toolSearch.*` and the per-server `includeTools` / `excludeTools` filters; and nothing in `tools.eager` can drop one of the rest: that takes a whole-tool `permissions.deny` rule, a `tools.disabled` entry, or `--exclude-tools` — and for the bridge pair, `tools.toolSearch.enabled: false` removes both halves at once. Removing a bridge half is not a neutral way out of the allowlist: every ordinary deferred tool (all MCP tools, the deferred Computer Use family) is then force-declared into every request, while a withheld tool stays out of top-level declarations and cannot be loaded back through the bridge. In direct and hybrid modes it stays registered, so a direct call by name still goes through normal approval. On the hybrid session surface, callable nested tools retain their schemas inside `exec` when the bridge is incomplete; withholding their top-level declarations does not save those schema tokens. - **`permissions.allow` saves nothing.** It is pure auto-approval: it never demotes, hides, or removes a tool. Neither do approval modes. -- **It needs both halves of the bridge.** A withheld tool is reviewed with `tool_search` and invoked through `tool_call`; `tool_search` alone can read a schema it cannot call. If either is unregistered — `tools.toolSearch.enabled: false` denies both; a `tool_search`/`tool_call` deny rule, an `--exclude-tools` entry, or a `tools.disabled` entry removes one — the allowlist still withholds the schemas, nothing can load them back, and the demoted tools are not offered to the model for that session (a warning is logged); they stay registered, so a direct call by name still goes through normal approval. Ordinary deferred tools fall back to being declared eagerly in that case; tools you demoted with `tools.eager` do not. +- **It needs both halves of the bridge.** A withheld tool is reviewed with `tool_search` and invoked through `tool_call`; `tool_search` alone can read a schema it cannot call. If either is unregistered — `tools.toolSearch.enabled: false` denies both; a `tool_search`/`tool_call` deny rule, an `--exclude-tools` entry, or a `tools.disabled` entry removes one — the allowlist still withholds top-level schemas and nothing can load them back through the bridge. In direct and hybrid modes a warning is logged; the tools stay registered, so a direct call by name still goes through normal approval. On the hybrid session surface, callable nested tools still carry their full schemas inside `exec`. Code Mode Only discovers deferred schemas through top-level `tool_search` and calls them through `exec`; when search is unavailable in the current scope, `exec` includes all allowed signatures. Only skips the bridge warning and budget preload, while `tools.eager` and `tools.visible` control its initial documented bindings. In Hybrid mode, AgentCore excludes tools still hidden by `tools.eager` from nested bindings. Ordinary deferred tools fall back to being declared eagerly in that case; tools you demoted with `tools.eager` do not. `tools.visible` is the escape hatch for one tool you want declared up front even though it is deferred by default. diff --git a/docs/users/features/sub-agents.md b/docs/users/features/sub-agents.md index 10f79f02c14..32e227dcb61 100644 --- a/docs/users/features/sub-agents.md +++ b/docs/users/features/sub-agents.md @@ -381,9 +381,11 @@ Do not modify any files. Use `tools` and `disallowedTools` to control which tools a subagent can access. -**`tools` (allowlist):** When specified, the subagent can only use the listed tools. When omitted, the subagent inherits all available tools from the parent session. +**`tools` (allowlist):** When specified, the list bounds direct tool calls and deferred-tool bridge targets. Code mode also supports the nested `exec` path described below. When omitted, the subagent inherits all available tools from the parent session. -The allowlist applies both to directly declared tools and to targets invoked through the `tool_search`/`tool_call` deferred-tool bridge. An explicit list of ordinary tools does not automatically include either bridge tool; name both `tool_search` and `tool_call` if the agent needs discovery and bridge invocation. A listed ordinary deferred target is declared directly, while a target demoted by `tools.eager` remains hidden unless a separate reveal rule applies. A hidden target must still be present in `tools` to be invoked through the bridge. `disallowedTools` and `permissions.deny` remain additional blocklists, and listing a tool does not bypass the subagent control-plane exclusions. Code Mode (`tools.codeModeOnly`) differs: `tool_call` is not available, an agent allowed any tool callable from `exec` gets `tool_search` automatically unless `disallowedTools` names it, and tools demoted by `tools.eager` stay callable through `exec`. +The allowlist applies both to directly declared tools and to targets invoked through the `tool_search`/`tool_call` deferred-tool bridge. An explicit list of ordinary tools does not automatically include either bridge tool; name both `tool_search` and `tool_call` if the agent needs discovery and bridge invocation. A listed ordinary deferred target is declared directly, while a target demoted by `tools.eager` remains hidden unless a separate reveal rule applies. A hidden target must still be present in `tools` to be invoked through the bridge. `disallowedTools` and `permissions.deny` remain additional blocklists, and listing a tool does not bypass the subagent control-plane exclusions. Code Mode (`tools.mode: "code_mode_only"`) differs: `tool_call` is not available, an agent allowed any tool callable from `exec` gets `tool_search` automatically unless `disallowedTools` names it, and tools demoted by `tools.eager` stay callable through `exec`. + +With `tools.mode: "code_mode"` or `"code_mode_only"`, `exec` provides a third path: nested `tools.(...)` calls. Explicitly listing `exec` also grants otherwise-admitted code-mode-callable bindings, including ordinary tools such as `write_file` and `run_shell_command` that are not individually listed. Thus `tools: [read_file, exec]` is not a read-only policy. Omit `exec` from the configured list to keep nested targets bounded by the listed names; the runtime still exposes the `exec` wrapper in code mode. `disallowedTools`, `permissions.deny`, and subagent exclusions continue to restrict access in both modes. Hybrid additionally excludes tools hidden by `tools.eager`. This nested grant does not add direct-call permission in hybrid mode. A fork's explicit `fork_tools: [exec]` grant also remains bounded by its parent's inherited execution policy. ``` --- @@ -412,7 +414,7 @@ disallowedTools: If both `tools` and `disallowedTools` are set, the allowlist is applied first, then the blocklist removes from that set. -**MCP tools** follow the same rules. If a subagent has no `tools` list, it inherits all MCP tools from the parent session. If a subagent has an explicit `tools` list, it only gets MCP tools that are explicitly named in that list. +**MCP tools** follow the same rules. Without a `tools` list, the subagent inherits MCP tools from the parent session. An explicit list ordinarily requires matching MCP entries. Granting `exec` also admits otherwise-available MCP bindings through the nested path when the list contains no MCP entries; once the list mentions MCP, nested MCP calls must match those entries. The blocklists and exclusions above still apply. The `disallowedTools` field supports MCP server-level patterns: diff --git a/docs/verification/resident-tool-prompt-assembly/README.md b/docs/verification/resident-tool-prompt-assembly/README.md index d5ef426b450..5b3c21077b5 100644 --- a/docs/verification/resident-tool-prompt-assembly/README.md +++ b/docs/verification/resident-tool-prompt-assembly/README.md @@ -133,7 +133,7 @@ done | 3 | Arena | 它直接调 `getCoreSystemPrompt`,没有快照 | `/arena --models a,b "简单任务"` | 同一构建内门控前后一致(无快照 = 不门控);跨版本比较时,提示词精简(如 #12546)带来的文案差异属预期 | | 4 | 子 agent | 定义型子 agent 渲染自己的提示词;fork 与恢复的后台 agent 逐字继承父级已渲染(已门控)的指令 | 用 `subagent_type: "fork"` 跑一次委派,与同一构建的父提示词对比 | 定义型子 agent 不受影响;fork / 恢复 agent 同一构建内与父提示词一致(继承是共享缓存前缀的前提),跨版本随父提示词有意变化 | | 5 | output style | `keepCodingInstructions: false` 与门控叠加 | 选一个自定义 style 再看提示词 | 两者各自生效,不互相吃掉 | -| 6 | code mode | 该模式下工具在 `exec` 内调用,实现里**明确不门控** | `tools.codeModeOnly: true` 起一个会话 | `tools.` 那些条目一条不少 | +| 6 | code mode | 该模式下工具在 `exec` 内调用,实现里**明确不门控** | `tools.mode: "code_mode_only"` 起一个会话 | `tools.` 那些条目一条不少 | | 7 | `QWEN_SYSTEM_MD` 覆盖 | 覆盖分支完全绕过默认提示词 | 设一个覆盖文件起会话 | 提示词就是该文件,门控不参与 | | 8 | 提示词缓存 | 静态前缀内容变了 | 同一会话连发 3 轮,看缓存命中 token | 命中率与改动前同阶;前缀只在会话开始时重写一次 | | 9 | 压缩 | 压缩走 `startChat`,会重算快照 | 触发一次 `/compress` | 压缩后提示词与压缩前一致(同一会话工具集没变) | diff --git a/packages/cli/src/acp-integration/session/Session.test.ts b/packages/cli/src/acp-integration/session/Session.test.ts index e0543cb0325..f5d1d432319 100644 --- a/packages/cli/src/acp-integration/session/Session.test.ts +++ b/packages/cli/src/acp-integration/session/Session.test.ts @@ -41706,6 +41706,25 @@ describe('Session', () => { expect(direct.parts[0].functionResponse?.response?.['error']).toContain( 'unavailable on this CodeModeOnly call surface', ); + + vi.mocked(mockConfig.getToolMode).mockReturnValue(core.ToolMode.CodeMode); + const hybridDirect = await ( + session as unknown as ToolCallInternals + ).runToolCalls( + new AbortController().signal, + 'prompt-hybrid-code-mode-direct', + [ + { + id: 'read-acp-hybrid-direct', + name: 'read_file', + args: { path: '/tmp/example.txt' }, + }, + ], + ); + expect(nestedExecute).toHaveBeenCalledTimes(2); + expect(hybridDirect.parts[0].functionResponse?.response?.['output']).toBe( + 'nested ACP output', + ); }); describe('Code Mode nested concurrency', () => { diff --git a/packages/cli/src/config/config.test.ts b/packages/cli/src/config/config.test.ts index e3651990b61..9c779035e31 100644 --- a/packages/cli/src/config/config.test.ts +++ b/packages/cli/src/config/config.test.ts @@ -21,6 +21,7 @@ import { Storage, SessionIdCaseConflictError, FatalConfigError, + ToolMode, } from '@qwen-code/qwen-code-core'; import { normalizeModelProposedGoals } from './config.js'; import { @@ -1646,7 +1647,7 @@ describe('loadCliConfig', () => { ); expect(mockConfigConstructorParams).toHaveBeenLastCalledWith( expect.objectContaining({ - codeModeOnly: false, + toolMode: ToolMode.Direct, disableAllHooks: true, mcpServers: {}, overrideExtensions: [], @@ -4394,22 +4395,102 @@ describe('mergeExcludeTools', () => { expect(config.getToolSearchThreshold()).toBe(0); }); - it('should enable CodeModeOnly only when explicitly configured', async () => { + it('should resolve the tool mode enum', async () => { process.argv = ['node', 'script.js']; const argv = await parseArguments(); const direct = await loadCliConfig({}, argv, undefined, []); const codeMode = await loadCliConfig( - { tools: { codeModeOnly: true } }, + { tools: { mode: ToolMode.CodeMode } }, + argv, + undefined, + [], + ); + const codeModeOnly = await loadCliConfig( + { tools: { mode: ToolMode.CodeModeOnly } }, argv, undefined, [], ); expect(direct.getCodeModeOnly()).toBe(false); - expect(codeMode.getCodeModeOnly()).toBe(true); + expect(codeMode.getCodeModeOnly()).toBe(false); + expect(codeModeOnly.getCodeModeOnly()).toBe(true); expect(direct.getToolMode()).toBe('direct'); - expect(codeMode.getToolMode()).toBe('code_mode_only'); + expect(codeMode.getToolMode()).toBe('code_mode'); + expect(codeModeOnly.getToolMode()).toBe('code_mode_only'); + }); + + it.each([ToolMode.CodeMode, ToolMode.CodeModeOnly])( + 'warns when SSH downgrades %s to direct', + async (mode) => { + sshWorkspaceProbe.mockReturnValueOnce({ + host: 'host', + directory: '/srv/project', + }); + process.argv = ['node', 'script.js']; + const argv = await parseArguments(); + await loadCliConfig({ tools: { mode } }, argv, undefined, []); + expect(mockConfigConstructorParams).toHaveBeenLastCalledWith( + expect.objectContaining({ + toolMode: ToolMode.Direct, + warnings: expect.arrayContaining([ + `SSH workspaces do not support tools.mode = "${mode}"; using direct tools for this session.`, + ]), + }), + ); + }, + ); + + it('should fail closed for an invalid tool mode', async () => { + process.argv = ['node', 'script.js']; + const argv = await parseArguments(); + const config = await loadCliConfig( + { tools: { mode: 'code-mode' } } as unknown as Settings, + argv, + undefined, + [], + ); + + expect(config.getToolMode()).toBe(ToolMode.Direct); + expect(config.getWarnings()).toContain( + 'Unrecognized tools.mode "code-mode"; falling back to direct.', + ); + }); + + it('should honor the legacy tools.codeModeOnly setting', async () => { + process.argv = ['node', 'script.js']; + const argv = await parseArguments(); + const config = await loadCliConfig( + { tools: { codeModeOnly: true } } as unknown as Settings, + argv, + undefined, + [], + ); + + expect(config.getToolMode()).toBe(ToolMode.CodeModeOnly); + expect(config.getWarnings()).toContain( + 'tools.codeModeOnly is deprecated; use tools.mode = "code_mode_only".', + ); + }); + + it.each([ + ['--safe-mode', ToolMode.CodeMode], + ['--safe-mode', ToolMode.CodeModeOnly], + ['--bare', ToolMode.CodeMode], + ['--bare', ToolMode.CodeModeOnly], + ] as const)('should force direct mode for %s with %s', async (flag, mode) => { + process.argv = ['node', 'script.js', flag]; + const argv = await parseArguments(); + const config = await loadCliConfig( + { tools: { mode } }, + argv, + undefined, + [], + ); + + expect(config.getCodeModeOnly()).toBe(false); + expect(config.getToolMode()).toBe('direct'); }); it('should only enable tools.freeform inside CodeModeOnly', async () => { @@ -4423,30 +4504,41 @@ describe('mergeExcludeTools', () => { [], ); const codeMode = await loadCliConfig( - { tools: { codeModeOnly: true, freeform: true } }, + { tools: { mode: ToolMode.CodeModeOnly, freeform: true } }, argv, undefined, [], ); + const hybrid = await loadCliConfig( + { tools: { mode: ToolMode.CodeMode, freeform: true } }, + argv, + undefined, + [], + ); + expect(hybrid.getFreeform()).toBe(false); + expect(direct.getCodeModeOnly()).toBe(false); expect(direct.getFreeform()).toBe(false); expect(codeMode.getFreeform()).toBe(true); }); it.each(['--safe-mode', '--bare'])( - 'should disable CodeModeOnly and Freeform in %s mode', + 'should force direct mode for the legacy setting and disable Freeform with %s', async (flag) => { process.argv = ['node', 'script.js', flag]; const argv = await parseArguments(); const config = await loadCliConfig( - { tools: { codeModeOnly: true, freeform: true } }, + { + tools: { codeModeOnly: true, freeform: true }, + } as unknown as Settings, argv, undefined, [], ); expect(config.getCodeModeOnly()).toBe(false); + expect(config.getToolMode()).toBe(ToolMode.Direct); expect(config.getFreeform()).toBe(false); }, ); diff --git a/packages/cli/src/config/config.ts b/packages/cli/src/config/config.ts index 62b12623a3e..facdbc30e31 100755 --- a/packages/cli/src/config/config.ts +++ b/packages/cli/src/config/config.ts @@ -59,6 +59,11 @@ import { type OutputStyleDefinition, validateModelProvidersConfig, } from '@qwen-code/qwen-code-core'; +import { + isToolMode, + ToolMode, + type ToolMode as ToolModeValue, +} from '@qwen-code/qwen-code-core/tools/code-mode.js'; import { AGENT_HOST_SESSION_SOURCE_TYPE, AGENT_SESSION_SOURCE_TYPE, @@ -159,6 +164,32 @@ function isSkillLevel(value: unknown): value is SkillLevel { return SKILL_LEVELS.includes(value as SkillLevel); } +function resolveToolModeSetting(tools: Settings['tools']): { + mode: ToolModeValue; + warning?: string; +} { + const mode: unknown = tools?.mode; + if (mode !== undefined) { + return isToolMode(mode) + ? { mode } + : { + mode: ToolMode.Direct, + warning: `Unrecognized tools.mode ${JSON.stringify(mode)}; falling back to direct.`, + }; + } + + const legacyCodeModeOnly = (tools as { codeModeOnly?: unknown } | undefined) + ?.codeModeOnly; + if (legacyCodeModeOnly === true) { + return { + mode: ToolMode.CodeModeOnly, + warning: + 'tools.codeModeOnly is deprecated; use tools.mode = "code_mode_only".', + }; + } + return { mode: ToolMode.Direct }; +} + export interface CliArgs { query: string | undefined; model: string | undefined; @@ -2226,6 +2257,10 @@ export async function loadCliConfig( selectedAuthType, env: process.env as Record, }); + const resolvedToolMode = + bareMode || safeMode + ? { mode: ToolMode.Direct } + : resolveToolModeSetting(settings.tools); const { model: resolvedModel } = resolvedCliConfig; @@ -2560,8 +2595,7 @@ export async function loadCliConfig( disabledTools: disabledTools.length > 0 ? disabledTools : undefined, visibleTools: visibleTools.length > 0 ? visibleTools : undefined, eagerTools, - codeModeOnly: - !bareMode && !safeMode && settings.tools?.codeModeOnly === true, + toolMode: resolvedToolMode.mode, freeform: settings.tools?.freeform === true, toolSearchThreshold: bareMode || safeMode ? 0 : settings.tools?.toolSearch?.threshold, @@ -2733,7 +2767,10 @@ export async function loadCliConfig( generationConfigSources: resolvedCliConfig.sources, generationConfig: resolvedCliConfig.generationConfig, initialModelRegistryBaseUrl: resolvedCliConfig.registryBaseUrl, - warnings: resolvedCliConfig.warnings, + warnings: [ + ...resolvedCliConfig.warnings, + ...(resolvedToolMode.warning ? [resolvedToolMode.warning] : []), + ], bareMode, safeMode, allowedHttpHookUrls: @@ -2910,7 +2947,12 @@ export async function loadCliConfig( customIgnoreFiles: configParams.fileFiltering?.customIgnoreFiles, }, ); - configParams.codeModeOnly = false; + if (resolvedToolMode.mode !== ToolMode.Direct) { + configParams.warnings?.push( + `SSH workspaces do not support tools.mode = "${resolvedToolMode.mode}"; using direct tools for this session.`, + ); + } + configParams.toolMode = ToolMode.Direct; configParams.disableAllHooks = true; configParams.mcpServers = {}; configParams.overrideExtensions = []; diff --git a/packages/cli/src/config/settings-tool-mode.test.ts b/packages/cli/src/config/settings-tool-mode.test.ts new file mode 100644 index 00000000000..f63e06962c8 --- /dev/null +++ b/packages/cli/src/config/settings-tool-mode.test.ts @@ -0,0 +1,58 @@ +/** + * @license + * Copyright 2026 Qwen + * SPDX-License-Identifier: Apache-2.0 + */ +import { describe, expect, it, vi } from 'vitest'; +import { + getDisplayValue, + getEffectiveValue, + saveModifiedSettings, +} from './settingsUtils.js'; +import { + SettingScope, + type LoadedSettings, + type Settings, +} from './settings.js'; + +describe('legacy tool mode settings', () => { + const legacy = { tools: { codeModeOnly: true } } as unknown as Settings; + it('displays the effective legacy mode and honors an explicit mode', () => { + expect(getEffectiveValue('tools.mode', legacy, legacy)).toBe( + 'code_mode_only', + ); + expect(getDisplayValue('tools.mode', legacy, legacy, new Set())).toBe( + 'Code Mode Only*', + ); + expect( + getEffectiveValue('tools.mode', legacy, { tools: { mode: 'direct' } }), + ).toBe('direct'); + expect( + getEffectiveValue('tools.mode', { tools: { mode: 'code_mode' } }, legacy), + ).toBe('code_mode'); + }); + it('persists Direct so restart cannot reactivate the legacy mode', () => { + const settings = structuredClone(legacy); + const setValue = vi.fn( + (_scope: SettingScope, _key: string, value: 'direct') => { + settings.tools = { ...settings.tools, mode: value }; + }, + ); + const loaded = { + forScope: () => ({ settings }), + setValue, + } as unknown as LoadedSettings; + saveModifiedSettings( + new Set(['tools.mode']), + { tools: { mode: 'direct' } }, + loaded, + SettingScope.User, + ); + expect(setValue).toHaveBeenCalledWith( + SettingScope.User, + 'tools.mode', + 'direct', + ); + expect(getEffectiveValue('tools.mode', settings, settings)).toBe('direct'); + }); +}); diff --git a/packages/cli/src/config/settingsSchema.test.ts b/packages/cli/src/config/settingsSchema.test.ts index 538eba51f9e..61d6fd54100 100644 --- a/packages/cli/src/config/settingsSchema.test.ts +++ b/packages/cli/src/config/settingsSchema.test.ts @@ -32,6 +32,24 @@ import { describe('SettingsSchema', () => { describe('getSettingsSchema', () => { + it('distinguishes hybrid eager exclusions from allowlist narrowing in both code modes', () => { + const tools = getSettingsSchema().tools.properties; + const descriptions = [ + tools.mode.description, + tools.eager.description, + tools.toolSearch.properties.threshold.description, + ]; + + for (const description of descriptions) { + expect(description).toContain( + 'In Hybrid mode, AgentCore excludes tools still hidden by tools.eager from nested bindings.', + ); + expect(description).toContain( + 'In both code modes, agent allowlists that do not grant exec narrow nested bindings.', + ); + } + }); + it('should describe prompt hooks supported by the runtime', () => { const hookProperties = getSettingsSchema().hooks.properties.PreToolUse.items.properties?.[ @@ -232,6 +250,23 @@ describe('SettingsSchema', () => { }); }); + it('should expose the tool mode enum', () => { + expect(getSettingsSchema().tools.properties.mode).toMatchObject({ + type: 'enum', + default: 'direct', + requiresRestart: true, + showInDialog: true, + options: [ + { value: 'direct', label: 'Default' }, + { value: 'code_mode', label: 'Code Mode' }, + { value: 'code_mode_only', label: 'Code Mode Only' }, + ], + }); + expect(getSettingsSchema().tools.properties).not.toHaveProperty( + 'codeModeOnly', + ); + }); + it('should expose cumulative tool result threshold in clearContextOnIdle', () => { const threshold = getSettingsSchema().context.properties.clearContextOnIdle.properties diff --git a/packages/cli/src/config/settingsSchema.ts b/packages/cli/src/config/settingsSchema.ts index 7390a7e3e6d..e9a61d443cb 100644 --- a/packages/cli/src/config/settingsSchema.ts +++ b/packages/cli/src/config/settingsSchema.ts @@ -35,6 +35,7 @@ import { REASONING_EFFORT_TIERS, SENSITIVE_SPAN_ATTRIBUTE_MAX_LENGTH_LIMIT, } from '@qwen-code/qwen-code-core'; +import { ToolMode } from '@qwen-code/qwen-code-core/tools/code-mode.js'; import type { CustomTheme } from '../ui/themes/theme.js'; import { getLanguageSettingsOptions } from '../i18n/languages.js'; import { MergeStrategy } from '../utils/deepMerge.js'; @@ -2903,15 +2904,20 @@ const SETTINGS_SCHEMA = { }, }, }, - codeModeOnly: { - type: 'boolean', - label: 'Code Mode Only (Experimental)', + mode: { + type: 'enum', + label: 'Tool Mode (Experimental)', category: 'Tools', requiresRestart: true, - default: false, + default: ToolMode.Direct, description: - 'Expose ordinary tools through the isolated exec JavaScript tool. Load deferred descriptions and schemas on demand with tool_search; if search is unavailable, include all allowed signatures in exec. Direct control tools remain available. Ignored in safe and bare modes.', + 'Choose how tools are exposed to the model. Direct uses ordinary tool calls; Code Mode also exposes the isolated exec JavaScript tool; Code Mode Only exposes ordinary tools only through exec. Safe and bare modes always use Direct. Container execution warns and uses direct tools for Code Mode, and rejects Code Mode Only. SSH workspaces warn and use Direct for either code mode. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by tools.eager from nested bindings. In both code modes, agent allowlists that do not grant exec narrow nested bindings. Inheriting or explicitly granting exec keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns.', showInDialog: true, + options: [ + { value: ToolMode.Direct, label: 'Default' }, + { value: ToolMode.CodeMode, label: 'Code Mode' }, + { value: ToolMode.CodeModeOnly, label: 'Code Mode Only' }, + ], }, freeform: { type: 'boolean', @@ -2920,7 +2926,7 @@ const SETTINGS_SCHEMA = { requiresRestart: true, default: false, description: - 'Use raw text input for the Code Mode exec tool on OpenAI Responses models. Effective only when tools.codeModeOnly is true and the selected model uses wireApi "responses". Enable only for endpoints that support Responses Custom Tools.', + 'Use raw text input for the Code Mode exec tool on OpenAI Responses models. Effective only when tools.mode is "code_mode_only" and the selected model uses wireApi "responses". Enable only for endpoints that support Responses Custom Tools.', showInDialog: false, }, sandbox: { @@ -3036,7 +3042,7 @@ const SETTINGS_SCHEMA = { requiresRestart: true, default: 0, description: - 'Context-window percentage used as the session-start budget for preloading ordinary deferred tools (bundled built-ins and MCP alike). Defaults to 0, which performs no threshold-based preload; ordinary deferred tools normally stay behind the stable ToolSearch + ToolCall bridge, at the cost of one tool_search round trip before first use. Raise it to N so that, when every eligible deferred schema fits within N% of the context window, all are declared upfront for direct calls with no bridge round trip; otherwise they stay behind the bridge while both bridge tools are registered. Tools demoted by tools.eager are excluded from this preload and stay reachable on demand through that bridge while it is registered; when either bridge tool is unregistered (tools.toolSearch.enabled false denies both; a tool_search or tool_call deny rule removes one) the demoted tools that remain hidden are not offered to the model and cannot be reached through the bridge for that session, and a warning is logged; these bridge and warning rules apply to direct tool mode. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager also reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed tool signatures. In direct mode they stay registered, so a direct call by their own name is still evaluated and approved normally. Separate paths can still declare deferred tools at 0: tools.visible; the live-history compatibility scan on every tool-set refresh (including resume, MCP discovery, first plan-mode entry, and subagent definition changes); the incomplete-bridge eager fallback; and daemon ACP late registration, which explicitly reveals and pins create_sub_session.', + 'Context-window percentage used as the session-start budget for preloading ordinary deferred tools (bundled built-ins and MCP alike). Defaults to 0, which performs no threshold-based preload; ordinary deferred tools normally stay behind the stable ToolSearch + ToolCall bridge, at the cost of one tool_search round trip before first use. Raise it to N so that, when every eligible deferred schema fits within N% of the context window, all are declared upfront for direct calls with no bridge round trip; otherwise they stay behind the bridge while both bridge tools are registered. Tools demoted by tools.eager are excluded from this preload and stay reachable on demand through that bridge while it is registered; when either bridge tool is unregistered (tools.toolSearch.enabled false denies both; a tool_search or tool_call deny rule removes one) the demoted tools that remain hidden are absent from top-level declarations and cannot be reached through the bridge for that session, and a warning is logged; these bridge and warning rules apply to direct and hybrid tool modes. In Code Mode, withheld tools with actual exec bindings also remain callable through exec; the warning names that subset. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by tools.eager from nested bindings. In both code modes, agent allowlists that do not grant exec narrow nested bindings. Inheriting or explicitly granting exec keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns. In direct and hybrid modes they stay registered, so a direct call by their own name is still evaluated and approved normally. Separate paths can still declare deferred tools at 0: tools.visible; the live-history compatibility scan on every tool-set refresh (including resume, MCP discovery, first plan-mode entry, and subagent definition changes); the incomplete-bridge eager fallback; and daemon ACP late registration, which explicitly reveals and pins create_sub_session.', showInDialog: true, // A percentage of the context window: values above 100 would set a // budget larger than the window and unconditionally preload every @@ -3207,7 +3213,7 @@ const SETTINGS_SCHEMA = { requiresRestart: true, default: undefined as string[] | undefined, description: - 'Deferred tool names made visible at startup without requiring the ToolSearch + ToolCall bridge. Listed tools appear alongside core tools in the initial session.', + 'Deferred tool names made visible at startup without requiring the ToolSearch + ToolCall bridge. Listed tools appear alongside core tools in the initial session. In Code Mode Only, listed tools have signatures in the initial exec description; other deferred schemas are discovered through top-level tool_search.', showInDialog: false, mergeStrategy: MergeStrategy.UNION, }, @@ -3218,7 +3224,7 @@ const SETTINGS_SCHEMA = { requiresRestart: true, default: undefined as string[] | undefined, description: - 'Allowlist of eager-by-default built-in tool names whose schemas remain eligible for the initial model request. Unlisted non-exempt tools are deferred but stay registered, listed in /tools, and reachable through the tool_search + tool_call bridge. Tools already deferred by default stay on demand even when listed; use tools.visible to surface one at startup. tool_search, tool_call, structured_output, plan-mode lifecycle tools, task_stop, MCP tools, and computer_use__* tools are unaffected. An explicitly empty list ([]) defers every non-exempt eager-by-default tool; omit the setting for no restriction. Pairs with the ToolSearch + ToolCall bridge: when either half is not registered — tools.toolSearch.enabled false (which denies both), a tool_search or tool_call deny rule, or a tools.disabled entry — the allowlist still withholds the schemas, but nothing can load them back, so the demoted tools that remain hidden are not offered to the model and cannot be reached through the bridge for that session, and a warning is logged; these bridge and warning rules apply to direct tool mode. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager also reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed tool signatures. In direct mode they stay registered, so a direct call by their own name is still evaluated and approved normally — except tools also listed in tools.visible, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares. Differs from tools.disabled, which removes tools entirely, and from permissions.allow, which only auto-approves calls.', + 'Allowlist of eager-by-default built-in tool names whose schemas remain eligible for the initial model request. Unlisted non-exempt tools are deferred but stay registered, listed in /tools, and reachable through the tool_search + tool_call bridge. Tools already deferred by default stay on demand even when listed; use tools.visible to surface one at startup. tool_search, tool_call, structured_output, plan-mode lifecycle tools, task_stop, MCP tools, and computer_use__* tools are unaffected. An explicitly empty list ([]) defers every non-exempt eager-by-default tool; omit the setting for no restriction. Pairs with the ToolSearch + ToolCall bridge: when either half is not registered — tools.toolSearch.enabled false (which denies both), a tool_search or tool_call deny rule, or a tools.disabled entry — the allowlist still withholds the schemas, but nothing can load them back, so the demoted tools that remain hidden are absent from top-level declarations and cannot be reached through the bridge for that session, and a warning is logged; these bridge and warning rules apply to direct and hybrid tool modes. In Code Mode, withheld tools with actual exec bindings also remain callable through exec; the warning names that subset. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by tools.eager from nested bindings. In both code modes, agent allowlists that do not grant exec narrow nested bindings. Inheriting or explicitly granting exec keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns. In direct and hybrid modes they stay registered, so a direct call by their own name is still evaluated and approved normally — except tools also listed in tools.visible, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares. Differs from tools.disabled, which removes tools entirely, and from permissions.allow, which only auto-approves calls.', showInDialog: false, }, approvalMode: { diff --git a/packages/cli/src/config/settingsUtils.ts b/packages/cli/src/config/settingsUtils.ts index fda3b694f57..1f035342e02 100644 --- a/packages/cli/src/config/settingsUtils.ts +++ b/packages/cli/src/config/settingsUtils.ts @@ -137,6 +137,20 @@ export function getEffectiveValue( return value as SettingsValue; } + if ( + key === 'tools.mode' && + (getNestedValue(mergedSettings as Record, [ + 'tools', + 'codeModeOnly', + ]) ?? + getNestedValue(settings as Record, [ + 'tools', + 'codeModeOnly', + ])) === true + ) { + return 'code_mode_only'; + } + // Return default value if no value is set anywhere return definition.default; } @@ -165,7 +179,7 @@ export function getAllSettingKeys(): string[] { const SETTINGS_DIALOG_ORDER: readonly string[] = [ // Workflow Control - most impactful setting 'tools.approvalMode', - 'tools.codeModeOnly', + 'tools.mode', // Localization - users often set this first 'general.language', @@ -609,7 +623,12 @@ export function saveModifiedSettings( const isDefaultValue = value === getDefaultValue(settingKey); - if (existsInOriginalFile || !isDefaultValue) { + // An explicit mode selection must override legacy codeModeOnly too. + if ( + existsInOriginalFile || + !isDefaultValue || + settingKey === 'tools.mode' + ) { loadedSettings.setValue(scope, settingKey, value); } }); @@ -636,8 +655,8 @@ export function getDisplayValue( // Show the value defined at the current scope if present value = getEffectiveValue(key, settings, {}); } else { - // Fall back to the schema default when the key is unset in this scope - value = getDefaultValue(key); + // Include legacy values while keeping the display scoped to this file. + value = getEffectiveValue(key, settings, {}); } let valueString = value === undefined ? t('(not set)') : String(value); diff --git a/packages/cli/src/i18n/locales/en.js b/packages/cli/src/i18n/locales/en.js index 1b143aafcab..3890534d7bc 100644 --- a/packages/cli/src/i18n/locales/en.js +++ b/packages/cli/src/i18n/locales/en.js @@ -760,7 +760,10 @@ export default { // ============================================================================ // Settings Labels // ============================================================================ - 'Code Mode Only (Experimental)': 'Code Mode Only (Experimental)', + 'Tool Mode (Experimental)': 'Tool Mode (Experimental)', + Default: 'Default', + 'Code Mode': 'Code Mode', + 'Code Mode Only': 'Code Mode Only', 'Vim Mode': 'Vim Mode', 'Attribution: commit': 'Attribution: commit', 'Terminal Bell Notification': 'Terminal Bell Notification', diff --git a/packages/cli/src/i18n/locales/zh-TW.js b/packages/cli/src/i18n/locales/zh-TW.js index e4291c3ff11..6097e8e79c9 100644 --- a/packages/cli/src/i18n/locales/zh-TW.js +++ b/packages/cli/src/i18n/locales/zh-TW.js @@ -719,7 +719,10 @@ export default { Settings: '設置', 'To see changes, Qwen Code must be restarted. Press r to exit and apply changes now.': '要查看更改,必須重啟 Qwen Code。按 r 退出並立即應用更改。', - 'Code Mode Only (Experimental)': '僅程式碼模式(實驗性)', + 'Tool Mode (Experimental)': '工具模式(實驗性)', + Default: '預設', + 'Code Mode': '程式碼模式', + 'Code Mode Only': '僅程式碼模式', 'Vim Mode': 'Vim 模式', 'Attribution: commit': '署名:提交', 'Terminal Bell Notification': '終端響鈴通知', diff --git a/packages/cli/src/i18n/locales/zh.js b/packages/cli/src/i18n/locales/zh.js index ac05d5f20d1..3bbb94f46b4 100644 --- a/packages/cli/src/i18n/locales/zh.js +++ b/packages/cli/src/i18n/locales/zh.js @@ -756,7 +756,10 @@ export default { // ============================================================================ // Settings Labels // ============================================================================ - 'Code Mode Only (Experimental)': '仅代码模式(实验性)', + 'Tool Mode (Experimental)': '工具模式(实验性)', + Default: '默认', + 'Code Mode': '代码模式', + 'Code Mode Only': '仅代码模式', 'Vim Mode': 'Vim 模式', 'Attribution: commit': '署名:提交', 'Terminal Bell Notification': '终端响铃通知', diff --git a/packages/cli/src/i18n/mustTranslateKeys.test.ts b/packages/cli/src/i18n/mustTranslateKeys.test.ts index ed10e8b3eb2..8f2a2bedba4 100644 --- a/packages/cli/src/i18n/mustTranslateKeys.test.ts +++ b/packages/cli/src/i18n/mustTranslateKeys.test.ts @@ -35,6 +35,13 @@ const FORK_COMMAND_REQUIRED_KEYS = [ 'Forked into a background agent. It inherits this conversation and runs without blocking — track it in the background tasks panel; it reports back when done.', ] as const; +const TOOL_MODE_TRANSLATION_KEYS = [ + 'Tool Mode (Experimental)', + 'Default', + 'Code Mode', + 'Code Mode Only', +] as const; + const NON_ENGLISH_LANGUAGES = SUPPORTED_LANGUAGES.filter( (language) => language.code !== 'en', ); @@ -119,6 +126,17 @@ describe('must-translate locale coverage', () => { SLOW_LOCALE_TEST_TIMEOUT_MS, ); + it.each(['zh', 'zh-TW'] as const)( + 'translates tool mode settings in %s', + async (language) => { + await setLanguageAsync(language); + + expect( + TOOL_MODE_TRANSLATION_KEYS.filter((key) => t(key) === key), + ).toEqual([]); + }, + ); + it.each(NON_ENGLISH_LANGUAGES)( 'translates built-in command descriptions in %s', async (language) => { diff --git a/packages/cli/src/serve/routes/workspace-settings.test.ts b/packages/cli/src/serve/routes/workspace-settings.test.ts index cbcc010a992..d1b21481959 100644 --- a/packages/cli/src/serve/routes/workspace-settings.test.ts +++ b/packages/cli/src/serve/routes/workspace-settings.test.ts @@ -271,6 +271,75 @@ function makeQualifiedApp( }; } +describe('GET tool mode settings', () => { + it.each([ + { legacy: true, mode: undefined, expected: 'code_mode_only' }, + { legacy: false, mode: undefined, expected: 'direct' }, + { legacy: 'true', mode: undefined, expected: 'direct' }, + { legacy: true, mode: 'direct', expected: 'direct' }, + { legacy: true, mode: 'code_mode', expected: 'code_mode' }, + ])( + 'serves the effective mode without migrating settings: $legacy/$mode', + async ({ legacy, mode, expected }) => { + const userSettings = { tools: { codeModeOnly: legacy } }; + const workspaceSettings = mode === undefined ? {} : { tools: { mode } }; + const { app, persistSetting } = makeApp({ + userSettings, + workspaceSettings, + }); + const response = await request(app).get('/workspace/settings'); + expect(response.status).toBe(200); + const toolMode = response.body.settings.find( + (row: { key: string }) => row.key === 'tools.mode', + ); + expect(toolMode.values.effective).toBe(expected); + expect(toolMode.values.user).toBeUndefined(); + expect(toolMode.values.workspace).toBe(mode); + expect(persistSetting).not.toHaveBeenCalled(); + expect(userSettings).toEqual({ tools: { codeModeOnly: legacy } }); + }, + ); + + it('uses the selected workspace legacy mode without reading or writing a primary fallback', async () => { + const { app, persistSetting } = makeQualifiedApp({ + workspaceCwd: '/selected-workspace', + }); + vi.mocked(loadSettings).mockImplementation( + (workspace) => + ({ + merged: { + tools: + workspace === '/selected-workspace' + ? { codeModeOnly: true } + : { mode: 'direct' }, + }, + user: { settings: {} }, + workspace: { settings: {} }, + forScope: vi.fn().mockReturnValue({ settings: {} }), + }) as never, + ); + vi.mocked(loadSettings).mockClear(); + const response = await request(app).get('/workspaces/primary/settings'); + expect(response.status).toBe(200); + expect( + response.body.settings.find( + (row: { key: string }) => row.key === 'tools.mode', + ).values.effective, + ).toBe('code_mode_only'); + expect(loadSettings).toHaveBeenCalledWith('/selected-workspace', { + skipLoadEnvironment: true, + skipWorkspaceSettings: false, + workspaceTrusted: true, + }); + expect( + vi + .mocked(loadSettings) + .mock.calls.every(([workspace]) => workspace === '/selected-workspace'), + ).toBe(true); + expect(persistSetting).not.toHaveBeenCalled(); + }); +}); + describe('POST /workspace/settings', () => { it('updates live sessions when Session Workflow changes', async () => { // Seeded as the post-write state: the route reads back the effective value. diff --git a/packages/cli/src/serve/routes/workspace-settings.ts b/packages/cli/src/serve/routes/workspace-settings.ts index 38f8087e5d5..7bb3f9f1dc3 100644 --- a/packages/cli/src/serve/routes/workspace-settings.ts +++ b/packages/cli/src/serve/routes/workspace-settings.ts @@ -19,6 +19,7 @@ import type { } from '../../config/settingsSchema.js'; import { getDialogSettingKeys, + getEffectiveValue, getNestedProperty, getSettingDefinition, validateSettingValue, @@ -190,10 +191,10 @@ function buildSettingsResponse( const def = getSettingDefinition(key); if (!def) continue; - const mergedEffective = getNestedProperty( - loaded.merged as Record, - key, - ); + const mergedEffective = + key === 'tools.mode' + ? getEffectiveValue(key, {}, loaded.merged) + : getNestedProperty(loaded.merged as Record, key); const userVal = getNestedProperty( loaded.user.settings as Record, key, diff --git a/packages/cli/src/ui/commands/config-command.test.ts b/packages/cli/src/ui/commands/config-command.test.ts index bdf74bc943a..cdce674473c 100644 --- a/packages/cli/src/ui/commands/config-command.test.ts +++ b/packages/cli/src/ui/commands/config-command.test.ts @@ -43,6 +43,32 @@ describe('configCommand', () => { expect(configCommand.description).toBeTruthy(); }); + it.each([ + { tools: { codeModeOnly: true }, expected: 'code_mode_only' }, + { tools: { codeModeOnly: false }, expected: 'direct' }, + { tools: { codeModeOnly: 'true' }, expected: 'direct' }, + { tools: { codeModeOnly: true, mode: 'direct' }, expected: 'direct' }, + { tools: { codeModeOnly: true, mode: 'code_mode' }, expected: 'code_mode' }, + ])( + 'reads effective tool mode without persisting legacy settings: $tools', + async ({ tools, expected }) => { + const { ctx, setValuesMock } = createMockContext({ tools }); + const before = structuredClone(ctx.services.settings.merged); + const listing = await configCommand.action!(ctx, '--help'); + const value = await configCommand.action!(ctx, 'tools.mode'); + + expect(listing).toMatchObject({ + type: 'message', + content: expect.stringMatching( + new RegExp(`^tools\\.mode\\s+enum\\s+${expected}`, 'm'), + ), + }); + expect(value).toMatchObject({ content: `tools.mode = ${expected}` }); + expect(setValuesMock).not.toHaveBeenCalled(); + expect(ctx.services.settings.merged).toEqual(before); + }, + ); + describe('set boolean value', () => { it('sets a boolean setting to true', async () => { const { ctx, setValuesMock } = createMockContext({ diff --git a/packages/cli/src/ui/commands/config-command.ts b/packages/cli/src/ui/commands/config-command.ts index 0b10c3062c9..1048a68f772 100644 --- a/packages/cli/src/ui/commands/config-command.ts +++ b/packages/cli/src/ui/commands/config-command.ts @@ -15,6 +15,7 @@ import type { SettingDefinition } from '../../config/settingsSchema.js'; import { t } from '../../i18n/index.js'; import { getAllSettingKeys, + getEffectiveValue, getFlattenedSchema, getDefaultValue, getNestedProperty, @@ -220,7 +221,10 @@ function listAllSettings(context: CommandContext): MessageActionReturn { const def = flattened[key]!; if (!SETTABLE_TYPES.has(def.type)) continue; - const current = getNestedProperty(merged as Record, key); + const current = + key === 'tools.mode' + ? getEffectiveValue(key, {}, merged) + : getNestedProperty(merged as Record, key); let displayCurrent: string; if (isSensitiveKey(key)) { displayCurrent = maskValue(current ?? def.default); @@ -283,10 +287,13 @@ export const configCommand: SlashCommand = { }; } - const currentValue = getNestedProperty( - context.services.settings.merged as Record, - key, - ); + const currentValue = + key === 'tools.mode' + ? getEffectiveValue(key, {}, context.services.settings.merged) + : getNestedProperty( + context.services.settings.merged as Record, + key, + ); if (isToggle && def.type !== 'boolean') { const display = isSensitiveKey(key) diff --git a/packages/cli/src/ui/commands/contextCommand.test.ts b/packages/cli/src/ui/commands/contextCommand.test.ts index 972b3501a17..17c5f4d120e 100644 --- a/packages/cli/src/ui/commands/contextCommand.test.ts +++ b/packages/cli/src/ui/commands/contextCommand.test.ts @@ -138,6 +138,90 @@ describe('collectContextData (contextCommand)', () => { expect(getFunctionDeclarationsSpy).toHaveBeenCalledWith(); }); + it('accounts for decorated CodeMode declarations in per-tool rows', async () => { + const execTool = { + name: 'exec', + schema: { name: 'exec', description: 'raw' }, + }; + const decoratedExec = { + name: 'exec', + description: 'decorated '.repeat(200), + }; + const config = { + ...mockConfig, + getToolRegistry: vi.fn().mockReturnValue({ + getAllTools: vi.fn().mockReturnValue([execTool]), + getFunctionDeclarations: vi.fn().mockReturnValue([decoratedExec]), + isDeferredAndHidden: vi.fn().mockReturnValue(false), + }), + } as unknown as Config; + + const data = await collectContextData(config, true); + + expect(data.builtinTools).toHaveLength(1); + expect(data.builtinTools[0]?.name).toBe('exec'); + expect(data.builtinTools[0]?.tokens).toBeGreaterThan(100); + }); + + it('counts only emitted declarations when ordinary tools are nested in exec', async () => { + const decoratedExec = { + name: 'exec', + description: 'tools.read_file and tools.skill '.repeat(100), + }; + const config = { + ...mockConfig, + getToolRegistry: vi.fn().mockReturnValue({ + getAllTools: () => + ['exec', 'read_file', 'skill'].map((name) => ({ + name, + schema: { name, description: 'raw schema '.repeat(20) }, + })), + getFunctionDeclarations: () => [decoratedExec], + isDeferredAndHidden: () => false, + }), + } as unknown as Config; + + const data = await collectContextData(config, true); + + expect(data.builtinTools.map((tool) => tool.name)).toEqual(['exec']); + expect(data.breakdown.skills).toBe(0); + expect(data.breakdown.builtinTools).toBe( + estimateContextTextTokens(JSON.stringify([decoratedExec])), + ); + }); + + it('charges the decorated Skill declaration once and leaves the residual in built-ins', async () => { + const decoratedExec = { name: 'exec', description: 'exec '.repeat(100) }; + const decoratedSkill = { + name: 'skill', + description: 'skill plus its nested declaration '.repeat(100), + }; + const declarations = [decoratedExec, decoratedSkill]; + const config = { + ...mockConfig, + getToolRegistry: vi.fn().mockReturnValue({ + getAllTools: () => + ['exec', 'skill'].map((name) => ({ + name, + schema: { name, description: 'raw' }, + })), + getFunctionDeclarations: () => declarations, + isDeferredAndHidden: () => false, + }), + } as unknown as Config; + + const data = await collectContextData(config, true); + const skillTokens = estimateContextTextTokens( + JSON.stringify(decoratedSkill), + ); + + expect(data.breakdown.skills).toBe(skillTokens); + expect(data.breakdown.builtinTools).toBe( + estimateContextTextTokens(JSON.stringify(declarations)) - skillTokens, + ); + expect(data.builtinTools.map((tool) => tool.name)).toEqual(['exec']); + }); + it('reads the per-session chat token count, not the process-global singleton (#5763)', async () => { // uiTelemetryService is a module-level singleton shared by every session // in a `serve` daemon. Reading it here would report whichever session most @@ -1001,13 +1085,10 @@ describe('collectContextData (contextCommand)', () => { expect(sumRows(data.breakdown)).toBe(100_000); }); - it('charges the builtin-clamp deficit to the mcp row, not to messages', async () => { - // Under `tools.codeModeOnly` the declarations collapse to a few control - // tools while an `alwaysLoadTools` MCP server still bills every schema - // the detail loop sees, so the billed tools exceed the declared ones and - // `displayBuiltinTools` clamps at 0. The clamp deficit must come out of - // the mcp row — the row whose billing overshoots the declarations; - // otherwise `attributedOverhead` silently takes it out of `messages`. + it('bills the declared mcp schema to the mcp row, not to messages', async () => { + // The detail loop bills only declared tools, so a declared MCP schema + // lands on the mcp row and the rows must still partition the provider + // total exactly: `messages` absorbs only the calibrated remainder. // Own value properties shadow the prototype's getters (Object.assign // would trip the setter-less `schema` accessor on DeclarativeTool). const mcpToolDouble = Object.defineProperties( @@ -1034,7 +1115,7 @@ describe('collectContextData (contextCommand)', () => { mcpToolDouble, { name: controlSchema.name, schema: controlSchema }, ]; - const declared = [skillToolSchema, controlSchema]; + const declared = [skillToolSchema, mcpToolDouble.schema, controlSchema]; const history = [prelude, ...conversation]; const unscaled = await collectContextData( @@ -1050,32 +1131,49 @@ describe('collectContextData (contextCommand)', () => { unscaled.breakdown.autocompactBuffer, ), ); - // The deficit comes out of the mcp row, not the built-in or skills rows. expect(unscaled.breakdown.mcpTools).toBe( + estimateContextTextTokens(JSON.stringify(mcpToolDouble.schema)), + ); + expect(unscaled.breakdown.builtinTools).toBe( estimateContextTextTokens(JSON.stringify(declared)) - - estimateContextTextTokens(JSON.stringify(skillToolSchema)), + estimateContextTextTokens(JSON.stringify(skillToolSchema)) - + estimateContextTextTokens(JSON.stringify(mcpToolDouble.schema)), ); - expect(unscaled.breakdown.mcpTools).toBeGreaterThan(0); - expect(unscaled.breakdown.builtinTools).toBe(0); - // With nothing declared the mcp row cannot absorb the whole deficit, so - // the rest is charged to skills and the window still adds up. + // The hidden MCP schema does not consume the declared control's budget. + const hiddenMcp = await collectContextData( + makeChatConfig({ + total: 0, + tools, + declared: [skillToolSchema, controlSchema], + history, + }), + false, + ); + expect(hiddenMcp.breakdown.mcpTools).toBe(0); + expect(hiddenMcp.breakdown.builtinTools).toBe( + estimateContextTextTokens( + JSON.stringify([skillToolSchema, controlSchema]), + ) - estimateContextTextTokens(JSON.stringify(skillToolSchema)), + ); + // With no declared schemas, only the listing is billed to skills. const undeclared = await collectContextData( makeChatConfig({ total: 0, tools, declared: [], history }), false, ); expect(undeclared.breakdown.mcpTools).toBe(0); expect(undeclared.breakdown.skills).toBe( - estimateContextTextTokens(listingReminder) + - estimateContextTextTokens(JSON.stringify([])), - ); - expect(undeclared.breakdown.freeSpace).toBe( - Math.max( - 0, - undeclared.contextWindowSize - - sumRows(undeclared.breakdown) - - undeclared.breakdown.autocompactBuffer, - ), + estimateContextTextTokens(listingReminder), ); + for (const estimate of [hiddenMcp, undeclared]) { + expect(estimate.breakdown.freeSpace).toBe( + Math.max( + 0, + estimate.contextWindowSize - + sumRows(estimate.breakdown) - + estimate.breakdown.autocompactBuffer, + ), + ); + } // The provider-side total: the measured overhead plus the 300-token // conversation, so exactly 300 tokens are left for `messages`. const total = @@ -1089,12 +1187,10 @@ describe('collectContextData (contextCommand)', () => { false, ); - // The fixture does put the clamp in force: billed skill definition plus - // mcp schemas exceed the declared tools. - expect( - estimateContextTextTokens(JSON.stringify(skillToolSchema)) + - estimateContextTextTokens(JSON.stringify(mcpToolDouble.schema)), - ).toBeGreaterThan(estimateContextTextTokens(JSON.stringify(declared))); + // The mcp schema is declared, so the guard bills it to the mcp row. + expect(data.breakdown.mcpTools).toBe( + estimateContextTextTokens(JSON.stringify(mcpToolDouble.schema)), + ); expect(data.breakdown.messages).toBe(300); expect(sumRows(data.breakdown)).toBe(total); diff --git a/packages/cli/src/ui/commands/contextCommand.ts b/packages/cli/src/ui/commands/contextCommand.ts index 6da6a5b0461..e3fcd022cad 100644 --- a/packages/cli/src/ui/commands/contextCommand.ts +++ b/packages/cli/src/ui/commands/contextCommand.ts @@ -424,6 +424,11 @@ export async function collectContextData( const toolDeclarations = toolRegistry ? toolRegistry.getFunctionDeclarations() : []; + const toolDeclarationsByName = new Map( + toolDeclarations + .filter((declaration) => declaration.name) + .map((declaration) => [declaration.name!, declaration] as const), + ); const toolsJsonStr = JSON.stringify(toolDeclarations); const allToolsTokens = estimateContextTextTokens(toolsJsonStr); @@ -440,7 +445,9 @@ export async function collectContextData( if (isMediaPolicyToolHiddenFromModel(config, tool)) { continue; } - const toolJsonStr = JSON.stringify(tool.schema); + const declaration = toolDeclarationsByName.get(tool.name); + if (!declaration) continue; + const toolJsonStr = JSON.stringify(declaration); const tokens = estimateContextTextTokens(toolJsonStr); if (tool instanceof DiscoveredMCPTool) { mcpTools.push({ @@ -471,8 +478,9 @@ export async function collectContextData( const memoryFilesTokens = memoryFiles.reduce((sum, f) => sum + f.tokens, 0); const skillTool = allTools.find((tool) => tool.name === ToolNames.SKILL); - const skillToolDefinitionTokens = skillTool - ? estimateContextTextTokens(JSON.stringify(skillTool.schema)) + const skillDeclaration = toolDeclarationsByName.get(ToolNames.SKILL); + const skillToolDefinitionTokens = skillDeclaration + ? estimateContextTextTokens(JSON.stringify(skillDeclaration)) : 0; const loadedContentNames: ReadonlyMap = diff --git a/packages/cli/src/ui/components/SettingsDialog.test.tsx b/packages/cli/src/ui/components/SettingsDialog.test.tsx index 2b06cbb3ee1..4d012246c2f 100644 --- a/packages/cli/src/ui/components/SettingsDialog.test.tsx +++ b/packages/cli/src/ui/components/SettingsDialog.test.tsx @@ -1241,6 +1241,55 @@ describe('SettingsDialog', () => { unmount(); }); + it.each(['\u0003', '\u000c'])( + 'persists resetting legacy Tool Mode with %j and requests restart', + async (key) => { + const settings = createMockSettings({ tools: { codeModeOnly: true } }); + const setValue = vi + .spyOn(settings, 'setValue') + .mockImplementation(() => {}); + const actual = await vi.importActual< + typeof import('../../config/settingsUtils.js') + >('../../config/settingsUtils.js'); + vi.mocked(saveModifiedSettings).mockImplementation( + actual.saveModifiedSettings, + ); + const onRestartRequest = vi.fn(); + const { stdin, lastFrame, unmount } = render( + + + , + ); + try { + await waitFor(() => expect(lastFrame()).toContain('Code Mode Only')); + stdin.write(TerminalKeys.DOWN_ARROW); + await wait(30); + stdin.write(key); + await waitFor(() => + expect(lastFrame()).toContain( + 'Press r to exit and apply changes now', + ), + ); + stdin.write('r'); + await waitFor(() => + expect(setValue).toHaveBeenCalledWith( + SettingScope.User, + 'tools.mode', + 'direct', + ), + ); + expect(onRestartRequest).toHaveBeenCalledOnce(); + } finally { + unmount(); + } + }, + ); + it('should handle Ctrl+C to reset current setting to default', async () => { const settings = createMockSettings({ vimMode: true }); // Start with vimMode enabled const onSelect = vi.fn(); diff --git a/packages/cli/src/ui/components/SettingsDialog.tsx b/packages/cli/src/ui/components/SettingsDialog.tsx index 548b2a20077..98c8f283dbb 100644 --- a/packages/cli/src/ui/components/SettingsDialog.tsx +++ b/packages/cli/src/ui/components/SettingsDialog.tsx @@ -1016,9 +1016,8 @@ export function SettingsDialog({ const scopeSettings = settings.forScope(selectedScope).settings; const resetChangesValue = - !isDefaultValue(currentSetting.value, scopeSettings) && getEffectiveValue(currentSetting.value, scopeSettings, {}) !== - defaultValue; + defaultValue; setModifiedSettings((prev) => { const updated = new Set(prev); if (resetChangesValue) updated.add(currentSetting.value); diff --git a/packages/cli/src/ui/components/__snapshots__/SettingsDialog.test.tsx.snap b/packages/cli/src/ui/components/__snapshots__/SettingsDialog.test.tsx.snap index 3237a4a8a50..0be7e032d63 100644 --- a/packages/cli/src/ui/components/__snapshots__/SettingsDialog.test.tsx.snap +++ b/packages/cli/src/ui/components/__snapshots__/SettingsDialog.test.tsx.snap @@ -10,7 +10,7 @@ exports[`SettingsDialog > Snapshot Tests > should render default state correctly │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ @@ -35,7 +35,7 @@ exports[`SettingsDialog > Snapshot Tests > should render focused on scope select │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ @@ -60,7 +60,7 @@ exports[`SettingsDialog > Snapshot Tests > should render with accessibility sett │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ @@ -85,7 +85,7 @@ exports[`SettingsDialog > Snapshot Tests > should render with all boolean settin │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ @@ -110,7 +110,7 @@ exports[`SettingsDialog > Snapshot Tests > should render with different scope se │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ @@ -135,7 +135,7 @@ exports[`SettingsDialog > Snapshot Tests > should render with different scope se │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ @@ -160,7 +160,7 @@ exports[`SettingsDialog > Snapshot Tests > should render with file filtering set │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ @@ -185,7 +185,7 @@ exports[`SettingsDialog > Snapshot Tests > should render with mixed boolean and │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ @@ -210,7 +210,7 @@ exports[`SettingsDialog > Snapshot Tests > should render with tools and security │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ @@ -235,7 +235,7 @@ exports[`SettingsDialog > Snapshot Tests > should render with various boolean se │ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ●︎ Tool Approval Mode Auto │ -│ Code Mode Only (Experimental) false │ +│ Tool Mode (Experimental) Default │ │ Language: UI Auto (detect from system) │ │ Language: Model Auto (follow user input) │ │ Theme Qwen Dark ▸ │ diff --git a/packages/cli/src/ui/opentui/dialogs-settings.test.ts b/packages/cli/src/ui/opentui/dialogs-settings.test.ts index b11cb051ea2..8006e3f89b4 100644 --- a/packages/cli/src/ui/opentui/dialogs-settings.test.ts +++ b/packages/cli/src/ui/opentui/dialogs-settings.test.ts @@ -121,7 +121,7 @@ describe('buildSettingsListItems', () => { expect(items.length).toBeGreaterThan(0); const keys = items.map((item) => item.key); expect(keys).toContain('ui.theme'); - expect(keys.indexOf('tools.codeModeOnly')).toBe( + expect(keys.indexOf('tools.mode')).toBe( keys.indexOf('tools.approvalMode') + 1, ); expect(keys).not.toContain('tools.freeform'); diff --git a/packages/core/src/agents/agent-transcript.ts b/packages/core/src/agents/agent-transcript.ts index 000c9a9b1ab..641e169c552 100644 --- a/packages/core/src/agents/agent-transcript.ts +++ b/packages/core/src/agents/agent-transcript.ts @@ -148,6 +148,8 @@ export interface AgentMeta { * exclusion; an empty list means deny-all. */ executionAllowedTools?: string[]; + /** Exact names allowed only inside exec; never grants direct tool calls. */ + nestedExecutionAllowedTools?: string[]; /** * Launch-time per-agent tool blocklist of a fork, persisted beside * `executionAllowedTools`: the resume path rebuilds the fork's toolConfig diff --git a/packages/core/src/agents/background-agent-resume.test.ts b/packages/core/src/agents/background-agent-resume.test.ts index 16e77d7f748..6fcdd189dca 100644 --- a/packages/core/src/agents/background-agent-resume.test.ts +++ b/packages/core/src/agents/background-agent-resume.test.ts @@ -188,6 +188,7 @@ describe('BackgroundAgentResumeService', () => { * negative rows pass vacuously. */ registeredToolNames?: string[]; + skillVisibility?: 'hidden' | 'visible'; toolMode?: ToolMode; hookSystem?: | { @@ -254,6 +255,8 @@ describe('BackgroundAgentResumeService', () => { .fn() .mockReturnValue(options.deferredToolSummary ?? []), isDeferredToolRevealed: vi.fn().mockReturnValue(false), + isPermissionDeferred: (name: string) => + name === ToolNames.SKILL && options.skillVisibility !== undefined, getMcpServerInstructions: vi .fn() .mockReturnValue(options.mcpServerInstructions ?? new Map()), @@ -294,6 +297,8 @@ describe('BackgroundAgentResumeService', () => { getStopHookBlockingCap: () => options.stopHookBlockingCap ?? 8, getApprovalMode: () => 'default', getToolMode: () => options.toolMode, + getVisibleTools: () => + new Set(options.skillVisibility === 'visible' ? [ToolNames.SKILL] : []), getModel: () => 'parent-model', getBareMode: () => false, getSandbox: () => undefined, @@ -1185,9 +1190,107 @@ describe('BackgroundAgentResumeService', () => { }, boolean, ToolMode?, + boolean?, + ('hidden' | 'visible')?, ] >([ ['inherits every tool', {}, true], + [ + 'CodeMode * with hidden eager Skill', + { tools: ['*'] }, + false, + ToolMode.CodeMode, + true, + 'hidden', + ], + [ + 'CodeMode * with visible eager Skill', + { tools: ['*'] }, + true, + ToolMode.CodeMode, + true, + 'visible', + ], + [ + 'CodeMode skill with hidden eager Skill', + { tools: ['skill'] }, + false, + ToolMode.CodeMode, + true, + 'hidden', + ], + [ + 'CodeMode skill with visible eager Skill', + { tools: ['skill'] }, + true, + ToolMode.CodeMode, + true, + 'visible', + ], + [ + 'CodeMode exec with hidden eager Skill', + { tools: ['exec'] }, + false, + ToolMode.CodeMode, + true, + 'hidden', + ], + [ + 'CodeMode exec with visible eager Skill', + { tools: ['exec'] }, + true, + ToolMode.CodeMode, + true, + 'visible', + ], + [ + 'CodeModeOnly * with hidden eager Skill', + { tools: ['*'] }, + true, + ToolMode.CodeModeOnly, + true, + 'hidden', + ], + [ + 'CodeModeOnly * with visible eager Skill', + { tools: ['*'] }, + true, + ToolMode.CodeModeOnly, + true, + 'visible', + ], + [ + 'CodeModeOnly skill with hidden eager Skill', + { tools: ['skill'] }, + true, + ToolMode.CodeModeOnly, + true, + 'hidden', + ], + [ + 'CodeModeOnly skill with visible eager Skill', + { tools: ['skill'] }, + true, + ToolMode.CodeModeOnly, + true, + 'visible', + ], + [ + 'CodeModeOnly exec with hidden eager Skill', + { tools: ['exec'] }, + true, + ToolMode.CodeModeOnly, + true, + 'hidden', + ], + [ + 'CodeModeOnly exec with visible eager Skill', + { tools: ['exec'] }, + true, + ToolMode.CodeModeOnly, + true, + 'visible', + ], [ 'disallows the Skill tool', { tools: ['*'], disallowedTools: [ToolNames.SKILL] }, @@ -1231,6 +1334,20 @@ describe('BackgroundAgentResumeService', () => { true, ToolMode.CodeModeOnly, ], + [ + 'names exec without skill under Hybrid with registered exec', + { tools: [ToolNames.EXEC] }, + true, + ToolMode.CodeMode, + true, + ], + [ + 'names exec without skill under Hybrid without exec', + { tools: [ToolNames.EXEC] }, + false, + ToolMode.CodeMode, + false, + ], // Same definition, Direct mode: no gateway, so no listing. Pins that the // row above is the tool mode and not the `exec` name doing the work. [ @@ -1240,7 +1357,14 @@ describe('BackgroundAgentResumeService', () => { ], ])( 'matches the launch-time skill listing when the definition %s', - async (_label, toolFields, expectListing, toolMode) => { + async ( + _label, + toolFields, + expectListing, + toolMode, + hasExec = false, + skillVisibility, + ) => { const sessionId = 'session-skill-listing'; const agentId = 'agent-skill-listing'; const metaPath = getAgentMetaPath(tempDir, sessionId, agentId); @@ -1296,11 +1420,15 @@ describe('BackgroundAgentResumeService', () => { }; const { service, subagentManager } = createService({ toolMode, + skillVisibility, // The session this resume runs in does have the Skill tool; the rows // below are about `subagentWillHaveSkillTool`, not about #12838's // registry gate. Omitting this made every row answer "no listing" for // the registry's reason and the two negative rows pass vacuously. - registeredToolNames: [ToolNames.SKILL], + registeredToolNames: [ + ToolNames.SKILL, + ...(hasExec ? [ToolNames.EXEC] : []), + ], skillManager: { listSkills: vi.fn().mockResolvedValue([ { @@ -1784,6 +1912,7 @@ describe('BackgroundAgentResumeService', () => { format: 'persisted deny-all execution policy', legacyCapabilities: {}, executionAllowedTools: [] as string[] | undefined, + nestedExecutionAllowedTools: ['read_file'], includeDisplayImage: false, deniedTool: 'Read', expectedExecutionAllowedTools: [], @@ -1841,6 +1970,7 @@ describe('BackgroundAgentResumeService', () => { async ({ legacyCapabilities, executionAllowedTools, + nestedExecutionAllowedTools, includeDisplayImage, deniedTool, expectedExecutionAllowedTools, @@ -1854,7 +1984,7 @@ describe('BackgroundAgentResumeService', () => { launchPrompt, [userText('bootstrap env'), modelText('bootstrap ack')], { - meta: { executionAllowedTools }, + meta: { executionAllowedTools, nestedExecutionAllowedTools }, payload: legacyCapabilities, reply: 'Working silently', }, @@ -1951,6 +2081,9 @@ describe('BackgroundAgentResumeService', () => { ToolNames.ASK_USER_QUESTION, ], executionAllowedTools: expectedExecutionAllowedTools, + ...(nestedExecutionAllowedTools !== undefined + ? { nestedExecutionAllowedTools } + : {}), }); expect(createArgs?.[9]).toBe(launchPrompt); expect(createArgs?.[10]).toBe(agentId); diff --git a/packages/core/src/agents/background-agent-resume.ts b/packages/core/src/agents/background-agent-resume.ts index b5c943796ae..8edee057866 100644 --- a/packages/core/src/agents/background-agent-resume.ts +++ b/packages/core/src/agents/background-agent-resume.ts @@ -75,8 +75,11 @@ import { buildInheritedForkExecutionToolNames, extractParentToolNames, } from './runtime/agent-core.js'; -import { toolConfigAllowsSkill } from './runtime/subagent-plan-tool-policy.js'; -import { ToolMode } from '../tools/code-mode.js'; +import { + hasAgentSkillExecBinding, + isAgentSkillEagerHidden, + toolConfigAllowsSkill, +} from './runtime/subagent-plan-tool-policy.js'; import { toolSearchBridgeSentence } from '../skills/bundled-reference.js'; import { ToolNames } from '../tools/tool-names.js'; import type { @@ -138,7 +141,8 @@ const CONTAINER_EXECUTION_BLOCKED_REASON = */ function subagentWillHaveSkillTool( subagentConfig: SubagentConfig | undefined, - codeModeOnly = false, + execBindingsAvailable = false, + skillEagerHidden = false, ): boolean { // Launch reads `config.tools?.length ? resolveToolNames(config.tools) : ['*']`, // and `resolveToolNames`' `for...of` walks a bare string per character, @@ -166,7 +170,8 @@ function subagentWillHaveSkillTool( ? disallowedTools : undefined, }, - codeModeOnly, + execBindingsAvailable, + skillEagerHidden, ); } @@ -1016,7 +1021,8 @@ export class BackgroundAgentResumeService { includeDeferredToolsReminder: false, includeAvailableSkillsReminder: subagentWillHaveSkillTool( target.subagentConfig, - activeAgentConfig.getToolMode?.() === ToolMode.CodeModeOnly, + hasAgentSkillExecBinding(activeAgentConfig), + isAgentSkillEagerHidden(activeAgentConfig), ), }) )[0], @@ -1075,6 +1081,7 @@ export class BackgroundAgentResumeService { meta.disallowedTools, meta.agentId, meta.description, + meta.nestedExecutionAllowedTools, ); } else { const resumeSubagentConfig = @@ -1868,6 +1875,7 @@ export class BackgroundAgentResumeService { disallowedTools?: string[], subagentId?: string, taskName?: string, + nestedExecutionAllowedTools?: string[], ): Promise { const promptConfig: PromptConfig = { renderedSystemPrompt: structuredClone(runtime.systemInstruction), @@ -1887,6 +1895,9 @@ export class BackgroundAgentResumeService { runtime.toolNames, ), ), + ...(nestedExecutionAllowedTools !== undefined + ? { nestedExecutionAllowedTools: [...nestedExecutionAllowedTools] } + : {}), // Restore the persisted blocklist beside the allowlist: the // invocation-level re-check is the only enforcement a wildcard // allowlist entry (e.g. mcp__*) cannot provide on its own. diff --git a/packages/core/src/agents/runtime/agent-core.fork-policy.test.ts b/packages/core/src/agents/runtime/agent-core.fork-policy.test.ts new file mode 100644 index 00000000000..e13b0870567 --- /dev/null +++ b/packages/core/src/agents/runtime/agent-core.fork-policy.test.ts @@ -0,0 +1,174 @@ +/** + * @license + * Copyright 2026 Qwen + * SPDX-License-Identifier: Apache-2.0 + */ +import { describe, expect, it, vi } from 'vitest'; +import { + AgentCore, + buildInheritedForkExecutionToolNames, +} from './agent-core.js'; +import { getCurrentAgentConfiguredToolAllowlist } from './agent-context.js'; +import type { ToolConfig } from './agent-types.js'; +import { makeFakeConfig } from '../../test-utils/config.js'; +import { MockTool } from '../../test-utils/mock-tool.js'; +import { ToolRegistry } from '../../tools/tool-registry.js'; +import { ExecTool } from '../../tools/exec.js'; +import { ToolMode } from '../../tools/code-mode.js'; + +describe('fork MCP policy inheritance', () => { + it.each([ToolMode.CodeMode, ToolMode.CodeModeOnly])( + 'does not turn the %s exec wrapper into an inherited grant', + async (mode) => { + const config = makeFakeConfig({ toolMode: mode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: 'read_file' })); + registry.registerTool(new MockTool({ name: 'write_file' })); + for (const separate of [false, true]) { + let tools: ToolConfig = { + tools: ['read_file'], + ...(separate ? { executionAllowedTools: ['read_file'] } : {}), + }; + for (let generation = 0; generation < 3; generation++) { + const core = new AgentCore( + 'fork', + config, + { systemPrompt: '' }, + { model: 'test' }, + { max_turns: 1 }, + tools, + ); + const declarations = await core.prepareTools(); + const exec = declarations.find((d) => d.name === 'exec')!; + expect(exec.description).toContain('read_file'); + expect(exec.description).not.toContain('write_file'); + let frame: readonly string[] | undefined; + await core.runInAgentFrames(async () => { + frame = getCurrentAgentConfiguredToolAllowlist(); + }); + const inherited = buildInheritedForkExecutionToolNames( + declarations.map((d) => d.name!), + registry.getAllToolNames(), + frame, + ); + expect(inherited).toEqual(['read_file']); + tools = { + tools: declarations.map((d) => d.name!), + executionAllowedTools: inherited, + }; + } + } + }, + ); + + it('keeps nested hybrid bindings separate from the inherited direct grant', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: 'read_file' })); + registry.registerTool(new MockTool({ name: 'write_file' })); + let tools: ToolConfig = { tools: ['read_file', 'exec'] }; + for (let generation = 0; generation < 3; generation++) { + const core = new AgentCore( + 'fork', + config, + { systemPrompt: '' }, + { model: 'test' }, + { max_turns: 1 }, + tools, + ); + const declarations = await core.prepareTools(); + expect( + declarations.find((d) => d.name === 'exec')!.description, + ).toContain('write_file'); + let frame: readonly string[] | undefined; + await core.runInAgentFrames(async () => { + frame = getCurrentAgentConfiguredToolAllowlist(); + }); + const inherited = buildInheritedForkExecutionToolNames( + declarations.map((d) => d.name!), + registry.getAllToolNames(), + frame, + ); + expect(inherited).toEqual(['read_file']); + tools = { + tools: declarations.map((d) => d.name!), + executionAllowedTools: inherited, + nestedExecutionAllowedTools: ['read_file', 'write_file'], + }; + } + }); + + it.each([ + [ToolMode.CodeMode, false], + [ToolMode.CodeMode, true], + [ToolMode.CodeModeOnly, false], + [ToolMode.CodeModeOnly, true], + ] as const)( + 'keeps %s restrictions with execution list=%s across generations', + async (mode, separate) => { + const config = makeFakeConfig({ toolMode: mode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + for (const [serverName, serverToolName] of [ + ['github', 'read_file'], + ['github', 'create_issue'], + ['payments', 'charge'], + ]) { + registry.registerTool( + Object.assign( + new MockTool({ + name: `mcp__${serverName}__${serverToolName}`, + }), + { serverName, serverToolName }, + ), + ); + } + for (const pattern of ['mcp__github__read_*', 'mcp__offline__read_*']) { + let tools: ToolConfig = { + tools: ['exec', pattern], + ...(separate ? { executionAllowedTools: ['exec', pattern] } : {}), + }; + for (let generation = 0; generation < 3; generation++) { + const core = new AgentCore( + 'fork', + config, + { systemPrompt: '' }, + { model: 'test' }, + { max_turns: 1 }, + tools, + ); + const declarations = await core.prepareTools(); + const exec = declarations.find((d) => d.name === 'exec')!; + expect(exec.description).not.toContain('mcp__payments__charge'); + expect(exec.description).not.toContain('mcp__github__create_issue'); + if (pattern.includes('github')) { + expect(exec.description).toContain('mcp__github__read_file'); + } + let frame: readonly string[] | undefined; + await core.runInAgentFrames(async () => { + frame = getCurrentAgentConfiguredToolAllowlist(); + }); + const inherited = buildInheritedForkExecutionToolNames( + declarations.map((d) => d.name!), + registry.getAllToolNames(), + frame, + ); + expect(inherited).not.toContain(pattern); + expect(inherited).not.toContain('exec'); + expect(inherited).not.toContain('mcp__payments__charge'); + if (pattern.includes('github')) + expect(inherited).toContain('mcp__github__read_file'); + tools = { + tools: declarations.map((d) => d.name!), + executionAllowedTools: inherited, + }; + } + } + }, + ); +}); diff --git a/packages/core/src/agents/runtime/agent-core.skill-eager-review.test.ts b/packages/core/src/agents/runtime/agent-core.skill-eager-review.test.ts new file mode 100644 index 00000000000..6434a63c90e --- /dev/null +++ b/packages/core/src/agents/runtime/agent-core.skill-eager-review.test.ts @@ -0,0 +1,90 @@ +/** + * @license + * Copyright 2026 Qwen + * SPDX-License-Identifier: Apache-2.0 + */ + +import { describe, expect, it, vi } from 'vitest'; +import { AgentCore } from './agent-core.js'; +import { makeFakeConfig } from '../../test-utils/config.js'; +import { MockTool } from '../../test-utils/mock-tool.js'; +import { ToolRegistry } from '../../tools/tool-registry.js'; +import { ToolNames } from '../../tools/tool-names.js'; +import { ToolMode } from '../../tools/code-mode.js'; +import { ExecTool } from '../../tools/exec.js'; + +describe('Skill announcement for eager-hidden subagent tools', () => { + it.each( + [ToolMode.CodeMode, ToolMode.CodeModeOnly].flatMap((mode) => + [false, true].flatMap((warm) => + ['hidden', 'visible', 'revealed'].flatMap((visibility) => + [false, true].flatMap((copyRegistry) => + [['*'], [ToolNames.SKILL], [ToolNames.EXEC]].map((tools) => ({ + mode, + warm, + visibility, + copyRegistry, + tools, + })), + ), + ), + ), + ), + )( + 'matches invocation for $mode tools=$tools warm=$warm visibility=$visibility copied=$copyRegistry', + async ({ mode, warm, visibility, copyRegistry, tools }) => { + const config = makeFakeConfig({ toolMode: mode }); + const makeRegistry = () => { + const registry = new ToolRegistry(config); + registry.registerTool(new ExecTool(config)); + registry.registerPermissionDeferredFactory( + ToolNames.SKILL, + async () => new MockTool({ name: ToolNames.SKILL }), + ); + return registry; + }; + let registry = makeRegistry(); + vi.spyOn(config, 'getToolRegistry').mockImplementation(() => registry); + if (warm || visibility === 'revealed') { + await registry.ensureTool(ToolNames.SKILL); + } + if (visibility === 'visible') { + vi.spyOn(config, 'getVisibleTools').mockReturnValue( + new Set([ToolNames.SKILL]), + ); + } else if (visibility === 'revealed') { + registry.revealDeferredTool(ToolNames.SKILL); + } + if (copyRegistry) { + const childRegistry = makeRegistry(); + childRegistry.copyDiscoveredToolsFrom(registry); + registry = childRegistry; + } + const core = new AgentCore( + 'skill-eager-review', + config, + { systemPrompt: '' }, + { model: 'test-model' }, + { max_turns: 1 }, + { tools }, + ); + const observed = core as unknown as { + willHaveSkillTool(): boolean; + canInvokeSkill(names: ReadonlySet): boolean; + }; + const announced = observed.willHaveSkillTool(); + const declared = new Set( + (await core.prepareTools()).map((declaration) => declaration.name), + ); + const invokable = observed.canInvokeSkill(declared); + const expected = + mode === ToolMode.CodeModeOnly || + visibility === 'visible' || + (visibility === 'revealed' && !copyRegistry); + + expect.soft(invokable).toBe(expected); + expect.soft(announced).toBe(expected); + expect(announced).toBe(invokable); + }, + ); +}); diff --git a/packages/core/src/agents/runtime/agent-core.skill-gate.test.ts b/packages/core/src/agents/runtime/agent-core.skill-gate.test.ts index 73dfd699991..ac0272f7d83 100644 --- a/packages/core/src/agents/runtime/agent-core.skill-gate.test.ts +++ b/packages/core/src/agents/runtime/agent-core.skill-gate.test.ts @@ -12,7 +12,10 @@ import { makeFakeConfig } from '../../test-utils/config.js'; import { ToolRegistry } from '../../tools/tool-registry.js'; import { ExecTool } from '../../tools/exec.js'; import { MockTool } from '../../test-utils/mock-tool.js'; +import { ToolMode } from '../../tools/code-mode.js'; +import { CoreToolScheduler } from '../../core/coreToolScheduler.js'; import { ToolSearchTool } from '../../tools/tool-search.js'; +import { hasAgentSkillExecBinding } from './subagent-plan-tool-policy.js'; // The skill-announcement gate asks whether the model can INVOKE a skill, and // that is two conditions, not one. @@ -93,7 +96,7 @@ describe('AgentCore skill-gate inputs', () => { toolConfig: ConstructorParameters[5], toolNames: string[], ) { - const config = makeFakeConfig({ codeModeOnly: true }); + const config = makeFakeConfig({ toolMode: 'code_mode_only' }); const registry = new ToolRegistry(config); vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); registry.registerTool(new ExecTool(config)); @@ -119,6 +122,35 @@ describe('AgentCore skill-gate inputs', () => { .codeModeAllowedToolNames; describe('declared', () => { + it.each([ToolMode.Direct, ToolMode.CodeMode, ToolMode.CodeModeOnly])( + 'keeps an explicit empty runtime tool list empty in %s', + async (toolMode) => { + const config = makeFakeConfig({ toolMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: ToolNames.READ_FILE })); + const core = new AgentCore( + 'no-tools', + config, + { systemPrompt: '' }, + { model: 'test-model' }, + { max_turns: 1 }, + { tools: [] }, + ); + + expect(await core.prepareTools()).toEqual([]); + expect(executable(core, ToolNames.EXEC)).toBe(false); + expect(executable(core, ToolNames.READ_FILE)).toBe(false); + const inherited = ( + core as unknown as { + getInheritedToolExecutionAllowlist: () => readonly string[]; + } + ).getInheritedToolExecutionAllowlist(); + expect(inherited).toEqual([]); + }, + ); + it('excludes a tool the disallowedTools blocklist removed', async () => { // `tools: ['*']` says "everything", so a `toolConfig`-based read reports // SKILL as available; the blocklist is applied after, to the list. @@ -266,6 +298,39 @@ describe('AgentCore skill-gate inputs', () => { ).toBe(true); }); + it.each([true, false])( + 'lists Hybrid exec-only skills when lazy exec is registered: %s', + (registered) => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + if (registered) { + registry.registerFactory( + ToolNames.EXEC, + async () => new ExecTool(config), + ); + } + expect(registry.getTool(ToolNames.EXEC)).toBeUndefined(); + registry.registerFactory( + ToolNames.SKILL, + async () => new MockTool({ name: ToolNames.SKILL }), + ); + const core = new AgentCore( + 'lazy-exec-skill-listing', + config, + { systemPrompt: '' }, + { model: 'test-model' }, + { max_turns: 1 }, + { tools: [ToolNames.EXEC] }, + ); + expect( + ( + core as unknown as { willHaveSkillTool: () => boolean } + ).willHaveSkillTool(), + ).toBe(registered); + }, + ); + it.each([ { tools: ['*'] }, { tools: [ToolNames.SKILL] }, @@ -277,6 +342,98 @@ describe('AgentCore skill-gate inputs', () => { expect(gate(core, declared)).toBe(true); }); + it.each( + [ToolMode.CodeMode, ToolMode.CodeModeOnly].flatMap((mode) => + [false, true].flatMap((warm) => + ['hidden', 'visible', 'revealed'].map((visibility) => ({ + mode, + warm, + visibility, + })), + ), + ), + )( + 'matches the Skill route in $mode with warm=$warm and $visibility', + async ({ mode, warm, visibility }) => { + const config = makeFakeConfig({ toolMode: mode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerPermissionDeferredFactory( + ToolNames.SKILL, + async () => new MockTool({ name: ToolNames.SKILL }), + ); + if (warm || visibility === 'revealed') { + await registry.ensureTool(ToolNames.SKILL); + } + if (visibility === 'visible') { + vi.spyOn(config, 'getVisibleTools').mockReturnValue( + new Set([ToolNames.SKILL]), + ); + } else if (visibility === 'revealed') { + registry.revealDeferredTool(ToolNames.SKILL); + } + expect(hasAgentSkillExecBinding(config)).toBe( + mode === ToolMode.CodeModeOnly || visibility !== 'hidden', + ); + const childRegistry = new ToolRegistry(config); + childRegistry.registerTool(new ExecTool(config)); + childRegistry.registerPermissionDeferredFactory( + ToolNames.SKILL, + async () => new MockTool({ name: ToolNames.SKILL }), + ); + childRegistry.copyDiscoveredToolsFrom(registry); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(childRegistry); + expect(childRegistry.isDeferredToolRevealed(ToolNames.SKILL)).toBe( + false, + ); + expect(hasAgentSkillExecBinding(config)).toBe( + mode === ToolMode.CodeModeOnly || visibility === 'visible', + ); + const core = new AgentCore( + 'skill-route', + config, + { systemPrompt: '' }, + { model: 'test-model' }, + { max_turns: 1 }, + { tools: [ToolNames.EXEC] }, + ); + const willHaveSkill = ( + core as unknown as { + willHaveSkillTool: () => boolean; + } + ).willHaveSkillTool(); + expect(willHaveSkill).toBe( + mode === ToolMode.CodeModeOnly || visibility === 'visible', + ); + const declared = await declaredNames(core); + expect(gate(core, declared)).toBe(willHaveSkill); + }, + ); + + it('opens for a hybrid nested-only skill', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: ToolNames.SKILL })); + const core = new AgentCore( + 'hybrid-nested-skill', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: [ToolNames.EXEC], + executionAllowedTools: [ToolNames.EXEC], + }, + ); + + const declared = await declaredNames(core); + expect(declared).toEqual(new Set([ToolNames.EXEC])); + expect(gate(core, declared)).toBe(true); + }); + it.each([ { tools: ['*'], disallowedTools: [ToolNames.SKILL] }, { tools: ['*'], disallowedTools: [ToolNames.EXEC] }, @@ -308,6 +465,46 @@ describe('AgentCore skill-gate inputs', () => { expect(gate(core, new Set([ToolNames.EXEC]))).toBe(false); }); + it('carries the inherited hybrid binding allowlist into exec requests', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: ToolNames.READ_FILE })); + registry.registerTool(new MockTool({ name: ToolNames.WRITE_FILE })); + const core = new AgentCore( + 'request-wiring-hybrid-code-mode', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: [ToolNames.EXEC], + executionAllowedTools: [ToolNames.READ_FILE], + }, + ); + const abortController = new AbortController(); + const schedule = vi + .spyOn(CoreToolScheduler.prototype, 'schedule') + .mockImplementation(async () => abortController.abort()); + + await core.processFunctionCalls( + [{ id: 'exec-call', name: ToolNames.EXEC, args: { source: '' } }], + abortController, + 'fork-prompt', + 1, + [{ name: ToolNames.EXEC }], + ); + + const scheduled = schedule.mock.calls[0]?.[0]; + expect(Array.isArray(scheduled) ? scheduled[0] : scheduled).toEqual( + expect.objectContaining({ + name: ToolNames.EXEC, + codeModeAllowedToolNames: [ToolNames.READ_FILE], + }), + ); + }); + it('closes for an unregistered skill', async () => { const core = makeCodeModeCore({ tools: ['*'] }, false); expect(gate(core, await declaredNames(core))).toBe(false); @@ -330,7 +527,7 @@ describe('AgentCore skill-gate inputs', () => { ])( 'offers scoped discovery for deferred Code Mode tools: %j', async (toolConfig) => { - const config = makeFakeConfig({ codeModeOnly: true }); + const config = makeFakeConfig({ toolMode: 'code_mode_only' }); const registry = new ToolRegistry(config); vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); registry.registerTool(new ExecTool(config)); @@ -367,7 +564,7 @@ describe('AgentCore skill-gate inputs', () => { ); it('falls back to scoped signatures when the agent disallows search', async () => { - const config = makeFakeConfig({ codeModeOnly: true }); + const config = makeFakeConfig({ toolMode: 'code_mode_only' }); const registry = new ToolRegistry(config); vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); registry.registerTool(new ExecTool(config)); @@ -430,6 +627,418 @@ describe('AgentCore skill-gate inputs', () => { expect(codeModeAllowed(core)).toEqual([ToolNames.READ_FILE]); }); + it('keeps direct and exec surfaces aligned for a narrowed CodeMode agent', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: ToolNames.READ_FILE })); + registry.registerTool(new MockTool({ name: ToolNames.WRITE_FILE })); + const core = new AgentCore( + 'restricted-hybrid-code-mode', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: [ToolNames.READ_FILE], + executionAllowedTools: [ToolNames.READ_FILE], + }, + ); + + const declarations = await core.prepareTools(); + const exec = declarations.find( + (declaration) => declaration.name === ToolNames.EXEC, + ); + const readFile = declarations.find( + (declaration) => declaration.name === ToolNames.READ_FILE, + ); + + expect(declarations.map((declaration) => declaration.name)).toEqual([ + ToolNames.EXEC, + ToolNames.READ_FILE, + ]); + expect(executable(core, ToolNames.EXEC)).toBe(true); + expect(exec?.description).toContain('"name":"read_file"'); + expect(exec?.description).not.toContain('"name":"write_file"'); + expect(readFile?.description).toContain( + 'declare const tools: { read_file(args:', + ); + expect( + ( + core as unknown as { + codeModeAllowedToolNames?: readonly string[]; + } + ).codeModeAllowedToolNames, + ).toEqual([ToolNames.READ_FILE]); + }); + + it('keeps inherited exec bindings off the direct CodeMode surface', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: ToolNames.READ_FILE })); + registry.registerTool(new MockTool({ name: ToolNames.WRITE_FILE })); + const core = new AgentCore( + 'exec-only-hybrid-code-mode', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: [ToolNames.EXEC], + executionAllowedTools: [ToolNames.EXEC], + }, + ); + + const declarations = await core.prepareTools(); + + expect(declarations.map((declaration) => declaration.name)).toEqual([ + ToolNames.EXEC, + ]); + expect(declarations[0]?.description).toContain('"name":"read_file"'); + expect(declarations[0]?.description).toContain('"name":"write_file"'); + expect(declarations[0]?.description).toContain('tools.read_file(args:'); + expect(declarations[0]?.description).toContain('tools.write_file(args:'); + expect( + ( + core as unknown as { + codeModeAllowedToolNames?: readonly string[]; + } + ).codeModeAllowedToolNames, + ).toEqual([ToolNames.READ_FILE, ToolNames.WRITE_FILE]); + }); + + it('keeps bounded fork bindings nested-only with an immutable policy', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + for (const name of [ + ToolNames.READ_FILE, + ToolNames.WRITE_FILE, + 'mcp__payments__charge', + ]) { + registry.registerTool(new MockTool({ name })); + } + const nested: string[] = [ToolNames.READ_FILE]; + const core = new AgentCore( + 'bounded-exec-fork', + config, + { systemPrompt: '' }, + { model: 'test-model' }, + { max_turns: 1 }, + { + tools: [ToolNames.EXEC, ToolNames.READ_FILE, ToolNames.WRITE_FILE], + executionAllowedTools: [], + nestedExecutionAllowedTools: nested, + }, + ); + nested.push(ToolNames.WRITE_FILE); + const declarations = await core.prepareTools(); + expect(declarations.map((tool) => tool.name)).toEqual([ToolNames.EXEC]); + expect(executable(core, ToolNames.READ_FILE)).toBe(false); + expect(declarations[0]?.description).toContain('tools.read_file(args:'); + expect(declarations[0]?.description).not.toContain( + 'tools.write_file(args:', + ); + expect(declarations[0]?.description).not.toContain( + 'mcp__payments__charge', + ); + }); + + it('honors MCP execution patterns in the hybrid nested binding set', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + const readTool = new MockTool({ name: 'mcp__github__read_file' }); + Object.assign(readTool, { + serverName: 'github', + serverToolName: 'read_file', + }); + const createTool = new MockTool({ name: 'mcp__github__create_issue' }); + Object.assign(createTool, { + serverName: 'github', + serverToolName: 'create_issue', + }); + registry.registerTool(readTool); + registry.registerTool(createTool); + const core = new AgentCore( + 'mcp-pattern-hybrid-code-mode', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: [ToolNames.EXEC, 'mcp__github__read_*'], + executionAllowedTools: [ToolNames.EXEC, 'mcp__github__read_*'], + }, + ); + + await core.prepareTools(); + + expect( + ( + core as unknown as { + codeModeAllowedToolNames?: readonly string[]; + } + ).codeModeAllowedToolNames, + ).toEqual(['mcp__github__read_file']); + }); + + it('honors MCP patterns from the tools list alone in the hybrid nested binding set', async () => { + // No `executionAllowedTools`: the agent-definition surface + // (`SubagentConfig.tools`) is the only allowlist an SDK integrator can + // set, so the configured-list branch must apply the same MCP narrowing + // as the execution-allowlist branch. + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + const readTool = new MockTool({ name: 'mcp__github__read_file' }); + Object.assign(readTool, { + serverName: 'github', + serverToolName: 'read_file', + }); + const createTool = new MockTool({ name: 'mcp__github__create_issue' }); + Object.assign(createTool, { + serverName: 'github', + serverToolName: 'create_issue', + }); + const chargeTool = new MockTool({ name: 'mcp__payments__charge' }); + Object.assign(chargeTool, { + serverName: 'payments', + serverToolName: 'charge', + }); + registry.registerTool(readTool); + registry.registerTool(createTool); + registry.registerTool(chargeTool); + const core = new AgentCore( + 'mcp-pattern-tools-only-hybrid-code-mode', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: [ToolNames.EXEC, 'mcp__github__read_*'], + }, + ); + + await core.prepareTools(); + + expect( + ( + core as unknown as { + codeModeAllowedToolNames?: readonly string[]; + } + ).codeModeAllowedToolNames, + ).toEqual(['mcp__github__read_file']); + }); + + it('honors MCP patterns from the tools list alone in the CodeModeOnly nested binding set', async () => { + // CodeModeOnly widens a configured exec grant into every + // code-mode-callable registry name; that expansion must not swallow + // the MCP names the configured list narrows by pattern. + const config = makeFakeConfig({ toolMode: ToolMode.CodeModeOnly }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + const readTool = new MockTool({ name: 'mcp__github__read_file' }); + Object.assign(readTool, { + serverName: 'github', + serverToolName: 'read_file', + }); + const createTool = new MockTool({ name: 'mcp__github__create_issue' }); + Object.assign(createTool, { + serverName: 'github', + serverToolName: 'create_issue', + }); + const chargeTool = new MockTool({ name: 'mcp__payments__charge' }); + Object.assign(chargeTool, { + serverName: 'payments', + serverToolName: 'charge', + }); + registry.registerTool(readTool); + registry.registerTool(createTool); + registry.registerTool(chargeTool); + const core = new AgentCore( + 'mcp-pattern-tools-only-code-mode-only', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: [ToolNames.EXEC, 'mcp__github__read_*'], + }, + ); + + await core.prepareTools(); + + expect( + ( + core as unknown as { + codeModeAllowedToolNames?: readonly string[]; + } + ).codeModeAllowedToolNames, + ).toEqual(['mcp__github__read_file']); + }); + + it('honors an exact MCP tool name in the hybrid nested binding set', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + const readTool = new MockTool({ name: 'mcp__github__read_file' }); + Object.assign(readTool, { + serverName: 'github', + serverToolName: 'read_file', + }); + const deleteTool = new MockTool({ name: 'mcp__github__delete_repo' }); + Object.assign(deleteTool, { + serverName: 'github', + serverToolName: 'delete_repo', + }); + registry.registerTool(readTool); + registry.registerTool(deleteTool); + const core = new AgentCore( + 'mcp-exact-hybrid-code-mode', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: [ToolNames.EXEC, 'mcp__github__read_file'], + executionAllowedTools: [ToolNames.EXEC, 'mcp__github__read_file'], + }, + ); + + await core.prepareTools(); + + expect( + ( + core as unknown as { + codeModeAllowedToolNames?: readonly string[]; + } + ).codeModeAllowedToolNames, + ).toEqual(['mcp__github__read_file']); + }); + + it('honors a server-level MCP name in the hybrid nested binding set', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + const readTool = new MockTool({ name: 'mcp__github__read_file' }); + Object.assign(readTool, { + serverName: 'github', + serverToolName: 'read_file', + }); + const createTool = new MockTool({ name: 'mcp__github__create_issue' }); + Object.assign(createTool, { + serverName: 'github', + serverToolName: 'create_issue', + }); + const chargeTool = new MockTool({ name: 'mcp__payments__charge' }); + Object.assign(chargeTool, { + serverName: 'payments', + serverToolName: 'charge', + }); + registry.registerTool(readTool); + registry.registerTool(createTool); + registry.registerTool(chargeTool); + const core = new AgentCore( + 'mcp-server-hybrid-code-mode', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: [ToolNames.EXEC, 'mcp__github'], + executionAllowedTools: [ToolNames.EXEC, 'mcp__github'], + }, + ); + + await core.prepareTools(); + + expect( + ( + core as unknown as { + codeModeAllowedToolNames?: readonly string[]; + } + ).codeModeAllowedToolNames, + ).toEqual(['mcp__github__read_file', 'mcp__github__create_issue']); + }); + + it('keeps inherited CodeMode declarations executable only when directly allowed', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: ToolNames.READ_FILE })); + registry.registerTool(new MockTool({ name: ToolNames.WRITE_FILE })); + const core = new AgentCore( + 'allowlisted-hybrid-code-mode', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { + tools: ['*'], + executionAllowedTools: [ToolNames.EXEC, ToolNames.READ_FILE], + }, + ); + + const declarations = await core.prepareTools(); + + expect(declarations.map((declaration) => declaration.name)).toEqual([ + ToolNames.EXEC, + ToolNames.READ_FILE, + ]); + expect(executable(core, ToolNames.READ_FILE)).toBe(true); + expect(executable(core, ToolNames.WRITE_FILE)).toBe(false); + expect( + ( + core as unknown as { + codeModeAllowedToolNames?: readonly string[]; + } + ).codeModeAllowedToolNames, + ).toEqual([ToolNames.READ_FILE, ToolNames.WRITE_FILE]); + }); + + it('inherits the complete registry in CodeMode', async () => { + const config = makeFakeConfig({ toolMode: ToolMode.CodeMode }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: ToolNames.READ_FILE })); + registry.registerTool(new MockTool({ name: ToolNames.WRITE_FILE })); + const core = new AgentCore( + 'inherited-hybrid-code-mode', + config, + { systemPrompt: '' } as never, + { model: 'test-model' } as never, + { max_turns: 1 } as never, + { tools: ['*'] }, + ); + + const declarations = await core.prepareTools(); + + expect(declarations.map((declaration) => declaration.name)).toEqual([ + ToolNames.EXEC, + ToolNames.READ_FILE, + ToolNames.WRITE_FILE, + ]); + expect( + ( + core as unknown as { + codeModeAllowedToolNames?: readonly string[]; + } + ).codeModeAllowedToolNames, + ).toEqual([ToolNames.READ_FILE, ToolNames.WRITE_FILE]); + }); + it('narrows an inherited fork exec surface to its execution allowlist', async () => { const core = codeModeCore( 'restricted-fork', diff --git a/packages/core/src/agents/runtime/agent-core.test.ts b/packages/core/src/agents/runtime/agent-core.test.ts index d046875d7e6..226f5215392 100644 --- a/packages/core/src/agents/runtime/agent-core.test.ts +++ b/packages/core/src/agents/runtime/agent-core.test.ts @@ -192,7 +192,9 @@ describe('AgentCore.runInAgentFrames', () => { const runConfig: RunConfig = { max_turns: 1 }; return new AgentCore( name, - {} as unknown as Config, + { + getToolRegistry: () => ({ getAllToolNames: () => [] }), + } as unknown as Config, promptConfig, modelConfig, runConfig, diff --git a/packages/core/src/agents/runtime/agent-core.ts b/packages/core/src/agents/runtime/agent-core.ts index 79ccead2ebc..428ba0516e7 100644 --- a/packages/core/src/agents/runtime/agent-core.ts +++ b/packages/core/src/agents/runtime/agent-core.ts @@ -120,7 +120,11 @@ import { AgentEventEmitter, AgentEventType } from './agent-events.js'; import { AgentStatistics, type AgentStatsSummary } from './agent-statistics.js'; import { canonicalToolName, ToolNames } from '../../tools/tool-names.js'; import type { ToolRegistry } from '../../tools/tool-registry.js'; -import { getToolExposure, ToolMode } from '../../tools/code-mode.js'; +import { + getToolExposure, + isCodeModeEnabled, + ToolMode, +} from '../../tools/code-mode.js'; import { DEFAULT_QWEN_MODEL } from '../../config/models.js'; import { type ContextState, templateString } from './agent-headless.js'; import { getResponseText } from '../../utils/partUtils.js'; @@ -144,6 +148,8 @@ import { isLeaderOnlyToolUnavailableInSubagent, isPlanLifecycleToolUnavailableInSubagent, isToolExcludedForCurrentContext, + hasAgentSkillExecBinding, + isAgentSkillEagerHidden, matchesAgentToolBlocklist, toolConfigAllowsSkill, } from './subagent-plan-tool-policy.js'; @@ -239,7 +245,9 @@ export function extractParentToolNames( * Build the executable fork surface shared by launch and resume. Deferred * tools are absent from the parent's declarations but remain reachable * through tool_search/tool_call, so the live registry is part of this - * surface. A configured positive allowlist remains the outer bound. + * surface. A configured positive allowlist remains the outer bound: inherit + * concrete bindings without the broad exec grant. The exec wrapper remains + * callable in code modes even when absent from the execution allowlist. */ export function buildInheritedForkExecutionToolNames( advertisedToolNames: readonly string[], @@ -252,7 +260,8 @@ export function buildInheritedForkExecutionToolNames( (toolName) => !EXCLUDED_TOOLS_FOR_SUBAGENTS.has(toolName) && (configuredToolAllowlist === undefined || - configuredToolAllowlist.includes(toolName)), + (toolName !== ToolNames.EXEC && + configuredToolAllowlist.includes(toolName))), ); } @@ -423,6 +432,7 @@ export class AgentCore { readonly runConfig: RunConfig; readonly toolConfig?: ToolConfig; private readonly executionAllowedTools?: readonly string[]; + private readonly nestedExecutionAllowedTools?: ReadonlySet; private readonly executionAllowedExactTools?: ReadonlySet; private readonly executionAllowedMcpPatterns?: readonly string[]; private readonly executionAllowlistErrorSummary?: string; @@ -514,6 +524,11 @@ export class AgentCore { this.modelConfig = modelConfig; this.runConfig = runConfig; this.toolConfig = toolConfig; + if (toolConfig?.nestedExecutionAllowedTools !== undefined) { + this.nestedExecutionAllowedTools = new Set( + toolConfig.nestedExecutionAllowedTools, + ); + } if (toolConfig?.executionAllowedTools !== undefined) { this.executionAllowedTools = Object.freeze([ ...toolConfig.executionAllowedTools, @@ -644,9 +659,19 @@ export class AgentCore { * disallowed `skill` was still shown every skill it could not load. */ private willHaveSkillTool(): boolean { + if ( + this.runtimeContext.getToolMode?.() === ToolMode.CodeMode && + !this.runtimeContext + .getToolRegistry() + .getAllToolNames() + .includes(ToolNames.SKILL) + ) { + return false; + } return toolConfigAllowsSkill( this.toolConfig, - this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly, + hasAgentSkillExecBinding(this.runtimeContext), + isAgentSkillEagerHidden(this.runtimeContext), ); } @@ -658,6 +683,10 @@ export class AgentCore { * array denies all tools. */ async prepareTools(): Promise { + if (this.toolConfig?.tools.length === 0) { + this.codeModeAllowedToolNames = Object.freeze([]); + return []; + } const toolRegistry = this.runtimeContext.getToolRegistry(); await toolRegistry.warmAll(); const toolsList: FunctionDeclaration[] = []; @@ -693,7 +722,7 @@ export class AgentCore { toolRegistry.isPermissionDeferred?.(name) === true && toolRegistry.isDeferredAndHidden?.(name) === true; - if (this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly) { + if (isCodeModeEnabled(this.runtimeContext.getToolMode?.())) { const stringTools = this.toolConfig?.tools.filter( (tool): tool is string => typeof tool === 'string', @@ -717,13 +746,17 @@ export class AgentCore { (name) => (!configuredNames || configuredNames.has(name) || + name === ToolNames.EXEC || (inheritsCodeModeBindings && getToolExposure(name) === 'code-mode-callable')) && !isExcluded(name) && + (this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly || + !isHiddenByEagerAllowList(name)) && !this.isToolDisallowedByAgentConfig(name, toolRegistry) && - this.isToolExecutionAllowed(name), + this.isToolExecutionAllowed(name, true), ); if ( + this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly && allowedNames.some( (name) => getToolExposure(name) === 'code-mode-callable', ) && @@ -739,8 +772,20 @@ export class AgentCore { (name) => getToolExposure(name) === 'code-mode-callable', ), ); - const declarations = - toolRegistry.getFunctionDeclarationsFiltered(allowedNames); + const declarationNames = + this.runtimeContext.getToolMode?.() === ToolMode.CodeMode + ? allowedNames.filter( + (name) => + (!configuredNames || + name === ToolNames.EXEC || + configuredNames.has(name)) && + this.isToolExecutionAllowed(name), + ) + : allowedNames; + const declarations = toolRegistry.getFunctionDeclarationsFiltered( + declarationNames, + new Set(this.codeModeAllowedToolNames), + ); declarations.push( ...inlineTools.filter( (tool) => @@ -964,7 +1009,7 @@ export class AgentCore { // parent's policy frame. const runWithToolPolicy = () => runWithAgentConfiguredToolAllowlist( - this.getConfiguredToolExecutionAllowlist(), + this.getInheritedToolExecutionAllowlist(), () => runWithAgentDisallowedTools( this.toolConfig?.disallowedTools, @@ -1820,11 +1865,11 @@ export class AgentCore { ): boolean { return ( (declaredToolNames.has(ToolNames.SKILL) || - (this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly && + (isCodeModeEnabled(this.runtimeContext.getToolMode?.()) && declaredToolNames.has(ToolNames.EXEC) && !!this.runtimeContext.getToolRegistry().getTool(ToolNames.SKILL) && this.codeModeAllowedToolNames?.includes(ToolNames.SKILL) === true)) && - this.isToolExecutionAllowed(ToolNames.SKILL) + this.isToolExecutionAllowed(ToolNames.SKILL, true) ); } @@ -1856,8 +1901,8 @@ export class AgentCore { } /** - * The finite positive allowlist configured for this agent. Wildcard/empty - * configurations inherit the registry and therefore return `undefined`. + * The finite positive allowlist configured for this agent. Wildcard or + * absent configurations inherit the registry and return `undefined`. * A separate execution allowlist also returns `undefined`: fork agents use * `toolConfig.tools` as a declaration snapshot while deliberately allowing * additional bridged targets through `executionAllowedTools`. @@ -1874,10 +1919,7 @@ export class AgentCore { .filter((tool): tool is FunctionDeclaration => typeof tool !== 'string') .map((tool) => tool.name) .filter((name): name is string => typeof name === 'string'); - if ( - stringTools.includes('*') || - (stringTools.length === 0 && inlineToolNames.length === 0) - ) { + if (stringTools.includes('*')) { return undefined; } @@ -1886,10 +1928,17 @@ export class AgentCore { this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly && allowed.has(ToolNames.EXEC) ) { + // A configured list that mentions MCP narrows MCP names through the + // raw-identity match in isToolExecutionAllowed; expanding them in here + // would short-circuit that narrowing via `includes`. + const mentionsMcp = [...allowed].some((name) => name.startsWith('mcp__')); for (const toolName of this.runtimeContext .getToolRegistry() .getAllToolNames()) { - if (getToolExposure(toolName) === 'code-mode-callable') { + if ( + getToolExposure(toolName) === 'code-mode-callable' && + !(mentionsMcp && toolName.startsWith('mcp__')) + ) { allowed.add(toolName); } } @@ -1897,37 +1946,98 @@ export class AgentCore { return [...allowed]; } - private isToolExecutionAllowed(toolName: string): boolean { + private getInheritedToolExecutionAllowlist(): readonly string[] | undefined { + const configured = + this.executionAllowedTools ?? this.getConfiguredToolExecutionAllowlist(); + if (configured === undefined) return undefined; + return Array.from( + new Set([ + ...configured, + ...this.runtimeContext + .getToolRegistry() + .getAllToolNames() + .filter( + (name) => + getToolExposure(name) === 'code-mode-callable' && + this.isToolExecutionAllowed(name), + ), + ]), + ); + } + + private isToolExecutionAllowed( + toolName: string, + forNestedBinding = false, + ): boolean { if (this.isToolDisallowedByAgentConfig(toolName)) { return false; } + if ( + forNestedBinding && + isCodeModeEnabled(this.runtimeContext.getToolMode?.()) && + getToolExposure(toolName) === 'code-mode-callable' && + this.nestedExecutionAllowedTools !== undefined + ) { + return this.nestedExecutionAllowedTools.has(toolName); + } if (this.executionAllowedTools === undefined) { - // Code mode declares exec unconditionally (getCodeModeFunctionDeclarations - // keeps exposure 'exec' regardless of the allowed set), and prepareTools - // adds tool_search beside it, so a finite configured list that omits - // them must not refuse the tools the model was shown — the same - // carve-out the executionAllowedTools branch applies below. Both - // gateways apply the agent's scoped nested-tool allowlist themselves. + // Non-empty code-mode configurations declare exec even when the + // finite list omits it. An explicit empty list declares nothing. if ( - (toolName === ToolNames.EXEC || toolName === ToolNames.TOOL_SEARCH) && - this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly + this.toolConfig?.tools.length !== 0 && + ((toolName === ToolNames.EXEC && + isCodeModeEnabled(this.runtimeContext.getToolMode?.())) || + (toolName === ToolNames.TOOL_SEARCH && + this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly)) ) { return true; } const configuredAllowlist = this.getConfiguredToolExecutionAllowlist(); return ( configuredAllowlist === undefined || - configuredAllowlist.includes(toolName) + configuredAllowlist.includes(toolName) || + (toolName.startsWith('mcp__') && + this.matchesMcpAllowlist( + toolName, + new Set(configuredAllowlist.filter((name) => !name.includes('*'))), + configuredAllowlist.filter((name) => name.includes('*')), + )) || + (forNestedBinding && + isCodeModeEnabled(this.runtimeContext.getToolMode?.()) && + configuredAllowlist.includes(ToolNames.EXEC) && + getToolExposure(toolName) === 'code-mode-callable' && + // Once the configured list mentions MCP at all, an MCP name must + // pass the raw-identity match instead of the exec carve-out — + // the same rule the executionAllowedTools branch applies below. + (!toolName.startsWith('mcp__') || + !configuredAllowlist.some((name) => name.startsWith('mcp__')) || + this.matchesMcpAllowlist( + toolName, + new Set( + configuredAllowlist.filter((name) => !name.includes('*')), + ), + configuredAllowlist.filter((name) => name.includes('*')), + ))) ); } if ( - (toolName === ToolNames.EXEC || toolName === ToolNames.TOOL_SEARCH) && - this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly + (toolName === ToolNames.EXEC && + isCodeModeEnabled(this.runtimeContext.getToolMode?.())) || + (toolName === ToolNames.TOOL_SEARCH && + this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly) ) { return true; } if ( - this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly && + forNestedBinding && + isCodeModeEnabled(this.runtimeContext.getToolMode?.()) && + // An MCP tool name must fall through to the exact-name and pattern + // checks below once the allowlist mentions MCP at all: they are the + // only place a server-level or exact-tool narrowing can be honored. + (!toolName.startsWith('mcp__') || + !this.executionAllowedTools?.some((name) => + name.startsWith('mcp__'), + )) && this.executionAllowedExactTools?.has(ToolNames.EXEC) && getToolExposure(toolName) === 'code-mode-callable' ) { @@ -1940,9 +2050,24 @@ export class AgentCore { return false; } - // Match MCP patterns against the registry's raw server/tool identity. - // Comparing provider-sanitized prefixes can merge distinct server names - // such as "repo.bad" and "repo/bad", so it is unsafe for an allowlist. + return this.matchesMcpAllowlist( + toolName, + this.executionAllowedExactTools!, + this.executionAllowedMcpPatterns!, + ); + } + + /** + * Matches an MCP tool name against allowlist entries using the registry's + * raw server/tool identity. Comparing provider-sanitized prefixes can merge + * distinct server names such as "repo.bad" and "repo/bad", so it is unsafe + * for an allowlist. + */ + private matchesMcpAllowlist( + toolName: string, + exact: ReadonlySet, + patterns: readonly string[], + ): boolean { const registeredTool = this.runtimeContext .getToolRegistry() .getTool(toolName) as @@ -1959,14 +2084,11 @@ export class AgentCore { const serverToolName = registeredTool.serverToolName; const serverPattern = `mcp__${serverName}`; const rawToolName = `${serverPattern}__${serverToolName}`; - if ( - this.executionAllowedExactTools?.has(serverPattern) || - this.executionAllowedExactTools?.has(rawToolName) - ) { + if (exact.has(serverPattern) || exact.has(rawToolName)) { return true; } - return this.executionAllowedMcpPatterns!.some((pattern) => { + return patterns.some((pattern) => { if (pattern === 'mcp__*') { return true; } @@ -2017,7 +2139,7 @@ export class AgentCore { }> { if ( this.codeModeAllowedToolNames === undefined && - this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly + isCodeModeEnabled(this.runtimeContext.getToolMode?.()) ) { await this.prepareTools(); } @@ -2566,13 +2688,18 @@ export class AgentCore { ...(toolCallArgumentsWereIncomplete(fc) ? { hadIncompleteArguments: true } : {}), - ...((toolName === ToolNames.EXEC || - toolName === ToolNames.TOOL_SEARCH) && - this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly - ? { - codeModeAllowedToolNames: this.codeModeAllowedToolNames ?? [], - } - : {}), + // The narrowed binding set rides on every code-mode agent request: + // exec dispatch gates on it, and the scheduler exposes it as the + // ambient allowlist so reachability hints resolve against the + // surface this agent actually received. + ...(this.codeModeAllowedToolNames !== undefined + ? { codeModeAllowedToolNames: this.codeModeAllowedToolNames } + : (toolName === ToolNames.EXEC && + isCodeModeEnabled(this.runtimeContext.getToolMode?.())) || + (toolName === ToolNames.TOOL_SEARCH && + this.runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly) + ? { codeModeAllowedToolNames: [] } + : {}), }; if (canonicalToolName(toolName) === ToolNames.TOOL_CALL) { diff --git a/packages/core/src/agents/runtime/agent-types.ts b/packages/core/src/agents/runtime/agent-types.ts index 8223bbd20a2..210df4fd3be 100644 --- a/packages/core/src/agents/runtime/agent-types.ts +++ b/packages/core/src/agents/runtime/agent-types.ts @@ -103,6 +103,8 @@ export interface ToolConfig { * Supports exact tool names and MCP server-level patterns. */ executionAllowedTools?: string[]; + /** Exact names allowed only inside exec; never grants direct tool calls. */ + nestedExecutionAllowedTools?: string[]; /** * Optional list of tool names to exclude from the agent's tool pool. diff --git a/packages/core/src/agents/runtime/subagent-plan-tool-policy.test.ts b/packages/core/src/agents/runtime/subagent-plan-tool-policy.test.ts index c0b902c8287..1bce8404d1f 100644 --- a/packages/core/src/agents/runtime/subagent-plan-tool-policy.test.ts +++ b/packages/core/src/agents/runtime/subagent-plan-tool-policy.test.ts @@ -252,7 +252,7 @@ describe('subagent plan tool policy', () => { expect(toolConfigAllowsSkill(toolConfig)).toBe(false); }); - it('credits the exec gateway only under CodeModeOnly', () => { + it('credits the exec gateway only when its bindings are available', () => { const execList = { tools: [ToolNames.EXEC, ToolNames.READ_FILE] }; expect(toolConfigAllowsSkill(execList, true)).toBe(true); expect(toolConfigAllowsSkill(execList, false)).toBe(false); diff --git a/packages/core/src/agents/runtime/subagent-plan-tool-policy.ts b/packages/core/src/agents/runtime/subagent-plan-tool-policy.ts index a607cdc7474..a66e60dac67 100644 --- a/packages/core/src/agents/runtime/subagent-plan-tool-policy.ts +++ b/packages/core/src/agents/runtime/subagent-plan-tool-policy.ts @@ -5,6 +5,7 @@ */ import { ToolNames } from '../../tools/tool-names.js'; +import { ToolMode } from '../../tools/code-mode.js'; import { matchesToolPattern } from '../../permissions/rule-parser.js'; import type { ToolResult } from '../../tools/tools.js'; import type { ToolConfig } from './agent-types.js'; @@ -88,6 +89,43 @@ export const EXCLUDED_TOOLS_FOR_SUBAGENTS: ReadonlySet = new Set([ ToolNames.MANAGE_MEMORY, ]); +type AgentSkillContext = Pick< + Config, + 'getToolMode' | 'getToolRegistry' | 'getVisibleTools' +>; + +export function isAgentSkillEagerHidden( + context: AgentSkillContext, + registryWillBeRebuilt = false, +): boolean { + const registry = context.getToolRegistry?.(); + // Lazy factories may not have a Tool instance yet. Registry rebuilds copy + // discovery metadata, but deliberately leave transient reveals behind. + return ( + context.getToolMode?.() === ToolMode.CodeMode && + registry?.isPermissionDeferred?.(ToolNames.SKILL) === true && + !context.getVisibleTools?.().has(ToolNames.SKILL) && + (registryWillBeRebuilt || + registry?.isDeferredToolRevealed?.(ToolNames.SKILL) !== true) + ); +} + +export function hasAgentSkillExecBinding( + context: AgentSkillContext, + registryWillBeRebuilt = false, +): boolean { + const mode = context.getToolMode?.(); + return ( + mode === ToolMode.CodeModeOnly || + (mode === ToolMode.CodeMode && + !!context + .getToolRegistry?.() + ?.getAllToolNames() + .includes(ToolNames.EXEC) && + !isAgentSkillEagerHidden(context, registryWillBeRebuilt)) + ); +} + /** * Whether an agent running with `toolConfig` is declared the Skill tool. * @@ -102,25 +140,13 @@ export const EXCLUDED_TOOLS_FOR_SUBAGENTS: ReadonlySet = new Set([ * listing and the pointer cannot disagree about whether a skill can actually * be loaded — the disagreement #12424 reports. * - * Tool mode is an input; registry state deliberately is not. A - * `permissions.deny` or `excludeTools` entry is a registry property, and - * `resolveBundledReferenceRoute` answers the route from it. The per-agent - * policy is the one input that resolver cannot see (#12424). - * - * Of the two `ToolMode.CodeModeOnly` arms, this predicate covers the `exec` - * gateway: `prepareTools()` additionally admits every `code-mode-callable` - * registry tool when the configured names include `exec` - * (`inheritsCodeModeBindings`, `agent-core.ts`), and `getToolExposure(SKILL)` - * is `code-mode-callable` because SKILL is in neither `HIDDEN_TOOLS` nor - * `DIRECT_ONLY_TOOLS`. So an agent whose finite list names `exec` but not - * `skill` reaches the Skill tool, and callers must pass the mode — with it - * omitted this answers `false` for that shape, which would withhold the manager - * and, through the `config.ts` registration guard, the Skill tool itself. - * - * The `exec` arm also holds when a `tools.eager` allowlist demotes `skill`: - * `prepareTools()` keeps eager-demoted tools in the code-mode allowlist - * (#12898), where they stay callable and discoverable through `tool_search`, - * so the pointer this answer leads to can be followed (#12809). + * Callers supply whether exec bindings are reachable in the current mode and + * registry. A finite list naming exec can therefore reach Skill without + * naming it directly. CodeModeOnly retains eager-deferred nested targets; + * The eager-hidden input vetoes every route in Hybrid, including wildcards + * and explicit Skill entries. Other permission bounds stay in prepareTools(). + * Registry deny/exclude rules remain the bundled-reference resolver's concern; + * this predicate supplies the per-agent policy that resolver cannot see. * * Matching is exact, as `prepareTools()`'s is: `SubagentManager` resolves * configured names to canonical tool names before they reach a `ToolConfig`. @@ -133,9 +159,10 @@ export const EXCLUDED_TOOLS_FOR_SUBAGENTS: ReadonlySet = new Set([ */ export function toolConfigAllowsSkill( toolConfig: ToolConfig | undefined, - codeModeOnly = false, + execBindingsAvailable = false, + skillEagerHidden = false, ): boolean { - if (EXCLUDED_TOOLS_FOR_SUBAGENTS.has(ToolNames.SKILL)) { + if (skillEagerHidden || EXCLUDED_TOOLS_FOR_SUBAGENTS.has(ToolNames.SKILL)) { return false; } // No per-agent config inherits the whole registry. @@ -153,9 +180,9 @@ export function toolConfigAllowsSkill( // list holding only inline declarations inherits: both take the explicit // branch there, which declares no registry tool. const inheritsRegistry = names.includes('*'); - // Under CodeModeOnly, naming `exec` inherits every code-mode-callable - // binding (`prepareTools()`), and `skill` is one of them. - const reachesThroughExec = codeModeOnly && names.includes(ToolNames.EXEC); + // Naming an available exec gateway can reach the nested Skill binding. + const reachesThroughExec = + execBindingsAvailable && names.includes(ToolNames.EXEC); return ( inheritsRegistry || names.includes(ToolNames.SKILL) || reachesThroughExec ); diff --git a/packages/core/src/code-mode/code-mode.test.ts b/packages/core/src/code-mode/code-mode.test.ts index d6f19a7d6fd..43578bba930 100644 --- a/packages/core/src/code-mode/code-mode.test.ts +++ b/packages/core/src/code-mode/code-mode.test.ts @@ -16,6 +16,7 @@ import { getToolExposure, isCodeModeToolCallAllowed, planCodeModeBindings, + ToolMode, type CodeModeBindingPlan, } from '../tools/code-mode.js'; import { executeCodeMode } from './host-client.js'; @@ -87,7 +88,7 @@ function registryWith( params?: Record, ) { const registry = new ToolRegistry( - makeFakeConfig(config ?? { codeModeOnly: true }), + makeFakeConfig(config ?? { toolMode: 'code_mode_only' }), ); for (const name of names) { registry.registerTool(new MockTool({ name, params })); @@ -190,32 +191,244 @@ describe('CodeModeOnly exposure', () => { expect(description).toContain('terminal update_goal'); }); - it('registers exec only when CodeModeOnly is enabled', async () => { + it('registers exec in both code modes', async () => { const direct = makeFakeConfig(); const directRegistry = await direct.createToolRegistry(undefined, { skipDiscovery: true, }); - const codeMode = makeFakeConfig({ codeModeOnly: true }); + const codeMode = makeFakeConfig({ toolMode: ToolMode.CodeMode }); const codeModeRegistry = await codeMode.createToolRegistry(undefined, { skipDiscovery: true, }); + const codeModeOnly = makeFakeConfig({ toolMode: ToolMode.CodeModeOnly }); + const codeModeOnlyRegistry = await codeModeOnly.createToolRegistry( + undefined, + { skipDiscovery: true }, + ); + const invalid = makeFakeConfig(); + vi.spyOn(invalid, 'getToolMode').mockReturnValue( + 'code-mode' as ReturnType, + ); + const invalidRegistry = await invalid.createToolRegistry(undefined, { + skipDiscovery: true, + }); expect(directRegistry.getAllToolNames()).not.toContain('exec'); expect(codeModeRegistry.getAllToolNames()).toContain('exec'); expect(codeModeRegistry.getAllToolNames()).toContain('tool_search'); + expect(codeModeOnlyRegistry.getAllToolNames()).toContain('exec'); + expect(invalidRegistry.getAllToolNames()).not.toContain('exec'); }); it('keeps Direct declarations unchanged', () => { const registry = new ToolRegistry(makeFakeConfig()); - for (const name of ['read_file', 'tool_search', 'agent']) { + const tools = ['read_file', 'tool_search', 'agent'].map( + (name) => new MockTool({ name }), + ); + for (const tool of tools) registry.registerTool(tool); + + expect(registry.getFunctionDeclarations()).toEqual( + tools + .sort((left, right) => left.name.localeCompare(right.name)) + .map((tool) => tool.schema), + ); + }); + + it('keeps ordinary tools direct and adds nested declarations in CodeMode', () => { + const registry = new ToolRegistry( + makeFakeConfig({ toolMode: ToolMode.CodeMode }), + ); + registry.registerTool( + new MockTool({ + name: 'read_file', + description: 'Reads a file from the filesystem.', + params: { + type: 'object', + properties: { path: { type: 'string' } }, + required: ['path'], + }, + }), + ); + for (const name of ['tool_search', 'tool_call', 'agent', 'exec']) + registry.registerTool(new MockTool({ name })); + + const declarations = registry.getFunctionDeclarations(); + expect(declarations.map((item) => item.name)).toEqual([ + 'agent', + 'exec', + 'read_file', + 'tool_call', + 'tool_search', + ]); + const readFileDescription = declarations.find( + (item) => item.name === 'read_file', + )?.description; + expect(readFileDescription).toContain('Reads a file from the filesystem.'); + expect(readFileDescription).toContain( + 'read_file(args: { "path": string })', + ); + expect( + declarations.find((item) => item.name === 'agent')?.description, + ).not.toContain('declare const tools:'); + const execDescription = declarations.find( + (item) => item.name === 'exec', + )?.description; + expect(execDescription).not.toContain('tools.read_file(args:'); + expect(execDescription).toContain('"name":"read_file"'); + expect(execDescription).toContain( + 'tool_search returns deferred parameter schemas without changing top-level declarations', + ); + }); + + it('keeps filtered CodeMode declarations and nested bindings in scope', () => { + const registry = new ToolRegistry( + makeFakeConfig({ toolMode: ToolMode.CodeMode }), + ); + for (const name of ['read_file', 'write_file', 'agent', 'exec']) { registry.registerTool(new MockTool({ name })); } - expect(registry.getFunctionDeclarations().map((item) => item.name)).toEqual( - ['agent', 'read_file', 'tool_search'], + const declarations = registry.getFunctionDeclarationsFiltered([ + 'read_file', + 'exec', + ]); + expect(declarations.map((item) => item.name)).toEqual([ + 'exec', + 'read_file', + ]); + expect(declarations[0]?.description).toContain('"name":"read_file"'); + expect(declarations[0]?.description).not.toContain('"name":"write_file"'); + // Every binding is already declared top-level, so no schema section + // follows: the description must not promise one. + expect(declarations[0]?.description).not.toContain('declared below'); + expect(declarations[1]?.description).toContain( + 'declare const tools: { read_file(args:', + ); + }); + + it('leaves direct declarations unchanged when exec is unavailable', () => { + const registry = new ToolRegistry( + makeFakeConfig({ toolMode: ToolMode.CodeMode }), + ); + const readFile = new MockTool({ name: 'read_file' }); + registry.registerTool(readFile); + + expect(registry.getFunctionDeclarations()).toEqual([readFile.schema]); + expect(registry.getFunctionDeclarationsFiltered(['read_file'])).toEqual([ + readFile.schema, + ]); + }); + + it('preserves deferred discovery in CodeMode', () => { + const registry = new ToolRegistry( + makeFakeConfig({ toolMode: ToolMode.CodeMode }), + ); + registry.registerTool(new MockTool({ name: 'exec' })); + registry.registerTool(new MockTool({ name: 'tool_search' })); + registry.registerTool(new MockTool({ name: 'tool_call' })); + registry.registerTool( + new MockTool({ name: 'deferred_tool', shouldDefer: true }), + ); + + const initial = registry.getFunctionDeclarations(); + expect(initial.map((item) => item.name)).toEqual([ + 'exec', + 'tool_call', + 'tool_search', + ]); + expect(initial[0]?.description).toContain('"name":"deferred_tool"'); + expect(initial[0]?.description).not.toContain('tools.deferred_tool(args:'); + expect(initial[0]?.description).toContain('use tool_call outside exec'); + expect(initial[0]?.description).toContain('returned by tool_search'); + expect(registry.getDeferredToolSummary().map((item) => item.name)).toEqual([ + 'deferred_tool', + ]); + + const filtered = registry.getFunctionDeclarationsFiltered( + ['exec', 'tool_call', 'tool_search'], + new Set(['deferred_tool']), + ); + expect(filtered[0]?.description).toContain('tools.deferred_tool(args:'); + expect(filtered[0]?.description).toContain('use tool_call outside exec'); + + const revealed = registry.getFunctionDeclarations({ + includeDeferred: true, + }); + expect( + revealed.find((item) => item.name === 'deferred_tool')?.description, + ).toContain('declare const tools: { deferred_tool(args:'); + }); + + it.each([ + { bridgeTools: [] }, + { bridgeTools: ['tool_search'] }, + { bridgeTools: ['tool_call'] }, + ])( + 'includes nested schemas with incomplete bridge $bridgeTools', + ({ bridgeTools }) => { + const registry = new ToolRegistry( + makeFakeConfig({ toolMode: ToolMode.CodeMode }), + ); + registry.registerTool(new MockTool({ name: 'exec' })); + for (const name of bridgeTools) { + registry.registerTool(new MockTool({ name })); + } + registry.registerTool( + new MockTool({ + name: 'deferred_tool', + shouldDefer: true, + params: { + type: 'object', + properties: { path: { type: 'string' } }, + required: ['path'], + }, + }), + ); + + const description = registry.getFunctionDeclarations()[0]?.description; + expect(description).toContain( + 'tools.deferred_tool(args: { "path": string })', + ); + expect(description).not.toContain('use tool_call outside exec'); + expect(description).not.toContain('returned by tool_search'); + expect(description).toContain('the tool has no nested binding'); + }, + ); + + it('describes an empty filtered CodeMode surface accurately', () => { + const registry = new ToolRegistry( + makeFakeConfig({ toolMode: ToolMode.CodeMode }), + ); + registry.registerTool(new MockTool({ name: 'exec' })); + registry.registerTool(new MockTool({ name: 'ask_user_question' })); + + const description = registry + .getFunctionDeclarationsFiltered(['exec', 'ask_user_question'], new Set()) + .find((declaration) => declaration.name === 'exec')?.description; + expect(description).toContain( + '// No ordinary tools are available in this context.', + ); + expect(description).not.toContain( + 'Nested tool declarations for directly exposed tools', ); }); + it('keeps filtered and unfiltered CodeMode declaration order aligned', () => { + const registry = new ToolRegistry( + makeFakeConfig({ toolMode: ToolMode.CodeMode }), + ); + for (const name of ['read_file', 'readFile', 'exec']) + registry.registerTool(new MockTool({ name })); + + const unfiltered = registry + .getFunctionDeclarations() + .map((declaration) => declaration.name); + const filtered = registry + .getFunctionDeclarationsFiltered(['read_file', 'readFile', 'exec']) + .map((declaration) => declaration.name); + expect(filtered).toEqual(unfiltered); + }); + it('exposes exec and direct controls while retaining ordinary and hidden tools', () => { const registry = registryWith([ 'read_file', @@ -251,6 +464,31 @@ describe('CodeModeOnly exposure', () => { expect(declarations[0]?.description).not.toContain('tools.write_file'); }); + it('keeps a permission-deferred binding on the session surface but not on a filtered one', async () => { + // The eager-deferral immunity the docs promise holds on the session + // surface only: an AgentCore surface (subagent, headless agent, arena) + // narrows the nested binding set by its own allowlist, so a tool demoted + // by `tools.eager` has no nested binding there. + const registry = new ToolRegistry( + makeFakeConfig({ toolMode: ToolMode.CodeModeOnly }), + ); + registry.registerTool(new MockTool({ name: 'exec' })); + registry.registerPermissionDeferredFactory( + 'write_file', + async () => new MockTool({ name: 'write_file' }), + ); + await registry.warmAll(); + + const sessionDescription = registry + .getFunctionDeclarations() + .find((item) => item.name === 'exec')?.description; + expect(sessionDescription).toContain('tools.write_file(args:'); + + const declarations = registry.getFunctionDeclarationsFiltered(['exec']); + expect(declarations.map((item) => item.name)).toEqual(['exec']); + expect(declarations[0]?.description).not.toContain('write_file'); + }); + it('keeps exec structured across Gemini, OpenAI, and Anthropic tool conversion', async () => { const registry = registryWith(['read_file', 'agent', 'exec'], undefined, { type: 'object', @@ -269,17 +507,17 @@ describe('CodeModeOnly exposure', () => { expect(anthropic[1]?.description).toContain('tools.read_file'); }); - it('builds stable declarations and resolves normalized-name collisions first-wins', () => { + it('builds stable declarations and prefers exact names over rewritten collisions', () => { const tools = [ + new MockTool({ name: 'z-tool' }), new MockTool({ - name: 'z-tool', + name: 'z_tool', params: { type: 'object', properties: { count: { type: 'integer' } }, required: ['count'], }, }), - new MockTool({ name: 'z_tool' }), new MockTool({ name: 'a-tool', shouldDefer: true, @@ -299,10 +537,10 @@ describe('CodeModeOnly exposure', () => { expect(first).toEqual(second); expect(first.bindings.map((item) => item.name)).toEqual([ 'a-tool', - 'z-tool', + 'z_tool', ]); expect(first.collisions).toEqual([ - { jsName: 'z_tool', kept: 'z-tool', omitted: 'z_tool' }, + { jsName: 'z_tool', kept: 'z_tool', omitted: 'z-tool' }, ]); for (const fragment of [ 'tools.z_tool(args: { "count": number })', @@ -316,9 +554,84 @@ describe('CodeModeOnly exposure', () => { ]) { expect(buildExecDescription(first)).toContain(fragment); } + expect(buildExecDescription(first, { codeModeOnly: false })).toContain( + 'an uncaught rejection aborts the program', + ); expect(buildExecDescription(first)).not.toContain('text(result.value)'); }); + it('describes normalized-name collisions on the hybrid surface', () => { + const registry = new ToolRegistry( + makeFakeConfig({ toolMode: ToolMode.CodeMode }), + ); + registry.registerTool(new MockTool({ name: 'exec' })); + registry.registerTool( + new MockTool({ + name: 'read-file', + params: { + type: 'object', + properties: { path: { type: 'string' } }, + }, + }), + ); + registry.registerTool(new MockTool({ name: 'read_file' })); + + const declarations = registry.getFunctionDeclarations(); + const execDescription = declarations.find( + (item) => item.name === 'exec', + )?.description; + const keptDescription = declarations.find( + (item) => item.name === 'read_file', + )?.description; + const omittedDescription = declarations.find( + (item) => item.name === 'read-file', + )?.description; + + expect(execDescription).toContain( + '- read-file is omitted because it collides with read_file as tools.read_file.', + ); + expect(keptDescription).toContain('declare const tools: { read_file(args:'); + expect(omittedDescription).not.toContain('declare const tools:'); + }); + + it('does not reserve an exact name excluded from the nested allowlist', () => { + const bindings = planCodeModeBindings( + [new MockTool({ name: 'get-data' }), new MockTool({ name: 'get_data' })], + () => false, + new Set(['get-data']), + ); + expect( + bindings.bindings.map(({ name, jsName }) => ({ name, jsName })), + ).toEqual([{ name: 'get-data', jsName: 'get_data' }]); + expect(bindings.collisions).toEqual([]); + }); + + it.each([ToolMode.CodeMode, ToolMode.CodeModeOnly])( + 'dispatches the exact canonical MCP name under a collision in %s', + async (toolMode) => { + const registry = new ToolRegistry(makeFakeConfig({ toolMode })); + for (const name of [ + 'exec', + 'mcp__my-server__fetch', + 'mcp__my_server__fetch', + ]) { + registry.registerTool(new MockTool({ name })); + } + const dispatched: string[] = []; + const result = await executeCodeMode( + 'text((await tools.mcp__my_server__fetch({})).name)', + registry.getCodeModeBindingPlan(), + runtime(async (name) => { + dispatched.push(name); + return { callId: name, name, status: 'success', output: name }; + }), + new AbortController().signal, + ); + expect(dispatched).toEqual(['mcp__my_server__fetch']); + expect(result.output).toBe('mcp__my_server__fetch'); + }, + ); + it('keeps deferred tool schemas when search is unavailable', () => { const deferredPlan = planCodeModeBindings( [ @@ -347,8 +660,30 @@ describe('CodeModeOnly exposure', () => { expect(description).toContain('"deferred":true'); }); + it('defers hybrid descriptions while retaining exact nested binding names', () => { + const config = makeFakeConfig({ toolMode: 'code_mode' }); + const registry = new ToolRegistry(config); + for (const name of ['exec', 'tool_search', 'tool_call', 'read_file']) { + registry.registerTool(new MockTool({ name })); + } + registry.registerTool( + new MockTool({ + name: 'remote_lookup', + shouldDefer: true, + description: 'PRIVATE_DEFERRED_DESCRIPTION', + }), + ); + const description = registry + .getFunctionDeclarations() + .find((d) => d.name === 'exec')!.description!; + expect(description).not.toContain('PRIVATE_DEFERRED_DESCRIPTION'); + expect(description).toContain('"name":"remote_lookup"'); + expect(description).toContain('"jsName":"remote_lookup"'); + expect(description).toContain('"deferred":true'); + }); + it('omits deferred metadata while retaining callable bindings and stable declarations', () => { - const config = makeFakeConfig({ codeModeOnly: true }); + const config = makeFakeConfig({ toolMode: 'code_mode_only' }); const registry = new ToolRegistry(config); for (const name of ['exec', 'tool_search', 'read_file']) { registry.registerTool(new MockTool({ name })); @@ -386,7 +721,7 @@ describe('CodeModeOnly exposure', () => { }); it('keeps visible deferred signatures and omits hidden collision names', () => { - const config = makeFakeConfig({ codeModeOnly: true }); + const config = makeFakeConfig({ toolMode: 'code_mode_only' }); vi.spyOn(config, 'getVisibleTools').mockReturnValue( new Set(['visible_tool']), ); diff --git a/packages/core/src/code-mode/goal-evidence.test.ts b/packages/core/src/code-mode/goal-evidence.test.ts index c82c71b9a0f..bfaef696529 100644 --- a/packages/core/src/code-mode/goal-evidence.test.ts +++ b/packages/core/src/code-mode/goal-evidence.test.ts @@ -21,6 +21,7 @@ import { ChatRecordingService } from '../services/chatRecordingService.js'; import { buildApiHistoryFromConversation } from '../services/session-api-history.js'; import { makeFakeConfig } from '../test-utils/config.js'; import { validateTranscriptRecord } from '../utils/transcript-records.js'; +import { ToolMode } from '../tools/code-mode.js'; import { ExecTool } from '../tools/exec.js'; import { ReadFileTool } from '../tools/read-file.js'; import { ToolRegistry } from '../tools/tool-registry.js'; @@ -46,7 +47,7 @@ describe('Code Mode Goal evidence', () => { targetDir: workspace, cwd: workspace, sessionId: randomUUID(), - codeModeOnly: true, + toolMode: ToolMode.CodeModeOnly, chatRecording: false, telemetry: { enabled: false }, deferTelemetryInitialization: true, diff --git a/packages/core/src/code-mode/omni-integration.test.ts b/packages/core/src/code-mode/omni-integration.test.ts index 3279e1d1420..94d5ffda7cd 100644 --- a/packages/core/src/code-mode/omni-integration.test.ts +++ b/packages/core/src/code-mode/omni-integration.test.ts @@ -12,6 +12,7 @@ import { makeFakeConfig } from '../test-utils/config.js'; import { MockTool } from '../test-utils/mock-tool.js'; import { ExecTool } from '../tools/exec.js'; import { ToolRegistry } from '../tools/tool-registry.js'; +import { ToolMode } from '../tools/code-mode.js'; import { Kind, type MediaPolicyToolDescriptor } from '../tools/tools.js'; class PolicyTool extends MockTool { @@ -26,7 +27,7 @@ class PolicyTool extends MockTool { function setup(omniEnabled = false) { const config = makeFakeConfig({ - codeModeOnly: true, + toolMode: ToolMode.CodeModeOnly, omniEnabled, approvalMode: ApprovalMode.DEFAULT, targetDir: '/tmp', diff --git a/packages/core/src/code-mode/scheduler.test.ts b/packages/core/src/code-mode/scheduler.test.ts index 658c3c0de78..2f352ed19f8 100644 --- a/packages/core/src/code-mode/scheduler.test.ts +++ b/packages/core/src/code-mode/scheduler.test.ts @@ -49,7 +49,7 @@ function setup( } = {}, ) { const config = makeFakeConfig({ - codeModeOnly: true, + toolMode: 'code_mode_only', approvalMode: ApprovalMode.DEFAULT, chatRecording: false, targetDir: '/tmp', @@ -106,7 +106,7 @@ const findUpdate = (updates: ToolCall[][], name: string, status?: string) => describe('CodeModeOnly scheduler dispatch', () => { it('searches a scoped deferred tool, executes it, and still validates arguments', async () => { const config = makeFakeConfig({ - codeModeOnly: true, + toolMode: 'code_mode_only', targetDir: '/tmp', cwd: '/tmp', }); @@ -470,7 +470,7 @@ describe('CodeModeOnly scheduler dispatch', () => { it('runs Code Mode Bash calls in one Promise.allSettled batch', async () => { const config = makeFakeConfig({ - codeModeOnly: true, + toolMode: 'code_mode_only', approvalMode: ApprovalMode.DEFAULT, targetDir: '/tmp', cwd: '/tmp', @@ -640,7 +640,7 @@ describe('CodeModeOnly scheduler dispatch', () => { it('delivers nested PreToolUse context with the exec result, not the script value', async () => { const config = makeFakeConfig({ - codeModeOnly: true, + toolMode: 'code_mode_only', approvalMode: ApprovalMode.DEFAULT, targetDir: '/tmp', cwd: '/tmp', @@ -722,7 +722,7 @@ describe('CodeModeOnly scheduler dispatch', () => { it('delivers nested PostToolUseFailure context with the exec result, not the script error', async () => { const config = makeFakeConfig({ - codeModeOnly: true, + toolMode: 'code_mode_only', approvalMode: ApprovalMode.DEFAULT, targetDir: '/tmp', cwd: '/tmp', @@ -889,6 +889,47 @@ describe('CodeModeOnly scheduler dispatch', () => { ); }); + it('allows an ordinary direct call on the hybrid CodeMode surface', async () => { + const config = makeFakeConfig({ + toolMode: 'code_mode', + approvalMode: ApprovalMode.DEFAULT, + targetDir: '/tmp', + cwd: '/tmp', + }); + const registry = new ToolRegistry(config); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + const execute = vi.fn().mockResolvedValue({ + llmContent: 'read ok', + returnDisplay: 'read ok', + }); + registry.registerTool( + new MockTool({ name: 'read_probe', kind: Kind.Read, execute }), + ); + const completed = vi.fn(); + const scheduler = new CoreToolScheduler({ + config, + onAllToolCallsComplete: async (calls) => completed(calls), + onToolCallsUpdate: vi.fn(), + getPreferredEditor: () => undefined, + onEditorClose: vi.fn(), + }); + + await scheduler.schedule( + { + callId: 'hybrid-direct-read', + name: 'read_probe', + args: {}, + isClientInitiated: false, + prompt_id: 'prompt-hybrid-direct-read', + }, + new AbortController().signal, + ); + + expect(execute).toHaveBeenCalledOnce(); + await vi.waitFor(() => expect(completed).toHaveBeenCalledOnce()); + expect(completed.mock.calls[0]?.[0][0].status).toBe('success'); + }); + it('enforces a restricted agent allowlist inside exec', async () => { const read = vi.fn().mockResolvedValue({ llmContent: 'read ok', diff --git a/packages/core/src/config/config-execution-environment.test.ts b/packages/core/src/config/config-execution-environment.test.ts index c9ab6641275..e252e55cd09 100644 --- a/packages/core/src/config/config-execution-environment.test.ts +++ b/packages/core/src/config/config-execution-environment.test.ts @@ -13,6 +13,7 @@ import { deriveWorktreeConfig, } from './config.js'; import { ToolNames } from '../tools/tool-names.js'; +import { ToolMode } from '../tools/code-mode.js'; import type { DebugLogger } from '../utils/debugLogger.js'; import { ExecutionCleanupError, @@ -203,17 +204,51 @@ describe('execution environment ownership', () => { ); it('rejects a code-mode-only container registry for direct derived Config callers', async () => { - const parent = new Config({ ...params, codeModeOnly: true }); + const parent = new Config({ + ...params, + toolMode: ToolMode.CodeModeOnly, + }); const child = deriveConfig(parent, { getExecutionEnvironment: () => ({}) as ExecutionEnvironment, }); await expect( child.createToolRegistry(undefined, { skipDiscovery: true }), - ).rejects.toThrow('tools.codeModeOnly'); + ).rejects.toThrow('tools.mode = "code_mode_only"'); expect(parent.getCodeModeOnly()).toBe(true); expect(parent.getExecutionEnvironment()).toBeUndefined(); }); + it('warns once per root session when hybrid container registries fall back to direct tools', async () => { + const parent = new Config({ ...params, toolMode: ToolMode.CodeMode }); + const child = deriveConfig(parent, { + getExecutionEnvironment: () => ({}) as ExecutionEnvironment, + }); + const warn = vi.spyOn(console, 'warn').mockImplementation(() => {}); + try { + const registry = await child.createToolRegistry(undefined, { + skipDiscovery: true, + }); + expect(registry.getAllToolNames()).toContain(ToolNames.READ_FILE); + expect(registry.getAllToolNames()).not.toContain(ToolNames.EXEC); + expect(warn).toHaveBeenCalledWith( + expect.stringContaining('continuing with direct tools'), + ); + await child.createToolRegistry(undefined, { skipDiscovery: true }); + await deriveConfig(parent, { + getExecutionEnvironment: () => ({}) as ExecutionEnvironment, + }).createToolRegistry(undefined, { skipDiscovery: true }); + expect(warn).toHaveBeenCalledTimes(1); + + await deriveConfig( + new Config({ ...params, toolMode: ToolMode.CodeMode }), + { getExecutionEnvironment: () => ({}) as ExecutionEnvironment }, + ).createToolRegistry(undefined, { skipDiscovery: true }); + expect(warn).toHaveBeenCalledTimes(2); + } finally { + warn.mockRestore(); + } + }); + it.each(['ready', 'starting'])( 'waits for %s environments during session shutdown', async (state) => { diff --git a/packages/core/src/config/config.ts b/packages/core/src/config/config.ts index 0ed025ce25e..a29f1492b4e 100644 --- a/packages/core/src/config/config.ts +++ b/packages/core/src/config/config.ts @@ -126,6 +126,8 @@ import { ToolRegistry, type ToolFactory } from '../tools/tool-registry.js'; import type { McpBudgetEvent } from '../tools/mcp-client-manager.js'; import { ToolNames } from '../tools/tool-names.js'; import { + isCodeModeEnabled, + isToolMode, ToolMode, type ToolMode as ToolModeValue, } from '../tools/code-mode.js'; @@ -1216,8 +1218,8 @@ export interface ConfigParameters { * auto-approval and never affects registration (#10075). */ eagerTools?: string[]; - /** Replace ordinary model-facing tools with the isolated exec bridge. */ - codeModeOnly?: boolean; + /** Select how model-facing tools are exposed. */ + toolMode?: ToolModeValue; /** Use Responses Custom Tool text input for exec in Code Mode Only. */ freeform?: boolean; /** @@ -3246,6 +3248,7 @@ export class Config { private readonly advisorUsage = { calls: 0 }; private readonly webSearchSettings?: WebSearchSettings; private webSearchNoticeEmitted = false; + private readonly codeModeWarnings = { containerFallback: false }; /** * Per-session web_search call count. An object that is never reassigned: * derived Configs (`deriveConfig` → `Object.create(base)`) must mutate the @@ -3614,9 +3617,11 @@ export class Config { this.bareMode = params.bareMode ?? false; this.safeMode = params.safeMode ?? isSafeModeEnv(); this.toolMode = - params.codeModeOnly && !this.bareMode && !this.safeMode - ? ToolMode.CodeModeOnly - : ToolMode.Direct; + this.bareMode || this.safeMode + ? ToolMode.Direct + : isToolMode(params.toolMode) + ? params.toolMode + : ToolMode.Direct; this.freeform = this.toolMode === ToolMode.CodeModeOnly && params.freeform === true; if (this.safeMode) { @@ -12410,6 +12415,7 @@ export class Config { this, this.eventEmitter, sendSdkMcpMessage, + options?.forSubAgent, ); // The registry refuses every other tool of a Managed session, but its // manager still connects a runtime-added server. @@ -12479,7 +12485,7 @@ export class Config { }; const registerExecIfEnabled = async (): Promise => { - if (this.getToolMode() !== ToolMode.CodeModeOnly) return; + if (!isCodeModeEnabled(this.getToolMode())) return; await registerLazy(ToolNames.EXEC, async () => { const { ExecTool } = await import('../tools/exec.js'); return new ExecTool(this); @@ -12538,7 +12544,17 @@ export class Config { if (environment) { if (this.getCodeModeOnly()) { throw new Error( - 'Container execution cannot be combined with tools.codeModeOnly.', + 'Container execution cannot be combined with tools.mode = "code_mode_only".', + ); + } + if ( + this.getToolMode() === ToolMode.CodeMode && + !this.codeModeWarnings.containerFallback + ) { + this.codeModeWarnings.containerFallback = true; + // eslint-disable-next-line no-console -- the fallback must be visible without debug logging + console.warn( + 'Container execution does not support exec; continuing with direct tools for tools.mode = "code_mode".', ); } const [{ createExecutionTools }, { wrapExecutionTool }] = diff --git a/packages/core/src/config/review-workflow-cache.test.ts b/packages/core/src/config/review-workflow-cache.test.ts index 4bae2b56697..fd0f8e935a9 100644 --- a/packages/core/src/config/review-workflow-cache.test.ts +++ b/packages/core/src/config/review-workflow-cache.test.ts @@ -38,7 +38,7 @@ describe('review workflow cache continuity', () => { targetDir: directory, model: 'test-model', debugMode: false, - codeModeOnly: true, + toolMode: 'code_mode_only', }); const permissions = new PermissionManager(config); permissions.initialize(); diff --git a/packages/core/src/core/client.test.ts b/packages/core/src/core/client.test.ts index e40a6800fdc..f6c477bd273 100644 --- a/packages/core/src/core/client.test.ts +++ b/packages/core/src/core/client.test.ts @@ -124,6 +124,8 @@ import { import { collectAvailableSkillEntries } from '../tools/skill-utils.js'; import type { AvailableSkillEntry } from '../tools/skill-utils.js'; import { ToolNames } from '../tools/tool-names.js'; +import { planCodeModeBindings, ToolMode } from '../tools/code-mode.js'; +import { MockTool } from '../test-utils/mock-tool.js'; import { DEFERRED_TOOL_CALL_CANCELLATION_PREFIX, DEFERRED_TOOL_CALL_REFUSAL_PREFIX, @@ -680,6 +682,7 @@ describe('Gemini Client (client.ts)', () => { /** The suite's tool-registry mock, typed so each stub is a `Mock`. */ const registryMock = () => vi.mocked(mockConfig.getToolRegistry)() as unknown as Record< + | 'getCodeModeBindingPlan' | 'warmAll' | 'getDeferredToolSummary' | 'getMcpServerInstructions' @@ -842,6 +845,9 @@ describe('Gemini Client (client.ts)', () => { // LlmClient's constructor starts an async startChat that needs a // fully-formed Config, so the whole Config is mocked. const mockToolRegistry = { + getCodeModeBindingPlan: vi + .fn() + .mockReturnValue({ bindings: [], collisions: [] }), warmAll: vi.fn().mockResolvedValue(undefined), ensureTool: vi.fn().mockResolvedValue(null), getFunctionDeclarations: vi.fn().mockReturnValue([]), @@ -3403,6 +3409,102 @@ describe('Gemini Client (client.ts)', () => { ); }); + it('reports normal direct approval when exec is absent in CodeMode', async () => { + const reg = registryMock(); + reg.getTool.mockReturnValue(null); + reg.getDeferredToolSummary.mockReturnValue([ + { name: 'write_file', description: 'write' }, + ]); + reg.isPermissionDeferred.mockReturnValue(true); + mockConfig.getToolMode = vi.fn().mockReturnValue(ToolMode.CodeMode); + vi.spyOn(client.getChat(), 'setTools').mockImplementation(() => {}); + const warnSpy = vi.spyOn(console, 'warn').mockImplementation(() => {}); + + await client.setTools(); + + expect(warnSpy).toHaveBeenCalledWith( + expect.stringContaining( + 'direct calls by name still use normal approval: write_file', + ), + ); + expect(warnSpy).not.toHaveBeenCalledWith( + expect.stringContaining('remain callable through exec'), + ); + warnSpy.mockRestore(); + }); + + it('reports only code-mode-callable withheld tools as reachable through exec', async () => { + const reg = registryMock(); + reg.getCodeModeBindingPlan.mockReturnValue( + planCodeModeBindings( + [ + new MockTool({ name: 'write_file' }), + new MockTool({ name: 'send_message' }), + ], + () => true, + ), + ); + reg.getTool.mockImplementation((name: string) => + name === ToolNames.EXEC ? ({} as never) : null, + ); + reg.getDeferredToolSummary.mockReturnValue([ + { name: 'write_file', description: 'write' }, + { name: 'send_message', description: 'send' }, + ]); + reg.isPermissionDeferred.mockReturnValue(true); + mockConfig.getToolMode = vi.fn().mockReturnValue(ToolMode.CodeMode); + vi.spyOn(client.getChat(), 'setTools').mockImplementation(() => {}); + const warnSpy = vi.spyOn(console, 'warn').mockImplementation(() => {}); + + await client.setTools(); + + expect(warnSpy).toHaveBeenCalledWith( + expect.stringContaining('remain callable through exec: write_file'), + ); + expect(warnSpy).toHaveBeenCalledWith( + expect.stringContaining( + 'direct calls by name still use normal approval: write_file, send_message', + ), + ); + warnSpy.mockRestore(); + }); + + it('does not advertise an omitted collision binding as reachable through exec', async () => { + const reg = registryMock(); + reg.getTool.mockImplementation((name: string) => + name === ToolNames.EXEC ? ({} as never) : null, + ); + reg.getCodeModeBindingPlan.mockReturnValue( + planCodeModeBindings( + [ + new MockTool({ name: 'get--data' }), + new MockTool({ name: 'get-_data' }), + ], + () => true, + ), + ); + reg.getDeferredToolSummary.mockReturnValue([ + { name: 'get-_data', description: 'omitted target' }, + ]); + reg.isPermissionDeferred.mockReturnValue(true); + mockConfig.getToolMode = vi.fn().mockReturnValue(ToolMode.CodeMode); + vi.spyOn(client.getChat(), 'setTools').mockImplementation(() => {}); + const warn = vi.spyOn(console, 'warn').mockImplementation(() => {}); + try { + await client.setTools(); + expect(warn).toHaveBeenCalledWith( + expect.stringContaining( + 'direct calls by name still use normal approval: get-_data', + ), + ); + expect(warn).not.toHaveBeenCalledWith( + expect.stringContaining('remain callable through exec'), + ); + } finally { + warn.mockRestore(); + } + }); + it('names the missing bridge half when only tool_call is excluded', async () => { // The guard withholds on EITHER missing half; the warning must not blame // the half that IS registered (tool_search present → tool_call missing). diff --git a/packages/core/src/core/client.ts b/packages/core/src/core/client.ts index 2d54ccdfbaf..7b6fe9880fd 100644 --- a/packages/core/src/core/client.ts +++ b/packages/core/src/core/client.ts @@ -2218,16 +2218,15 @@ export class LlmClient { * revealed here so they land in the declaration list. Tools explicitly * demoted by `tools.eager` stay hidden unless the history-reveal pass above * already re-exposed one for a resumed session — that schema stays in the - * declarations, so it is not counted unreachable below. Skipping this for + * declarations, so it is not counted as bridge-hidden below. Skipping this for * ordinary deferred tools would leave them both off the declarations AND * off the deferred-summary list * (since `undefined` is returned in that branch) — a silent disappearance. * - * Returns `undefined` when the bridge is incomplete (ToolSearch or - * ToolCall unregistered): reminders must not advertise tools the model has - * no way to invoke on demand. Tools held back by `tools.eager` in that - * state are unreachable for the session, which is warned about once per - * session. + * Returns `undefined` when the ToolSearch + ToolCall bridge is incomplete. + * Held-back schemas stay out of top-level declarations, but registered tools use normal + * approval when called by name. Hybrid exec also retains its actual nested + * bindings. The warning identifies the missing bridge half and those bindings. */ private resolveDeferredToolsForReminder( deferredSummary: readonly DeferredToolSummary[], @@ -2266,6 +2265,18 @@ export class LlmClient { } if (withheld.length > 0 && !this.warnedAboutUnreachableEagerTools) { this.warnedAboutUnreachableEagerTools = true; + const hybridCodeMode = + this.config.getToolMode?.() === ToolMode.CodeMode; + const boundNames = new Set( + hybridCodeMode && toolRegistry.getTool(ToolNames.EXEC) + ? toolRegistry + .getCodeModeBindingPlan() + .bindings.map((binding) => binding.name) + : [], + ); + const nestedReachable = new Set( + withheld.filter((name) => boundNames.has(name)), + ); const missingHalves: string[] = []; if (!toolRegistry.getTool(ToolNames.TOOL_SEARCH)) { missingHalves.push(ToolNames.TOOL_SEARCH); @@ -2277,8 +2288,11 @@ export class LlmClient { console.warn( `tools.eager is holding back ${withheld.length} tool(s) in a session where the ` + `ToolSearch + ToolCall bridge is incomplete (${missingHalves.join(' and ')} not registered), ` + - `so they are not offered to the model and cannot be loaded through the bridge until restart; ` + + `so they are absent from top-level declarations and cannot be loaded through the bridge until restart; ` + `they remain registered and direct calls by name still use normal approval: ${withheld.join(', ')}. ` + + (nestedReachable.size > 0 + ? `These tools also retain their schemas and remain callable through exec: ${[...nestedReachable].join(', ')}. ` + : '') + `Enable tools.toolSearch.enabled (which registers both bridge tools) and drop any ` + `tool_search/tool_call deny rule, --exclude-tools entry, or tools.disabled entry to keep them loadable, ` + `list them in tools.eager to send their schemas upfront, or use permissions.deny if removal was the intent.`, diff --git a/packages/core/src/core/coreToolScheduler.test.ts b/packages/core/src/core/coreToolScheduler.test.ts index 12f81dacd8b..a28fdfcb5c1 100644 --- a/packages/core/src/core/coreToolScheduler.test.ts +++ b/packages/core/src/core/coreToolScheduler.test.ts @@ -13390,6 +13390,7 @@ describe('CoreToolScheduler activation wiring', () => { toolResult?: ToolResult; containerExecution?: boolean; withheldFromConfig?: boolean; + toolMode?: 'direct' | 'code_mode' | 'code_mode_only'; }; /** The single read_file request most activation cases schedule. */ @@ -13428,6 +13429,7 @@ describe('CoreToolScheduler activation wiring', () => { { getExecutionEnvironment: () => opts.containerExecution ? {} : undefined, + getToolMode: () => opts.toolMode, addInlineAnnouncedSkillKeys, ...(opts.withheldFromConfig ? { @@ -13459,6 +13461,24 @@ describe('CoreToolScheduler activation wiring', () => { }; } + it('uses the declared invocation surface for nested CodeModeOnly skill activation', async () => { + const { responseText, completedCall } = await runWithSkillManager( + { + matchAndActivateByPaths: vi.fn().mockResolvedValue(['tsx-helper']), + skillToolPresent: true, + toolMode: 'code_mode_only', + }, + { ...readRequest('/proj/src/App.tsx'), source: 'code_mode' }, + ); + expect(completedCall().status).toBe('success'); + expect(responseText()).toContain( + 'Load a skill by name using the tool interface declared in this session', + ); + expect(responseText()).not.toContain( + 'pass its name to the top-level Skill tool', + ); + }); + /** runWithSkillManager (App.tsx read) whose activation yields tsx-helper. */ async function runTsxActivation( opts: Omit, @@ -13522,7 +13542,9 @@ describe('CoreToolScheduler activation wiring', () => { expect(matchAndActivateByPaths).toHaveBeenCalledWith(['/proj/src/App.tsx']); expect(completedCall().status).toBe('success'); expect(responseText()).toContain('tsx-helper'); - expect(responseText()).toContain('became available via the Skill tool'); + expect(responseText()).toContain( + 'Load a skill by name using the tool interface declared in this session', + ); }); it('stays silent when SkillTool is registered but was never declared', async () => { @@ -13538,7 +13560,9 @@ describe('CoreToolScheduler activation wiring', () => { }); expect(completedCall().status).toBe('success'); - expect(responseText()).not.toContain('became available via the Skill tool'); + expect(responseText()).not.toContain( + 'Load a skill by name using the tool interface declared in this session', + ); expect(responseText()).not.toContain('tsx-helper'); // The half that starves the parent: moving this call outside the gate // (text still inside) passes everything else, yet the orchestrator's @@ -13557,7 +13581,9 @@ describe('CoreToolScheduler activation wiring', () => { }); expect(matchAndActivateByPaths).toHaveBeenCalledWith(['/proj/src/App.tsx']); expect(completedCall().status).toBe('success'); - expect(responseText()).not.toContain('became available via the Skill tool'); + expect(responseText()).not.toContain( + 'Load a skill by name using the tool interface declared in this session', + ); expect(addInlineAnnouncedSkillKeys).not.toHaveBeenCalled(); }); @@ -13571,7 +13597,9 @@ describe('CoreToolScheduler activation wiring', () => { declaredHasSkillTool: true, }); - expect(responseText()).toContain('became available via the Skill tool'); + expect(responseText()).toContain( + 'Load a skill by name using the tool interface declared in this session', + ); // …and the announcement IS consumed, so the parent does not repeat it; // the pair makes the negative assertion above mean "not consumed". expect(addInlineAnnouncedSkillKeys).toHaveBeenCalled(); diff --git a/packages/core/src/core/coreToolScheduler.ts b/packages/core/src/core/coreToolScheduler.ts index c4680d21f8e..cbab49a37af 100644 --- a/packages/core/src/core/coreToolScheduler.ts +++ b/packages/core/src/core/coreToolScheduler.ts @@ -242,6 +242,7 @@ import { runWithToolCallSource, type CodeModeToolResult, } from '../code-mode/tool-call-runtime.js'; +import { runWithCodeModeAllowedNames } from '../utils/code-mode-allowed-names.js'; import { isCodeModeToolCallAllowed, ToolMode } from '../tools/code-mode.js'; const debugLogger = createDebugLogger('TOOL_SCHEDULER'); @@ -1466,15 +1467,16 @@ interface CoreToolSchedulerOptions { /** Lets an outer owner suppress a scheduler result it already emitted. */ shouldObserveProducer?: (callId: string) => boolean; /** - * Whether the model this scheduler serves was DECLARED the Skill tool. + * Whether the model this scheduler serves can invoke a skill. * * The skill-activation reminder must not announce a skill to a model that - * cannot invoke one, and the registry cannot answer that: `SKILL` stays - * registered while a `tools.eager` allowlist defers its schema, and a - * subagent running an explicit `tools` list may never have it declared — - * nor is being declared sufficient, since a fork can keep a declaration it - * is forbidden to execute. An owner that filters either passes its own - * predicate here. + * cannot invoke one, and the registry cannot answer that: `SKILL` is + * registered unconditionally, including for subagents, while a subagent + * running an explicit `tools` list may never have it declared — nor is + * being declared sufficient, since a fork can keep a declaration it is + * forbidden to execute. In code mode, `exec` may instead expose Skill as a + * nested binding without a top-level declaration. An owner that filters + * either surface passes its own predicate here. * * It is NOT the predicate behind the startup `` snapshot, * and the two are independent rather than ordered. The snapshot is decided @@ -5863,27 +5865,31 @@ export class CoreToolScheduler { setPromoteAbortControllerCallback, canPromoteForegroundShell, ); - return scheduledCall.request.name === ToolNames.EXEC || - scheduledCall.request.name === ToolNames.TOOL_SEARCH - ? runWithToolCallRuntime( - { - parentCallId: callId, - allowedToolNames: - scheduledCall.request.codeModeAllowedToolNames, - dispatch: (name, args, nestedSignal, onResult) => - this.dispatchCodeModeTool( - name, - args, - scheduledCall.request, - nestedSignal, - onResult, - ), - }, - execute, - ) - : scheduledCall.request.source === 'code_mode' - ? runWithToolCallSource({ kind: 'code_mode' }, execute) - : execute(); + return runWithCodeModeAllowedNames( + scheduledCall.request.codeModeAllowedToolNames, + () => + scheduledCall.request.name === ToolNames.EXEC || + scheduledCall.request.name === ToolNames.TOOL_SEARCH + ? runWithToolCallRuntime( + { + parentCallId: callId, + allowedToolNames: + scheduledCall.request.codeModeAllowedToolNames, + dispatch: (name, args, nestedSignal, onResult) => + this.dispatchCodeModeTool( + name, + args, + scheduledCall.request, + nestedSignal, + onResult, + ), + }, + execute, + ) + : scheduledCall.request.source === 'code_mode' + ? runWithToolCallSource({ kind: 'code_mode' }, execute) + : execute(), + ); }), ); } else { @@ -5905,27 +5911,31 @@ export class CoreToolScheduler { liveOutputCallback, shellExecutionConfig, ); - return scheduledCall.request.name === ToolNames.EXEC || - scheduledCall.request.name === ToolNames.TOOL_SEARCH - ? runWithToolCallRuntime( - { - parentCallId: callId, - allowedToolNames: - scheduledCall.request.codeModeAllowedToolNames, - dispatch: (name, args, nestedSignal, onResult) => - this.dispatchCodeModeTool( - name, - args, - scheduledCall.request, - nestedSignal, - onResult, - ), - }, - execute, - ) - : scheduledCall.request.source === 'code_mode' - ? runWithToolCallSource({ kind: 'code_mode' }, execute) - : execute(); + return runWithCodeModeAllowedNames( + scheduledCall.request.codeModeAllowedToolNames, + () => + scheduledCall.request.name === ToolNames.EXEC || + scheduledCall.request.name === ToolNames.TOOL_SEARCH + ? runWithToolCallRuntime( + { + parentCallId: callId, + allowedToolNames: + scheduledCall.request.codeModeAllowedToolNames, + dispatch: (name, args, nestedSignal, onResult) => + this.dispatchCodeModeTool( + name, + args, + scheduledCall.request, + nestedSignal, + onResult, + ), + }, + execute, + ) + : scheduledCall.request.source === 'code_mode' + ? runWithToolCallSource({ kind: 'code_mode' }, execute) + : execute(), + ); }), ); } @@ -6356,13 +6366,13 @@ export class CoreToolScheduler { const activatedSkills = await skillManager?.matchAndActivateByPaths(candidatePaths); if (activatedSkills && activatedSkills.length > 0 && skillManager) { - // Gate on whether SkillTool was DECLARED to the model — the - // registry cannot answer that. See `hasSkillTool` in - // `CoreToolSchedulerOptions` for the mechanism and the reason. - const hasSkillTool = this.hasSkillToolOverride + // Gate on whether the model can invoke Skill through any declared + // surface — the registry cannot answer that. See `hasSkillTool` + // in `CoreToolSchedulerOptions` for the mechanism and the reason. + const canInvokeSkill = this.hasSkillToolOverride ? this.hasSkillToolOverride() : !!this.toolRegistry.getTool(ToolNames.SKILL); - if (hasSkillTool) { + if (canInvokeSkill) { // Render the just-activated skills with their description/whenToUse // (the full listing is no longer in the tool description, so the // model needs enough here to decide whether to invoke them). Source @@ -6394,7 +6404,7 @@ export class CoreToolScheduler { } if (activatedEntries.length > 0) { reminderBlocks.push( - `${SKILLS_ACTIVATED_OPENER}; invoke a skill by passing its name to the Skill tool:\n\n${renderAvailableSkillsBlock( + `${SKILLS_ACTIVATED_OPENER}. Load a skill by name using the tool interface declared in this session:\n\n${renderAvailableSkillsBlock( activatedEntries, )}\n`, ); diff --git a/packages/core/src/core/turn.ts b/packages/core/src/core/turn.ts index 8b6fe0b2e00..dc4851a46e9 100644 --- a/packages/core/src/core/turn.ts +++ b/packages/core/src/core/turn.ts @@ -206,7 +206,11 @@ export interface ToolCallRequestInfo { /** Parent model tool call for a programmatically dispatched child call. */ parentCallId?: string; source?: 'model' | 'code_mode'; - /** Exact tools an exec call may dispatch for a restricted agent. */ + /** + * The owning agent's narrowed nested binding set: exec dispatch gates on + * it, and the scheduler exposes it as the ambient allowlist for + * reachability hints. + */ codeModeAllowedToolNames?: readonly string[]; } diff --git a/packages/core/src/goals/goal-tools.test.ts b/packages/core/src/goals/goal-tools.test.ts index b8062cab023..409559096d0 100644 --- a/packages/core/src/goals/goal-tools.test.ts +++ b/packages/core/src/goals/goal-tools.test.ts @@ -597,7 +597,7 @@ describe('UpdateGoalTool', () => { [new UpdateGoalTool(makeConfig({}))], () => true, ); - const declaration = buildExecDescription(plan, true); + const declaration = buildExecDescription(plan, { searchAvailable: true }); expect(declaration).toContain( 'Deferred tool signatures and descriptions are omitted below', ); @@ -1145,7 +1145,7 @@ describe('ProposeGoalTool', () => { ); // A default code-mode session has the bridge, so the signature moves out // of the exec declaration and the discovery path is announced instead. - const bridged = buildExecDescription(plan, true); + const bridged = buildExecDescription(plan, { searchAvailable: true }); expect(bridged).toContain( 'Deferred tool signatures and descriptions are omitted below', ); diff --git a/packages/core/src/index.ts b/packages/core/src/index.ts index 674d26c3245..3dc80610712 100644 --- a/packages/core/src/index.ts +++ b/packages/core/src/index.ts @@ -829,6 +829,7 @@ export { } from './code-mode/tool-call-runtime.js'; export { getToolExposure, + isCodeModeEnabled, isCodeModeToolCallAllowed, ToolMode, type ToolExposure, diff --git a/packages/core/src/skills/bundled-reference.test.ts b/packages/core/src/skills/bundled-reference.test.ts index 86687e133bb..1f8f20eaaad 100644 --- a/packages/core/src/skills/bundled-reference.test.ts +++ b/packages/core/src/skills/bundled-reference.test.ts @@ -25,7 +25,15 @@ import { describe, expect, it, vi } from 'vitest'; import { AGENT_DELEGATION_SKILL_NAME } from './agent-delegation-skill.js'; -import { readBundledReference } from './bundled-reference.js'; +import { + readBundledReference, + resolveBundledReferenceRoute, + isToolHiddenBehindToolSearch, +} from './bundled-reference.js'; +import type { Config } from '../config/config.js'; +import { makeFakeConfig } from '../test-utils/config.js'; +import { ToolNames } from '../tools/tool-names.js'; +import { PermissionManager } from '../permissions/permission-manager.js'; import { WORKFLOW_AUTHORING_SKILL_NAME } from './workflow-authoring-skill.js'; /** First body line of each reference, after the frontmatter is stripped. */ @@ -37,6 +45,75 @@ const REFERENCES = [ [AGENT_DELEGATION_SKILL_NAME, DELEGATION_ANCHOR, WORKFLOW_ANCHOR], ] as const; +describe('hybrid reference routes', () => { + it.each([ToolNames.TOOL_SEARCH, ToolNames.TOOL_CALL])( + 'inlines an eager-hidden Skill for an agent missing %s', + async (missingBridge) => { + const config = makeFakeConfig({ + toolMode: 'code_mode', + eagerTools: [ToolNames.READ_FILE], + disabledTools: [missingBridge], + }); + const permissions = new PermissionManager(config); + permissions.initialize(); + vi.spyOn(config, 'getPermissionManager').mockReturnValue(permissions); + vi.spyOn(config, 'getSkillManager').mockReturnValue({} as never); + const registry = await config.createToolRegistry(undefined, { + skipDiscovery: true, + forSubAgent: true, + }); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); + try { + expect(registry.getAllToolNames()).toContain(ToolNames.EXEC); + expect(registry.getAllToolNames()).toContain(ToolNames.SKILL); + expect(registry.isPermissionDeferred(ToolNames.SKILL)).toBe(true); + expect(resolveBundledReferenceRoute(config, 'agent-delegation')).toBe( + 'inline', + ); + } finally { + await registry.stop(); + } + }, + ); + + it.each([ + [[], 'skill'], + [['tool_search'], 'skill'], + [['tool_call'], 'skill'], + [['tool_search', 'tool_call'], 'skill-via-tool-search'], + ] as const)('uses a reachable skill with bridge %j', (bridge, route) => { + const config = { + getToolMode: () => 'code_mode', + getSkillManager: () => ({}), + getToolRegistry: () => ({ + getAllToolNames: () => ['exec', 'skill', ...bridge], + isPermissionDeferred: () => true, + isDeferredToolRevealed: () => false, + }), + } as unknown as Config; + expect(resolveBundledReferenceRoute(config, 'agent-delegation')).toBe( + route, + ); + expect(isToolHiddenBehindToolSearch(config, 'skill')).toBe( + bridge.length === 2, + ); + }); + + it('inlines a hidden skill when both exec and the bridge are unavailable', () => { + const config = { + getToolMode: () => 'code_mode', + getSkillManager: () => ({}), + getToolRegistry: () => ({ + getAllToolNames: () => ['skill'], + isPermissionDeferred: () => true, + }), + } as unknown as Config; + expect(resolveBundledReferenceRoute(config, 'agent-delegation')).toBe( + 'inline', + ); + }); +}); + describe('readBundledReference', () => { /** * Both orders, because a single-slot cache is wrong in whichever order it is diff --git a/packages/core/src/skills/bundled-reference.ts b/packages/core/src/skills/bundled-reference.ts index 4f3d6ea97fc..d7a25c7850d 100644 --- a/packages/core/src/skills/bundled-reference.ts +++ b/packages/core/src/skills/bundled-reference.ts @@ -207,11 +207,26 @@ export function resolveBundledReferenceSurface( * Whether a registered tool's schema can be withheld from the request: * permission-deferred by a `tools.eager` allowlist and not listed in * `tools.visible`. A ToolSearch reveal is not consulted, because `/clear` - * drops it — a decision recorded once has to ask this. CodeModeOnly invokes - * deferred tools through `exec` instead of the Direct-mode bridge. + * drops it — a decision recorded once has to ask this. CodeModeOnly discovers + * schemas through top-level search and invokes deferred tools through exec. + * Session Hybrid exec also carries their schemas when the bridge is incomplete. + * AgentCore filters eager-hidden tools from its nested bindings, so an agent + * registry cannot use that session fallback, even before prepareTools runs. */ function isToolDeferredBehindToolSearch(config: Config, name: string): boolean { - if (config.getToolMode?.() === ToolMode.CodeModeOnly) return false; + const mode = config.getToolMode?.(); + if (mode === ToolMode.CodeModeOnly) return false; + const names = config.getToolRegistry?.()?.getAllToolNames?.() ?? []; + if ( + mode === ToolMode.CodeMode && + !config.getToolRegistry?.()?.forSubAgent && + names.includes(ToolNames.EXEC) && + !( + names.includes(ToolNames.TOOL_SEARCH) && + names.includes(ToolNames.TOOL_CALL) + ) + ) + return false; if (!config.getToolRegistry?.()?.isPermissionDeferred?.(name)) return false; return !config.getVisibleTools?.()?.has(name); } diff --git a/packages/core/src/subagents/subagent-manager.skill-eager-review.test.ts b/packages/core/src/subagents/subagent-manager.skill-eager-review.test.ts new file mode 100644 index 00000000000..9bb756095f1 --- /dev/null +++ b/packages/core/src/subagents/subagent-manager.skill-eager-review.test.ts @@ -0,0 +1,130 @@ +/** + * @license + * Copyright 2026 Qwen + * SPDX-License-Identifier: Apache-2.0 + */ + +import { describe, expect, it, vi } from 'vitest'; +import type { Config } from '../config/config.js'; +import type { ToolConfig } from '../agents/runtime/agent-types.js'; +import { makeFakeConfig } from '../test-utils/config.js'; +import { PermissionManager } from '../permissions/permission-manager.js'; +import { SubagentManager } from './subagent-manager.js'; +import { AgentHeadless } from '../agents/runtime/agent-headless.js'; +import { AgentCore } from '../agents/runtime/agent-core.js'; +import { resolveAgentDelegationSurface } from '../skills/agent-delegation-skill.js'; +import { ToolMode } from '../tools/code-mode.js'; + +describe('nested agents retain their actual Skill availability', () => { + it.each([ + { mode: ToolMode.CodeMode, visible: false }, + { mode: ToolMode.CodeMode, visible: true }, + { mode: ToolMode.CodeModeOnly, visible: false }, + ])( + 'preserves Skill policy in $mode visible=$visible', + async ({ mode, visible }) => { + const config = makeFakeConfig({ + targetDir: process.cwd(), + cwd: process.cwd(), + toolMode: mode, + eagerTools: ['exec', 'agent', ...(visible ? ['skill'] : [])], + coreTools: ['exec', 'agent', 'skill', 'tool_search', 'tool_call'], + }); + const skills = { + addChangeListener: () => () => {}, + listSkills: async () => [], + getCachedSkills: () => [], + hasDiscoveryErrors: () => false, + isSkillActive: () => true, + }; + const permissions = new PermissionManager(config); + permissions.initialize(); + vi.spyOn(config, 'getPermissionManager').mockReturnValue(permissions); + vi.spyOn(config, 'getSkillManager').mockReturnValue(skills as never); + const manager = new SubagentManager(config); + vi.spyOn(manager, 'listSubagents').mockResolvedValue([]); + vi.spyOn(config, 'getSubagentManager').mockReturnValue(manager); + const rootRegistry = await config.createToolRegistry(undefined, { + skipDiscovery: true, + }); + vi.spyOn(config, 'getToolRegistry').mockReturnValue(rootRegistry); + + let capture: + | { context: Config; tools: ToolConfig | undefined } + | undefined; + const createSpy = vi + .spyOn(AgentHeadless, 'create') + .mockImplementation( + async (_name, context, _prompt, _model, _run, tools) => { + capture = { context, tools }; + return { + getCore: () => ({ subagentId: 'skill-eager-review' }), + } as unknown as AgentHeadless; + }, + ); + const handles: Array<{ dispose(): Promise }> = []; + let parent = config; + const expected = mode === ToolMode.CodeModeOnly || visible; + try { + for (let depth = 1; depth <= 3; depth++) { + handles.push( + await manager.createAgentHeadless( + { + name: `skill-eager-review-${depth}`, + description: 'Isolated nested Skill test', + systemPrompt: 'Inspect Skill availability only.', + level: 'session', + ...(depth === 2 ? { tools: ['exec', 'agent'] } : {}), + }, + parent, + ), + ); + expect(capture).toBeDefined(); + const { context, tools } = capture!; + const core = new AgentCore( + 'skill-eager-review', + context, + { systemPrompt: 'Inspect Skill availability only.' }, + { model: 'test-model' }, + { max_turns: 1 }, + tools, + ); + const observed = core as unknown as { + willHaveSkillTool(): boolean; + canInvokeSkill(names: ReadonlySet): boolean; + }; + const announced = observed.willHaveSkillTool(); + const declared = new Set( + (await core.prepareTools()).map((declaration) => declaration.name), + ); + expect + .soft(!!context.getSkillManager(), `depth ${depth}`) + .toBe(expected); + expect + .soft(announced, `announcement at depth ${depth}`) + .toBe(expected); + expect + .soft( + observed.canInvokeSkill(declared), + `invocation at depth ${depth}`, + ) + .toBe(expected); + if (!expected) { + expect + .soft( + resolveAgentDelegationSurface(context), + `pointer at depth ${depth}`, + ) + .toBe('inline'); + } + parent = context; + } + } finally { + createSpy.mockRestore(); + for (const handle of handles.reverse()) await handle.dispose(); + await rootRegistry.stop(); + vi.restoreAllMocks(); + } + }, + ); +}); diff --git a/packages/core/src/subagents/subagent-manager.test.ts b/packages/core/src/subagents/subagent-manager.test.ts index cf2cc6df958..2713765b9c3 100644 --- a/packages/core/src/subagents/subagent-manager.test.ts +++ b/packages/core/src/subagents/subagent-manager.test.ts @@ -14,7 +14,9 @@ import { SubagentError, SubagentErrorCode, } from './types.js'; -import type { ToolRegistry } from '../tools/tool-registry.js'; +import { ToolRegistry } from '../tools/tool-registry.js'; +import { ExecTool } from '../tools/exec.js'; +import { MockTool } from '../test-utils/mock-tool.js'; import type { Config } from '../config/config.js'; import { ApprovalMode } from '../config/approval-mode.js'; import { makeFakeConfig } from '../test-utils/config.js'; @@ -3204,25 +3206,68 @@ describe('SubagentManager', () => { // skills through the exec gateway and must keep its manager. The parent // is a real CodeModeOnly Config: dropping the tool-mode argument at the // createAgentHeadless call site turns this case red. - it('keeps the manager for an exec-only agent under CodeModeOnly', async () => { - const codeModeParent = makeFakeConfig({ codeModeOnly: true }); - vi.spyOn(codeModeParent, 'getSkillManager').mockReturnValue( - sessionManager, - ); - vi.spyOn(codeModeParent, 'getSubagentManager').mockReturnValue(manager); - vi.spyOn(codeModeParent, 'getToolRegistry').mockReturnValue( - mockToolRegistry, - ); + it.each(['code_mode_only', 'code_mode'] as const)( + 'keeps the manager and registered Skill for an exec-only agent in %s', + async (toolMode) => { + const codeModeParent = makeFakeConfig({ toolMode }); + const codeModeRegistry = new ToolRegistry(codeModeParent); + codeModeRegistry.registerFactory( + ToolNames.EXEC, + async () => new ExecTool(codeModeParent), + ); + expect(codeModeRegistry.getTool(ToolNames.EXEC)).toBeUndefined(); + vi.spyOn(codeModeParent, 'getSkillManager').mockReturnValue( + sessionManager, + ); + vi.spyOn(codeModeParent, 'getSubagentManager').mockReturnValue( + manager, + ); + vi.spyOn(codeModeParent, 'getToolRegistry').mockReturnValue( + codeModeRegistry, + ); - const context = await launch( - { tools: [ToolNames.EXEC] }, - codeModeParent, - ); - expect(context.getSkillManager()).toBe(sessionManager); - expect(context.getToolRegistry().getAllToolNames()).toContain( - ToolNames.SKILL, - ); - }); + const context = await launch( + { tools: [ToolNames.EXEC] }, + codeModeParent, + ); + expect(context.getSkillManager()).toBe(sessionManager); + expect(context.getToolRegistry().getAllToolNames()).toContain( + ToolNames.SKILL, + ); + }, + ); + + it.each([false, true])( + 'withholds the manager for an eager-hidden Hybrid Skill even if parent revealed it: %s', + async (revealed) => { + const parent = makeFakeConfig({ + toolMode: 'code_mode', + eagerTools: [ToolNames.READ_FILE], + }); + const registry = new ToolRegistry(parent); + registry.registerTool(new ExecTool(parent)); + registry.registerPermissionDeferredFactory( + ToolNames.SKILL, + async () => new MockTool({ name: ToolNames.SKILL }), + ); + await registry.ensureTool(ToolNames.SKILL); + if (revealed) registry.revealDeferredTool(ToolNames.SKILL); + vi.spyOn(parent, 'getSkillManager').mockReturnValue(sessionManager); + vi.spyOn(parent, 'getSubagentManager').mockReturnValue(manager); + vi.spyOn(parent, 'getToolRegistry').mockReturnValue(registry); + const { context, dispose } = await launchHandle( + { tools: [ToolNames.EXEC, ToolNames.AGENT] }, + parent, + ); + try { + expect(context.getSkillManager()).toBeNull(); + expect(resolveAgentDelegationSurface(context)).toBe('inline'); + } finally { + await dispose(); + await registry.stop(); + } + }, + ); // The rebuilt registry's tools are per-subagent instances: the nested // Agent tool subscribes to the *shared session* SubagentManager in its diff --git a/packages/core/src/subagents/subagent-manager.ts b/packages/core/src/subagents/subagent-manager.ts index 3f66282b188..2422f63812b 100644 --- a/packages/core/src/subagents/subagent-manager.ts +++ b/packages/core/src/subagents/subagent-manager.ts @@ -88,8 +88,12 @@ import { hasRebuiltToolRegistry, rebuildToolRegistryOnOverride, } from '../tools/agent/agent.js'; -import { toolConfigAllowsSkill } from '../agents/runtime/subagent-plan-tool-policy.js'; import { ToolMode } from '../tools/code-mode.js'; +import { + hasAgentSkillExecBinding, + isAgentSkillEagerHidden, + toolConfigAllowsSkill, +} from '../agents/runtime/subagent-plan-tool-policy.js'; const AGENT_CONFIG_DIR = 'agents'; @@ -1154,6 +1158,13 @@ export class SubagentManager { ], } : {}), + ...(configuredToolConfig?.nestedExecutionAllowedTools !== undefined + ? { + nestedExecutionAllowedTools: [ + ...configuredToolConfig.nestedExecutionAllowedTools, + ], + } + : {}), disallowedTools: Array.from( new Set([ ...(configuredToolConfig?.disallowedTools ?? []), @@ -1175,9 +1186,37 @@ export class SubagentManager { modelConfig.reasoningEffort, ); + const skillRegistryWillBeRebuilt = + !!Object.keys(config.mcpServers ?? {}).length || + !hasRebuiltToolRegistry(runtimeContext) || + sessionSkillManager(runtimeContext) !== + runtimeContext.getSkillManager(); + let skillEagerHidden = isAgentSkillEagerHidden( + runtimeContext, + skillRegistryWillBeRebuilt, + ); + if ( + skillRegistryWillBeRebuilt && + runtimeContext.getToolMode?.() === ToolMode.CodeMode && + !runtimeContext + .getToolRegistry() + .getAllToolNames() + .includes(ToolNames.SKILL) + ) { + // A parent without a SkillManager omits the Skill factory entirely. + // Recover its registration policy before restoring skills to a child. + const status = await runtimeContext + .getPermissionManager?.() + ?.getToolRegistrationStatus(ToolNames.SKILL); + skillEagerHidden = + status === 'disabled' || + (status === 'deferred' && + !runtimeContext.getVisibleTools().has(ToolNames.SKILL)); + } const skillsAvailable = toolConfigAllowsSkill( toolConfig, - runtimeContext.getToolMode?.() === ToolMode.CodeModeOnly, + hasAgentSkillExecBinding(runtimeContext, skillRegistryWillBeRebuilt), + skillEagerHidden, ); const { context: subagentContext, cleanup } = await this.buildSubagentContextOverride( diff --git a/packages/core/src/tools/agent/agent.code-mode-review.test.ts b/packages/core/src/tools/agent/agent.code-mode-review.test.ts new file mode 100644 index 00000000000..38e0fadfa01 --- /dev/null +++ b/packages/core/src/tools/agent/agent.code-mode-review.test.ts @@ -0,0 +1,62 @@ +/** + * @license + * Copyright 2026 Qwen + * SPDX-License-Identifier: Apache-2.0 + */ + +import { describe, expect, it } from 'vitest'; +import { buildChildMessage } from './fork-subagent.js'; +import { makeFakeConfig } from '../../test-utils/config.js'; +import { MockTool } from '../../test-utils/mock-tool.js'; +import { ToolRegistry } from '../tool-registry.js'; +import { ExecTool } from '../exec.js'; +import { ToolMode } from '../code-mode.js'; +import { ToolNames } from '../tool-names.js'; + +describe('Code-mode fork restriction matches declared call surfaces', () => { + it.each([ToolMode.CodeMode, ToolMode.CodeModeOnly])( + 'only advertises possible direct calls in %s', + (mode) => { + const config = makeFakeConfig({ toolMode: mode }); + const registry = new ToolRegistry(config); + registry.registerTool(new ExecTool(config)); + registry.registerTool(new MockTool({ name: ToolNames.READ_FILE })); + const declarationNames = new Set( + registry + .getFunctionDeclarationsFiltered([ + ToolNames.EXEC, + ToolNames.READ_FILE, + ]) + .map((declaration) => declaration.name), + ); + const message = buildChildMessage( + 'Inspect the implementation', + [ToolNames.READ_FILE], + undefined, + [ToolNames.READ_FILE], + mode, + ); + const directClaim = message.match(/direct-call allowlist: (\[[^\n]*\])/); + const claimedDirectTools: string[] = directClaim + ? JSON.parse(directClaim[1]) + : []; + + expect(declarationNames.has(ToolNames.EXEC)).toBe(true); + expect(message).toContain( + 'Inside exec, only these exact nested tool names are permitted: ["read_file"]', + ); + for (const name of claimedDirectTools) { + expect( + declarationNames.has(name), + `${name} is not directly declared`, + ).toBe(true); + } + if (mode === ToolMode.CodeMode) { + expect(claimedDirectTools).toContain(ToolNames.READ_FILE); + } else { + expect(declarationNames.has(ToolNames.READ_FILE)).toBe(false); + expect(claimedDirectTools).not.toContain(ToolNames.READ_FILE); + } + }, + ); +}); diff --git a/packages/core/src/tools/agent/agent.test.ts b/packages/core/src/tools/agent/agent.test.ts index 68b14739c59..f8bca0822cb 100644 --- a/packages/core/src/tools/agent/agent.test.ts +++ b/packages/core/src/tools/agent/agent.test.ts @@ -4,6 +4,7 @@ * SPDX-License-Identifier: Apache-2.0 */ +import { runWithCodeModeAllowedNames } from '../../utils/code-mode-allowed-names.js'; import { goalTurnContext } from '../../goals/goal-turn-context.js'; import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'; import { @@ -644,9 +645,9 @@ describe('AgentTool', () => { it('rejects code mode before starting a container, including when validation is bypassed', async () => { config.getCodeModeOnly = () => true; expect(agentTool.validateToolParams(params)).toContain( - 'tools.codeModeOnly', + 'tools.mode = "code_mode_only"', ); - await expectRefused('tools.codeModeOnly', true); + await expectRefused('tools.mode = "code_mode_only"', true); }); it('rejects nested launch rather than defaulting to host tools', async () => { @@ -3497,6 +3498,140 @@ describe('AgentTool', () => { ); }); + describe('inherited execution policy persistence', () => { + function preparePolicyRuntime(mode: 'code_mode' | 'code_mode_only') { + config.getToolMode = vi.fn().mockReturnValue(mode); + config.getToolRegistry().getTool = vi.fn(); + const emitter = new AgentEventEmitter(); + mockAgent.getCore().getEventEmitter = () => emitter; + mockAgent.setExternalMessageProvider = vi.fn(); + mockAgent.setExternalMessageWaiter = vi.fn(); + mockAgent.setExternalMessageWaitPredicate = vi.fn(); + vi.spyOn(transcript, 'attachJsonlTranscriptWriter').mockReturnValue({ + cleanup: vi.fn(), + }); + } + afterEach(() => vi.restoreAllMocks()); + it.each(['code_mode', 'code_mode_only'] as const)( + 'bounds and persists an explicit exec fork in %s', + async (mode) => { + preparePolicyRuntime(mode); + const names = [ + 'exec', + 'read_file', + 'mcp__github__read_file', + 'mcp__payments__charge', + ]; + vi.mocked(config.getToolRegistry().getAllToolNames).mockReturnValue( + names, + ); + vi.mocked(config.getLlmClient).mockReturnValue({ + getHistory: vi.fn().mockReturnValue([]), + getChat: vi.fn().mockReturnValue({ + getGenerationConfig: vi.fn().mockReturnValue({ + systemInstruction: 'parent system', + tools: [{ functionDeclarations: [{ name: 'exec' }] }], + }), + }), + } as unknown as ReturnType); + const writeMetaSpy = vi + .spyOn(transcript, 'writeAgentMeta') + .mockImplementation(() => {}); + const invocation = ( + agentTool as AgentToolWithProtectedMethods + ).createInvocation({ + description: 'inherit bounded exec', + prompt: 'read the implementation', + subagent_type: 'fork', + fork_tools: ['exec'], + run_in_background: true, + }); + const result = await runWithAgentConfiguredToolAllowlist( + [ + 'exec', + 'read_file', + 'mcp__github__read_*', + 'mcp__github__read_file', + ], + () => invocation.execute(), + ); + expect(result.error).toBeUndefined(); + const tools = vi.mocked(AgentHeadless.create).mock.calls[0]?.[5]; + expect(tools?.executionAllowedTools).toEqual([]); + expect(tools?.nestedExecutionAllowedTools).toEqual([ + 'read_file', + 'mcp__github__read_file', + ]); + expect(writeMetaSpy.mock.calls[0]?.[1]).toMatchObject({ + executionAllowedTools: [], + nestedExecutionAllowedTools: [ + 'read_file', + 'mcp__github__read_file', + ], + }); + writeMetaSpy.mockRestore(); + }, + ); + + it.each([false, true])( + 'bounds a default fork and persists its background sidecar (background=%s)', + async (background) => { + preparePolicyRuntime('code_mode'); + vi.mocked(config.getToolRegistry().getAllToolNames).mockReturnValue([ + 'exec', + 'read_file', + 'write_file', + ]); + vi.mocked(config.getLlmClient).mockReturnValue({ + getHistory: vi.fn().mockReturnValue([]), + getChat: vi.fn().mockReturnValue({ + getGenerationConfig: vi.fn().mockReturnValue({ + systemInstruction: 'parent system', + tools: [{ functionDeclarations: [{ name: 'exec' }] }], + }), + }), + } as unknown as ReturnType); + const writeMetaSpy = vi + .spyOn(transcript, 'writeAgentMeta') + .mockImplementation(() => {}); + const invocation = ( + agentTool as AgentToolWithProtectedMethods + ).createInvocation({ + description: 'inherit bounded default', + prompt: 'read the implementation', + subagent_type: 'fork', + run_in_background: background, + }); + const result = await runWithAgentConfiguredToolAllowlist( + ['read_file'], + () => + runWithCodeModeAllowedNames(['read_file', 'write_file'], () => + invocation.execute(), + ), + ); + expect(result.error).toBeUndefined(); + expect( + vi.mocked(AgentHeadless.create).mock.calls[0]?.[5] + ?.executionAllowedTools, + ).toEqual(['read_file']); + expect( + vi.mocked(AgentHeadless.create).mock.calls[0]?.[5] + ?.nestedExecutionAllowedTools, + ).toEqual(['read_file', 'write_file']); + if (background) { + expect(writeMetaSpy.mock.calls[0]?.[1]).toMatchObject({ + executionAllowedTools: ['read_file'], + nestedExecutionAllowedTools: ['read_file', 'write_file'], + }); + } else { + // Interactive forks return a placeholder and have no sidecar. + expect(writeMetaSpy).not.toHaveBeenCalled(); + } + writeMetaSpy.mockRestore(); + }, + ); + }); + it('preserves fork_tools deny-all inside a configured parent allowlist', async () => { parentTools(readSearchCall, [...bridged()]); await runFork( diff --git a/packages/core/src/tools/agent/agent.ts b/packages/core/src/tools/agent/agent.ts index 19ad0d2916e..5c958d6c1fd 100644 --- a/packages/core/src/tools/agent/agent.ts +++ b/packages/core/src/tools/agent/agent.ts @@ -10,6 +10,8 @@ import { randomUUID } from 'node:crypto'; import { realpath } from 'node:fs/promises'; import { BaseDeclarativeTool, BaseToolInvocation, Kind } from '../tools.js'; import { ToolNames, ToolDisplayNames } from '../tool-names.js'; +import { getCurrentCodeModeAllowedNames } from '../../utils/code-mode-allowed-names.js'; +import { getToolExposure, isCodeModeEnabled } from '../code-mode.js'; import { buildInheritedForkExecutionToolNames, EXCLUDED_TOOLS_FOR_SUBAGENTS, @@ -320,7 +322,7 @@ function getExecutionBackendError( return 'Container execution is not enabled by this host.'; } if (config.getCodeModeOnly?.()) { - return 'Container execution cannot be combined with tools.codeModeOnly.'; + return 'Container execution cannot be combined with tools.mode = "code_mode_only".'; } if (params.name !== undefined || !isTopLevelSession()) { return 'Container execution is available only for top-level regular subagents.'; @@ -1881,6 +1883,7 @@ class AgentToolInvocation extends BaseToolInvocation { }; const buildParentBoundExecutionAllowlist = ( fallbackTools: readonly string[], + forNestedBinding = false, ): string[] => { if (parentConfiguredToolAllowlist === undefined) { return buildForkExecutionAllowlist( @@ -1895,6 +1898,12 @@ class AgentToolInvocation extends BaseToolInvocation { !EXCLUDED_TOOLS_FOR_SUBAGENTS.has(toolName) && keepOffParentBlocklist(toolName) && (isRequestedByFork(toolName) || + (forNestedBinding && + isCodeModeEnabled(agentConfig.getToolMode?.()) && + requestedTools?.includes(ToolNames.EXEC) && + getToolExposure(toolName) === 'code-mode-callable' && + (!toolName.startsWith('mcp__') || + !requestedTools.some((name) => name.startsWith('mcp__')))) || (requestedTools !== undefined && requestedTools.length > 0 && (toolName === ToolNames.TOOL_SEARCH || @@ -1902,8 +1911,16 @@ class AgentToolInvocation extends BaseToolInvocation { parentToolNames.includes(toolName))), ); }; + const nestedExecutionAllowedTools = + parentConfiguredToolAllowlist !== undefined && + isCodeModeEnabled(agentConfig.getToolMode?.()) + ? buildParentBoundExecutionAllowlist( + getCurrentCodeModeAllowedNames() ?? defaultExecutionToolNames, + true, + ) + : undefined; const requestedExecutionAllowedTools = - requestedTools === undefined + requestedTools === undefined && nestedExecutionAllowedTools === undefined ? undefined : resolveForkExecutionAllowedTools( parentToolNames, @@ -1963,6 +1980,8 @@ class AgentToolInvocation extends BaseToolInvocation { lastMessage, requestedExecutionAllowedTools, profilePromptHint, + nestedExecutionAllowedTools, + agentConfig.getToolMode?.(), ); if (forkedMessages.length > 0) { // Model had function calls: append tool responses + directive, @@ -1996,6 +2015,8 @@ class AgentToolInvocation extends BaseToolInvocation { this.params.prompt, requestedExecutionAllowedTools, profilePromptHint, + nestedExecutionAllowedTools, + agentConfig.getToolMode?.(), ); } @@ -2025,6 +2046,9 @@ class AgentToolInvocation extends BaseToolInvocation { parentToolNames, buildParentBoundExecutionAllowlist(defaultExecutionToolNames), ), + ...(nestedExecutionAllowedTools !== undefined + ? { nestedExecutionAllowedTools } + : {}), ...(parentDisallowedTools?.length ? { disallowedTools: [...parentDisallowedTools] } : {}), @@ -2040,6 +2064,9 @@ class AgentToolInvocation extends BaseToolInvocation { parentToolNames, buildParentBoundExecutionAllowlist(defaultExecutionToolNames), ), + ...(nestedExecutionAllowedTools !== undefined + ? { nestedExecutionAllowedTools } + : {}), ...(parentDisallowedTools?.length ? { disallowedTools: [...parentDisallowedTools] } : {}), @@ -3756,7 +3783,8 @@ class AgentToolInvocation extends BaseToolInvocation { resolvedApprovalMode, ...(isFork && (this.params.fork_tools !== undefined || - this.forkProfile !== undefined) && + this.forkProfile !== undefined || + getCurrentAgentConfiguredToolAllowlist() !== undefined) && bgToolConfig?.executionAllowedTools !== undefined ? { executionAllowedTools: [...bgToolConfig.executionAllowedTools], @@ -3765,6 +3793,13 @@ class AgentToolInvocation extends BaseToolInvocation { // Unlike the allowlist above, the blocklist persists whenever the // fork carries one — it also bounds plain forks whose allowlist is // rebuilt from the live parent surface on resume. + ...(isFork && bgToolConfig?.nestedExecutionAllowedTools !== undefined + ? { + nestedExecutionAllowedTools: [ + ...bgToolConfig.nestedExecutionAllowedTools, + ], + } + : {}), ...(isFork && bgToolConfig?.disallowedTools?.length ? { disallowedTools: [...bgToolConfig.disallowedTools] } : {}), @@ -4678,7 +4713,8 @@ class AgentToolInvocation extends BaseToolInvocation { resolvedApprovalMode, ...(isFork && (this.params.fork_tools !== undefined || - this.forkProfile !== undefined) && + this.forkProfile !== undefined || + getCurrentAgentConfiguredToolAllowlist() !== undefined) && toolConfig?.executionAllowedTools !== undefined ? { executionAllowedTools: [...toolConfig.executionAllowedTools], @@ -4687,6 +4723,13 @@ class AgentToolInvocation extends BaseToolInvocation { // Unlike the allowlist above, the blocklist persists whenever the // fork carries one — it also bounds plain forks whose allowlist is // rebuilt from the live parent surface on resume. + ...(isFork && toolConfig?.nestedExecutionAllowedTools !== undefined + ? { + nestedExecutionAllowedTools: [ + ...toolConfig.nestedExecutionAllowedTools, + ], + } + : {}), ...(isFork && toolConfig?.disallowedTools?.length ? { disallowedTools: [...toolConfig.disallowedTools] } : {}), diff --git a/packages/core/src/tools/agent/fork-subagent.ts b/packages/core/src/tools/agent/fork-subagent.ts index 892641e51fd..166e9608e03 100644 --- a/packages/core/src/tools/agent/fork-subagent.ts +++ b/packages/core/src/tools/agent/fork-subagent.ts @@ -4,6 +4,7 @@ import type { Config } from '../../config/config.js'; import type { SubagentConfig } from '../../subagents/types.js'; import { BUBBLE_APPROVAL_MODE } from '../../subagents/types.js'; import { ToolNames } from '../tool-names.js'; +import { ToolMode } from '../code-mode.js'; import { getStartupContextLength, isSystemReminderContent, @@ -309,6 +310,8 @@ export function buildForkedMessages( assistantMessage: Content, executionAllowedTools?: readonly string[], promptHint?: string, + nestedExecutionAllowedTools?: readonly string[], + toolMode?: ToolMode, ): Content[] { const toolUseParts = assistantMessage.parts?.filter((part) => part.functionCall) || []; @@ -356,7 +359,13 @@ export function buildForkedMessages( parts: [ ...toolResultParts, { - text: buildChildMessage(directive, executionAllowedTools, promptHint), + text: buildChildMessage( + directive, + executionAllowedTools, + promptHint, + nestedExecutionAllowedTools, + toolMode, + ), }, ], }; @@ -409,14 +418,25 @@ export function buildChildMessage( directive: string, executionAllowedTools?: readonly string[], promptHint?: string, + nestedExecutionAllowedTools?: readonly string[], + toolMode?: ToolMode, ): string { const executionRestriction = - executionAllowedTools === undefined - ? '' - : executionAllowedTools.length === 0 - ? `\n\nTOOL EXECUTION RESTRICTION: + nestedExecutionAllowedTools !== undefined + ? `\n\nTOOL EXECUTION RESTRICTION: +${ + toolMode === ToolMode.CodeModeOnly + ? `You may call exec and, when declared, tool_search. Ordinary tools are reachable only inside exec; other declared direct-only control tools must also match this allowlist: ${JSON.stringify(executionAllowedTools ?? [])}.` + : `You may call exec and declared tools matched by this direct-call allowlist: ${JSON.stringify(executionAllowedTools ?? [])}.` +} +Inside exec, only these exact nested tool names are permitted: ${JSON.stringify(nestedExecutionAllowedTools)}. +A nested allowance alone does not authorize a direct call; direct calls must satisfy the mode and direct-call allowlist above.` + : executionAllowedTools === undefined + ? '' + : executionAllowedTools.length === 0 + ? `\n\nTOOL EXECUTION RESTRICTION: You may not execute any tools, even though tool declarations remain visible. Do not attempt tool calls.` - : `\n\nTOOL EXECUTION RESTRICTION: + : `\n\nTOOL EXECUTION RESTRICTION: You may execute only tools matched by this allowlist: ${JSON.stringify(executionAllowedTools)}. Other visible tool declarations are unavailable to you. Do not call them.`; const profileGuidance = promptHint diff --git a/packages/core/src/tools/code-mode.ts b/packages/core/src/tools/code-mode.ts index af6acd0fb3d..11a70791b67 100644 --- a/packages/core/src/tools/code-mode.ts +++ b/packages/core/src/tools/code-mode.ts @@ -20,11 +20,20 @@ export type ToolExposure = export const ToolMode = { Direct: 'direct', + CodeMode: 'code_mode', CodeModeOnly: 'code_mode_only', } as const; export type ToolMode = (typeof ToolMode)[keyof typeof ToolMode]; +export function isToolMode(mode: unknown): mode is ToolMode { + return Object.values(ToolMode).some((candidate) => candidate === mode); +} + +export function isCodeModeEnabled(mode: unknown): boolean { + return mode === ToolMode.CodeMode || mode === ToolMode.CodeModeOnly; +} + const HIDDEN_TOOLS = new Set(['tool_call']); const DIRECT_ONLY_TOOLS = new Set([ @@ -92,15 +101,21 @@ export function planCodeModeBindings( const bindings: CodeModeToolBinding[] = []; const collisions: CodeModeBindingPlan['collisions'] = []; const claimed = new Map(); - const sorted = [...tools].sort((a, b) => - a.name < b.name ? -1 : a.name > b.name ? 1 : 0, - ); + const sorted = tools + .filter( + (tool) => + getToolExposure(tool.name) === 'code-mode-callable' && + (!allowedNames || allowedNames.has(tool.name)), + ) + .sort((a, b) => (a.name < b.name ? -1 : a.name > b.name ? 1 : 0)); + const canonicalNames = new Set(sorted.map((tool) => tool.name)); for (const tool of sorted) { - if (getToolExposure(tool.name) !== 'code-mode-callable') continue; - if (allowedNames && !allowedNames.has(tool.name)) continue; const jsName = normalizeCodeModeToolName(tool.name); - const kept = claimed.get(jsName); + // An exact name owns its binding even if a rewritten name sorts first. + const kept = + claimed.get(jsName) ?? + (tool.name !== jsName && canonicalNames.has(jsName) ? jsName : undefined); if (kept) { collisions.push({ jsName, kept, omitted: tool.name }); continue; @@ -193,15 +208,48 @@ function schemaToType(schema: unknown): string { return 'unknown'; } +function bindingSignature(binding: CodeModeToolBinding): string { + return `${binding.jsName}(args: ${schemaToType(binding.parametersJsonSchema)}): Promise`; +} + export function describeCodeModeBinding(binding: CodeModeToolBinding): string { - const params = schemaToType(binding.parametersJsonSchema); - return `tools.${binding.jsName}(args: ${params}): Promise;`; + return `tools.${bindingSignature(binding)};`; +} + +export function augmentDeclarationForCodeMode( + declaration: FunctionDeclaration, + binding: CodeModeToolBinding, +): FunctionDeclaration { + return { + ...declaration, + description: `${declaration.description ?? ''} + +exec tool declaration: +\`\`\`ts +declare const tools: { ${bindingSignature(binding)}; }; +\`\`\``, + }; +} + +export interface ExecDescriptionOptions { + codeModeOnly?: boolean; + searchAvailable?: boolean; + topLevelBindingNames?: ReadonlySet; + canSearchDeferredSchemas?: boolean; + hasToolCallBridge?: boolean; } export function buildExecDescription( plan: CodeModeBindingPlan, - searchAvailable = false, + options: ExecDescriptionOptions = {}, ): string { + const { + codeModeOnly = true, + topLevelBindingNames = new Set(), + canSearchDeferredSchemas = false, + hasToolCallBridge = canSearchDeferredSchemas, + } = options; + const searchAvailable = codeModeOnly && (options.searchAvailable ?? false); const visibleBindings = plan.bindings.filter( (binding) => !searchAvailable || !binding.deferred, ); @@ -209,7 +257,7 @@ export function buildExecDescription( ({ name, jsName, description, deferred }) => ({ name, jsName, - description, + ...(!deferred || !canSearchDeferredSchemas ? { description } : {}), deferred, }), ); @@ -219,17 +267,42 @@ export function buildExecDescription( `- ${omitted} is omitted because it collides with ${kept} as tools.${jsName}.`, ) .join('\n'); - const declarations = visibleBindings.map(describeCodeModeBinding).join('\n'); + const uncoveredBindings = plan.bindings.filter( + (binding) => + !topLevelBindingNames.has(binding.name) && + (!binding.deferred || !canSearchDeferredSchemas), + ); + const declarations = + plan.bindings.length === 0 + ? '' + : codeModeOnly + ? visibleBindings.map(describeCodeModeBinding).join('\n') + : [ + 'Nested tool declarations for directly exposed tools are included in their top-level tool descriptions.', + hasToolCallBridge + ? 'With both tool_search and tool_call available, tool_search returns deferred parameter schemas without changing top-level declarations. Match the returned schema name exactly to ALL_TOOLS.name, then call tools[entry.jsName] with arguments shaped by that schema. If no entry matches, do not normalize or guess a binding; use tool_call outside exec, or an available direct tool, subject to normal validation and approval.' + : uncoveredBindings.length === 0 + ? '' + : 'Other nested parameter schemas are declared below. Match tool names exactly to ALL_TOOLS.name, then call tools[entry.jsName] with arguments shaped by that schema. If no entry matches, the tool has no nested binding; do not normalize or guess one.', + uncoveredBindings.map(describeCodeModeBinding).join('\n'), + ] + .filter(Boolean) + .join('\n'); + const toolsDescription = codeModeOnly + ? 'all registered code-mode-callable tool functions permitted in this context, including deferred tools.' + : `the code-mode-callable functions listed in ALL_TOOLS. Parameter schemas are in their top-level tool descriptions or below${hasToolCallBridge ? ', or returned by tool_search' : ''}; use the exact ALL_TOOLS name-to-jsName mapping.`; return `Execute JavaScript in a fresh isolated runtime and wait for it to finish. Use async/await and call registered tools through tools.(args), using the exact JavaScript name documented for the tool. Calls use the same validation, permissions, approvals, hooks, telemetry, cancellation, concurrency, and output limits as direct tool calls. Batch independent searches and reads with await Promise.allSettled([...]) and inspect every result: emit fulfilled output with text(result.value.output) and rejected reasons with text(String(result.reason)). Safe calls run concurrently within the runtime limit; a rejected promise does not discard sibling results. Keep dependent actions, mutations, and approvals sequential. User cancellation still stops unfinished calls. Await every tool promise; unawaited calls are cancelled when the script finishes. The exec tool, direct control tools, tool_search, and tool_call are not callable through tools. ${searchAvailable ? '\nDeferred tool signatures and descriptions are omitted below. Invoke tool_search as a separate top-level tool call, outside exec; it is not a JavaScript global or a tools binding. Search by keywords to obtain the registered name, full schema, and jsName. Use select: only with a registered name, including the full MCP prefix. Read the result before writing a later exec call using tools.(args). Reuse schemas already in the current context. If a schema is missing, including after context compression, search again before constructing arguments. Search results do not change this declaration.\n' : ''} +A denied or failed nested call rejects its promise; an uncaught rejection aborts the program. Catch expected failures if execution should continue. Pass values you need to inspect or return to text(); assigning a result does not include it in the exec output. + Nested tool results stay in JavaScript values and are not automatically added to the exec response. Use text(value) to return text, image(value) or audio(value) to return media, and generatedImage(value) for the result of tools.image_gen(...). Only explicit helper calls are returned; bare return values and successful script completion produce no output. Read loaded skill instructions before taking dependent actions in a later exec call. A terminal update_goal result ends the script and prevents further tool calls. When Omni is enabled, uploaded media and its resource metadata use the native media transport. Available globals: -- tools: all registered code-mode-callable tool functions permitted in this context, including deferred tools. +- tools: ${toolsDescription} - ALL_TOOLS: frozen metadata for every function in tools. - text(value): append bounded text output. Non-string values are JSON-stringified when possible. - image(imageUrlOrItem: string | ImageContent): append an image from a base64 data URL or Qwen MCP ImageContent. To return a nested MCP image, emit available items with result.content?.forEach(image). @@ -255,10 +328,10 @@ ${collisionText ? `\nName collisions:\n${collisionText}` : ''}`; export function buildExecDeclaration( execTool: AnyDeclarativeTool, plan: CodeModeBindingPlan, - searchAvailable = false, + options: ExecDescriptionOptions = {}, ): FunctionDeclaration { return { ...execTool.schema, - description: buildExecDescription(plan, searchAvailable), + description: buildExecDescription(plan, options), }; } diff --git a/packages/core/src/tools/exec-context-budget.test.ts b/packages/core/src/tools/exec-context-budget.test.ts index 097773fe8b2..4a9b703e082 100644 --- a/packages/core/src/tools/exec-context-budget.test.ts +++ b/packages/core/src/tools/exec-context-budget.test.ts @@ -27,7 +27,7 @@ describe('exec context output budget', () => { userMemory: '', memoryFileCount: 0, approvalMode: ApprovalMode.YOLO, - codeModeOnly: true, + toolMode: 'code_mode_only', disableAllHooks: true, truncateToolOutputThreshold: 200_000, toolOutputBatchBudget: 200_000, diff --git a/packages/core/src/tools/exec-context-tools.test.ts b/packages/core/src/tools/exec-context-tools.test.ts index 02af65ae3bd..b09b9f4f47a 100644 --- a/packages/core/src/tools/exec-context-tools.test.ts +++ b/packages/core/src/tools/exec-context-tools.test.ts @@ -34,7 +34,7 @@ function setup( userMemory: '', memoryFileCount: 0, approvalMode: ApprovalMode.DEFAULT, - codeModeOnly: true, + toolMode: 'code_mode_only', }); const registry = new ToolRegistry(config); vi.spyOn(config, 'getToolRegistry').mockReturnValue(registry); diff --git a/packages/core/src/tools/tool-registry.code-mode-budget-review.test.ts b/packages/core/src/tools/tool-registry.code-mode-budget-review.test.ts new file mode 100644 index 00000000000..96b1eb29dbf --- /dev/null +++ b/packages/core/src/tools/tool-registry.code-mode-budget-review.test.ts @@ -0,0 +1,96 @@ +/** + * @license + * Copyright 2026 Qwen + * SPDX-License-Identifier: Apache-2.0 + */ + +import { describe, expect, it } from 'vitest'; +import { makeFakeConfig } from '../test-utils/config.js'; +import { MockTool } from '../test-utils/mock-tool.js'; +import { CHARS_PER_TOKEN } from '../services/tokenEstimation.js'; +import { ToolRegistry } from './tool-registry.js'; +import { ToolMode } from './code-mode.js'; +import { ToolNames } from './tool-names.js'; +import { ExecTool } from './exec.js'; + +const target = () => + new MockTool({ + name: 'deferred', + description: 'Budget-controlled deferred tool. '.repeat(80), + shouldDefer: true, + params: { + type: 'object', + properties: { path: { type: 'string' } }, + required: ['path'], + }, + }); + +function fixture(mode: ToolMode, hiddenExec = false) { + const config = makeFakeConfig({ toolMode: mode }); + const registry = new ToolRegistry(config); + registry.registerTool(new MockTool({ name: ToolNames.TOOL_SEARCH })); + registry.registerTool(new MockTool({ name: ToolNames.TOOL_CALL })); + if (mode === ToolMode.CodeMode) { + if (hiddenExec) { + registry.registerPermissionDeferredFactory( + ToolNames.EXEC, + async () => new ExecTool(config), + ); + } else { + registry.registerTool(new ExecTool(config)); + } + } + const deferred = target(); + registry.registerTool(deferred); + return { registry, deferred }; +} + +describe('Code-mode preload follows actual declaration growth', () => { + it('keeps Direct preload at the raw-schema budget', () => { + const { registry, deferred } = fixture(ToolMode.Direct); + const rawTokens = Math.ceil( + JSON.stringify(deferred.schema).length / CHARS_PER_TOKEN, + ); + expect(registry.preloadDeferredToolsWithinBudget(rawTokens - 1)).toBe(0); + expect(registry.preloadDeferredToolsWithinBudget(rawTokens)).toBe(1); + }); + + it('keeps real exec declared even from a permission-deferred factory', async () => { + const { registry, deferred } = fixture(ToolMode.CodeMode, true); + await registry.warmAll(); + const rawTokens = Math.ceil( + JSON.stringify(deferred.schema).length / CHARS_PER_TOKEN, + ); + const before = registry.getFunctionDeclarations(); + expect(registry.isPermissionDeferred(ToolNames.EXEC)).toBe(true); + expect(before.some((item) => item.name === ToolNames.EXEC)).toBe(true); + + expect(registry.preloadDeferredToolsWithinBudget(rawTokens)).toBe(0); + expect(registry.isDeferredToolRevealed(deferred.name)).toBe(false); + }); + + it('rejects a budget that fits the target declaration but not the actual prompt growth', () => { + const measuring = fixture(ToolMode.CodeMode); + const before = measuring.registry.getFunctionDeclarations(); + measuring.registry.revealDeferredTool(measuring.deferred.name); + const after = measuring.registry.getFunctionDeclarations(); + const targetDeclaration = after.find( + (item) => item.name === measuring.deferred.name, + ); + const ownSchemaTokens = Math.ceil( + JSON.stringify(targetDeclaration).length / CHARS_PER_TOKEN, + ); + const actualGrowthTokens = Math.ceil( + (JSON.stringify(after).length - JSON.stringify(before).length) / + CHARS_PER_TOKEN, + ); + expect(actualGrowthTokens).toBeGreaterThan(ownSchemaTokens); + + const { registry, deferred } = fixture(ToolMode.CodeMode); + expect( + registry.preloadDeferredToolsWithinBudget(ownSchemaTokens), + `Own schema budget ${ownSchemaTokens}, actual growth ${actualGrowthTokens}`, + ).toBe(0); + expect(registry.isDeferredToolRevealed(deferred.name)).toBe(false); + }); +}); diff --git a/packages/core/src/tools/tool-registry.test.ts b/packages/core/src/tools/tool-registry.test.ts index 4a8e6a60259..dd1fbd8a773 100644 --- a/packages/core/src/tools/tool-registry.test.ts +++ b/packages/core/src/tools/tool-registry.test.ts @@ -389,7 +389,7 @@ describe('ToolRegistry', () => { const baseUrl = 'https://images.example/v1'; const config = new Config({ ...baseConfigParams, - codeModeOnly: true, + toolMode: 'code_mode_only', experimentalZedIntegration: true, modelProvidersConfig: { openai: [ @@ -744,7 +744,7 @@ describe('ToolRegistry', () => { 'applies modelAccess=%s to CodeModeOnly bindings', (enabled) => { const registry = registryFor({ - codeModeOnly: true, + toolMode: 'code_mode_only', omniPolicyTools: { omni_compress_image: { modelAccess: { enabled } }, }, @@ -1057,6 +1057,42 @@ describe('ToolRegistry', () => { expect(toolRegistry.isDeferredToolRevealed(toolB.name)).toBe(false); }); + it('counts CodeMode declaration decoration toward the budget', () => { + const directTool = new MockTool({ + name: 'deferred', + shouldDefer: true, + params: { + type: 'object', + properties: { path: { type: 'string' } }, + required: ['path'], + }, + }); + const codeModeTool = new MockTool({ + name: 'deferred', + shouldDefer: true, + params: { + type: 'object', + properties: { path: { type: 'string' } }, + required: ['path'], + }, + }); + const directRegistry = new ToolRegistry(new Config(baseConfigParams)); + const codeModeRegistry = new ToolRegistry( + new Config({ ...baseConfigParams, toolMode: 'code_mode' }), + ); + directRegistry.registerTool(directTool); + codeModeRegistry.registerTool(new MockTool({ name: 'exec' })); + codeModeRegistry.registerTool(codeModeTool); + const rawBudget = tokensFor(directTool); + + expect(directRegistry.preloadDeferredToolsWithinBudget(rawBudget)).toBe( + 1, + ); + expect( + codeModeRegistry.preloadDeferredToolsWithinBudget(rawBudget), + ).toBe(0); + }); + it('excludes visible deferred tools from the preload budget', () => { const visibleTool = deferred('visible', { description: 'x'.repeat(CHARS_PER_TOKEN * 10), @@ -1116,7 +1152,7 @@ describe('ToolRegistry', () => { it('getDeferredToolSummary is empty in CodeModeOnly', () => { // Code Mode discovers tools without a startup catalog or Direct-mode // bridge reminders. - const registry = registryFor({ codeModeOnly: true }); + const registry = registryFor({ toolMode: 'code_mode_only' }); register(registry, deferred('deferred'), cronList()); expect(registry.getDeferredToolSummary()).toEqual([]); @@ -1293,6 +1329,28 @@ describe('ToolRegistry', () => { // `tool_call` bridge while their schemas stay out of the eager model // request (#9827). describe('permission-deferred tools (#10075)', () => { + it('keeps deferred direct-control tools out of CodeMode bindings', async () => { + const registry = new ToolRegistry( + new Config({ + ...baseConfigParams, + toolMode: ToolMode.CodeMode, + }), + ); + registry.registerTool(new MockTool({ name: 'exec' })); + registry.registerPermissionDeferredFactory( + 'send_message', + async () => new MockTool({ name: 'send_message' }), + ); + await registry.warmAll(); + + expect(registry.isDeferredAndHidden('send_message')).toBe(true); + expect( + registry + .getCodeModeBindingPlan() + .bindings.some((binding) => binding.name === 'send_message'), + ).toBe(false); + }); + const HIDDEN = 'hidden_by_allowlist'; async function addPermissionDeferred(registry = toolRegistry) { diff --git a/packages/core/src/tools/tool-registry.ts b/packages/core/src/tools/tool-registry.ts index 1d95aabf3c3..2603ea429e7 100644 --- a/packages/core/src/tools/tool-registry.ts +++ b/packages/core/src/tools/tool-registry.ts @@ -38,6 +38,7 @@ import { CHARS_PER_TOKEN } from '../services/tokenEstimation.js'; import { getCurrentAgentChat } from '../agents/runtime/agent-context.js'; import type { LlmChat } from '../core/llm-chat.js'; import { + augmentDeclarationForCodeMode, buildExecDeclaration, getToolExposure, planCodeModeBindings, @@ -277,6 +278,7 @@ export class ToolRegistry { config: Config, eventEmitter?: EventEmitter, sendSdkMcpMessage?: SendSdkMcpMessage, + readonly forSubAgent = false, ) { this.config = config; // options-bag @@ -1034,11 +1036,12 @@ export class ToolRegistry { getFunctionDeclarations(options?: { includeDeferred?: boolean; }): FunctionDeclaration[] { - if (this.config.getToolMode?.() === ToolMode.CodeModeOnly) { + const toolMode = this.config.getToolMode?.(); + if (toolMode === ToolMode.CodeModeOnly) { return this.getCodeModeFunctionDeclarations(); } const includeDeferred = options?.includeDeferred === true; - return Array.from(this.tools.values()) + const declarations = Array.from(this.tools.values()) .filter((tool) => this.isToolAvailable(tool.name)) .filter((tool) => this.isToolDeclared(tool.name)) .filter((tool) => this.isMemoryRecallToolDeclared(tool.name)) @@ -1049,8 +1052,50 @@ export class ToolRegistry { tool.alwaysLoad || !this.isDeferredAndHidden(tool.name), ) - .sort(ToolRegistry.compareToolsByDeclarationName) - .map((tool) => tool.schema); + .sort(ToolRegistry.compareToolsByDeclarationName); + if (toolMode !== ToolMode.CodeMode) { + return declarations.map((tool) => tool.schema); + } + return this.decorateCodeModeDeclarations(declarations); + } + + private decorateCodeModeDeclarations( + tools: AnyDeclarativeTool[], + allowedNames?: ReadonlySet, + canSearchDeferredSchemas = tools.some( + (tool) => tool.name === ToolNames.TOOL_SEARCH, + ) && tools.some((tool) => tool.name === ToolNames.TOOL_CALL), + ): FunctionDeclaration[] { + if (!tools.some((tool) => tool.name === ToolNames.EXEC)) { + return tools.map((tool) => tool.schema); + } + const hasToolCallBridge = + tools.some((tool) => tool.name === ToolNames.TOOL_SEARCH) && + tools.some((tool) => tool.name === ToolNames.TOOL_CALL); + const plan = this.getCodeModeBindingPlan(allowedNames); + const bindings = new Map( + plan.bindings.map((binding) => [binding.name, binding]), + ); + const topLevelBindingNames = new Set( + tools + .filter((tool) => tool.name !== ToolNames.EXEC) + .filter((tool) => bindings.has(tool.name)) + .map((tool) => tool.name), + ); + return tools.map((tool) => { + if (tool.name === ToolNames.EXEC) { + return buildExecDeclaration(tool, plan, { + codeModeOnly: false, + topLevelBindingNames, + canSearchDeferredSchemas, + hasToolCallBridge, + }); + } + const binding = bindings.get(tool.name); + return binding + ? augmentDeclarationForCodeMode(tool.schema, binding) + : tool.schema; + }); } /** @@ -1087,7 +1132,7 @@ export class ToolRegistry { .sort(ToolRegistry.compareCodeModeTools) .map((tool) => tool.name === ToolNames.EXEC - ? buildExecDeclaration(tool, plan, searchAvailable) + ? buildExecDeclaration(tool, plan, { searchAvailable }) : tool.schema, ); } @@ -1387,6 +1432,17 @@ export class ToolRegistry { preloadDeferredToolsWithinBudget(budgetTokens: number): number { const candidates: string[] = []; let totalChars = 0; + const execTool = this.tools.get(ToolNames.EXEC); + const declaredNames = + this.config.getToolMode?.() === ToolMode.CodeMode + ? new Set(this.getFunctionDeclarations().map((tool) => tool.name)) + : undefined; + const codeModePlan = declaredNames?.has(ToolNames.EXEC) + ? this.getCodeModeBindingPlan() + : undefined; + const codeModeBindings = codeModePlan + ? new Map(codeModePlan.bindings.map((binding) => [binding.name, binding])) + : undefined; for (const tool of this.tools.values()) { if (!this.isToolAvailable(tool.name)) continue; if (!this.isEffectivelyDeferred(tool) || tool.alwaysLoad) continue; @@ -1399,7 +1455,52 @@ export class ToolRegistry { if (this.permissionDeferred.has(tool.name)) continue; if (this.config.getVisibleTools().has(tool.name)) continue; candidates.push(tool.name); - totalChars += JSON.stringify(tool.schema).length; + const binding = codeModeBindings?.get(tool.name); + totalChars += JSON.stringify( + binding + ? augmentDeclarationForCodeMode(tool.schema, binding) + : tool.schema, + ).length; + } + if (codeModePlan && execTool && declaredNames) { + const candidateNames = new Set(candidates); + // Compare complete exec declarations without changing registry state. + // Treat prior reveals as hidden in the baseline so repeated preloads + // keep charging their footprint instead of ratcheting past the budget. + const execChars = (revealed: boolean): number => { + const topLevelBindingNames = new Set( + [...declaredNames].filter( + (name): name is string => + name !== undefined && !candidateNames.has(name), + ), + ); + if (revealed) { + for (const name of candidates) topLevelBindingNames.add(name); + } + const hasToolCallBridge = + topLevelBindingNames.has(ToolNames.TOOL_SEARCH) && + topLevelBindingNames.has(ToolNames.TOOL_CALL); + return JSON.stringify( + buildExecDeclaration( + execTool, + { + ...codeModePlan, + bindings: codeModePlan.bindings.map((binding) => + candidateNames.has(binding.name) + ? { ...binding, deferred: !revealed } + : binding, + ), + }, + { + codeModeOnly: false, + topLevelBindingNames, + canSearchDeferredSchemas: hasToolCallBridge, + hasToolCallBridge, + }, + ), + ).length; + }; + totalChars += Math.max(0, execChars(true) - execChars(false)); } const estimatedTokens = Math.ceil(totalChars / CHARS_PER_TOKEN); if (candidates.length === 0) { @@ -1436,12 +1537,17 @@ export class ToolRegistry { /** * Retrieves a filtered list of tool schemas based on a list of tool names. * @param toolNames - An array of tool names to include. + * @param codeModeAllowedNames - Optional nested binding allowlist when the + * direct and exec surfaces differ. * @returns An array of FunctionDeclarations for the specified tools. * @remarks Requires all tool factories to be resolved first. Call * {@link warmAll} before invoking this method, otherwise factory-registered * tools that have not yet been loaded will be silently omitted. */ - getFunctionDeclarationsFiltered(toolNames: string[]): FunctionDeclaration[] { + getFunctionDeclarationsFiltered( + toolNames: string[], + codeModeAllowedNames?: ReadonlySet, + ): FunctionDeclaration[] { if (toolNames.length === 0) return []; if (this.factories.size > 0) { debugLogger.warn( @@ -1449,9 +1555,26 @@ export class ToolRegistry { `tool factories. Call warmAll() first to avoid incomplete results.`, ); } - if (this.config.getToolMode?.() === ToolMode.CodeModeOnly) { + const toolMode = this.config.getToolMode?.(); + if (toolMode === ToolMode.CodeModeOnly) { return this.getCodeModeFunctionDeclarations(new Set(toolNames)); } + if (toolMode === ToolMode.CodeMode) { + const allowedNames = new Set(toolNames); + const tools = Array.from(this.tools.values()) + .filter( + (tool) => + allowedNames.has(tool.name) && + this.isToolAvailable(tool.name) && + this.isToolDeclared(tool.name), + ) + .sort(ToolRegistry.compareToolsByDeclarationName); + return this.decorateCodeModeDeclarations( + tools, + codeModeAllowedNames ?? allowedNames, + false, + ); + } const declarations: FunctionDeclaration[] = []; for (const name of toolNames) { const tool = this.getTool(name); diff --git a/packages/core/src/tools/tool-search.test.ts b/packages/core/src/tools/tool-search.test.ts index e6b158835e9..b821faa5a10 100644 --- a/packages/core/src/tools/tool-search.test.ts +++ b/packages/core/src/tools/tool-search.test.ts @@ -61,7 +61,7 @@ function makeConfigWithRegistry( const { withToolCall = true, params } = options; const config = new Config({ ...baseConfigParams, - codeModeOnly: options.codeModeOnly, + toolMode: options.codeModeOnly ? 'code_mode_only' : 'direct', ...params, }); const registry = new ToolRegistry(config); @@ -335,10 +335,14 @@ describe('Code Mode discovery', () => { new MockTool({ name: 'remote_fetch', shouldDefer: true }), ); const result = await new ToolSearchTool(config) - .build({ query: 'select:remote_fetch,tool_search,exec' }) + .build({ query: 'select:remote-fetch,tool_search,exec' }) .execute(new AbortController().signal); expect(String(result.llmContent)).not.toContain(''); expect(String(result.llmContent)).toContain('Not found:'); + const winner = await new ToolSearchTool(config) + .build({ query: 'select:remote_fetch' }) + .execute(new AbortController().signal); + expect(String(winner.llmContent)).toContain('"name":"remote_fetch"'); }); it('uses the scoped collision winner for both search and execution', async () => { @@ -346,24 +350,22 @@ describe('Code Mode discovery', () => { registry.registerTool( new MockTool({ name: 'remote_fetch', - description: 'scoped winner', + description: 'global winner', shouldDefer: true, }), ); const result = await runWithToolCallRuntime( { parentCallId: 'search', - allowedToolNames: ['remote_fetch'], + allowedToolNames: ['remote-fetch'], dispatch: vi.fn(), }, () => new ToolSearchTool(config) - .build({ query: 'select:remote_fetch' }) + .build({ query: 'select:remote-fetch' }) .execute(new AbortController().signal), ); - expect(String(result.llmContent)).toContain( - '"description":"scoped winner"', - ); + expect(String(result.llmContent)).toContain('"name":"remote-fetch"'); expect(String(result.llmContent)).toContain('"jsName":"remote_fetch"'); }); diff --git a/packages/core/src/utils/code-mode-allowed-names.ts b/packages/core/src/utils/code-mode-allowed-names.ts new file mode 100644 index 00000000000..b25e9de19fd --- /dev/null +++ b/packages/core/src/utils/code-mode-allowed-names.ts @@ -0,0 +1,28 @@ +/** + * @license + * Copyright 2026 Qwen + * SPDX-License-Identifier: Apache-2.0 + */ + +import { AsyncLocalStorage } from 'node:async_hooks'; + +const storage = new AsyncLocalStorage(); + +// Nested-binding reachability is per-agent: the session-wide binding plan +// can hold bindings the calling agent's exec surface does not. An `undefined` +// store leaves the ambient set untouched so tools scheduled from a parent +// exec keep its allowlist. +export function runWithCodeModeAllowedNames( + allowedNames: readonly string[] | undefined, + callback: () => T, +): T { + return allowedNames === undefined + ? callback() + : storage.run(allowedNames, callback); +} + +export function getCurrentCodeModeAllowedNames(): + | readonly string[] + | undefined { + return storage.getStore(); +} diff --git a/packages/core/src/utils/fileUtils.test.ts b/packages/core/src/utils/fileUtils.test.ts index 6cff1ea428f..777c34e8814 100644 --- a/packages/core/src/utils/fileUtils.test.ts +++ b/packages/core/src/utils/fileUtils.test.ts @@ -42,6 +42,8 @@ import { decodeBufferWithEncodingInfo } from '../services/sync-file-encoding.js' import { iconvEncode } from './iconvHelper.js'; import { LargeNonUtf8TextError } from './read-text-range.js'; import type { Config } from '../config/config.js'; +import { ToolMode } from '../tools/code-mode.js'; +import { runWithCodeModeAllowedNames } from './code-mode-allowed-names.js'; import { StandardFileSystemService } from '../services/fileSystemService.js'; import { ToolErrorType } from '../tools/tool-error.js'; import { @@ -192,6 +194,7 @@ describe('fileUtils', () => { getFunctionDeclarations: () => [ { name: 'read_file' }, { name: 'tool_search' }, + { name: 'tool_call' }, ], getDeferredToolSummary: () => [{ name: 'zoom_image' }], getCodeModeBindingPlan: () => ({ @@ -1181,9 +1184,11 @@ describe('fileUtils', () => { it.each<{ codeModeOnly: boolean; + toolMode?: ToolMode; declared: string[]; deferred: string[]; bindings: string[]; + allowedNames?: string[]; hint: string; }>([ { @@ -1195,7 +1200,7 @@ describe('fileUtils', () => { }, { codeModeOnly: false, - declared: ['read_file', 'tool_search'], + declared: ['read_file', 'tool_search', 'tool_call'], deferred: ['zoom_image'], bindings: [], hint: @@ -1223,24 +1228,106 @@ describe('fileUtils', () => { bindings: ['zoom_image'], hint: ' If details are too small, call tools.zoom_image with coordinates normalized from 0 to 1000.', }, + ...[[], ['tool_search'], ['tool_call'], ['tool_search', 'tool_call']].map( + (bridgeTools) => ({ + codeModeOnly: false, + toolMode: ToolMode.CodeMode, + declared: ['read_file', 'exec', ...bridgeTools], + deferred: ['zoom_image'], + bindings: ['zoom_image'], + hint: + bridgeTools.length === 2 + ? ' If details are too small, review zoom_image with tool_search and invoke it through tool_call, with coordinates normalized from 0 to 1000.' + : ' If details are too small, call tools.zoom_image with coordinates normalized from 0 to 1000.', + }), + ), + ...[[], ['tool_search'], ['tool_call']].map((bridgeTools) => ({ + codeModeOnly: false, + toolMode: ToolMode.CodeMode, + declared: ['read_file', ...bridgeTools], + deferred: ['zoom_image'], + bindings: ['zoom_image'], + hint: '', + })), + { + // A declared zoom_image is directly callable: the guidance must not + // route it through exec. + codeModeOnly: false, + toolMode: ToolMode.CodeMode, + declared: ['read_file', 'exec', 'zoom_image'], + deferred: [], + bindings: ['zoom_image'], + hint: ' If details are too small, call zoom_image with coordinates normalized from 0 to 1000.', + }, + { + // The session plan binds zoom_image, but the calling agent's + // narrowed plan does not: the guidance must not advertise it. + codeModeOnly: false, + toolMode: ToolMode.CodeMode, + declared: ['read_file', 'exec'], + deferred: ['zoom_image'], + bindings: ['zoom_image'], + allowedNames: ['read_file'], + hint: '', + }, + { + // Same narrowing through the bridge route: the session declares + // tool_search and tool_call, but the narrowed agent has neither + // half, so the guidance must not advertise the bridge. + codeModeOnly: false, + toolMode: ToolMode.CodeMode, + declared: ['read_file', 'tool_search', 'tool_call'], + deferred: ['zoom_image'], + bindings: ['zoom_image'], + allowedNames: ['read_file'], + hint: '', + }, + ...[ + ['read_file', 'exec', 'tool_search', 'tool_call'], + ['read_file', 'exec', 'zoom_image'], + ['read_file', 'exec'], + ].flatMap((declared) => + [[], ['read_file'], ['read_file', 'zoom_image']].map( + (allowedNames) => ({ + codeModeOnly: false, + toolMode: ToolMode.CodeMode, + declared, + deferred: ['zoom_image'], + bindings: ['zoom_image'], + allowedNames, + hint: '', + }), + ), + ), ])( 'uses only exposed tools for image guidance: $declared, code mode $codeModeOnly', - async ({ codeModeOnly, declared, deferred, bindings, hint }) => { + async ({ + codeModeOnly, + toolMode, + declared, + deferred, + bindings, + allowedNames, + hint, + }) => { await writePng(testImageFilePath, 20, 10); mockMimeGetType.mockReturnValue('image/png'); - const result = await read( - testImageFilePath, - withConfig({ + const result = await runWithCodeModeAllowedNames(allowedNames, () => + processSingleFileContent(testImageFilePath, { + ...mockConfig, getCodeModeOnly: () => codeModeOnly, + getToolMode: () => toolMode, getToolRegistry: () => ({ getFunctionDeclarations: () => declared.map((name) => ({ name })), getDeferredToolSummary: () => deferred.map((name) => ({ name })), - getCodeModeBindingPlan: () => ({ - bindings: bindings.map((name) => ({ name })), + getCodeModeBindingPlan: (allowed?: ReadonlySet) => ({ + bindings: bindings + .filter((name) => !allowed || allowed.has(name)) + .map((name) => ({ name })), collisions: [], }), }), - }), + } as unknown as Config), ); const parts = result.llmContent as Part[]; expect(parts[0]).toEqual({ diff --git a/packages/core/src/utils/fileUtils.ts b/packages/core/src/utils/fileUtils.ts index cbbaa962e11..ab482a19625 100644 --- a/packages/core/src/utils/fileUtils.ts +++ b/packages/core/src/utils/fileUtils.ts @@ -35,6 +35,7 @@ import { shouldRequirePDFPageRange, } from './pdf.js'; import { VISION_BRIDGE_MAX_IMAGES } from './vision-bridge-constants.js'; +import { getCurrentCodeModeAllowedNames } from './code-mode-allowed-names.js'; import type { VisionBridgePdfContinuation } from '../services/visionBridge/vision-bridge-service.js'; import { extensionForMimeType, @@ -1661,29 +1662,36 @@ export async function processSingleFileContent( .map((declaration) => declaration.name), ); const zoomDeclared = declaredTools.has('zoom_image'); - // Reachability, not existence (#12271): advertising a tool the - // session cannot call costs the model a turn. Checking - // `tool_search` alone is enough even though invoking a deferred - // tool also needs `tool_call` — when either bridge tool is - // unavailable, deferred tools are declared eagerly, which - // `zoomDeclared` already covers. - const zoomAvailable = codeModeOnly - ? registry - ?.getCodeModeBindingPlan() - .bindings.some((binding) => binding.name === 'zoom_image') - : zoomDeclared || - (declaredTools.has('tool_search') && - registry - ?.getDeferredToolSummary() - .some((tool) => tool.name === 'zoom_image')); + const hasToolCallBridge = + declaredTools.has('tool_search') && + declaredTools.has('tool_call'); + const useNestedZoom = + codeModeOnly || + (config.getToolMode?.() === 'code_mode' && + declaredTools.has('exec') && + !zoomDeclared && + !hasToolCallBridge); + const ambientAllowedNames = getCurrentCodeModeAllowedNames(); + const zoomAvailable = + ambientAllowedNames === undefined && + (useNestedZoom + ? registry + ?.getCodeModeBindingPlan() + .bindings.some((binding) => binding.name === 'zoom_image') + : zoomDeclared || + (hasToolCallBridge && + registry + ?.getDeferredToolSummary() + .some((tool) => tool.name === 'zoom_image'))); let zoomHint = ''; + // An agent's target allowlist does not identify its declared + // direct, bridge, or exec routes. Do not advertise session routes. if (zoomAvailable) { - // CodeModeOnly calls zoom_image through `exec`, never `tool_call`; - // if its schema is deferred, exec's own guidance points the model - // at tool_search. - const toolName = codeModeOnly ? 'tools.zoom_image' : 'zoom_image'; + const toolName = useNestedZoom + ? 'tools.zoom_image' + : 'zoom_image'; zoomHint = - codeModeOnly || zoomDeclared + useNestedZoom || zoomDeclared ? ` If details are too small, call ${toolName} with coordinates normalized from 0 to 1000.` : ' If details are too small, review zoom_image with tool_search and invoke it through tool_call, with coordinates normalized from 0 to 1000.'; } diff --git a/packages/sdk-typescript/README.md b/packages/sdk-typescript/README.md index 44580db924c..49a5bd67921 100644 --- a/packages/sdk-typescript/README.md +++ b/packages/sdk-typescript/README.md @@ -52,27 +52,27 @@ Creates a new query session with the Qwen Code. #### QueryOptions -| Option | Type | Default | Description | -| ------------------------ | -------------------------------------------------------------------- | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `cwd` | `string` | `process.cwd()` | The working directory for the query session. Determines the context in which file operations and commands are executed. | -| `model` | `string` | - | The AI model to use (e.g., `'qwen-max'`, `'qwen-plus'`, `'qwen-turbo'`). Takes precedence over `OPENAI_MODEL` and `QWEN_MODEL` environment variables. | -| `pathToQwenExecutable` | `string` | Bundled CLI | Path to the Qwen Code executable. Supports multiple formats: `'qwen'` (native binary from PATH), `'/path/to/qwen'` (explicit path), `'/path/to/cli.js'` (Node.js bundle), `'node:/path/to/cli.js'` (force Node.js runtime), `'bun:/path/to/cli.js'` (force Bun runtime). If not provided, the SDK uses the bundled CLI included with the package. | -| `permissionMode` | `'default' \| 'plan' \| 'auto-edit' \| 'auto' \| 'yolo'` | `'default'` | Permission mode controlling tool execution approval. See [Permission Modes](#permission-modes) for details. | -| `canUseTool` | `CanUseTool` | - | Custom permission handler for tool execution approval. Invoked when a tool requires confirmation. Must respond within 60 seconds or the request will be auto-denied. See [Custom Permission Handler](#custom-permission-handler). | -| `env` | `Record` | - | Environment variables to pass to the Qwen Code process. Merged with the current process environment. | -| `systemPrompt` | `string \| QuerySystemPromptPreset` | - | System prompt configuration for the main session. Use a string to fully override the built-in Qwen Code system prompt, or a preset object to keep the built-in prompt and append extra instructions. | -| `mcpServers` | `Record` | - | MCP (Model Context Protocol) servers to connect. Supports external servers (stdio/SSE/HTTP) and SDK-embedded servers. External servers are configured with transport options like `command`, `args`, `url`, `httpUrl`, etc. SDK servers use `{ type: 'sdk', name: string, instance: Server }`. | -| `abortController` | `AbortController` | - | Controller to cancel the query session. Call `abortController.abort()` to terminate the session and cleanup resources. | -| `debug` | `boolean` | `false` | Enable debug mode for verbose logging from the CLI process. | -| `maxSessionTurns` | `number` | `-1` (unlimited) | Maximum number of conversation turns before the session automatically terminates. Must be an integer. A turn consists of a user message and an assistant response. | -| `coreTools` | `string[]` | - | Uses the legacy `coreTools` / CLI `--core-tools` allowlist semantics. If specified, only matching core tools are registered for the session. This is the only allowlist-style option that restricts built-in tool registration; a whole-tool `permissions.deny` / `excludeTools` rule (and `tools.disabled` in settings.json) also removes a tool from the registry. `permissions.allow` in settings.json is pure auto-approval and never removes, demotes, or hides a tool (#10075). To keep a tool's schema out of the initial model request, use `tools.eager` in settings.json (requires restart, #9827) — `tool_search`, `tool_call`, `structured_output`, plan-mode lifecycle tools, `task_stop`, `mcp__*` and `computer_use__*` tools are exempt from that allowlist and keep their normal loading; tools demoted this way stay registered and reachable through `tool_search` + `tool_call` while both bridge tools are registered — when either is unregistered (`tools.toolSearch.enabled: false` denies both; a `tool_search` or `tool_call` deny rule, or a `tools.disabled` entry removes one) the demoted tools that remain hidden are not offered to the model and cannot be reached through the bridge for that session, and a warning is written to the CLI process's stderr (SDK forwards it only with piped stderr and effective `debug` logging; an explicit `logLevel` always wins). These bridge and warning rules apply to direct tool mode. CodeModeOnly finds deferred schemas with top-level `tool_search` and calls them through `exec`; `tools.eager` also shrinks the initial `exec` description. Without search in scope, `exec` lists each allowed tool's signature. In direct mode they stay registered, so a direct call by their own name is still evaluated and approved normally — except tools also listed in `tools.visible`, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares.; to remove a tool entirely, use a whole-tool `excludeTools` / `permissions.deny` rule — a rule with a specifier (such as `'Bash(rm *)'`) only denies matching invocations at runtime. MCP tools are exempt from deny-based removal: hide them with the per-server `excludeTools` / `tools.disabled` filters instead (deny still blocks their calls at runtime). Example: `['read_file', 'edit', 'run_shell_command']`. | -| `excludeTools` | `string[]` | - | Equivalent to `permissions.deny` in settings.json. Excluded tools return a permission error immediately. Takes highest priority over all other permission settings. Supports tool name aliases and pattern matching: tool name (`'write_file'`), shell command prefix (`'Bash(rm *)'`), or path patterns (`'Read(.env)'`, `'Edit(/src/**)'`). | -| `allowedTools` | `string[]` | - | Equivalent to `permissions.allow` in settings.json for auto-approval. Matching tools bypass `canUseTool` callback and execute automatically. Only applies when tool requires confirmation. Like `permissions.allow`, this is pure auto-approval and never affects which tools are registered or which schemas are sent (#10075). Supports same pattern matching as `excludeTools`. Example: `['Bash(git status)', 'Bash(npm test)']`. | -| `authType` | `'openai' \| 'anthropic' \| 'qwen-oauth' \| 'gemini' \| 'vertex-ai'` | - | Authentication type for the AI service. When provided, the SDK forwards it to the CLI as `--auth-type`. | -| `agents` | `SubagentConfig[]` | - | Configuration for subagents that can be invoked during the session. Subagents are specialized AI agents for specific tasks or domains. | -| `includePartialMessages` | `boolean` | `false` | When `true`, the SDK emits incomplete messages as they are being generated, allowing real-time streaming of the AI's response. | -| `resume` | `string` | - | Resume a previous session by providing its session ID. Equivalent to CLI's `--resume` flag. | -| `sessionId` | `string` | - | Specify a session ID for the new session. Ensures SDK and CLI use the same ID without resuming history. Equivalent to CLI's `--session-id` flag. | +| Option | Type | Default | Description | +| ------------------------ | -------------------------------------------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `cwd` | `string` | `process.cwd()` | The working directory for the query session. Determines the context in which file operations and commands are executed. | +| `model` | `string` | - | The AI model to use (e.g., `'qwen-max'`, `'qwen-plus'`, `'qwen-turbo'`). Takes precedence over `OPENAI_MODEL` and `QWEN_MODEL` environment variables. | +| `pathToQwenExecutable` | `string` | Bundled CLI | Path to the Qwen Code executable. Supports multiple formats: `'qwen'` (native binary from PATH), `'/path/to/qwen'` (explicit path), `'/path/to/cli.js'` (Node.js bundle), `'node:/path/to/cli.js'` (force Node.js runtime), `'bun:/path/to/cli.js'` (force Bun runtime). If not provided, the SDK uses the bundled CLI included with the package. | +| `permissionMode` | `'default' \| 'plan' \| 'auto-edit' \| 'auto' \| 'yolo'` | `'default'` | Permission mode controlling tool execution approval. See [Permission Modes](#permission-modes) for details. | +| `canUseTool` | `CanUseTool` | - | Custom permission handler for tool execution approval. Invoked when a tool requires confirmation. Must respond within 60 seconds or the request will be auto-denied. See [Custom Permission Handler](#custom-permission-handler). | +| `env` | `Record` | - | Environment variables to pass to the Qwen Code process. Merged with the current process environment. | +| `systemPrompt` | `string \| QuerySystemPromptPreset` | - | System prompt configuration for the main session. Use a string to fully override the built-in Qwen Code system prompt, or a preset object to keep the built-in prompt and append extra instructions. | +| `mcpServers` | `Record` | - | MCP (Model Context Protocol) servers to connect. Supports external servers (stdio/SSE/HTTP) and SDK-embedded servers. External servers are configured with transport options like `command`, `args`, `url`, `httpUrl`, etc. SDK servers use `{ type: 'sdk', name: string, instance: Server }`. | +| `abortController` | `AbortController` | - | Controller to cancel the query session. Call `abortController.abort()` to terminate the session and cleanup resources. | +| `debug` | `boolean` | `false` | Enable debug mode for verbose logging from the CLI process. | +| `maxSessionTurns` | `number` | `-1` (unlimited) | Maximum number of conversation turns before the session automatically terminates. Must be an integer. A turn consists of a user message and an assistant response. | +| `coreTools` | `string[]` | - | Uses the legacy `coreTools` / CLI `--core-tools` allowlist semantics. If specified, only matching core tools are registered for the session. This is the only allowlist-style option that restricts built-in tool registration; a whole-tool `permissions.deny` / `excludeTools` rule (and `tools.disabled` in settings.json) also removes a tool from the registry. `permissions.allow` in settings.json is pure auto-approval and never removes, demotes, or hides a tool (#10075). To keep a tool's schema out of the initial model request, use `tools.eager` in settings.json (requires restart, #9827) — `tool_search`, `tool_call`, `structured_output`, plan-mode lifecycle tools, `task_stop`, `mcp__*` and `computer_use__*` tools are exempt from that allowlist and keep their normal loading; tools demoted this way stay registered and reachable through `tool_search` + `tool_call` while both bridge tools are registered — when either is unregistered (`tools.toolSearch.enabled: false` denies both; a `tool_search` or `tool_call` deny rule, or a `tools.disabled` entry removes one) the demoted tools that remain hidden are absent from top-level declarations and cannot be reached through the bridge for that session, and a warning is written to the CLI process's stderr (SDK forwards it only with piped stderr and effective `debug` logging; an explicit `logLevel` always wins). These bridge and warning rules apply to direct and hybrid code modes. On the session surface in hybrid mode, while `exec` itself is registered (container and SSH execution warn and fall back to direct tools without it), exec retains callable nested bindings; their schemas are included in exec when either bridge tool is unavailable. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by tools.eager from nested bindings. In both code modes, agent allowlists that do not grant `exec` narrow nested bindings. Inheriting or explicitly granting `exec` keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns. In direct and hybrid modes they stay registered, so a direct call by their own name is still evaluated and approved normally — except tools also listed in `tools.visible`, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares.; to remove a tool entirely, use a whole-tool `excludeTools` / `permissions.deny` rule — a rule with a specifier (such as `'Bash(rm *)'`) only denies matching invocations at runtime. MCP tools are exempt from deny-based removal: hide them with the per-server `excludeTools` / `tools.disabled` filters instead (deny still blocks their calls at runtime). Example: `['read_file', 'edit', 'run_shell_command']`. | +| `excludeTools` | `string[]` | - | Equivalent to `permissions.deny` in settings.json. Excluded tools return a permission error immediately. Takes highest priority over all other permission settings. Supports tool name aliases and pattern matching: tool name (`'write_file'`), shell command prefix (`'Bash(rm *)'`), or path patterns (`'Read(.env)'`, `'Edit(/src/**)'`). | +| `allowedTools` | `string[]` | - | Equivalent to `permissions.allow` in settings.json for auto-approval. Matching tools bypass `canUseTool` callback and execute automatically. Only applies when tool requires confirmation. Like `permissions.allow`, this is pure auto-approval and never affects which tools are registered or which schemas are sent (#10075). Supports same pattern matching as `excludeTools`. Example: `['Bash(git status)', 'Bash(npm test)']`. | +| `authType` | `'openai' \| 'anthropic' \| 'qwen-oauth' \| 'gemini' \| 'vertex-ai'` | - | Authentication type for the AI service. When provided, the SDK forwards it to the CLI as `--auth-type`. | +| `agents` | `SubagentConfig[]` | - | Configuration for subagents that can be invoked during the session. Subagents are specialized AI agents for specific tasks or domains. | +| `includePartialMessages` | `boolean` | `false` | When `true`, the SDK emits incomplete messages as they are being generated, allowing real-time streaming of the AI's response. | +| `resume` | `string` | - | Resume a previous session by providing its session ID. Equivalent to CLI's `--resume` flag. | +| `sessionId` | `string` | - | Specify a session ID for the new session. Ensures SDK and CLI use the same ID without resuming history. Equivalent to CLI's `--session-id` flag. | > [!tip] > If you need to configure `coreTools`, `excludeTools`, or `allowedTools`, it is **strongly recommended** to read the [permissions configuration documentation](../../docs/users/configuration/settings.md#permissions) first, especially the **Tool name aliases** and **Rule syntax examples** sections. Rule patterns such as `Bash(git *)`, `Read(.env)`, and `Edit(/src/**)` apply to `excludeTools` and `allowedTools`; `coreTools` accepts aliases but strips invocation specifiers. diff --git a/packages/sdk-typescript/src/types/types.ts b/packages/sdk-typescript/src/types/types.ts index 493b76a6357..b06459d34bf 100644 --- a/packages/sdk-typescript/src/types/types.ts +++ b/packages/sdk-typescript/src/types/types.ts @@ -418,7 +418,7 @@ export interface QueryOptions { * registered; when either is unregistered (`tools.toolSearch.enabled: false` * denies both; a `tool_search` or `tool_call` deny rule, or a * `tools.disabled` entry removes one) the - * demoted tools that remain hidden are not offered to the model and cannot + * demoted tools that remain hidden are absent from top-level declarations and cannot * be reached through the bridge for that session, and a warning is * written to the CLI process's stderr. The SDK forwards it only when * stderr is piped (`debug: true` or a `stderr` handler) and the effective @@ -432,11 +432,16 @@ export interface QueryOptions { * direct call to a still-hidden demoted tool, which any tool-set * refresh (resume, MCP discovery, the first plan-mode entry in a * session, a subagent definition change) re-declares. - * These bridge and warning rules describe direct tool mode. CodeModeOnly - * discovers deferred schemas through top-level tool_search and invokes them - * through exec. It skips deferred preload and startup catalogs; tools.eager - * also reduces the initial exec description. When search is unavailable in - * the current scope, exec includes all allowed tool signatures. + * These bridge and warning rules apply to direct and hybrid code modes. + * On the session surface in hybrid mode, while `exec` is registered (container or SSH execution + * warns and falls back to direct tools without it), exec retains callable + * nested bindings; their schemas are included in exec when either bridge + * tool is unavailable. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by tools.eager + * from nested bindings. In both code modes, agent allowlists that do not grant `exec` narrow nested bindings. + * Inheriting or explicitly granting `exec` keeps all otherwise admitted + * ordinary code-mode-callable bindings. An execution allowlist that + * mentions any MCP tool additionally restricts MCP bindings to matching + * exact names or server patterns. * Tools already deferred by default remain * on demand even when listed; `tools.visible` surfaces one at startup. The * allowlist does not affect MCP tools, the `--json-schema` @@ -490,7 +495,7 @@ export interface QueryOptions { * when either is unregistered (`tools.toolSearch.enabled: false` denies * both; a `tool_search` or `tool_call` deny rule, or a * `tools.disabled` entry removes one) the - * demoted tools that remain hidden are not offered to the model and cannot + * demoted tools that remain hidden are absent from top-level declarations and cannot * be reached through the bridge for that session, and a warning is * written to the CLI process's stderr. The SDK forwards it only when * stderr is piped (`debug: true` or a `stderr` handler) and the effective @@ -504,12 +509,21 @@ export interface QueryOptions { * contains a direct call to a still-hidden demoted tool, which any * tool-set refresh (resume, MCP discovery, the first plan-mode * entry in a session, a subagent definition change) re-declares (#9827). - * These bridge and warning rules describe direct tool mode. CodeModeOnly - * discovers deferred schemas through top-level tool_search and invokes - * them through exec. It skips deferred preload and startup catalogs; - * tools.eager also reduces the initial exec description. When search is - * unavailable in the current scope, exec includes all allowed tool - * signatures. + * These bridge and warning rules apply to direct and hybrid code modes. + * On the session surface in hybrid mode, while `exec` is registered (container or SSH + * execution warns and falls back to direct tools without it), exec + * retains callable nested bindings; their schemas are included in + * exec when either bridge tool is unavailable. CodeModeOnly discovers deferred schemas through + * top-level tool_search and invokes them through exec. It skips deferred + * preload and startup catalogs; tools.eager reduces the initial exec + * description. When search is unavailable in the current scope, exec + * includes all allowed signatures. + * In Hybrid mode, AgentCore excludes tools still hidden by tools.eager + * from nested bindings. In both code modes, agent allowlists that do not grant `exec` narrow nested bindings. + * Inheriting or explicitly granting `exec` keeps all otherwise admitted + * ordinary code-mode-callable bindings. An execution allowlist that + * mentions any MCP tool additionally restricts MCP bindings to matching + * exact names or server patterns. * * **Pattern matching:** * - Tool name: `'write_file'` diff --git a/packages/vscode-ide-companion/schemas/settings.schema.json b/packages/vscode-ide-companion/schemas/settings.schema.json index 661f8e787ac..fccccd5c236 100644 --- a/packages/vscode-ide-companion/schemas/settings.schema.json +++ b/packages/vscode-ide-companion/schemas/settings.schema.json @@ -1421,13 +1421,17 @@ }, "description": "Linux tool execution confinement. Operator scopes only; workspace settings cannot override it. Model/auth/session traffic stays on the host." }, - "codeModeOnly": { - "description": "Expose ordinary tools through the isolated exec JavaScript tool. Load deferred descriptions and schemas on demand with tool_search; if search is unavailable, include all allowed signatures in exec. Direct control tools remain available. Ignored in safe and bare modes.", - "type": "boolean", - "default": false + "mode": { + "description": "Choose how tools are exposed to the model. Direct uses ordinary tool calls; Code Mode also exposes the isolated exec JavaScript tool; Code Mode Only exposes ordinary tools only through exec. Safe and bare modes always use Direct. Container execution warns and uses direct tools for Code Mode, and rejects Code Mode Only. SSH workspaces warn and use Direct for either code mode. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by tools.eager from nested bindings. In both code modes, agent allowlists that do not grant exec narrow nested bindings. Inheriting or explicitly granting exec keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns. Options: direct, code_mode, code_mode_only", + "enum": [ + "direct", + "code_mode", + "code_mode_only" + ], + "default": "direct" }, "freeform": { - "description": "Use raw text input for the Code Mode exec tool on OpenAI Responses models. Effective only when tools.codeModeOnly is true and the selected model uses wireApi \"responses\". Enable only for endpoints that support Responses Custom Tools.", + "description": "Use raw text input for the Code Mode exec tool on OpenAI Responses models. Effective only when tools.mode is \"code_mode_only\" and the selected model uses wireApi \"responses\". Enable only for endpoints that support Responses Custom Tools.", "type": "boolean", "default": false }, @@ -1491,7 +1495,7 @@ "minimum": 0, "maximum": 100, "default": 0, - "description": "Context-window percentage used as the session-start budget for preloading ordinary deferred tools (bundled built-ins and MCP alike). Defaults to 0, which performs no threshold-based preload; ordinary deferred tools normally stay behind the stable ToolSearch + ToolCall bridge, at the cost of one tool_search round trip before first use. Raise it to N so that, when every eligible deferred schema fits within N% of the context window, all are declared upfront for direct calls with no bridge round trip; otherwise they stay behind the bridge while both bridge tools are registered. Tools demoted by tools.eager are excluded from this preload and stay reachable on demand through that bridge while it is registered; when either bridge tool is unregistered (tools.toolSearch.enabled false denies both; a tool_search or tool_call deny rule removes one) the demoted tools that remain hidden are not offered to the model and cannot be reached through the bridge for that session, and a warning is logged; these bridge and warning rules apply to direct tool mode. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager also reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed tool signatures. In direct mode they stay registered, so a direct call by their own name is still evaluated and approved normally. Separate paths can still declare deferred tools at 0: tools.visible; the live-history compatibility scan on every tool-set refresh (including resume, MCP discovery, first plan-mode entry, and subagent definition changes); the incomplete-bridge eager fallback; and daemon ACP late registration, which explicitly reveals and pins create_sub_session." + "description": "Context-window percentage used as the session-start budget for preloading ordinary deferred tools (bundled built-ins and MCP alike). Defaults to 0, which performs no threshold-based preload; ordinary deferred tools normally stay behind the stable ToolSearch + ToolCall bridge, at the cost of one tool_search round trip before first use. Raise it to N so that, when every eligible deferred schema fits within N% of the context window, all are declared upfront for direct calls with no bridge round trip; otherwise they stay behind the bridge while both bridge tools are registered. Tools demoted by tools.eager are excluded from this preload and stay reachable on demand through that bridge while it is registered; when either bridge tool is unregistered (tools.toolSearch.enabled false denies both; a tool_search or tool_call deny rule removes one) the demoted tools that remain hidden are absent from top-level declarations and cannot be reached through the bridge for that session, and a warning is logged; these bridge and warning rules apply to direct and hybrid tool modes. In Code Mode, withheld tools with actual exec bindings also remain callable through exec; the warning names that subset. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by tools.eager from nested bindings. In both code modes, agent allowlists that do not grant exec narrow nested bindings. Inheriting or explicitly granting exec keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns. In direct and hybrid modes they stay registered, so a direct call by their own name is still evaluated and approved normally. Separate paths can still declare deferred tools at 0: tools.visible; the live-history compatibility scan on every tool-set refresh (including resume, MCP discovery, first plan-mode entry, and subagent definition changes); the incomplete-bridge eager fallback; and daemon ACP late registration, which explicitly reveals and pins create_sub_session." } } }, @@ -1577,14 +1581,14 @@ } }, "visible": { - "description": "Deferred tool names made visible at startup without requiring the ToolSearch + ToolCall bridge. Listed tools appear alongside core tools in the initial session.", + "description": "Deferred tool names made visible at startup without requiring the ToolSearch + ToolCall bridge. Listed tools appear alongside core tools in the initial session. In Code Mode Only, listed tools have signatures in the initial exec description; other deferred schemas are discovered through top-level tool_search.", "type": "array", "items": { "type": "string" } }, "eager": { - "description": "Allowlist of eager-by-default built-in tool names whose schemas remain eligible for the initial model request. Unlisted non-exempt tools are deferred but stay registered, listed in /tools, and reachable through the tool_search + tool_call bridge. Tools already deferred by default stay on demand even when listed; use tools.visible to surface one at startup. tool_search, tool_call, structured_output, plan-mode lifecycle tools, task_stop, MCP tools, and computer_use__* tools are unaffected. An explicitly empty list ([]) defers every non-exempt eager-by-default tool; omit the setting for no restriction. Pairs with the ToolSearch + ToolCall bridge: when either half is not registered — tools.toolSearch.enabled false (which denies both), a tool_search or tool_call deny rule, or a tools.disabled entry — the allowlist still withholds the schemas, but nothing can load them back, so the demoted tools that remain hidden are not offered to the model and cannot be reached through the bridge for that session, and a warning is logged; these bridge and warning rules apply to direct tool mode. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager also reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed tool signatures. In direct mode they stay registered, so a direct call by their own name is still evaluated and approved normally — except tools also listed in tools.visible, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares. Differs from tools.disabled, which removes tools entirely, and from permissions.allow, which only auto-approves calls.", + "description": "Allowlist of eager-by-default built-in tool names whose schemas remain eligible for the initial model request. Unlisted non-exempt tools are deferred but stay registered, listed in /tools, and reachable through the tool_search + tool_call bridge. Tools already deferred by default stay on demand even when listed; use tools.visible to surface one at startup. tool_search, tool_call, structured_output, plan-mode lifecycle tools, task_stop, MCP tools, and computer_use__* tools are unaffected. An explicitly empty list ([]) defers every non-exempt eager-by-default tool; omit the setting for no restriction. Pairs with the ToolSearch + ToolCall bridge: when either half is not registered — tools.toolSearch.enabled false (which denies both), a tool_search or tool_call deny rule, or a tools.disabled entry — the allowlist still withholds the schemas, but nothing can load them back, so the demoted tools that remain hidden are absent from top-level declarations and cannot be reached through the bridge for that session, and a warning is logged; these bridge and warning rules apply to direct and hybrid tool modes. In Code Mode, withheld tools with actual exec bindings also remain callable through exec; the warning names that subset. CodeModeOnly discovers deferred schemas through top-level tool_search and invokes them through exec. It skips deferred preload and startup catalogs; tools.eager reduces the initial exec description. When search is unavailable in the current scope, exec includes all allowed signatures. In Hybrid mode, AgentCore excludes tools still hidden by tools.eager from nested bindings. In both code modes, agent allowlists that do not grant exec narrow nested bindings. Inheriting or explicitly granting exec keeps all otherwise admitted ordinary code-mode-callable bindings. An execution allowlist that mentions any MCP tool additionally restricts MCP bindings to matching exact names or server patterns. In direct and hybrid modes they stay registered, so a direct call by their own name is still evaluated and approved normally — except tools also listed in tools.visible, which are declared upfront, and sessions whose live history contains a direct call to a still-hidden demoted tool, which any tool-set refresh (resume, MCP discovery, the first plan-mode entry in a session, a subagent definition change) re-declares. Differs from tools.disabled, which removes tools entirely, and from permissions.allow, which only auto-approves calls.", "type": "array", "items": { "type": "string" diff --git a/packages/web-shell/client/settings.test.ts b/packages/web-shell/client/settings.test.ts index fce17350c19..a19f030f01d 100644 --- a/packages/web-shell/client/settings.test.ts +++ b/packages/web-shell/client/settings.test.ts @@ -43,6 +43,14 @@ describe('settings presentation aliases', () => { ).toBe(false); expect(WEB_SHELL_SETTING_ITEM_IDS).toContain('setting:workflow-name-only'); }); + it('keeps the stable code-mode alias for the tool mode row', () => { + expect( + isSettingVisible('tools.mode', { + excludeItems: ['setting:code-mode-only'], + }), + ).toBe(false); + expect(WEB_SHELL_SETTING_ITEM_IDS).toContain('setting:code-mode-only'); + }); it('matches published builtin ids by direct membership', () => { expect( isItemVisible('builtin:model-management', { @@ -241,7 +249,7 @@ describe('settings presentation aliases', () => { 'setting:respect-git-ignore': 'context.fileFiltering.respectGitIgnore', 'setting:respect-qwen-ignore': 'context.fileFiltering.respectQwenIgnore', 'setting:fuzzy-file-search': 'context.fileFiltering.enableFuzzySearch', - 'setting:code-mode-only': 'tools.codeModeOnly', + 'setting:code-mode-only': 'tools.mode', 'setting:web-search': 'tools.webSearch.enabled', 'setting:web-search-model': 'tools.webSearch.model', 'setting:web-extractor': 'tools.webSearch.webExtractor', diff --git a/packages/web-shell/client/settings.ts b/packages/web-shell/client/settings.ts index 3222fd743bc..0484b88f038 100644 --- a/packages/web-shell/client/settings.ts +++ b/packages/web-shell/client/settings.ts @@ -33,7 +33,7 @@ const SETTING_KEYS = { 'setting:respect-git-ignore': 'context.fileFiltering.respectGitIgnore', 'setting:respect-qwen-ignore': 'context.fileFiltering.respectQwenIgnore', 'setting:fuzzy-file-search': 'context.fileFiltering.enableFuzzySearch', - 'setting:code-mode-only': 'tools.codeModeOnly', + 'setting:code-mode-only': 'tools.mode', 'setting:web-search': 'tools.webSearch.enabled', 'setting:web-search-model': 'tools.webSearch.model', 'setting:web-extractor': 'tools.webSearch.webExtractor', diff --git a/packages/web-shell/client/settings/messages.ts b/packages/web-shell/client/settings/messages.ts index 1c449101bd8..c474f4065c5 100644 --- a/packages/web-shell/client/settings/messages.ts +++ b/packages/web-shell/client/settings/messages.ts @@ -347,9 +347,9 @@ export const SETTINGS_MESSAGES_ZH: Record = { 'settings.label.tools.listDirectory.enabled': '启用 ListDirectory', 'settings.description.tools.listDirectory.enabled': '启用内置 list_directory 工具。默认关闭;当它被显式列入 coreTools 白名单(--core-tools / tools.core)时会自动启用。', - 'settings.label.tools.codeModeOnly': '仅代码模式(实验性)', - 'settings.description.tools.codeModeOnly': - '普通工具只通过隔离的 exec JavaScript 工具暴露给模型。直接控制类工具仍然可用。在 safe 和 bare 模式下忽略。', + 'settings.label.tools.mode': '工具模式(实验性)', + 'settings.description.tools.mode': + '选择工具向模型暴露的方式。Direct 使用普通工具调用;Code Mode 额外提供隔离的 exec JavaScript 工具;Code Mode Only 仅通过 exec 暴露普通工具。safe 和 bare 模式始终使用 Direct。容器执行时,Code Mode 会警告并使用直接工具,Code Mode Only 则被拒绝。SSH 工作区会警告并将两种代码模式回退为 Direct。Code Mode Only 通过顶层 tool_search 按需发现 schema,再通过 exec 调用;跳过预算预加载与启动工具清单。tools.eager 会减少初始 exec 声明,当前范围无法搜索时则保留所有允许工具的签名。Hybrid 模式下,AgentCore 不会将仍被 tools.eager 隐藏的工具加入嵌套绑定;未授予 exec 的智能体白名单会收窄嵌套绑定。继承或显式授予 exec 会保留其他规则允许的所有普通代码模式工具绑定。执行白名单只要提及任一 MCP 工具,就会进一步将 MCP 绑定限制为匹配的精确名称或服务器模式。', 'settings.label.tools.todoWrite.enabled': '启用 Todo Write', 'settings.description.tools.todoWrite.enabled': '启用内置 todo_write 工具及其系统提示词引导。',