Skip to content

feat(workflows): let the model invoke extension workflows that declare whenToUse - #11957

Merged
qqqys merged 1 commit into
QwenLM:mainfrom
qqqys:feat/workflow-model-invocable
Sep 15, 2026
Merged

qqqys merged 1 commit into
QwenLM:mainfrom
qqqys:feat/workflow-model-invocable

Conversation

@qqqys

@qqqys qqqys commented Sep 15, 2026

Copy link
Copy Markdown
Collaborator

What this PR does

An extension workflow whose script declares meta.whenToUse is now listed for the model. Its command appears in the model's skill listing with the description and that condition, so the model can start the workflow when a request matches, through the Skill tool and then Workflow({ name }). Every run still goes through the workflow approval. A workflow without whenToUse stays out of the model's listing, and project and user workflows are unchanged.

whenToUse is carried on ExtensionWorkflowDefinition and SavedWorkflowEntry, trimmed and shortened to 500 characters like the description; a blank value counts as absent.

Outside the interactive UI a {type:'tool'} command return cannot run. That covers the model invoking the command through the Skill tool, headless qwen -p, and ACP. There, a saved-workflow command now expands to a prompt: Run the "<name>" workflow., the description, the condition, and Invoke: Workflow({ name: "<name>", args: … }). The prompt uses the qualified workflow name, so a command renamed on a collision with its extension's skill still runs the right workflow. Typing /<name> in the interactive UI still dispatches the tool with the script path, unchanged. Because every mode can now run the command, it no longer restricts itself to interactive mode.

The Workflow tool's opt-in rule now reads: "A skill or slash command that ran — invoked by the user, or by you through the Skill tool — instructs you to use this tool."

Why it's needed

An extension that ships a workflow had no way to let the model choose it. The command was never model-invocable, and the three model-invocable command executors accept only submit_prompt, so even a listed command would have come back as "not found". The opt-in rule also counted only commands the user invoked. The only way to get the model to run an extension workflow was an instruction in QWEN.md telling it to, which the opt-in rule does not accept and which fires on every matching request regardless of what the extension author intended.

Claude Code wraps every workflow as a prompt command with its whenToUse, expanding to "Run the … workflow … Invoke: Workflow({name})". This follows the same shape, but lists only extension workflows whose author wrote the condition: the author says when it applies, the user consented to the extension, workflows are enabled, and the run is approved.

Reviewer Test Plan

How to verify

Enable workflows, install or link an extension with workflows/analysis.js declaring whenToUse, and a second script without it. Start a session and ask a question that matches the condition. The model should call the Skill tool with the workflow's name, receive the "Run the … workflow" text, and call Workflow({ name }), which opens the usual approval with Saved workflow: <name>. The workflow without whenToUse does not appear in the model's skill listing.

Type /dataworks:analysis in the interactive UI: it starts the workflow directly, as before. Run qwen -p '/dataworks:analysis': the command is no longer rejected as unsupported, and the model runs the workflow by name.

Unit tests:

cd packages/core && npx vitest run src/agents/runtime/workflow-extension.test.ts src/agents/runtime/workflow-saved.test.ts src/tools/workflow/workflow.test.ts src/tools/workflow/workflow-description.test.ts src/skills/bundled/workflow-authoring/SKILL.test.ts
cd packages/cli && npx vitest run src/services/saved-workflow-loader.test.ts src/services/CommandService.test.ts src/services/commandMetadata.test.ts src/nonInteractiveCliCommands.test.ts

330 core and 137 cli tests pass locally, as do the core type check and ESLint on the changed files. The cli type check is left to CI: locally it reads the other workspace packages' built output, which predates this change.

Evidence (Before & After)

Output from the real modules on this branch: an extension discovered from disk with loadExtensionWorkflows, adapted by SavedWorkflowLoader, registered in CommandService, and rendered with the skill-listing renderer. Before this change the listing had no entry for either workflow, the command action returned the tool dispatch in every context (which the executors turn into "not found" and headless turns into "Tool execution from slash commands is not supported in non-interactive mode."), and both commands were listed for the interactive mode only.

--- <available_skills> entries from the model-invocable commands ---
<skill>
<name>
dataworks:analysis
</name>
<description>
Answers a data question end to end — When the user asks a question that needs querying, analysing and charting warehouse data
</description>
</skill>

--- Skill({ skill: "dataworks:analysis", args: '{"table":"orders"}' }) through the executor (non_interactive context) ---
Run the "dataworks:analysis" workflow.

Answers a data question end to end

When the user asks a question that needs querying, analysing and charting warehouse data

Invoke: Workflow({ name: "dataworks:analysis", args: {"table":"orders"} })

--- /dataworks:analysis typed in the interactive UI ---
{"type":"tool","toolName":"workflow","toolArgs":{"scriptPath":"/tmp/ev/ext/workflows/analysis.js"}}

--- commands per mode ---
interactive: dataworks:analysis, dataworks:export
non_interactive: dataworks:analysis, dataworks:export
acp: dataworks:analysis, dataworks:export

--- Workflow tool opt-in rule, third bullet ---
- A skill or slash command that ran — invoked by the user, or by you through the Skill tool — instructs you to use this tool.

Tested on

OS Status
🍏 macOS ⚠️
🪟 Windows ⚠️
🐧 Linux ✅

Environment (optional)

Unit tests, the core type check, and a throwaway vitest file driving the real modules. No interactive TUI session was run.

Risk & Scope

  • Main risk or tradeoff: the model can start a third-party extension's multi-agent run without the user typing the command. That needs all of: the author declared whenToUse, the user consented to the extension, tools.workflowsEnabled is on (off by default), and the user approves the run. whenToUse lives in the script, so changing it changes the content digest and the next update asks for consent again. Headless and ACP /<name> now go through one model turn instead of being refused, which matches how skills and file commands behave there.
  • Not validated: an interactive session with a real model deciding to invoke the workflow; the unit tests and the evidence cover the listing, the expansion and the dispatch.
  • Out of scope: model invocation of project and user workflows (their discovery does not read the script, so there is no whenToUse to go on); a disableModelInvocation meta field; switching the interactive /<name> path from scriptPath to name, which would stop existing Workflow(scriptPath:…) rules from matching it.

Linked Issues / Bugs

Part of #11013


中文说明

这个 PR 做了什么

脚本 meta 声明了 whenToUse 的扩展 workflow 现在会列给模型。它的命令带着描述和适用条件出现在模型的技能列表里,请求匹配时模型可以经 Skill 工具、再调用 Workflow({ name }) 发起运行。每次运行仍走 workflow 审批。没写 whenToUse 的 workflow 不进模型的列表,project 和 user workflow 不变。

ExtensionWorkflowDefinition 与 SavedWorkflowEntry 带上 whenToUse,去掉首尾空白,并与描述一样截到 500 字符;空白值视为未声明。

在交互界面之外,命令返回的 {type:'tool'} 无法执行,包括模型经 Skill 工具调起命令、headless qwen -p 和 ACP。这些场景下,已保存 workflow 的命令现在展开为一段提示:Run the "<name>" workflow.、描述、适用条件,以及 Invoke: Workflow({ name: "<name>", args: … })。提示里用的是限定名,所以命令和同扩展的同名 skill 撞名被改名后,仍然运行正确的 workflow。在交互界面里敲 /<name> 仍按脚本路径直接调度工具,行为不变。既然所有模式都能运行,命令也不再只声明交互模式。

Workflow 工具显式请求规则的第 3 条改为:"A skill or slash command that ran — invoked by the user, or by you through the Skill tool — instructs you to use this tool."

为什么需要

扩展随附的 workflow 没有办法让模型自行选用。命令从来不是模型可调用的;三处模型可调用命令执行器只接受 submit_prompt,即便列出来也会返回"找不到";规则也只承认用户调起的命令。想让模型跑扩展 workflow,只能在 QWEN.md 里要求它这么做,而这既不被显式请求规则承认,也不管扩展作者的意图,匹配就触发。

Claude Code 把每个 workflow 包装成带 whenToUse 的 prompt 命令,展开为 "Run the … workflow … Invoke: Workflow({name})"。本 PR 采用相同形态,但只列出作者写了适用条件的扩展 workflow:作者说明何时适用、用户同意安装扩展、workflows 已开启、运行经过审批。

验证方式

开启 workflows,安装或链接一个扩展:workflows/analysis.js 声明 whenToUse,另一个脚本不声明。开一个会话,问一个符合条件的问题。模型应以 workflow 名调用 Skill 工具,拿到 "Run the … workflow" 文本,再调用 Workflow({ name }),弹出常规审批,标题为 Saved workflow: <name>。没写 whenToUse 的 workflow 不出现在模型的技能列表里。

在交互界面敲 /dataworks:analysis:和之前一样直接启动。运行 qwen -p '/dataworks:analysis':命令不再被判为不支持,而是由模型按名运行 workflow。

单元测试命令见英文部分。本地 core 330 个、cli 137 个测试通过,core 类型检查与改动文件的 ESLint 通过。cli 类型检查交给 CI:本地它读取其他 workspace 包的旧构建产物。

证据

见英文部分代码块:用本分支真实模块从磁盘发现扩展、经 SavedWorkflowLoader 适配、注册到 CommandService,并用技能列表渲染器输出。改动之前,列表里两个 workflow 都没有条目;命令在任何上下文都返回工具调度,执行器会把它变成"找不到",headless 会报 "Tool execution from slash commands is not supported in non-interactive mode.";两个命令都只在交互模式下列出。

测试平台

Linux ✅;macOS ⚠️;Windows ⚠️。

风险与范围

  • 主要风险/取舍:模型可以在用户没敲命令的情况下发起第三方扩展的多 agent 运行。前提是同时满足:作者声明了 whenToUse、用户同意安装扩展、tools.workflowsEnabled 已开启(默认关闭)、用户批准这次运行。whenToUse 写在脚本里,改动它会改变内容摘要,下次更新会重新征求同意。headless 与 ACP 的 /<name> 现在经模型一轮运行,而不是被拒绝,与 skill 和文件命令在这些场景的行为一致。
  • 未验证:真实模型在交互会话中自行决定调用 workflow;单元测试与证据覆盖了列表、展开和调度。
  • 范围之外:project 和 user workflow 的模型可调用(其发现阶段不读脚本,没有 whenToUse 可依据);disableModelInvocation meta 字段;把交互式 /<name> 从 scriptPath 改为 name,那会让已有的 Workflow(scriptPath:…) 规则不再匹配。

关联 Issue

Part of #11013

…e whenToUse

An extension workflow whose script declares `meta.whenToUse` is now listed
for the model, with its description and that condition, so the model can
start it when a request matches. Every run still goes through the workflow
approval; a workflow without `whenToUse` stays out of the model's listing.

- Carry `whenToUse` on `ExtensionWorkflowDefinition` and
  `SavedWorkflowEntry`, trimmed and shortened like the description.
- `SavedWorkflowLoader` marks such a command model-invocable and puts the
  condition in its model description.
- Outside the interactive UI (headless, ACP, and the model invoking the
  command through the Skill tool) a tool dispatch cannot run, so the command
  expands to a prompt asking the model to call `Workflow({ name })` with the
  qualified name. The interactive `/name` path is unchanged, and the command
  no longer restricts itself to interactive mode.
- The Workflow tool's opt-in rule counts a skill or command the model itself
  ran through the Skill tool.

@yiliang114 yiliang114 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM at 2e37ae3 — no blocking issues; first review pass. The exposure is scoped exactly right: only extension workflows with an explicit meta.whenToUse become model-invocable (user-saved ones never do), every run still goes through the approval dialog with the content-pinned grant from #11943, and whenToUse lives inside the digested script so editing it re-triggers install consent. The prompt expansion for headless/ACP names the qualified workflow name, and a collision-renamed command still runs the right one. CI green.

@chiga0 chiga0 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed for blocking issues only. None found.

  • Skill listing correctly gated on extension source + whenToUse presence; project/user workflows excluded.
  • Prompt expansion uses the qualified workflow name (not the collision-renamed command name) and JSON.stringifies it in the tool call.
  • Mode restriction removal is safe: action branches on executionMode, all non-interactive paths still end at the Workflow tool's approval gate.
  • whenToUse is trimmed, blank = absent, clamped to 500 code points.
  • No approval bypass path introduced.

@wenshao

wenshao commented Sep 15, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /review

@github-actions

Copy link
Copy Markdown
Contributor

Qwen Code review request accepted. Review is queued for an available runner; follow the workflow run for progress. A command-triggered review is not listed under the checks of this PR; the result is posted here as a review when it finishes.

@qqqys
qqqys added this pull request to the merge queue Sep 15, 2026
Merged via the queue into QwenLM:main with commit 5ca13b6 Sep 15, 2026
111 of 112 checks passed
wenshao added a commit to wenshao/qwen-code that referenced this pull request Sep 15, 2026
@wenshao

wenshao commented Sep 15, 2026

Copy link
Copy Markdown
Collaborator

Local verification on macOS — real build, real extension, real TUI

This is the hands-on run the PR's own Risk & Scope and the triage gate both listed as missing: an interactive session where the model reaches the extension workflow on its own, gets the expansion, calls Workflow({ name }) and lands on the approval dialog. It merged while the run was in progress, so this is a post-merge confirmation rather than a merge gate — with one documentation item still open on main.

Tree tested: PR head e52fe8d, built with npm run build && npm run bundle, against a base build at the merge base 9f8cb17. All 11 files of the squash commit 5ca13b61 are byte-identical to the tree I tested (git diff e52fe8d 5ca13b61 -- <the 11 files> is empty), so everything below applies to what landed.

Rig: macOS 15 (darwin 25.6.0), Node 24.18.1, isolated QWEN_HOME, tools.workflowsEnabled: true. A real extension linked with qwen extensions link (the consent prompt listed all four workflows). Fixtures: dataworks:analysis (declares whenToUse), dataworks:export (none), dataworks:blank (whitespace-only), dataworks:long (600-char condition), plus a project workflow .qwen/workflows/deep-research.js that does declare whenToUse. Controls that exist on the base build too: a project file command tidy-imports and an extension command audit-hint. The model is the repo's integration-tests/fake-openai-server.ts, scripted so the decision points are mine and every request body is recorded; the TUI runs in a real pty through the repo's terminal-capture harness.

1 · The model-invoked path works end to end

Asked a question matching the condition, the model calls the Skill tool with the workflow's qualified name:

model invokes the extension workflow through the Skill tool

It receives the expansion and calls Workflow({ name, args }), which opens the usual approval — with extension provenance, the phases, the agent() structure, the args and the script digest:

workflow approval reached from the model-invoked path

Approved, the run dispatches its subagent and the turn finishes on the workflow's result:

run result

2 · The typed command is unchanged

/dataworks:analysis typed in the TUI still dispatches the tool directly — 0 model requests in that step, and the always-allow keys show the scriptPath spelling, not the name:

typed command approval

3 · A/B against the base build

Same rig, same scripted model, base build: the Skill call is refused and the headless command is rejected outright.

base build rejects the same call

4 · What the model actually received

Every string here comes from the recorded HTTP bodies, not from the terminal:

wire evidence

5 · Full matrix

verification matrix

Also re-run on macOS, which CI skips for this PR: core 330 passed / cli 137 passed (the author's numbers reproduce exactly), npm run typecheck across all workspaces on a fully built tree exit 0 — that closes the "cli typecheck left to CI" caveat locally — and ESLint on the changed files exit 0.

Follow-ups for the maintainer

  1. The design doc still contradicts the shipped behaviour, now on main. docs/design/2026-09-14-extension-workflow-distribution.md:33 — "The slash command … is not model-invocable"; zh-CN :33 — "并且不能被模型调用"; zh-CN :17 still describes the colliding skill as keeping "模型可见的命令列表". The triage gate raised this before merge and it was not addressed in the squash. Worth a small doc PR that records the reversal deliberately.
  2. The listing reaches the model only once extension loading finishes. The session-start <available_skills> snapshot never carries the workflow entry; it arrives with the per-turn delta reminder a moment later. Measured on the head build: a question submitted within ~2 s of the input box appearing goes out without the entry (4/4 sessions); after a ≥5 s pause it is there (4/4 sessions, in the delta block). The pre-existing extension command audit-hint behaves identically, so this is not introduced here — but "ask a matching question as your first message" is exactly the flow the feature invites, and it can silently miss on that first turn. The headless registration path calls skillManager.notifyConfigChanged() after attaching the provider; the interactive one (slashCommandProcessor.ts) does not.
  3. Two wording details. The PR description says the approval opens with Saved workflow: <name>; for an extension workflow the real dialog reads Extension workflow: dataworks:analysis (the extension-provenance branch). And "always allow" is keyed differently on the two paths — name:<name>,sha256:<digest> from the model path vs scriptPath:<path> from the typed command — so an allow-always granted by typing the command does not cover the model-initiated call. Defensible, but undocumented.
  4. Third-party text is structurally safe. A whenToUse containing </description></skill><skill>… reaches the model XML-escaped, so a hostile extension cannot forge listing entries; the 500-char clamp is exactly 500 code points ending in … on the wire. The install consent still lists only name + description, as triage noted.

Not verified: a real LLM deciding by itself to pick the workflow — the listing, the expansion, the dispatch, the approval and the run are all real, but the choice to call Skill was scripted; Windows; the update-consent flow when whenToUse changes.

中文说明

macOS 本地实测:真实构建 + 真实扩展 + 真实 TUI

这正是 PR 自己的 Risk & Scope 和 triage 门禁都标为"未验证"的那条链路:在交互会话里,模型自己触达扩展 workflow、拿到展开文本、调用 Workflow({ name }),并落到审批对话框。 验证过程中 PR 已被合入,所以这是一份合入后的确认,而不是合并前的门禁——但有一条文档问题仍留在 main 上。

被测树: PR head e52fe8d,npm run build && npm run bundle 出真实产物;对照臂是合并基 9f8cb17 的同样产物。squash 提交 5ca13b61 的全部 11 个文件与我实测的树逐字节一致(git diff e52fe8d 5ca13b61 -- <11 个文件> 为空),因此下述结论适用于已合入的代码。

装置: macOS 15(darwin 25.6.0)、Node 24.18.1、隔离 QWEN_HOME、tools.workflowsEnabled: true。用 qwen extensions link 真实安装一个扩展(同意提示里列出了四个 workflow)。夹具:dataworks:analysis(声明 whenToUse)、dataworks:export(未声明)、dataworks:blank(仅空白)、dataworks:long(600 字符条件),外加一个声明了 whenToUse 的 project workflow .qwen/workflows/deep-research.js。基线上同样存在的对照项:project 文件命令 tidy-imports 与扩展命令 audit-hint。模型用仓库自带的 integration-tests/fake-openai-server.ts 脚本化,决策点由我控制并记录全部请求体;TUI 通过仓库的 terminal-capture 在真实 pty 中运行。

1 · 模型发起的链路端到端成立。 问一个符合条件的问题后,模型以限定名调用 Skill 工具(图 1),拿到展开文本后调用 Workflow({ name, args }),弹出常规审批——带扩展归属、phases、agent() 结构、args 与脚本摘要(图 2);批准后子 agent 派发、运行结束并给出结果(图 3)。

2 · 交互式输入命令行为不变。 在 TUI 敲 /dataworks:analysis 仍直接调度工具,该步骤零模型请求,且"始终允许"的键是 scriptPath: 形态而非名字(图 4)。

3 · 与基线构建的 A/B。 同一装置、同一脚本模型,在基线上 Skill 调用被判"not found",headless 命令直接被拒(图 5)。

4 · 模型实际收到的字节(全部取自记录的 HTTP 报文,非终端回显)见图 6;完整矩阵见图 7。

另外在 CI 对本 PR 跳过的 macOS 上复跑:core 330 通过 / cli 137 通过(与作者数字一致);在完整构建产物上跑全 workspace npm run typecheck exit 0——这补上了作者"cli 类型检查交给 CI"的缺口;改动文件的 ESLint exit 0。

给维护者的后续项

  1. 设计文档与已发布行为相反,且现在就在 main 上。 docs/design/2026-09-14-extension-workflow-distribution.md:33 写着"is not model-invocable";中文版 :33 写着"并且不能被模型调用";中文版 :17 仍把撞名的 skill 描述为保留"模型可见的命令列表"。triage 在合入前提过,squash 里没有处理。建议补一个小 PR,把这次反转记录清楚。
  2. 列表只有在扩展加载完成后才到达模型。 会话起始的 <available_skills> 快照始终不含该条目,它是随后由逐轮增量提醒送达的。在 head 构建上实测:输入框出现后约 2 秒内提交的问题,那一轮不带该条目(4/4 次会话);停顿 ≥5 秒再提交则带上(4/4 次,出现在增量块里)。既有的扩展命令 audit-hint 表现完全相同,所以不是本 PR 引入——但"开局第一句就问一个匹配的问题"恰恰是这个功能鼓励的用法,那一轮会静默落空。headless 的注册路径在挂上 provider 后会调 skillManager.notifyConfigChanged(),交互路径(slashCommandProcessor.ts)没有。
  3. 两处措辞细节。 PR 描述说审批标题是 Saved workflow: <name>;扩展 workflow 的真实对话框是 Extension workflow: dataworks:analysis(走扩展归属分支)。另外两条路径的"始终允许"键不同——模型路径是 name:<名字>,sha256:<摘要>,敲命令是 scriptPath:<路径>——因此敲命令时授予的"始终允许"并不覆盖模型发起的调用。这大概率是合理设计,但没有写进文档。
  4. 第三方文本在结构上是安全的。 含 </description></skill><skill>… 的 whenToUse 到达模型时已被 XML 转义,恶意扩展无法伪造列表条目;线上实测条件长度正好 500 个码点并以 … 结尾。安装同意提示仍只列名称与描述,与 triage 的说法一致。

未验证: 真实 LLM 自行决定选用该 workflow——列表、展开、调度、审批与运行都是真实的,但"调用 Skill"这一步是脚本化的;Windows;whenToUse 变化时的更新同意流程。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants