Skip to content

Commit 55f02ec

Browse files
chore(desktop-relay): merge main and keep the MCP App sandbox origin gate
Both sides edited the same `if (webShellDir)` block in createServeApp: this branch passes the desktop-relay CSP flag to mountWebShellAssets, while main (QwenLM#12258) rewired mountMcpAppSandbox to take an origin-allow callback and return a disposer that run-qwen-serve stops on shutdown. The two changes are adjacent, not overlapping, so keep both: the 4-arg mountWebShellAssets call and main's sandbox wiring with its app.locals capture.
2 parents 6cf0efc + 81c260b commit 55f02ec

352 files changed

Lines changed: 45944 additions & 2055 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.github/workflows/ci.yml‎

Lines changed: 14 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1277,9 +1277,22 @@ jobs:
12771277
# test:ci` (vitest) does not collect `node:test` files — so run them here
12781278
# too, or a compositor/publisher change could pass CI without its
12791279
# regression tests. Linux-only (they're platform-independent).
1280+
# One retry, the unit lane's contention doctrine (#10868): the battery
1281+
# spawns real subprocess loops on the shared ECS pool, and a main push
1282+
# has no flaky-rerun patrol — a single host hiccup failed the lane on a
1283+
# CSS-only commit (#12772). A real break still fails both attempts.
1284+
# Attempt 1 tees to a log so the warning can name the failed suites;
1285+
# pipefail (defaults.run.shell: bash) is what keeps node's status
1286+
# through that pipe. The warning is gated on the retry's success, so
1287+
# only an absorbed flake is annotated — a double failure reddens the
1288+
# lane directly (the #11364 doctrine from e2e.yml).
12801289
- name: 'Run .github/scripts helper tests'
12811290
if: "${{ needs.classify_pr.outputs.skip_ci != 'true' && steps.ci_profile.outputs.ci_profile == 'full' }}"
1282-
run: 'node --test --test-concurrency=1 ${{ env.HELPER_TESTS }}'
1291+
run: |-
1292+
node --test --test-concurrency=1 ${{ env.HELPER_TESTS }} 2>&1 | tee "${RUNNER_TEMP}/helper-attempt1.log" || {
1293+
node --test --test-concurrency=1 ${{ env.HELPER_TESTS }} &&
1294+
echo "::warning::helper tests failed once and the retry absorbed it; attempt-1 failures: $(grep -E '^not ok' "${RUNNER_TEMP}/helper-attempt1.log" | head -5 | paste -sd ';' -) — count this toward the #12772 helper-battery recurrence threshold"
1295+
}
12831296
12841297
# State-at-failure dump for the 19 substantive steps after install: the
12851298
# install sampler is reaped by that step's `trap … EXIT`, so a death in

‎.github/workflows/e2e.yml‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -286,6 +286,16 @@ jobs:
286286
npm config set fetch-retries 5
287287
npm config set fetch-timeout 300000
288288
289+
# Fail fast on a saturated pool host instead of dying on ENOSPC
290+
# mid-job (#10035): run 36241138248 lost this leg with no failed step
291+
# and no test result when the runner's own worker died writing its
292+
# diag log on a full disk (#12764). ci.yml and the release lane already
293+
# gate on this script; a gate failure is a clear, retryable signal and
294+
# a re-run lands on a host with headroom.
295+
- name: 'Disk floor gate (self-hosted)'
296+
if: "${{ runner.environment == 'self-hosted' }}"
297+
run: 'bash .github/scripts/check-disk-floor.sh "${GITHUB_WORKSPACE}" "${RUNNER_TEMP:-/tmp}"'
298+
289299
- name: 'Install dependencies'
290300
env:
291301
# `npm ci` runs the `prepare` script, which builds and bundles the

‎.github/workflows/sdk-java.yml‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -305,6 +305,13 @@ jobs:
305305
-Dqwen.cli.entry="${GITHUB_WORKSPACE}/dist/cli.js"
306306
-Dmysql.url="jdbc:mysql://127.0.0.1:${MYSQL_PORT}/hosted_harness_test?allowPublicKeyRetrieval=true&useSSL=false"
307307
-Dmysql.user=root -Dmysql.password=hosted-fixture verify checkstyle:check
308+
- name: 'Run Runtime Broker fault gates'
309+
timeout-minutes: 10
310+
env:
311+
MAVEN_ARGS: '--settings ${{ runner.temp }}/setup-java-m2/settings.xml --toolchains ${{ runner.temp }}/setup-java-m2/toolchains.xml'
312+
run: >-
313+
mvn --batch-mode --no-transfer-progress -f packages/sdk-java/runtime-broker/pom.xml
314+
-Pfault-gates -Dqwen.cli.entry="${GITHUB_WORKSPACE}/dist/cli.js" test
308315
- name: 'Upload Hosted process reports'
309316
if: 'always()'
310317
uses: 'actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a' # v7.0.1
Lines changed: 45 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,45 @@
1+
# Saved workflow slash-command completion
2+
3+
[English](2026-09-21-workflow-slash-completion.md) | [简体中文](2026-09-21-workflow-slash-completion.zh-CN.md)
4+
5+
Status: implementation under review in [#12415](https://github.com/QwenLM/qwen-code/pull/12415), addressing [#12176](https://github.com/QwenLM/qwen-code/issues/12176).
6+
7+
## Problem
8+
9+
A saved-workflow slash command starts a client-initiated tool call. Its foreground result does not continue the model turn, and the registry emits completion notifications only for background runs. The result can exist on disk without reaching the conversation.
10+
11+
## Decision
12+
13+
Keep saved slash commands in the foreground. Separate completion delivery from execution mode: the scheduler passes an internal notification flag to notification-aware invocations when `executionOrigin.kind` is `client`. Missing or model provenance does not opt in, including nested code-mode calls whose results return to their parent. The overloaded `isClientInitiated` flag does not determine delivery. The Workflow invocation passes the delivery flag to its runner and registry entry. It is not a model-facing tool parameter or a persisted execution-mode setting. Both initial scheduling and rebuilding an invocation after a `PermissionRequest` hook edits its arguments preserve the flag.
14+
15+
A completed or failed client-started run uses the existing completion callback and notification queue. Model-started foreground runs retain their tool-result return path; background runs retain their existing notification path. Cancellation produces no completion notification. Terminal-state guards prevent a run from reporting twice.
16+
17+
The runner refreshes the tool card with its terminal state before execution settles. After Esc cancellation, the card retains its run ID and phase history with `status: "cancelled"`, replacing stale running progress. The active-progress guidance disappears once the run settles. This display refresh does not enqueue a completion notification.
18+
19+
Foreground progress stays in the live tool card, with a run ID and guidance pointing to that card. `/workflows <runId>` can inspect the run after it settles; the command is queued during foreground execution. Its completion notice displays the run ID, status, a result preview or error capped at 4,096 characters per block, including the truncation marker, recorded subagent failures, and explicit lines for nonempty top-level `failed`, `errors`, or `error` values on structured object results. Read these fields from the original object, including when another field cannot be serialized or its getter throws. The shared failure formatter limits each reported field and recorded dispatch error to 400 characters, including any truncation marker, without splitting Unicode code points. Empty objects, arrays, Maps, Sets, and text containing only whitespace or terminal controls do not create reported-failure entries. Native Error values render as `Name: message` without runtime stack frames, including sandbox-created Errors and Errors nested in structured results. Nonempty Maps render as lists of key/value pairs and Sets as value lists, preserving their contents and nested Error reasons within the existing preview limits. The same reported lines are included in a separate `<reported-failures>` model element outside the result preview; they do not change runtime dispatch counts. Plain strings are not parsed as JSON failure reports. These fields are displayed as reported data, not used to reinterpret a completed run as a runtime failure. Arbitrary application-specific result schemas remain uninterpreted. A run without a return value displays `(workflow returned no value)` and sends no model result element.
20+
21+
Foreground and background completion notifications cap the model result preview at 25,000 XML-escaped characters, independent of tool-output settings. The model-facing tool result, terminal tool card, and completion notification share Error and collection serialization semantics while retaining their existing pretty and compact formatting. If a value cannot be serialized, the tool card retains the run metadata and replaces only the affected top-level field with a placeholder. Result text, reported failure fields, and recorded dispatch errors share ANSI/control normalization, including Unicode bidirectional embedding, override, and isolate controls, preserving newlines and expanding tabs to spaces. Labels and thrown errors pass through the same normalization, and both complete notification projections are sanitized before delivery; workflow messages and approval script excerpts use the same normalization helper. Escaping is limited to the preview; truncation never splits a Unicode code point or XML entity, using the same bounded XML helper as sub-session notifications. A truncation notice carries the absolute snapshot destination provided by storage, independent of the journal. It states that the snapshot is available after finalization if persistence succeeds: notification emission precedes the snapshot write. The notice explains that snapshots use plain JSON and do not preserve native Error, Map, or Set contents, and that result and reported-failure previews may also be truncated. Neither the snapshot nor the previews promise recovery of every omitted detail. Without a destination, it directs the reader to inspect the settled run without promising a full result or inventing a relative file path. Existing diagnostics retain the per-agent journal path. The registry result and persisted snapshot are not shortened. Background summaries retain their original wrapper quotes and escaped labels.
22+
23+
The TUI displays this foreground notice when the callback arrives, before waiting for model admission or a response. The same notification is then delivered through the existing queue and recorded in the conversation. A later question has the returned result in model context. The slash-command dispatch record may still have empty `outputHistoryItems`: the completion has its own notification record.
24+
25+
## Alternatives and tradeoff
26+
27+
Removing the foreground guard for every run would add notifications to model-started runs that already return a tool result. Forcing slash commands into background mode would change inline progress and foreground cancellation. The internal delivery flag preserves those behaviors while reusing the notification channel.
28+
29+
Letting client tools through the normal tool-result continuation would also require handling their missing model-authored tool call, turn ownership, and cancellation. Expanding the command into a model prompt would add an invocation before execution and give the model control over the requested arguments. Neither is needed for this fix.
30+
31+
The existing notification channel starts a model response and consumes tokens; this proposal preserves that report-back behavior requested by the issue. A configurable silent-context update or optional automatic summary is a separate feature, not implied by preserving foreground execution. The visible result does not depend on the model successfully generating a summary.
32+
33+
## Scope
34+
35+
In the ink renderer, interactive saved workflow commands keep their script-path/name-only choice, arguments, approvals, foreground progress, and cancellation. Headless and ACP commands retain prompt expansion. Code-mode nested calls return to their parent; ordinary model calls do not gain notifications. No public tool schema, configuration option, or persisted-schema migration is added. Snapshot serialization is unchanged: native Error, Map, and Set contents are not preserved by its plain JSON serializer. This repair preserves those values in notifications and model-facing and displayed tool results; it does not add snapshot persistence for their contents.
36+
37+
OpenTUI's separate unwired `schedule_tool` handler and a new persisted-result viewer in workflow history remain outside this change. Both authoring skills scope slash-command guidance to ink and advise asking the model to call `Workflow({ name: '<name>' })` in OpenTUI.
38+
39+
## Validation
40+
41+
Use an isolated interactive session and a recording synthetic model endpoint. Verify foreground progress before a controlled agent finishes, visible completion data and real subagent failures, exactly one result notification, persistence of JSON-compatible result data, and result retention on a follow-up question. Include a large XML-heavy result with the failed field beyond the preview: the explicit failure line survives in both the display and model context, and the bounded model preview, absolute snapshot destination, and separate failure element remain in follow-up context. Exercise thrown script errors, cancellation, and approval refusal. Use model-initiated foreground execution as a control: one tool response and no completion notification. Preserve loader coverage for arguments, name-only dispatch, headless, and ACP. Share the existing 90-second read-only `/about` readiness probe with the command-output visibility test.
42+
43+
Unit regressions cover authoritative provenance during initial scheduling and hook argument rebuilds, JSON-looking strings, non-serializable objects with valid failure fields, all three conventional error keys, sandbox-created and nested Errors, multiple Error reasons within one failure field, populated Maps/Sets, matching card/model result content, per-field degradation for BigInt/circular values, empty failure values, multiline failure text, no-return results, background summary escaping, throwing field getters, ANSI-heavy multiline results, intact Unicode/XML boundaries, and snapshot paths, persistence conditions, and missing-content qualifications for foreground/background runs with and without journals. The built-CLI large-result case verifies that the missing-content qualification reaches the model and remains in follow-up context.
44+
45+
Run targeted unit and terminal regressions, build/type checks, and checks for changed files. No local repository-wide test suite is part of this revision. Evidence belongs in the PR verification report and must distinguish synthetic transport checks from real-model summary quality and the reporter's original nine-agent workload.

0 commit comments

Comments
 (0)