Skip to content

docs(serve): fix stale permission-route note and document pending-prompt events - #13281

Merged
wenshao merged 1 commit into
QwenLM:mainfrom
LUOSENGWA:docs/serve-docs-corrections
Oct 7, 2026
Merged

wenshao merged 1 commit into
QwenLM:mainfrom
LUOSENGWA:docs/serve-docs-corrections

Conversation

@LUOSENGWA

Copy link
Copy Markdown
Contributor

What this PR does

Fixes a stale user-guide note and fills an event-reference gap, both surfaced by the triage on #13279. (1) docs/users/qwen-serve.md: the old note still described session-scoped permission routing as Stage 1.5 future work and pointed integrators at the legacy process-global vote route — it has shipped since Wave 4 / F3. The note now names POST /session/:id/permission/:requestId as the recommended form for new multi-workspace integrations, spells out the legacy route's 404 semantics (primary-bound; the same body as a lost vote for a session owned by another runtime), and keeps a standing shared-bearer caveat; the "mutable over HTTP" bullet is updated in the same direction. (2) docs/developers/daemon/09-event-schema.md: adds a "Pending prompt queue" subsection documenting pending_prompt_added, pending_prompt_started, and pending_prompt_completed — payload fields, the state: 'completed' | 'removed' union, the idle-session no-event caveat (a first prompt on an idle session emits no queue events), and the "queue-view bookkeeping, not a turn terminal" guidance (correlate turn_complete / turn_error by promptId for completion).

Why it's needed

An integrator following the user guide is actively pointed at the legacy permission-vote shape, and the three queue-lifecycle events are public SDK surface (packages/sdk-typescript/src/daemon/events.ts) without a reference entry — a completion detector keyed on pending_prompt_completed silently never fires for the common create-single-prompt-wait shape. Both fixes are concrete and independent of the external-supervisor recipe-page scope question in #13279; this PR covers the two items the triage marked as doable now.

Reviewer Test Plan

How to verify

Docs-only change. (1) The rewritten note in docs/users/qwen-serve.md renders and its link to the REST reference's #permissions anchor resolves; (2) the new subsection in docs/developers/daemon/09-event-schema.md renders as a table and its claims match packages/acp-bridge/src/session-control-plane.ts publish conditions (isQueued gating for added; head-of-FIFO promotion for started; isQueued && !removed for completed); (3) payload fields match the three DaemonPendingPrompt*Data interfaces in packages/sdk-typescript/src/daemon/events.ts.

Evidence (Before & After)

N/A — docs-only. Before: the user guide described the session-scoped permission route as future work, and the event reference had no pending-prompt entries. After: the note reflects the shipping route (with legacy-route 404 semantics), and the three events have reference entries with their conditions.

Tested on

OS Status
🍏 macOS N/A
🪟 Windows N/A
🐧 Linux N/A

N/A — docs-only; there is no runtime behavior to test across operating systems.

Environment (optional)

N/A — verifying documentation requires no runtime environment.

Risk & Scope

Linked Issues

References #13279 (without a closing keyword — this PR covers the two docs items from its triage; the recipe-page discussion stays on the issue).

中文说明

本 PR 做了什么

修复一处用户指南过期说明 + 补一处事件参考缺口(均来自 #13279 的 triage):(1)docs/users/qwen-serve.md 里旧的权限告示仍把 session 级权限路由写成 Stage 1.5 未来工作、并把集成方引向 legacy 全局投票路由——该路由自 Wave 4 / F3 已 ship。现在:POST /session/:id/permission/:requestId 写明为新多工作区集成的推荐形态;legacy 路由的 404 语义写明(primary-bound;对属于其他 runtime 的会话返回与"票已丢"相同的 404 体);保留共享 bearer 的安全告示;"HTTP 可变面"清单同步更新。(2)docs/developers/daemon/09-event-schema.md 新增 "Pending prompt queue" 子节:补全 pending_prompt_added / pending_prompt_started / pending_prompt_completed 的触发条件、载荷字段、state: 'completed' | 'removed' 联合,以及两个关键点——空转会话首条不产生队列事件;这是队列视图记账事件、不是回合终端(完成判定用 turn_complete / turn_error 按 promptId 关联)。

为什么需要

按用户指南走的集成方会被指到 legacy 权限票形态;而三个队列事件虽是公开 SDK 面(packages/sdk-typescript/src/daemon/events.ts),却无参考条目——以 pending_prompt_completed 做完成判定的监督器,在"新建→单发→等"这一最常见形态下会永远不触发。两项修复都具体且独立于 #13279 中待定的 recipe 页面范围问题;本 PR 覆盖 triage 标为"可立即做"的两项。

审查要点

纯文档。(1)改写段渲染正常、#permissions 锚点可达;(2)新表与 packages/acp-bridge/src/session-control-plane.ts 的发布条件一致(added 由 isQueued 门控;started 在队列头部晋升时发出;completed 在 isQueued && !removed 时发出);(3)载荷字段与 packages/sdk-typescript/src/daemon/events.ts 的三个 DaemonPendingPrompt*Data 类型一致。

风险与范围

低——纯文档、无行为变化。未含:外部监督者 recipe 页面(归属与形态属维护者范围决定;站点导航归 qwen-code-docs);若获准,可按新页面或 rest-api-integration.md 章节两种形态渲染。无破坏性变更。

关联

References #13279(不带关闭关键字——本 PR 仅覆盖其 triage 中的两个文档项;recipe 页面讨论留在 issue)。

…mpt events

docs/users/qwen-serve.md still described session-scoped permission routing
as Stage 1.5 future work and pointed integrators at the legacy process-global
vote route. Rewrite the note: the session-scoped route ships (Wave 4 / F3) and
is the recommended form for new multi-workspace integrations; the legacy route
remains for pre-F3 single-workspace clients with its 404 semantics spelled out
(the same body as a lost vote for a session owned by another runtime).

Add event-reference entries for pending_prompt_added / pending_prompt_started /
pending_prompt_completed to docs/developers/daemon/09-event-schema.md: payload
fields, the state union, the idle-session no-event caveat, and the "queue-view
bookkeeping, not a turn terminal" guidance (correlate turn_complete /
turn_error by promptId for completion).

Derived from the field notes and triage in QwenLM#13279.
@wenshao

wenshao commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /triage

@wenshao

wenshao commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

Verdict: findings — maintainer local verification, round 1. 39/39 scripted assertions passed at head e5c2cf8b9b2d5c47d3341e22470cc584064c043e (base 66a6279894), including a live two-workspace qwen serve E2E that confirms every load-bearing claim in both edits. Two non-blocking documentation nits: one sentence in the new event table is falsified by live behavior (state: 'removed' also fires for a prompt that already started), and the new table is not Prettier-normalized (base file was clean).

中文摘要

结论:findings(不阻塞,两条文档级建议) — 39/39 条脚本断言全部通过。

A/B 与实测定论:纯文档 PR,运行时代码无变化,故 A/B 的维度是"文档声明 vs 代码事实"。在 head 构建上启动了真实双工作区 qwen serve(假模型端点做线控预言机),实测:

  • 空闲会话首条 prompt 确实不产生任何 pending_prompt_* 事件(文档关键告诫成立);
  • 排队 prompt 依次发出 added{sessionId,promptId,text,queuedAt} → started{sessionId,promptId,text} → completed{state:'completed'};执行错误的 prompt 仍报 completed 且 turn_error 携带同一 promptId(文档的关联指引成立);
  • 权限路由:session 级路由把次要工作区的票路由到属主 runtime(200),legacy 全局路由对其他 runtime 的会话返回与"丢票"逐字节相同的 404 体(文档语义成立)。

两条建议(见 Findings):

  1. 新表格称 state:'removed' 表示"从未运行"——实测被删除的运行中 prompt 同样发出 completed{state:'removed'}(在 started 之后)。triage 机器人的静态核查漏掉了这一点(它只核对了 :11681 的 completed 发布点,没核对 :14729 的 removed 发布点无 isQueued 门控)。
  2. 新表格未过 Prettier(base 干净、head 失败;本 PR 走 docs_only 通道不会挂 CI,但合并后 main 变脏,下一个 full 通道 PR 的 prettier 门禁会失败)。

未覆盖:晋升后立即运行的 mid-turn 消息的 started 分支(时序敏感,仅静态核实);removed 路径中"从未排队的运行中首条 prompt 被删也会发 completed"(仅代码层核实,无 isQueued 门控);prettier 全仓检查在本机超时(>5min,未跑完)——docs/ 目录级 A/B 已足够定位问题文件。

Central claim + evidence

This PR is docs-only; runtime code is unchanged, so the A/B dimension is documented claim vs. code/live behavior rather than head-vs-base runtime. Every claim was verified against the head build statically and — for all load-bearing behavioral claims — by driving a real qwen serve daemon (dist/cli.js bundled from the head worktree, two --workspace runtimes, loopback fake OpenAI endpoint as wire oracle, SSE GET /session/:id/events capture, REST votes and deletes). Capture: 01-live-e2e-queue-and-permission.png.

Claim table — docs/developers/daemon/09-event-schema.md (new "Pending prompt queue" subsection)

Doc claim Verified against Result
added payload sessionId, promptId, text, queuedAt DaemonPendingPromptAddedData (events.ts:1036) + live SSE frame exact, live
started payload sessionId, promptId, text DaemonPendingPromptStartedData (events.ts:1043) + live frame (no queuedAt) exact, live
completed payload + state: 'completed' | 'removed' DaemonPendingPromptCompletedData (events.ts:1049) + live frames of both states exact, live
added only when genuinely queued; first prompt on idle session emits nothing isQueued = pendingPromptCount > 1 (:11011), publish under if (isQueued) (:11136); live: P1 emitted zero pending_prompt_* events exact, live
started at head-of-FIFO promotion; skipped when aborted before promotion :11215 publish, :11188 abort-throw; live: promoted P2 emitted started, deleted-queued P3 never did exact, live
execution errors still report completed; correlate turn_complete/turn_error by promptId publish in result.finally (:11666); live: queued prompt whose model stream errored emitted turn_error{promptId} and completed{state:'completed'} with the same promptId exact, live
state:'removed' = "removed from the queue and never ran"; completed "only emitted for genuinely queued prompts" falsified live — see Finding 1 mismatch

Claim table — docs/users/qwen-serve.md (rewritten permission note + mutable-over-HTTP bullet)

Doc claim Verified against Result
POST /session/:id/permission/:requestId ships; routes to the owning runtime, never falls back to primary routes/permission.ts:47 + requireSessionRuntime (not_found → 404, no primary fallback); live: vote to a session in the secondary workspace runtime returned 200 and executed the tool there exact, live
pre-flight caps.features.session_permission_vote capabilities.ts:91; live: GET /capabilities advertises it exact, live
legacy POST /permission/:requestId is primary-bound registered with bridge: primaryBridge (server.ts:3784); live: vote for a secondary-runtime session → 404 exact, live
legacy 404 for another runtime's session has "the same body as a lost vote under first-responder" live: unknown-id, lost-vote (2nd vote), and cross-runtime votes all returned {error:"No pending permission request", requestId} — identical shape exact, live
first-responder is the default policy opts.permissionPolicy ?? 'first-responder' (session-control-plane.ts:4438) exact, static
link ../developers/daemon-rest-api-reference.md#permissions resolves file exists; ## Permissions at :83 → #permissions; its table independently classifies the two routes live-session-owner vs legacy-primary resolves, consistent
"(Wave 4 / F3)" label corroborated by docs/developers/daemon/00-index.md (F3: multi-client permission coordination) and qwen-serve.md:47 (session-scoped permission routing (Wave 4 PRs)) consistent with project naming

Correction (to the triage stage-2 review comment)

The stage-2 bot comment's verification table has one row that checks the wrong publish site: "completed only when genuinely queued and not already removed — if (isQueued && !pendingEntry.removed) (:11681)". That guard covers only the state:'completed' publish. The state:'removed' publish at session-control-plane.ts:14729 (removePendingPrompt) has no isQueued guard — any listed prompt, queued or running, emits it on removal. The live run demonstrates the consequence: a prompt emitted pending_prompt_started and then pending_prompt_completed{state:'removed'}. The review's conclusion row is therefore incomplete, and the doc sentence it endorsed ("never ran") is inaccurate on that path. This is a correction of the review record — the code behaves sensibly (queue-view bookkeeping); it is the doc sentence that needs one clause.

Findings

1. (Suggestion) pending_prompt_completed row: "state: 'removed' (removed from the queue and never ran)" is falsified by the running-removal path

Repro (live, head build): with a session running prompt P2 (already added + started), DELETE /session/:id/pending-prompts/<p2> returns {removed:true} and the session's SSE stream emits pending_prompt_completed {state:'removed'} — for a prompt that demonstrably ran. Observed sequence in the harness: pending_prompt_added(P2) → pending_prompt_started(P2) → pending_prompt_completed(P2, state:'removed') (see 01-live-e2e-queue-and-permission.png, and the daemon log line removing promptId=… state=running in the raw log).

Code: removePendingPrompt publishes state:'removed' unconditionally after the queued/running branch (session-control-plane.ts:14727-14738); the running branch (target.removed = true) exists precisely so a removed running prompt keeps its list slot until settle (comment at :11670-11673 confirms the design). The same path means a never-queued running first prompt also emits completed{state:'removed'} on removal, which additionally falsifies the row's "Only emitted for prompts that were genuinely queued (an added was published)" clause (static only — not exercised live).

The sibling protocol doc already describes removal-aborts-running (qwen-serve-protocol.md:3296), so the new table also disagrees with its own sibling. Suggested rewording, consistent with both live observations and the sibling doc:

state: 'removed' (deleted by a client before settling: a queued prompt never runs; a running prompt is aborted where it stands)

and scope the "genuinely queued" sentence to state: 'completed'.

2. (Suggestion) The new table is not Prettier-normalized — base file was clean

npx prettier --check docs/ passes at base and fails at head on exactly one file (docs/developers/daemon/09-event-schema.md); --write would only re-pad the table columns and rewrite *genuinely queued* as _genuinely queued_ (02-prettier-ab-base-clean-head-fails.png). This PR classifies docs_only, so CI runs no lint lane and it merges green; but scripts/lint.js --prettier checks the whole repo (.prettierignore does not exclude docs/), so after merge the next full-profile PR's Prettier gate fails on lines it didn't author. One npx prettier --write docs/developers/daemon/09-event-schema.md fixes it. (The triage stage-3 comment already flagged "a formatter run" as a nit; this confirms it with the A/B and names the downstream cost.)

Not covered

  • The "promoted mid-turn message that starts immediately" started branch (isPromotedMidTurn, :11215): timing-sensitive to reproduce honestly over REST; verified statically only.
  • The removed-without-prior-added case (never-queued running prompt deleted): read from the unguarded publish at :14729, not exercised live.
  • Repo-wide prettier --check . timed out on this machine (>5 min); the docs/-scoped A/B above is the exact subset needed to attribute the regression to this PR, but I did not prove the rest of the tree clean at base.
  • Per-CI-lane behavior of the docs_only profile was read from ci.yml + classify-pr-profile.sh, not executed.
  • The "Wave 4 / F3" label is project-history naming; verified as consistent with the daemon docs' own usage, not against the original planning issues.

Methodology

Environment: Orange Pi (Linux aarch64), Node v24.14.0. Head worktree at e5c2cf8b9b (single commit on base 66a6279894); dependencies installed via node scripts/setup-worktree.js (pinned pnpm, --frozen-lockfile), then npm run build && npm run bundle. The live harness (pure Node stdlib) spawns node dist/cli.js serve --http-bridge --no-web --port 0 --token … --workspace wsA --workspace wsB with an isolated HOME/QWEN_HOME/trust store, points the model at a loopback fake OpenAI server (per-request hold gates released over an admin endpoint; marker-keyed tool_call/error responses scoped to the newest text block of the last user message), subscribes to both sessions' SSE streams with Last-Event-ID: 0, and drives the queue (3 prompts with one queued removal and one running removal, plus a mid-stream model failure on a queued prompt) and permission scenarios (unknown-id legacy vote, session-scoped vote, primary-owned legacy vote, lost vote, cross-runtime legacy vote, cross-runtime session-scoped vote). Every assertion is a scripted check printed as [PASS]/[FAIL]. Captures produced by scripts/verify-capture.mjs from the head tree. The identical harness passed 37/37 on three consecutive runs (35/35 before the error-case block was added).

Evidence

01-live-e2e-queue-and-permission.png — the full live run, 37/37:

01-live-e2e-queue-and-permission

02-prettier-ab-base-clean-head-fails.png — Prettier A/B (base exit 0, head exit 1, and the exact reformat --write would apply):

02-prettier-ab-base-clean-head-fails

@yiliang114 yiliang114 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Verified against the serve implementation: pending_prompt_added only fires for genuinely queued prompts, pending_prompt_started is skipped when the queued prompt aborts before promotion, pending_prompt_completed carries state: 'completed' | 'removed' with turn_error/turn_complete as the turn terminal, and POST /session/:id/permission/:requestId plus the session_permission_vote capability exist as documented. The replaced Stage-1.5 note was indeed stale.

@wenshao
wenshao added this pull request to the merge queue Oct 7, 2026
Merged via the queue into QwenLM:main with commit 8505479 Oct 7, 2026
47 checks passed
yiliang114 added a commit that referenced this pull request Oct 7, 2026
docs/developers/daemon/09-event-schema.md reached main unformatted in
8505479 (#13281), whose docs-only CI profile skipped the Prettier lane.
scripts/lint.js --prettier runs `prettier --check .` across the whole repo, so
every branch with that commit as an ancestor now fails Lint & Static on a file
it never touched. Normalize the emphasis markers and table padding so this
branch's own gate can pass.

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-conflict/jmuxqta15bw
yiliang114 added a commit that referenced this pull request Oct 7, 2026
docs/developers/daemon/09-event-schema.md reached main unformatted in
8505479 (#13281), whose docs-only CI profile skipped the Prettier lane.
scripts/lint.js --prettier runs `prettier --check .` across the whole repo, so
every branch with that commit as an ancestor fails Lint & Static on a file it
never touched. Normalize the emphasis markers and table padding so this
branch's own gate can pass.

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-conflict/jmuxqta15bw
wenshao added a commit that referenced this pull request Oct 7, 2026
…ier gate

main's #13281 added the pending-prompt queue table unformatted, so once
the branch merged main the repo-wide `node scripts/lint.js --prettier`
gate turned the Lint & Static lane red on `09-event-schema.md` — the
same fix H4b already carries as `e72c63a11c` upstream of its line.
Pure formatting; content unchanged.

Refs #13548
yiliang114 added a commit that referenced this pull request Oct 7, 2026
docs/developers/daemon/09-event-schema.md reached main unformatted in
8505479 (#13281), whose docs-only CI profile skipped the Prettier lane.
scripts/lint.js --prettier runs `prettier --check .` across the whole repo, so
every branch with that commit as an ancestor fails Lint & Static on a file it
never touched. Normalize the emphasis markers and table padding so this
branch's own gate can pass.

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-conflict/jmuxqta15bw
shpaker pushed a commit to shpaker/qwen-code that referenced this pull request Oct 7, 2026
…wenLM#13548)

* feat(managed-agent): H5a channel route and delivery record contract

Land slice H5a of the Managed Agent extension runtime (QwenLM#12827, stage H of
QwenLM#12380): the durable record contract for Channels. Define the
managed-channel_route and managed-channel_delivery record bodies — the two
domains H0b registered but left unbodied — with one shared fixture corpus
(83 shape cases, 54 successor cases) replayed identically by TypeScript and
Java, register both bodies with taskKind null, teach the Session authority
their resource closures, and keep both domains refused for submission. The
H5 design document (both languages) moves from contract direction to the
pinned byte-level contract this slice freezes.

Also widen ManagedExtensionRecords' value-based comparator to the package
so the channel records share the sibling JSON.stringify-compatible numeric
semantics instead of a double comparison that splits -0.0 from 0.0.

Refs QwenLM#12827 QwenLM#12380

* fix(managed-agent): close the two H5a review findings from QwenLM#13548

Two follow-ups from the wenshao review on the H5a record contract:

- The segment plan parser iterated with Array.prototype.map, so a sparse
  plan emptied its holes past every check and then failed its own durable
  round-trip once the holes serialized as null. Walk every index so a hole
  refuses as a non-object entry, pin a TypeScript regression for the
  sparse shapes, and add the explicit null-element corpus case both
  languages replay (84 shape cases, 54 successors).
- The Java store's resource closure only covered the MCP, Hook,
  child-run and monitor bodies, so a channel journal could persist a
  record the Session authority cannot reopen. Enumerate the channel
  route's policyRef and the delivery's contentRef, every segment
  contentRef and every non-null receipt proofRef in the commit
  transaction, with positive commits (planned, sending with a proofed
  receipt) and one refusal per missing resource kind.

Refs QwenLM#12827 QwenLM#13548

* fix(managed-agent): close the review F1/F2 follow-ups on H5a

- Leaving `unknown` requires proof: `unknown → partial` must settle a
  segment the unknown revision did not, in both validators, or the next
  `sending` revision could re-send a segment the provider may already
  hold (decision 5; the previous corpus composed the two-step bypass
  byte-for-byte). Corpus re-derived: `delivery-unknown-proves-partial`
  gains the third segment with a real new receipt, and
  `delivery-unknown-partial-without-proof` pins the refusal.
- Corpus strength: pin the ten rules both languages enforced without a
  fixture (thread scope with a senderId, single with a chatId, unknown
  scope kind with all-null carriers, non-boolean cancelRequested,
  unknown/rejected with every segment receipted, route accountId drift,
  delivery routeId drift, malformed proofRef, malformed segment
  contentRef), and correct `delivery-cancelled-with-receipt` so the
  refusal rides the zero-receipt conjunct it names.
- Design docs (both languages) record the leaving-unknown rule and the
  updated 92/57 corpus size.

Refs QwenLM#12827 QwenLM#13548

* fix(managed-agent): repair the store-test splice after the H4a merge

The H4a/H5a registry merge put the closure test and helper methods into a
splice that lost channelDelivery's closing brace and interleaved the H4a
tests mid-method. Re-laid the file on top of origin/main with the closure
test and helpers inserted before childRun, restoring compilation and the
full 25-test store suite.

Refs QwenLM#12827 QwenLM#13548

* fix(managed-agent): restore @test on the channel resource-closure regression

The store-test splice after the H4a merge dropped the annotation while
re-laying the file, so normal test runs silently skipped the regression
(the review-5438970962 Suggestion). Restored it; the suite runs 26 tests
including all positive commits and the four missing-resource refusals.

Refs QwenLM#13548

* chore(docs): format the daemon event-schema queue table for the Prettier gate

main's QwenLM#13281 added the pending-prompt queue table unformatted, so once
the branch merged main the repo-wide `node scripts/lint.js --prettier`
gate turned the Lint & Static lane red on `09-event-schema.md` — the
same fix H4b already carries as `e72c63a11c` upstream of its line.
Pure formatting; content unchanged.

Refs QwenLM#13548

* fix(managed-agent): widen RECORD_BODIES past Map.of's ten-pair cap

H6a's schedule and automation_run bodies bring the registry to eleven
bodies once this branch merges with main; Map.of constructor overloads
stop at ten. Convert to Map.ofEntries, mirroring the entries layout
that both contract surfaces and the projection fixtures drive.

Refs QwenLM#12827 QwenLM#13548
wenshao added a commit that referenced this pull request Oct 7, 2026
Upstream #13281 landed the doc with 4 badly wrapped table lines;
its own lint lane ran a 15 s light profile that never scanned it,
so the violation reached main unnoticed and began failing every PR
merge ref computed after 2026-10-07T04:45Z (this PR's run on
13d1735 inclusive). Format-only change: `prettier --write` (4
lines); the whole-tree `--check .` now passes, and upstream #13536's
version of the same file is already clean, so the next main merge
stays green by construction.
zjgzx1988 pushed a commit to zjgzx1988/qwen-code that referenced this pull request Oct 7, 2026
…LM#13573) (QwenLM#13580)

Main CI's Run Prettier step failed at 28512f1 because 8505479 (QwenLM#13281)
landed the pending-prompt table with asterisk emphasis that Prettier
normalizes to underscores, with the separator row re-padded to match.
Apply the normalization; no content change.

Co-authored-by: Qwen Autofix <[email protected]>
Co-authored-by: Qwen-Coder <[email protected]>
kvnloo pushed a commit to kvnloo/qwen-code that referenced this pull request Oct 7, 2026
…3574)

* fix(acp): preserve branch checkpoints after Code Mode turns

Co-authored-by: Qwen-Coder <[email protected]>

* ci(docs): format the daemon event-schema table for the Prettier gate

docs/developers/daemon/09-event-schema.md reached main unformatted in
8505479 (QwenLM#13281), whose docs-only CI profile skipped the Prettier lane.
scripts/lint.js --prettier runs `prettier --check .` across the whole repo, so
every branch with that commit as an ancestor fails Lint & Static on a file it
never touched. Normalize the emphasis markers and table padding so this
branch's own gate can pass.

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-conflict/jmuxqta15bw

* test(acp): accept code_mode_tool_result options arg in native-metadata test

The branch's checkpoint preservation now records nested Code Mode tool
results with a third options argument ({subtype: 'code_mode_tool_result'})
— a deliberate contract change already pinned by the sibling test at
~41694. The older 'allows top-level discovery' assertion still used a
two-argument toHaveBeenCalledWith, which no longer matched the nested
read_file call. Extend the matcher minimally to expect the new options
argument.

* test(acp): pin the Goal turn stamp on nested Code Mode originals

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-closeout/jmuxwj1f62z

* refactor(acp): keep subtype out of the queue callback input type

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-closeout/jmuxwj1f62z

---------

Co-authored-by: Qwen-Coder <[email protected]>
Co-authored-by: yiliang114 <[email protected]>
kvnloo pushed a commit to kvnloo/qwen-code that referenced this pull request Oct 7, 2026
…dger (M5c) (QwenLM#13352)

* feat(managed-agent): prove Shell process-group stops with a worker ledger (M5c)

M5a settled a cancel on the worker's word: a Shell member that ignored
SIGTERM survived its leader, and a worker that crashed left the Shell
process groups it started unnamed to the child's snapshot-only registry.

Each Runtime worker now keeps an incarnated ledger of its Shell process
groups, rewritten atomically before any settled step. A settled cancel
waits for the whole group to die before the call journals; the host sweeps
the ledger of a dead worker, and every new child sweeps older ledgers,
neither one trusted. Identity is judged from the live process table with a
leader-dated, one-sided start-time proof: a young live leader or a live
pid the table cannot name is never signalled, and a group nothing can
prove is a quarantine that blocks new Managed sessions until a reaper
proves it. Registers and enables nothing; Windows stays on its documented
liveness-only shape pending real-machine verification before M6.

Refs QwenLM#12380

* docs(managed-agent): cite the M5c pull request in the engine design

* fix(managed-agent): harden the M5c worker ledger after review round 1

Identity: judge ages inside the boot-clock domain on Linux — ps etime is
boot-derived and a wall-clock step must not make a live leader read as a
recycled impostor — resolve a recycled id only once no old-enough member
of the recorded group remains, probe hostPid as a process whose EPERM
holds rather than sweeps a sibling, refresh the identity snapshot after
the worker's proof wait, and date every proven group with its own budget
rather than one shared serial deadline.

Containment: ledger filesystem failures no longer escape into fatal
states — an addGroup failure stops the newborn group and fails the call
loudly, close and the settle-evidence block contain errors instead of
rejecting the shutdown or losing the journal's terminal state, an
unreadable ledger retires aside once it outlives the debris age instead
of quarantining forever, the .tmp debris unlink joins the per-entry
containment, and the 1 Hz watchdog prunes by liveness alone instead of
forking a blocking ps each tick.

Quarantine: a dedicated managed_engine_quarantined error kind mapped to
HTTP 503 with the reason kept at both daemon mappers, replacing the
resume-conflict 409 rewrite; a startup sweep failure now has its report
test; the reaper backs off to a 30-second cadence. The worker also
scrubs the ledger path from the environment before any Shell command
inherits it.

* fix(managed-agent): close the M5c quarantine verdict and witnessed-retry holes

Round-2 review follow-up:

- Sweep verdicts: sweepWorkerLedger/sweepStaleLedgers now answer
  proven/absent/held/retired, and startLedgerReaper lifts only on the
  sweep's own proof. A retired unreadable ledger never fires onProven,
  and a vanished ledger re-probes the groups the last failure named
  before any lift.
- A witnessed sweep consults the live process table again: the exit
  witness is fresh only at the close it names, so a reaper's retry hours
  later resolves a provably recycled id without signalling its new
  holder, while a genuine or undatable survivor is still killed.
- recordAgeMs refuses a negative boot-domain age: a stamp ahead of this
  boot (a ledger that outlived a reboot) falls back to the wall clock
  instead of inverting the identity guard into an ownership assertion.
- addGroup derives the default uptime stamp from the caller's startedAt,
  so a backdated record keeps both clocks telling one story (fixes the
  recycled-id test that ran red on Linux).
- A quarantine refusal resets the deferred Managed-conversation
  activation instead of poisoning it, so an existing session can enter
  its workspace after the reaper proves the stop; terminal refusals stay
  poisoned.
- The 503 refusal summarizes the quarantine client-safely (ledger
  basename and unproven count) instead of republishing the sweep's
  absolute paths and live pgids.
- An orphaned worker's record omits hostPid (init's pid 1 is no host and
  the reader refuses it), and waitForGroupExit drops only the record it
  judged, never one re-recorded on the same id.
- close() leaves a ledger its reaper already owns to the reaper.
- Design docs (EN+zh-CN) realigned: the quarantine error kind, the
  reaper's backoff, the retirement path, the liveness-only prune, the
  fourth accepted hole, and the daemon mapper files in the M5c row.

* test(cli): pin POSIX shim scripts to CommonJS against up-tree package.json

Deterministic verification rejected the M5c round-2 commit: two
packages/cli tests failed — the review command's git-shim case and the
chafa caching case. Both write an extensionless `#!/usr/bin/env node`
shim whose body uses require()/__dirname__ into os.tmpdir(). A
package.json with "type": "module" in tmpdir's up-tree (this runner's
/tmp/package.json) makes Node load those shims as ES modules, so they
crash before arming or counting. The mermaid fake-mmdc/fake-chafa
helpers share the pattern and fail the same way in that environment.

Each shim directory now gets a {"type":"commonjs"} package.json, pinning
the interpretation the shims were written for. Verified: without the pin
the named tests fail under the polluted tmpdir and pass under a clean
one; with the pin they pass under both, and the full packages/cli suite
is green (38595 passed, 92 skipped).

* fix(cli): close the managed runtime sweep's false-proof and never-prove holes

Review round for QwenLM#13352:

- the startup reaper now accepts its own sweep's proof: the directory
  sweep returns 'proven' only for a ledger it judged clean itself, and
  'absent' (nothing judged) falls through to the last-failure liveness
  probe, so a deleted ledger still holds the quarantine until the groups
  it named die while a sweep-proven clean bill lifts it.
- a retired ledger now rejects with LedgerSweepRetiredError instead of
  resolving 'retired', so every production caller — close, the exit hook,
  the failed-launch sweep, the startup sweep — quarantines it, and both
  reaper closures map the rejection (seeded from the arming one, since
  the aside is judged by no later pass) to 'terminal'.
- a bookkeeping unlink that fails (aged .tmp debris, or the post-proof
  ledger removal) is logged and skipped rather than escalated into an
  unprovable stop.
- queryProcessTable pins COLUMNS=4096 and passes -ww so procps never
  truncates the args column the identity markers are matched against.
- the reaper's retries no longer carry the exit witness: the witness is
  fresh only at the exit it names, and a retry minutes later must judge
  by the live table alone rather than SIGKILL a recycled id's new holder.
- a pid-1 host is recorded as-is and told from init by the live table's
  argv, never by liveness alone; the reader accepts hostPid >= 1, and the
  hold never applies to a ledger this process itself parented.
- the single-ledger reaper maps an empty last-named set to 'terminal':
  an unreadable ledger that is then deleted is no proof of anything.
- the reaper backoff doubles on every unresolved retry, thrown or
  resolved-'unproven' alike; the group judgement re-reads the process
  table only after a proof wait aged the first snapshot, and keeps the
  paid-for snapshot when the re-read fails.

Each new guard carries a mutation-probed witness test; the design doc
records the pid-1 reading, the undatable-group hole, and the Windows
attribution qualifier in both languages.

Co-authored-by: Qwen-Coder <[email protected]>

* fix(cli): pin the managed runtime reaper's lift and harden ledger reads

Review round for QwenLM#13352:

- the startup reaper's lift is pinned to the files that armed the
  quarantine: a directory-level 'proven' an unrelated ledger earned no
  longer lifts, a file gone without a judgement falls back to the
  liveness of every group any failure named (accumulated across
  retries, never replaced), and a retirement ends the reaper
  terminally only after the remaining provable ledgers settle.
- recordAgeMs judges a previous-boot record as unmatchable
  (Number.POSITIVE_INFINITY) instead of wall-aging it into a negative
  age that matched every live process — the identity guard inverted
  into an ownership assertion; both identity matchers also refuse a
  record whose stamp lies in the future.
- a read-side failure (EMFILE/EIO/ENOMEM) no longer retires a valid
  aged ledger: reader and bytes are judged apart, the file stays in
  place, and the stop stays unproven for the next pass.
- the liftable-quarantine refusal guard no longer dereferences a null
  `data`, which wedged provisional activations on a TypeError thrown
  inside the catch.
- the pid-1 ledger test sweeps through a sys seam instead of against
  the runner's own pid (a Windows fork killer), and the child-kill
  process test now asserts the ledger file is gone end to end (M5c.3).
- dropped the duplicate CommonJS shim pins main already carries.

Each new guard carries a mutation-probed witness test.

Co-authored-by: Qwen-Coder <[email protected]>

* style(cli): fix Prettier spacing in managed-runtime-tool-executor test import (QwenLM#13352)

* refactor(cli): drop unreachable terminal verdict in startup ledger reaper (QwenLM#13352)

Co-authored-by: Qwen-Coder <[email protected]>

* fix(managed-agent): hold the quarantine for a ledger that vanished unjudged

The startup reaper lifted the engine's quarantine when a never-readable
ledger was deleted from outside the sweep, while its single-ledger
sibling holds the same fact terminal: a file nobody ever judged proves
nothing by disappearing, so the stop stays unprovable. The directory
reaper now returns terminal on that shape like its sibling, and the
shipped startup-sweep test is re-pinned to the held quarantine, with the
positive lift case next door. The design records the terminal rule in
both languages.

From qqqys's review comment, answered there with the ruling.

* fix(managed-agent): sweep the ledger's final truth and judge boot-stamped records in their own clock domain (QwenLM#13352)

Co-authored-by: Qwen-Coder <[email protected]>

* test(cli): align ACP bridge refusal assertions

Match the current named validation guidance and refusal prefix in the ACP
regression tests. The production bridge contract is unchanged.

Co-authored-by: Qwen-Coder <[email protected]>

* fix(managed-agent): address review round 2 on the M5c physical-stop slice

Identity and write discipline: ledger addGroup writes durably before the
in-memory mutation so a failed write never leaves a memory-only group the
host sweep cannot see; staging is uniquely named per write so concurrent
writers can no longer clobber a shared temporary; ps elapsed parsing
rejects procps's negative-wraparound spelling; proof deadlines live on the
monotonic clock; killOutstanding shares one process-table read across all
its waiters.

Reaper and sweep ownership: the single-ledger reaper accumulates named
groups like its sibling instead of replacing, records retirement, and
holds the quarantine terminal on the same facts; close() and the
launch-failure path join the exit hook's in-flight sweep instead of
starting a second one; two environment creations in the same window share
one directory sweep; judged ledgers accumulate across reaper ticks so a
pass's proofs stay visible to the next.

Admissions: the quarantine refusal fires only where the call is a Managed
admission (executionEngine), never on transcript replay or fork-copy
paths, and refusals are counted and shaped by cause, with shapeless sweep
failures named as sweep-level; conversations no longer freezes a runtime
where a retry is followed by a liftable Managed quarantine refusal —
bindAndRelease surfaces it as a retryable managed_engine_quarantined
instead.

* test(managed-agent): pin the r2 witnesses for elapsed bounds and addGroup write-fail

Two probes the review asked for: procps's negative-wraparound etime
spelling is rejected outright beside the four-digit ceiling the parser
accepts, and a write that cannot reach the ledger leaves neither the
in-memory map nor the file claiming anything was recorded.

* fix(managed-agent): address review round 3 on the M5c physical-stop slice

Sweep-vs-live-writer race closed for good: a worker the sweep cannot
prove stopped keeps its ledger file to itself — the sweep never reads,
merges, and writes behind a live writer, so a durable addGroup landing
in any window can no longer be lost with the group exiting unsignalled.
The final-truth re-read now adopts a same-pgid record the worker
published over the swept snapshot (new call, new stamps), so a young
replacement group is judged by its own record instead of being read as
'recycled' against the stamps the group its pgid used to name carried.

Quarantine lift is pinned per ledger: the directory reaper attributes
every named group to the file that named it, so a sibling ledger's
proof can no longer lift a vanished never-read ledger whose own truth
was never judged — the mixed-failure batch ends terminal exactly like
the single-ghost sibling.

Witnesses: an inode-stability probe for the no-rewrite law, a
same-pgid replacement adoption probe, and a mixed ghost+sibling reaper
run against a real process group; all three go red under a false flip
and green restored.

* fix(managed-agent): keep undatable process rows instead of dropping them

Review round 5 measured that the wrapped-etime clamp traded a false
'ours' for a false 'gone': parseProcessTable dropped the whole row, so
a group whose only live member carried an undatable etime read as
empty — silently resolved 'gone', ledger deleted, group left running,
witnessed sweep included. The row now stays in the table with its age
undatable: it still counts as a member (never 'gone'), but no age
judgement can be made from it — it cannot prove 'ours', it cannot
prove 'recycled' via a young leader, and a worker or host row whose
age cannot be read is held ('unknown' / hold) rather than decided.

The R1-36 witness that made writes fail with chmod(555) now fails them
through the writeFileSync mock instead — root runners ignore mode bits,
so the same red ran on their clean tree.

Witnesses: parser keeps undatable rows with their age undefined;
an unwitnessed sweep holds (never signals, never unlinks) a group
whose only member is undatable; a witnessed sweep still SIGKILLs it;
the row-drop mutation turns the parser witness red before restore.

* test(managed-agent): pin the r5 witnesses the review asked for

Five behaviours shipped in the round-2 batch now have their own
probes, each red under its reverting mutation and green restored:

- killOutstanding consults the process table once for its whole
  fanout: a blocking-ps counter proves three outstanding groups cost
  the prune's consult plus exactly one shared read, never one per
  waiter (R1-12);
- the sweep's proof deadline rides the monotonic clock: a 10 s
  backward wall step inside the first proof poll leaves the measured
  wait at its 600 ms budget instead of stretching to ~10.6 s (R1-48);
- close() joins the sweep the exit hook is already running over the
  same ledger: the wrapped sweep records exactly one pass for the two
  triggers (R1-45, the mock gains a per-call hold so the overlap is
  scripted, not raced);
- two environment creations in one sweep window share the
  stale-ledger pass: directory passes are counted at their own entry
  point because file-level calls never cross the module boundary, and
  the seeded ledger is judged exactly once (R1-8);
- the single-ledger reaper keeps the groups an earlier failure named
  across a later nameless one: armed on garbage, named on a valid
  held-group ledger, silenced on garbage again, then the file vanishes
  and the group dies — only the accumulated name lifts, where the
  replace-instead-of-accumulate mutation ends terminal (R1-5).

* style(docs): format 09-event-schema.md so the Prettier lane passes again

Upstream QwenLM#13281 landed the doc with 4 badly wrapped table lines;
its own lint lane ran a 15 s light profile that never scanned it,
so the violation reached main unnoticed and began failing every PR
merge ref computed after 2026-10-07T04:45Z (this PR's run on
13d1735 inclusive). Format-only change: `prettier --write` (4
lines); the whole-tree `--check .` now passes, and upstream QwenLM#13536's
version of the same file is already clean, so the next main merge
stays green by construction.

* fix(managed-agent): address review round R2 on the M5c physical-stop slice

Classification and surfaces: a liftable Managed-engine quarantine
refusal is translated once, at the throw site in bindAndRelease (both
the commit and the release arm), so create and restore alike see the
retryable managed_engine_quarantined StandaloneSessionServiceError —
restore no longer falls through to working_directory_compromised and
its HTTP-409-shaped "retrying is pointless". The create path's own
catch passes the already-translated error through instead of
rewrapping it as a creation-outcome failure. Both StandaloneSession-
ServiceError status ladders (acp-http dispatch and server
error-response) now map the code to 503 like the raw child refusal,
never the default conflict tier, so 503=retry-later stays the daemon's
own convention on both surfaces.

Quarantine summary: an unreadable ledger in an aggregate no longer
erases the names and counts of every sibling the same sweep
identified — concrete entries keep their basename and group count, and
the unreadable clause is appended; the unreadable-only shape is the
same clause alone, never "ghost.json: 0 process group(s)".

Witnesses, each red under its reverting mutation and green restored:
a restore-path load/resume whose commit arm refuses twice is reported
code managed_engine_quarantined, retryable: true, with the refusal as
cause (raw-rethrow flip reds); a 503 case on each mapper (tier-removal
flip reds); the named-entries-survive and unreadable-only summary
wordings (clause-removal flip reds both); the failed-launch inline
sweep — a worker that boots, writes a ledger naming a live group, then
fails attestation — leaves the group dead and the file gone (sweep
deletion reds); and the launchedLedgerPaths wiring — a second
environment creation leaves the first session's live marker-bearing
worker and its ledger untouched (empty skip set reds).

---------

Co-authored-by: Qwen Autofix <qwen-code@localhost>
Co-authored-by: qwen-code-dev-bot <[email protected]>
Co-authored-by: qwen-code-ci-bot <[email protected]>
Co-authored-by: Qwen-Coder <[email protected]>
Co-authored-by: qwen-code-dev-bot <[email protected]>
Co-authored-by: yiliang114 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants