Skip to content

docs(managed-agent): Stage H4/H5/H6 slice designs - #13499

Merged
wenshao merged 5 commits into
mainfrom
docs/managed-agent-h4-h5-h6-slice-designs
Oct 6, 2026
Merged

wenshao merged 5 commits into
mainfrom
docs/managed-agent-h4-h5-h6-slice-designs

Conversation

@wenshao

@wenshao wenshao commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

What this PR does

This PR adds the Stage H4/H5/H6 slice design documents (English + Chinese twins): 2026-10-04-managed-child-agents.{md,zh-CN.md} (H4 child agents/workflows/teams), 2026-10-04-managed-channels.{md,zh-CN.md} (H5 Channels), and 2026-10-04-managed-automation.{md,zh-CN.md} (H6 automation), following the house design-doc format of the H0a/H0b/H0c pairs. Each document states its slice plan (a/b/c with explicit exit gates), record-body direction, validation plan (fixture parity, mutation checks, fault-injection E2E), acceptance criteria mapped to the reference design's §14 items, and its open questions — all marked proposed-design with nothing implemented and no domain enabled for submission.

Why it's needed

The Stage H tracker (#12827) assigns each capability its own admission/receipt/cancel/quota/recovery checks, and the pinned extension-runtime design defines the target semantics for children, channels, and automation. These documents turn those into sliceable work: each records verified current-state facts (against main@5ddfacc9d4, with naming corrections to the tracker's older snapshot — monitor_run already in the v1 domain index, H1/H2 bodies already enabled, no WakeIntent record by design) so later slices stop re-investigating, and narrows the first vertical slice explicitly (H4: read-only snapshot workspace only, deferring independent worktree/teams/mailbox; H5: email as the reference adapter; H6: scanner lease single-claim, delivery without model rerun).

Reviewer Test Plan

How to verify

Docs-only. The English and Chinese twins have identical section structure (11 sections per pair, verified), reciprocal language links, date conventions matching the house design dirs, and every factual claim about current code is cited with a verified path/source (no claim presented as existing unless it reads out of main). The readme-style pair check in the review guides covers whether both language versions stay in sync per the design-doc README.

Evidence (Before & After)

N/A (documentation). The documents distinguish "verified on main" content from "designed, not yet implemented" content in their status lines and non-goals; factual corrections to the tracker/current snapshot are documented inline (four-part ingress identity records; attachment staging; occurrenceKey shapes).

Tested on

N/A (documentation).

Risk & Scope

  • Main risk or tradeoff: none; documentation only.
  • Not validated / out of scope: no contracts or code in this PR — contract slices land separately (channel contract + envelope contract ship in their own PRs; the child/automation record contracts are currently blocked on reconciling child_run with the H3-registered shell body on main).
  • Breaking changes / migration notes: none.

Linked Issues

Part of #12827 (Stage H design slices) · Refs #12380

中文说明

本 PR 做了什么

本 PR 新增 Stage H4/H5/H6 的双语设计文档:2026-10-04-managed-child-agents.{md,zh-CN.md}(H4 子 Agent/workflow/team)、2026-10-04-managed-channels.{md,zh-CN.md}(H5 Channels)、2026-10-04-managed-automation.{md,zh-CN.md}(H6 自动化),沿用 H0a/H0b/H0c 的格式——切片计划带退出门槛、记录体方向、验证计划、对照参考设计 §14 的验收、以及 open questions;全部标注 proposed-design、无实现、无 domain 放开。

为什么需要

#12827 要求各能力各自具备准入/回执/取消/配额/恢复检查;文档把 pinned 设计落成可拆工作,并记录按 main@5ddfacc9d4 核验过的现状纠正(monitor_run 已在 v1 索引、H1/H2 body 已开、无 WakeIntent 记录),首个垂直切片显式收窄(H4 只读 snapshot、推迟 worktree/team/mailbox;H5 email 参照 adapter;H6 扫描者单 claim、交付不重跑模型)。

评审者验证计划

如何验证

纯文档。每对标题结构 11 节逐一对齐、互链存在、日期符合仓惯例,每项关于现状代码的断言都有已核验出处,未实现项明确标注。

证据(前后对照)

N/A(纯文档)。现状纠正内嵌记录。

已测试

N/A。

风险与范围

  • 纯文档。
  • 未验证/范围外:契约与代码在各自 PR;子/agent 契约当前堵于与 main 已注册 shell 形态的 child_run 对齐。
  • 破坏性:无。

关联 Issue

#12827(阶段 H 设计切片)· #12380

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

评审基于 dad188a936c0b54ece3c10c9c2d141379e4d5ce7,范围为全部六份 H4/H5/H6 双语设计文档,并对照本提交中的 TS/Java 契约及文档固定的外部参考设计。

发现 3 项 P2 设计问题,已逐行评论:H4 与已注册 Shell 正文的兼容方案缺失、前台 child 的结果接受流程与通知约束冲突、H6 定义更新缺少跨修订的调度覆盖水位。建议在据此冻结实现契约前修正;这些是设计层问题,本 PR 未修改生产逻辑。

另有一项非阻塞建议:H4 决策 1 把 child_acceptance 放在 child authority,父侧 accepted/consumed 则放进 child_run;固定参考设计 §2/§5把 child_acceptance 定义为父 authority 的接受回执与消费位置。若这是有意调整,请说明偏离及 relay ACK/资源保留的对应关系;否则建议沿用参考设计的领域归属。

验证:三对文档的标题层级一致、相对链接均可解析;全文对照未发现实质性的中英文翻译遗漏;git diff --check 通过。提交前未发现待核对的已有 Critical。仅做文档与源码静态核验,未运行构建、单元测试或运行时 E2E,也不将上述设计推演当作运行时复现。

Comment thread docs/design/2026-10-04-managed-child-agents.md
Comment thread docs/design/2026-10-04-managed-child-agents.md Outdated
Comment thread docs/design/2026-10-04-managed-automation.md
…n, foreground/background forks, cross-revision watermark

Co-authored-by: Qwen-Coder <[email protected]>

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复审提交 a33b30a202412fa221a6765a9112842f7b6200c8(增量 dad188a → a33b30a),已核对四份更新文档、双语上下文及固定参考设计;H5 两份文档没有变化。

上轮问题 本轮结论
H4 与现有 Shell 正文兼容 已修正设计方向:保留 v1,新增 v2 kind-union,明确各 kind 的身份、任务投影及 TS/Java 版本准入矩阵。
H4 前台结果误走通知输入 部分修正:决策 4 已禁止前台通知;但新文字把结果受理与父侧消费合并,仍有 1 项 P2,见行内评论。
H6 定义更新重跑已执行时隙 上轮反例已修正:跨修订 latestAdmittedSlot 参与 claim/catch-up,旧修订已提交的 slot 不再被新 revision 重新点火。

本轮剩余 1 项 P2。上一轮关于 child_acceptance 领域归属的非阻塞建议尚未调整,仍保持建议级别,不重复新增行内问题。

验证:新改动的中英文语义对齐;三对文档标题层级匹配、相对文件链接有效;全 PR 的 git diff --check 通过。以上为设计与源码静态核验,未运行构建、单元测试或运行时 E2E。

Comment thread docs/design/2026-10-04-managed-child-agents.md Outdated
… the answer, consumed as its own later revision

Co-authored-by: Qwen-Coder <[email protected]>

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复审提交 2c8cfd7466bd83be3758694e4b0d197278f8d98b(增量 a33b30a → 2c8cfd7)。本轮未发现新增 P2 及以上问题;此前报告的 P2 设计问题均已修正。

已按最新正文核实上一轮问题,而非仅依据评论线程的 resolved 标记:H4 决策 4现在将原工具结果与 accepted 一起持久提交,consumed 则由父 authority 在对应消费进度成立后单独推进。两者之间崩溃时,恢复保留“已接受、待消费”;前台仍不生成 notification input 或 wake。H4b 的范围与退出门槛已同步该语义,中英文一致。

同时确认,前两轮补充的 Shell v1/v2 兼容设计与 H6 跨修订 latestAdmittedSlot 规则仍然保留,原先的反例没有回归。本次仅改动 H4 的中英文文档,H5/H6 与上轮评审版本一致。

此前关于 child_acceptance 领域归属与固定参考设计不一致的意见仍为非阻塞建议,建议在 H4 契约定稿时说明归属及 relay ACK/资源保留关系;本轮不重复新增问题。

验证:检查了完整增量及相关决策、切片门槛和验收上下文;三对文档标题层级匹配、相对文件链接有效;全 PR 的 git diff --check 通过。这是文档与契约的静态复审,未运行构建、单元测试或运行时 E2E。

@wenshao

wenshao commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@wenshao

wenshao commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@qwen-code-review-bot

qwen-code-review-bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

✅ Qwen Triage finished — view run. See the stage comments in this thread for the result.

✅ Qwen Triage 已完成 —— 查看运行。结果见本线程中的各阶段评论。

@qwen-code-review-bot

Copy link
Copy Markdown
Collaborator

Sandboxed verification: ⚠️ not run — n/a - workflow run

This PR changes documentation/assets only — there is no code to execute, so a sandboxed verification has nothing to verify.

中文 — 判定:⚠️ 未运行 · 不适用

该 PR 仅改动文档/静态资源,没有可执行的代码,沙箱验证没有验证对象。

— Qwen Code · sandboxed verification

@qwen-code-review-bot

Copy link
Copy Markdown
Collaborator

Thanks for the PR!

Template looks good ✓ — every required heading is present, and the design-doc rule the template calls out (link both language versions, keep them synchronized) is satisfied: all three pairs exist with reciprocal links, and all 19 relative .md links in the six new files resolve at the head commit.

Problem: this isn't a bug fix, so there's no reproduction to ask for — the driver is the open tracker #12827, which assigns each Stage H capability its own admission/receipt/cancel/quota/recovery checks, and AGENTS.md requires a design doc for non-trivial work before the code lands. That's an observed, tracked need rather than theoretical hardening.

Direction: aligned. Stage H is an active roadmap item (#12827 is open and carries roadmap/multi-agent, roadmap/hooks-events, roadmap/background-automation), and these three docs sit squarely inside it. CHANGELOG gives a genuine supporting signal here: upstream is dense with fixes for exactly the failure class these designs target — subagents and scheduled tasks losing state across resume, compaction, restart and sleep, and background agents stalling after a host wake. The premise that an in-process callback, roster or PID is not a recovery credential is well corroborated by how often that breaks upstream.

Size: not applicable. All six files are new docs under docs/design/; no core paths are touched, so there are 0 production logic lines. The 1210 added lines are documentation, which the Stage 0 size gate deliberately does not count, and the author is a maintainer in any case.

Approach: the scope feels right, and — worth saying, because it's the less common direction — these docs actively narrow rather than expand. H4 defers independent worktrees, teams and the mailbox and admits only read_only_snapshot; H5 ports one reference adapter (email) instead of the twelve; H6 refuses webhook triggers at admission until their slice lands. Each capability gets explicit non-goals and per-slice exit gates, which is what makes a tracker like #12827 sliceable. Keeping the three in one PR is defensible since H6 consumes contracts from both H4 and H5.

I did verify the factual claims rather than take them on trust, and they hold up unusually well — the registered-but-not-enabled domain state, the absent record bodies for schedule/automation_run/channel_route/channel_delivery, the missing WakeIntent, the legacy scheduled-task-run.ts path, the public TaskKind enum, the two runtime-broker flags, the twelve-adapter inventory, and the email state store's exact bounds (33 pending / 64 outbound / 1024 dedupe / 256 routes) all read out of the tree as described. One stale item turned up in H4's current-state section; I've written it up in the review comment rather than here, since it's a code-review finding and not a gate concern.

Risk: no elevated risk signals — the Stage 1e high-risk path scan matched nothing.

Moving on to code review. 🔍

中文说明

感谢贡献!

模板完整 ✓ —— 所有必需标题齐备,模板点名的设计文档规则(附上双语版本链接并保持同步)也满足:三对文档均存在且互链,六个新文件里的 19 个相对 .md 链接在 head 提交上全部可解析。

问题:这不是 bug fix,因此无需复现——驱动来源是仍开放的 tracker #12827,它要求 H 阶段每个能力各自具备准入/回执/取消/配额/恢复检查;AGENTS.md 也要求非平凡工作在落代码前先有设计文档。这是已观测、已被跟踪的需求,不是理论性加固。

方向:对齐。H 阶段是活跃路线图项(#12827 开放,带 roadmap/multi-agent、roadmap/hooks-events、roadmap/background-automation 标签),这三篇文档正落在其中。CHANGELOG 也给出真实支撑信号:上游密集出现同类失效的修复——子 agent 与计划任务在 resume、压缩、重启、休眠后丢状态,后台 agent 在宿主唤醒后卡死。"进程内回调、roster 或 PID 不是恢复凭据"这一前提,被上游的故障频率充分印证。

规模:不适用。六个文件均为 docs/design/ 下新增文档,未触及核心路径,生产逻辑行数为 0。新增 1210 行属文档,Stage 0 规模门槛本就不计入;且作者为维护者。

方案:范围合理。更值得说的是——这批文档在主动收窄而非扩张:H4 推迟独立 worktree、team 与 mailbox,只准入 read_only_snapshot;H5 只移植一个参照 adapter(email)而非十二个;H6 在准入处拒绝 webhook 触发直到其切片落地。每个能力都有明确非目标与逐切片退出门槛,这正是让 #12827 这类 tracker 可拆分的关键。三篇放在一个 PR 里也说得通,因为 H6 要消费 H4 与 H5 的契约。

我确实核验了文中事实性断言而非直接采信,结果准确度相当高——domain "已注册但未开放"的状态、schedule/automation_run/channel_route/channel_delivery 缺失记录正文、没有 WakeIntent、legacy scheduled-task-run.ts 路径、公开 TaskKind 枚举、两个 runtime-broker 开关、十二个 adapter 清单,以及 email 状态存储的精确上限(33 pending / 64 outbound / 1024 去重 / 256 路由)全部与代码一致。H4 现状小节有一处过期,我写在代码审查评论里而非此处,因为那属于 review 发现而非 gate 问题。

风险:无升级风险信号——Stage 1e 高风险路径扫描无命中。

进入代码审查 🔍

— Qwen Code · qwen3.8-max-2026-09-02

Reviewed at 5001174978b55388ca52cc408051ee18d32013b6 · re-run with @qwen-code /triage

@qwen-code-review-bot

Copy link
Copy Markdown
Collaborator

Code review

Docs-only, so the review is about whether the documents are correct and self-consistent rather than whether code behaves. I read all six files at the head commit and checked the factual claims against the tree instead of taking them on trust — about twenty of them, including the ones that are easiest to get subtly wrong. They hold up remarkably well, and in a few places better than I expected: the email state store's bounds are quoted exactly right (pending ≤ 33, outboundPending 64, recent 1024, routes ≤ 256), the twelve-adapter inventory under packages/channels matches the directory listing entry for entry, and the claim that child_agent/automation_run are frozen in the public TaskKind is correct — it refers to the generated API contract's enum, not the unrelated internal TaskKind in agents/tasks/types.ts. The runtime-broker.trusted-local-reboot-recovery flag is real too; it lives as a Java property, not in the TypeScript tree.

One thing did turn up, and it's worth fixing before merge because of where it sits.

H4's current-state section is stale on child_run, and contradicts the same document. The "Domain index" bullet says that of child_run, child_acceptance and the four team_* names, "none of them has a record body in MANAGED_EXTENSION_RECORD_BODIES". On today's main that is false for child_run — it has a body at managed-extension-projection.ts:151, with taskKind: 'background_shell' and recordId = record.shellId. I checked both refs to be sure this is drift rather than an error: at the baseline the doc pins (5ddfacc9d4) the bodies were only the four MCP/Hook ones plus monitor_run, so the bullet was accurate when written; H3 landed the shell body afterwards. The document's own "Record bodies (H4a contract direction)" section then states the current fact correctly — "managed-child_run is already registered on main with the H3, single-kind body" — and the H3 module header it cites really does say "H4 extends the domain to the other child kinds under its own body version". So the file asserts both things, each true against a different commit.

That would be a harmless inconsistency in most sections, but the current-state section is precisely the part whose stated job is to stop later slices re-investigating — and the registered shell body is the single most consequential constraint on H4a. It is the whole reason H4a must define body version 2 as a kind-union with kind: 'shell' kept byte-identical, instead of writing a v1 body. A reader who trusts current state alone would design H4a the wrong way and break replay for every committed Shell. The Chinese twin carries the identical line (and the identical correct statement further down), so this is baseline drift in both languages rather than a translation-sync gap — the fix needs to land in both.

Two smaller items of the same kind:

  • "monitor_run has a body and still stays disabled until H3" — it does have a body and it is still absent from MANAGED_SESSION_ENABLED_DOMAINS, so both halves are right, but "until H3" now reads as though H3 were future work when H3 has landed.
  • "a Flyway migration after main's V34" appears in all three validation plans (six places across both languages). V34 was the highest migration at the pinned baseline; main is at V46 now. Since these lines will be read when H4a/H5a/H6a actually land, the number understates by twelve. The historical citation in H4 ("O4 retirement (feat(managed-agent): Protect Session-owned tool output retirement #13084, V30) and collection (V34)") is fine as-is — that records what V34 did rather than instructing anyone.

The cheapest honest fix is to keep the pinned baseline and add a short note where the doc already speaks about "main" unscoped, so current state and the record-body section agree. Re-verifying and refreshing the baseline wholesale would also work but is more churn than the problem needs.

None of this blocks: there is no runtime surface here, every document is explicitly marked proposed-design with nothing implemented and no domain enabled, and the H4a body-version-2 design that the stale line obscures is itself correct and well reasoned. Structural requirements are all met — eleven sections per pair in identical order and heading levels across all three pairs, reciprocal language links, and all nineteen relative links resolving. The byte volumes also confirm the Chinese versions are full translations rather than summaries (17.5 KB against 19.9 KB, 17.9 against 20.2, 21.2 against 24.9 — the lopsided line counts are just the English being hard-wrapped at ~76 columns while the Chinese runs one line per paragraph), which is what the design README requires.

Testing evidence

This is an unattended CI run, so per the gate rules I did not build or execute anything from the PR — the evidence below is the PR's own CI, read through the API for the reviewed commit. Nothing in this diff is user-visible (six new Markdown files, no code, no TUI surface), so real-scenario terminal testing is N/A, and there is no behavioural claim for a sandboxed lane to settle — @qwen-code /verify and @qwen-code /tmux would have nothing to observe. Not verified: the documents' forward-looking design decisions (body version 2, the four-part ingress identity, the scanner lease) are proposals, so there is nothing to test yet; that is by design and the docs say so.

The PR's Qwen Code CI workflow run completed successfully, so no checks were pending at review time. The skips are the path-filtered jobs — macOS/Windows unit legs, CLI integration and the web-shell E2E smoke — which is the expected shape for a docs-only diff; the Linux unit and lint legs that do run are green. review-pr was still in progress, but that is bot orchestration on the comment event, not PR CI, and does not gate this.

Check Conclusion
Test (ubuntu-latest, Node 22.x) success
Lint & Static (ubuntu-latest, Node 22.x) success
Integration Tests (no-AK, No Sandbox) success
Desktop Shell (ubuntu-22.04) success
Desktop Shell (windows-2022) success
Test (macos-latest, Node 22.x) skipped
Test (windows-latest, Node 22.x) skipped
Integration Tests (CLI, No Sandbox) skipped
web-shell E2E Smoke (ubuntu-latest, Node 22.x) skipped
Classify PR / assign / authorize / label success
review-pr in_progress (bot orchestration, not PR CI)
中文说明

代码审查

纯文档,因此审查重点是文档是否正确、是否自洽,而非代码行为。我在 head 提交上读完六个文件,并把事实性断言逐条对照代码树核验,而不是直接采信——大约二十条,包含最容易出细微错误的几条。结果准确度很高,某些地方甚至超出预期:email 状态存储的上限引用完全正确(pending ≤ 33、outboundPending 64、recent 1024、routes ≤ 256);packages/channels 下十二个 adapter 的清单与目录逐项一致;"child_agent/automation_run 冻结在公开 TaskKind 中"也成立——指的是生成的 API 契约枚举,而非 agents/tasks/types.ts 里那个同名的内部类型。runtime-broker.trusted-local-reboot-recovery 开关同样真实存在,只是它是 Java 侧属性,不在 TypeScript 树里。

但确实查出一处,且因为它所在的位置而值得在合入前修掉。

H4 现状小节关于 child_run 已过期,并与同一篇文档自相矛盾。"Domain 索引"那条写着:child_run、child_acceptance 与四个 team_* 名字"在 MANAGED_EXTENSION_RECORD_BODIES 中都没有记录正文"。在今天的 main 上,这对 child_run 是错的——它在 managed-extension-projection.ts:151 有正文,taskKind: 'background_shell'、recordId = record.shellId。为确认这是漂移而非笔误,我核对了两个 ref:在文档钉住的基线(5ddfacc9d4)上,正文只有四个 MCP/Hook 加 monitor_run,所以那条当时是准确的;H3 之后才落地 shell 正文。而同一篇文档的"记录正文(H4a 契约方向)"小节又正确陈述了当前事实——"managed-child_run 已在 main 上按 H3 的单 kind 正文注册"——它引用的 H3 模块头注释也确实写着"H4 用自己的 body 版本把该 domain 扩展到其他 child 种类"。于是这份文件同时断言了两件事,各自对一个不同的提交成立。

放在别的小节这只是一处无害的不一致,但现状小节恰恰是其声明职责为"让后续切片不必重复调查"的那部分——而已注册的 shell 正文正是 H4a 最关键的一条约束。它就是 H4a 必须把正文版本 2 定义成 kind-union、并让 kind: 'shell' 逐字节保持不变的全部理由,而不是去写一个 v1 正文。只信现状小节的读者会把 H4a 设计错,从而破坏每个已提交 Shell 的重放。中文版本带有完全相同的那句(以及下方同样正确的陈述),因此这是两种语言共同的基线漂移,而非翻译不同步——修复需要同时落在两侧。

另有两处同类小问题:

  • "monitor_run 有正文,但在 H3 之前同样不开放"——它有正文、也确实仍不在 MANAGED_SESSION_ENABLED_DOMAINS 中,两半都对,但"在 H3 之前"如今读起来像 H3 仍是未来工作,而 H3 已落地。
  • "main 的 V34 之后的一个 Flyway 迁移"出现在三份验证计划里(双语共六处)。V34 是钉住基线上的最高迁移号;main 现在是 V46。既然这几行会在 H4a/H5a/H6a 真正落地时才被读到,这个数字少了十二。H4 里的历史引用("O4 引退(feat(managed-agent): Protect Session-owned tool output retirement #13084,V30)与回收(V34)")保持原样即可——那记录的是 V34 做了什么,不是给谁的指令。

最省事的诚实修法是保留钉住的基线,只在文档已经不限定地说"main"的地方补一句短说明,让现状小节与记录正文小节一致。整体重新核验并刷新基线也可行,但对这个问题的规模来说是多余的改动。

这些都不构成阻塞:这里没有运行时面,每篇文档都明确标注为设计提案、无实现、无 domain 开放;而被那句过期陈述遮住的 H4a 正文版本 2 设计本身是正确且论证充分的。结构性要求全部满足——三对文档每对十一节、顺序与标题层级完全一致,互链齐备,十九个相对链接全部可解析。字节量也证明中文版是完整翻译而非摘要(17.5 KB 对 19.9 KB、17.9 对 20.2、21.2 对 24.9——行数悬殊只是因为英文按约 76 列硬折行,而中文一段一行),这正是设计文档 README 的要求。

测试证据

这是无人值守的 CI 运行,因此按 gate 规则我没有构建或执行 PR 里的任何内容——下面的证据是 PR 自己的 CI,通过 API 针对被审提交读取。本 diff 没有任何用户可见面(六个新 Markdown 文件,无代码、无 TUI),因此真实场景终端测试为 N/A;也没有需要沙箱通道来判定的行为性断言——@qwen-code /verify 与 @qwen-code /tmux 在此无可观测对象。未验证:文档中前瞻性的设计决策(正文版本 2、四段式 ingress 身份、扫描者租约)都是提案,因此尚无可测之物;这是设计使然,文档也已写明。

PR 的 Qwen Code CI workflow run 已成功完成,因此审查时没有待决检查。跳过的是按路径过滤的任务——macOS/Windows 单测腿、CLI 集成、web-shell E2E smoke——这对纯文档 diff 是预期形态;确实运行的 Linux 单测与 lint 腿均为绿。review-pr 仍在进行,但那是 comment 事件上的机器人编排任务,不属于 PR CI,不构成门禁。

— Qwen Code · qwen3.8-max-2026-09-02

Reviewed at 5001174978b55388ca52cc408051ee18d32013b6 · re-run with @qwen-code /triage

@qwen-code-review-bot

Copy link
Copy Markdown
Collaborator

Confidence: 4/5 — accurate, well-scoped design groundwork that meets every structural requirement; the one real defect is a stale current-state line that is non-blocking and cheap to fix.

Before reading the diff I wrote down what I'd expect from a tracker-driven design slice: one document per capability, bilingual twins, a current-state section pinned to a named commit so nobody re-investigates, a slice plan with exit gates that actually gate, and non-goals that say what's being deferred. This PR matches all of that, and it exceeds my baseline in one respect I would not have anticipated — the H4a body-version-2 kind-union. Discovering that managed-child_run already carries an H3 single-kind body, and designing so that kind: 'shell' stays byte-identical while kind: 'child_agent' gets the new body with recordId = childRunId, is the most valuable thing in the diff. You only find that by reading the tree rather than the tracker, and it stops H4a from being designed in a way that would break replay for every committed Shell. The domain.committed payload version widening to 1 | 2 for that one domain only, with every other domain held at version == 1, is the right amount of care.

The scope discipline is what I'd point at as the reason this doesn't read as volume. Three capabilities could easily have become three speculative platform designs; instead H4 admits only read_only_snapshot and refuses the other two isolation names at admission until their slice lands, H5 ports one adapter out of twelve and names why email is the right first one (its legacy state already separates mailbox generation, event identity, in-flight uncertainty and reply routes — the four concepts the contracts need), and H6 refuses webhook triggers rather than half-designing them. Every doc says plainly that nothing in it is implemented and no domain is enabled. Six files, zero deletions, no drive-by edits, nothing unrelated — the diff is exactly the stated goal.

My reservation is the one I wrote up in the review: H4's current-state bullet still says child_run has no record body, which was true at the pinned baseline 5ddfacc9d4 and is false on main now that H3 has landed, while the same file's record-body section states the current fact correctly. I'd normally weight a self-contradiction more heavily, but here the mechanism is benign and the correction is two lines per language — and notably the section that's stale is stale in the safe direction, since the accurate section is the one that carries the design consequence. The monitor_run ... until H3 phrasing and the six "after main's V34" references are the same drift and want the same touch. Worth folding into whichever PR picks up H4a rather than respinning this one.

On the questions I'm supposed to ask myself: this is groundwork for an open, labelled roadmap tracker, so the "does anyone want this" test is answered by #12827 existing. In six months the pinned baselines are what will make these documents trustworthy rather than archaeological, which is exactly why the one stale line is worth naming. I checked whether I was being worn down by volume or approving for want of objections — I found a genuine defect, traced it to the commit where it became wrong, and it doesn't block, so the approval is a judgement rather than a shrug. CI is green on the reviewed commit with nothing pending.

Approving. The three staleness items above are non-blocking follow-ups, not conditions.

中文说明

Confidence: 4/5 —— 准确、范围克制的设计前置工作,满足全部结构性要求;唯一的实质缺陷是一处过期的现状陈述,不阻塞且修复成本很低。

在读 diff 之前我先写下了自己对"tracker 驱动的设计切片"的预期:每个能力一篇文档、双语孪生、现状小节钉在具体提交上以免后人重复调查、带真正能拦住的退出门槛的切片计划,以及说明推迟了什么 non-goals。这个 PR 全部符合,并且在一点上超出我的基线——H4a 的正文版本 2 kind-union。发现 managed-child_run 已带有 H3 的单 kind 正文,进而设计成让 kind: 'shell' 逐字节不变、kind: 'child_agent' 以 recordId = childRunId 携带新正文,是这份 diff 里最有价值的部分。这只有去读代码树而不是读 tracker 才能发现,而它阻止了 H4a 被设计成破坏每个已提交 Shell 重放的样子。domain.committed payload 版本只对这一个 domain 放开到 1 | 2、其余全部维持 version == 1,分寸也拿得对。

范围克制正是它读起来不像"堆量"的原因。三个能力很容易写成三份投机的平台设计;但 H4 只准入 read_only_snapshot,另外两个隔离名在其切片落地前于准入处拒绝;H5 在十二个 adapter 里只移植一个,并说清为何 email 是正确的第一个(它的 legacy 状态已经分开了邮箱代数、事件身份、in-flight 不确定性与回复路由——契约需要的四个概念);H6 拒绝 webhook 触发,而不是半个设计。每篇都明确写着:本文无任何实现、无 domain 开放。六个文件、零删除、无顺手改动、无无关内容——diff 恰好等于所述目标。

我的保留意见就是审查里写的那条:H4 现状那条仍说 child_run 没有记录正文,这在钉住的基线 5ddfacc9d4 上成立,在 H3 已落地的 main 上不成立,而同一文件的记录正文小节又正确陈述了当前事实。我通常会更重视自相矛盾,但这里成因无害、修正是每语言两行——而且值得注意的是,过期的小节是朝安全方向过期的,因为承载设计后果的那一节是准确的。monitor_run ... 在 H3 之前的措辞与六处"main 的 V34 之后"属同一类漂移,需要同一顺手修正。更适合并入接手 H4a 的那个 PR,而不是为此重做本 PR。

关于我该自问的那些问题:这是一个开放且带标签的路线图 tracker 的前置工作,所以"有没有人要"由 #12827 的存在回答。六个月后,正是这些钉住的基线会让文档可信而非变成考古材料——这也正是要点名那一处过期陈述的原因。我也检查了自己是否被数量磨平、或是否因为找不出反对意见就放行——我发现了一个真实缺陷,追到了它开始出错的那个提交,而它不构成阻塞,所以这次放行是判断而非敷衍。被审提交上 CI 为绿,无待决检查。

予以批准。上述三项过期问题是非阻塞的后续项,不是合入条件。

— Qwen Code · qwen3.8-max-2026-09-02

Reviewed at 5001174978b55388ca52cc408051ee18d32013b6 · re-run with @qwen-code /triage

@qwen-code-review-bot qwen-code-review-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, looks ready to ship. ✅

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed head: 5001174978b55388ca52cc408051ee18d32013b6 (base main).

Approve. No historical blocking finding exists on this PR, and the Critical-only scan found nothing blocking. For a docs-only change the only Critical class that applies is mis-certification — a document asserting that current code behaves a way it does not — so that is what I checked, against live main rather than against the PR's own prose.

Historical blocking findings

None to re-verify. The PR carries eight reviews: seven COMMENTED author notes across the intermediate commits and one qwen-code-review-bot APPROVE filed against this exact head. All four review threads are resolved, there is no review ledger and no [Critical] inline comment anywhere on the PR, and reviewDecision is REVIEW_REQUIRED — no CHANGES_REQUESTED was ever filed. The bot's approval was not substituted for verification; the claims below were checked independently.

Critical-only scan

The diff is six new files under docs/design/ (+1210/−0): three English documents and their Chinese twins. No production code, no contract, no schema, no migration, no workflow — so there is no runtime surface that could regress, and the reviewable question is whether the documents are accurate about the code they describe.

The structural claim holds exactly. The PR asserts the twins have "identical section structure (11 sections per pair, verified)". At this head all six files carry one H1 and ten H2 headings, and the ten correspond one-to-one in order across every pair: Problem/问题, Current state/现状, Goals/目标, Non-goals/非目标, Decisions/决策, Record bodies (H4a/H5a/H6a contract direction)/记录正文, Slice plan/切片计划, Validation plan/验证计划, Acceptance criteria/验收标准, Open questions/未决问题. The reciprocal language link sits on line 3 of all six files and points at both twins. The three-way size asymmetry (283/294/337 English lines against 92/101/103 Chinese) is a summarising twin, not a divergent one — the section structure the claim is about is identical.

The "verified on main" claims check out against live main. These are the load-bearing facts, because a later slice is meant to stop re-investigating by trusting them:

  • monitor_run is already in the v1 domain index. Confirmed — it is in MANAGED_SESSION_DOMAINS in packages/core/src/managed-runtime/managed-session-records.ts, and that file's own header comment records it as parsing in readers since #12837. The document's correction to the tracker's older snapshot is right.
  • H1/H2 bodies are already enabled. Confirmed — MANAGED_SESSION_ENABLED_DOMAINS contains mcp_configuration, mcp_operation, hook_registration and hook_execution.
  • No H4/H5/H6 domain is enabled for submission. Confirmed — child_run, child_acceptance, team_state, team_task, team_message, team_plan, channel_route, channel_delivery, schedule and automation_run are all present in the closed registry and all absent from the enabled set. This is enforced, not merely conventional: assertManagedSessionDomainEnabled throws domain … is registered but not enabled for submission, and the source comment states the distinction the documents rely on ("recognising a name never means the capability is implemented or admitted, so submission is gated separately"). So the documents' repeated "proposed-design, nothing implemented, no domain enabled" status line is accurate.
  • Cited paths exist. packages/cli/src/runtime/scheduled-task-run.ts and packages/core/src/managed-runtime/managed-session-inbox.ts both resolve on main.
  • No WakeIntent record. Confirmed — no such name appears in the records module, matching the document's statement that there is none by design.

One incidental note that is not a finding and does not affect this approval: the automation_run string also appears in MANAGED_SESSION_ACTION_SOURCES, a separate list of action sources. It is not a second enabled-domain entry, and nothing in the documents conflates the two.

Nothing in the documents could mis-certify shipped behaviour, because each one states its own status line as proposed-design and scopes its non-goals explicitly. The risk this class of PR usually carries — a design document describing a planned capability in the present tense, so a reader believes it ships — is avoided here by the status lines and by the fact that the registry/enabled split in code independently enforces the inertness the documents claim.

Not scanned — disclosed, not asserted clean

I verified the structural claim across all six files and the factual claims about main listed above. I did not read all 1,210 lines line by line, did not verify the acceptance criteria against §14 of the reference design, and did not fetch the pinned external reference-design commit the documents cite, so I am not attesting that the section mapping to that reference is complete or that the pinned link resolves. I did not audit the slice-plan exit gates for achievability — that is a planning judgement, not a correctness one. I report no Critical in those areas because I found none where I looked, not because I proved absence.

CI

No failing checks at this head: 12 pass, 27 skipped, 1 pending. The passing set includes Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox) and Desktop Shell on both ubuntu and windows. The single pending job is the automatic review-pr, which per this channel's rules is summarised and never treated as a gate. Nothing to attribute to this PR.

Scope note

Approval is bound to commit 50011749. It covers the documents as documentation: their internal consistency, their twin parity, and the accuracy of the current-state facts they assert about main. It is not a judgement that the proposed H4/H5/H6 designs are the right decomposition, and it does not pre-approve any contract or implementation slice that follows from them.

@wenshao
wenshao added this pull request to the merge queue Oct 6, 2026
Merged via the queue into main with commit 4ce35fc Oct 6, 2026
58 of 59 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants