Repository navigation
feat(managed-agent): Implement private Hosted MCP runtime (H1) - #12946
Conversation
E2E and pre-commit verificationVerified on macOS with Node.js 22.22.2, JDK 21.0.11 and an isolated MariaDB 10.11.19 instance. The locally built Hosted CLI used the Java Broker and provisioned Runtime with real stdio and Streamable HTTP MCP servers; only the model was a deterministic fixture. Credentials and test records were synthetic.
The first expanded run used a stale Broker dependency embedded in the Spring JAR and failed the revoked-access close case. After reinstalling the Broker dependency and a clean server package, all 81 embedded Broker classes matched the current compiled classes; the complete E2E above passed. Final CLI SHA-256: Additional checks passed: root build/typecheck/bundle; 904 CLI + 346 Core + 157 Broker + 43 Java server tests; changed-file ESLint/Prettier; staged diff checks. Independent failure reproductions verified both dispatch/cancel race orderings, durable replay, real 25-second timeout followed by native cancellation and late settlement, discovery invalidation during listing, original-owner recovery and grant/effect identity isolation. Repeated open-ended and targeted audits found defects that were fixed before commit. The final two consecutive passes found no remaining actionable issues in snapshot Windows/Linux and the full production restart matrix were not exercised locally. Full SQL E2E covers stdio and Streamable HTTP; SSE, replacement, quotas and cancellation have focused transport/contract coverage. Production profile enablement and SDK reverse clients remain outside this private H1 slice. |
Real-stack verification — H1 Hosted MCP runtime @
|
| # | Trigger (ordinary event) | Observed | Scope |
|---|---|---|---|
| F1 | An MCP Session stays attached between turns | It keeps the Workspace execution lease, so other Sessions in the Workspace can't run tools | Workspace |
| F2 | */list_changed from the server, or an SSE server restart |
Every later MCP call gets managed_mcp_catalog_conflict, 0 effects, until someone reconfigures by hand |
Session MCP |
| F3 | Cancel of a running resource_read or prompt_get |
outcome_unknown for good; prompt 409, detach 503, lease held |
Session + Workspace |
| F4 | Prompt cancel during an MCP tool, or a tool taking more than 25 s | Broker execution UNKNOWN; a late success doesn't heal it; detach 503 |
Session + Workspace |
| F5 | Streamable HTTP server unreachable at close | detach 503, and still 503 after the server comes back | Session + Workspace |
| F6 | Each reconfiguration | The previous connection stays open; the 17th reconfigure hits the quota | Session MCP |
| F7 | First install refused (Workspace busy) | Detach 503 until another prompt is sent | Session |
F1 — An attached MCP Session holds the Workspace lease between turns. HostedWorkspaceToolTurn skips broker.release() for the MCP profile, and HostedMcpSession releases only in close(). I put three Sessions in one Workspace; the outcome is the same on both databases:
- B (MCP profile), while A sits idle but attached:
POST /prompt→ 503hosted_prompt_admission_failed(the Broker acquire returned 409). - C (files profile):
turn_error, and nothing is written. - After A detaches, B works, and then B blocks C in the same way.
So one attached MCP Session makes its Workspace unusable for tools in every other Session. Combined with F3–F5, one stuck Session locks the Workspace indefinitely.
F3/F4 — Cancel, or a call longer than 25 s, leaves the Session and its Workspace stuck. Per the MCP spec, the SDK server drops the response to a request it was told to cancel. I checked the success and error paths in shared/protocol.js. So after notifications/cancelled no settlement ever comes back.
| Case | Server ledger | Result |
|---|---|---|
Cancel a resource_read (HTTP on MySQL, SSE on MariaDB) |
Handler aborted within 1.5 s | outcome_unknown 12 s later. Status recoveryBlocked, next prompt 409, detach 503, lease held, a fresh Session in the same Workspace 503. |
| Prompt cancel while an MCP tool runs | Aborted 0.5 s after the cancel | The turn gets no terminal event. Same 409 / 503 / held / 503. |
| A 32 s tool, no cancel | Server answered at 41 s | The Runtime timed out at 25 s (TIMEOUT_MS, hard-coded; the manifest rejects a timeout field). |
The late reply in the 32 s case does settle the Runtime's own view. But qwen_tool_execution.execution_state stays UNKNOWN and is never re-queried, so the Session never heals. Close runs mcp-release 200, then tool-sessions/mcp:<sid>:release returns 409, and detach returns 503. For comparison, ordinary MCP uses MCP_DEFAULT_TIMEOUT_MSEC = 10 min, configurable per server.
F2 / F5 / F6 / F7 — connection lifecycle after the first turn.
-
F2. A
list_changednotification forresourcesortoolsbumps the singlecatalogRevisionthat every call kind is checked against, and the Harness never dispatchesmcp-discover. From then on:- the next tool call in the same turn gets
catalog_conflict; - so does every later turn;
- so does
prompt_get; - and the Harness catalog still says revision 1,
complete.
An SSE server restart has the same effect, because
onerrorretires the connection. The SDK'sMcpServeremitslist_changedwhenever a tool is registered, enabled or disabled after connect, so this trigger is not exotic. - the next tool call in the same turn gets
-
F5. Detach returns 503 within 1 s: 678 ms on MySQL, 228 ms on MariaDB. After the server comes back, detach returns 503 again. The configuration stays
releasingand the lease stays held.terminateSession()throws when the DELETE cannot be sent, although that connection had 0 pending operations. -
F6. Each
POST /mcp/configurationsleft the previous stdio process running: 1 → 16 live processes for one server in one Session. The 17th reconfiguration got 409. Only detach reaped them, and that took 6.9 s. This is stage-2 point 4 of the triage bot's review, now measured. Together with F2, the documented recovery path works about 15 times per Session. -
F7. After a refused first install,
statusreportsrecoveryBlocked:truewhile/promptstill admits. Detach stays 503 even after the Workspace frees up. It succeeds only after a later prompt dispatches the leftover configuration intent.
Possible directions (the choice is the maintainers'):
- F1: either stop holding the storage lease between turns for MCP owners, or document "one attached MCP Session per Workspace" as an H1 limit.
- F2: dispatch the existing
mcp-discoveroncatalog_conflict/stale, and reconnect with a new generation after a transport error. - F3: for
resource_read/prompt_get, settle a cancelled operation as a non-blockingcancelled, since its result is data only. - F4:
- make the timeout a manifest field;
- either don't send
notifications/cancelledfor tool calls, or bound the wait and settle them; - re-query the execution status on close or load, so a late success is reconciled.
- F5: with 0 pending operations, treat a failed or
405DELETE as released. - F6: close a retired connection once its pending set drains.
- F7: have
close()locally settle configuration intents that are still atexecution: intent, asnot_started_proven, the same way operations are handled.
Gates
- CI.
- At
a390b5d5,Test (ubuntu)was red:process-env-guard.test.tsflaggedQWEN_MANAGED_MCP_CONFIG,PATHandSystemRoot. - The author fixed it in
e3ecba66, and it is now green: 22 pass, 2 pending. - Locally the guard passes 4 of 5 runs; the one failure was at a load average near 100.
- At
- Main. The PR is
CONFLICTINGafter feat(serve): add gated Hosted foreground Shell turns #12848 landed: 8 files, 17 hunks in the 6 production files (8 of them inhosted-workspace-tool-turn.ts). That file is where F1 and F4 live, so the post-merge head needs a re-run of the relevant scenarios. - Merge order with feat(managed-agent): Add durable remote Shell result delivery #12894 (
a7c0a391).- Both PRs add
V19(V19__managed_mcp_records.sqlandV19__managed_tool_publication.sql), and git reports no conflict on those files. - The trial-merged jar on MySQL 8.4 fails at startup with
Found more than one migration with version 19. - The trial merge also has 8 textual conflicts.
- Whichever PR lands second has to renumber its migration.
- Both PRs add
- Local suites at
a390b5d5.- cli focused: 79/79. core focused: 230/230. managed-agent-server: 180/180. runtime-broker: 395/396.
- The one Broker failure,
HttpRuntimeTransportTest#closesTheConnectionOnTheDeadlineAndOnCallerCancel, also fails onmainbd45b95fon this host under load, and passes 5/5 alone.
- Mutation. Java: 8/8 killed. TS: 13/16 killed. Three survivors point at untested behaviour:
- M7: a
-32601from a tools-only server should mean an emptycompletelist. - M11: the prompt route's
hasPendingOperations()check. This guard returned the 409 in F3, yet no unit test pins it. - M13: configurations owned by two Runtime Sessions.
- M7: a
Smaller notes
- Identity conflicts on
/mcp/operationsand/mcp/configurationsreturn 503, which looks retryable. A 409 would fit better. - Model-facing tool names are
mcp_<32 hex>. They include the catalog revision and connection generation, so every name changes on each load. After a reload, the transcript refers to tools that are no longer advertised. - The Runtime's error receipts reach the model as bare codes, e.g.
managed_mcp_catalog_conflict. - stdio servers run with
HOMEset to the Workspace directory (observed:HOME = cwd = <ws>/child). Anything a server caches under~therefore lands in the user's Workspace.
Not verified
- Windows and Linux.
- A real model; like the author, I used a fixture model.
- Whether a restarted Harness can recover a Session stuck by F3–F5.
- The 32 in-flight quota.
- Interaction with feat(serve): implement generic Broker provider controls #12868, which the triage bot flagged.
- The post-merge head, after resolving feat(serve): add gated Hosted foreground Shell turns #12848.
Harness, scenarios, mutants and raw results: wenshao/qwen-code@a5ef4914/pr12946
中文版
真实环境验证 —— H1 Hosted MCP runtime @ e3ecba66
合并参考(一段话): 在本地构建的真实环境中,本 PR 声称的内容在 MySQL 8.4.11 与 MariaDB 10.11.18 上均成立。其中包括 PR 只有单测覆盖的 SSE,也包括作者 E2E 说明里的每一条恢复主张(均独立重跑)。我在 a390b5d5 发现的 CI 红腿(process-env guard)已由作者在我验证过程中修复。目前仍不建议直接合入,原因有三:
- 现在与
main冲突。 feat(serve): add gated Hosted foreground Shell turns #12848 已合入,与本 PR 在 8 个serve/文件上冲突,其中 6 个生产文件共 17 个冲突块。 - 与在飞的 feat(managed-agent): Add durable remote Shell result delivery #12894 撞 Flyway
V19。 git 不会提示,合并后的服务端无法启动。 - 五个行为问题(F1–F5)。 普通事件就会让 Session 卡死或 MCP 不可用,经由 F1 还会锁住整个 Workspace。这些事件包括:用户取消、超过 25 s 的工具、MCP 服务端宕机或重启、
list_changed通知。
该 profile 是私有且需显式启用的,维护者可以选择把 F1–F5 记为 H1 已知限制并另开后续 issue。但我认为 F3–F5 不应带入任何生产启用。
跑了什么
- 被测构建。
a390b5d5的 bundle(dist/cli.js)加内嵌 Broker 的 Spring jar。验证途中 head 前进到e3ecba66,只在process-env-guard.test.ts增加了 16 行,其余路径 diff 为空,因此以下运行时结果全部适用于e3ecba66。 - 链路。 打包后的 Hosted Harness(
qwen serve --profile hosted-harness)→ 内嵌 Java Broker(开启 session store)→ 由部署 manifest(QWEN_MANAGED_MCP_CONFIG)拉起的 Runtime worker → 基于仓库自带@modelcontextprotocol/sdk1.30.0 的真实 MCP 服务端:- stdio,作为 worker 的子进程;
- Streamable HTTP,带 Bearer 头;
- SSE,带 Bearer 头。
- 观测手段。 只有模型是假的。Broker 流量经计数、可注入故障的代理;每个物理 MCP 请求都写入 JSONL 台账。测试专用部分只有一个 Java main,用于预置 Workspace 并提供目录接口所需的 actor。
- 环境。 MySQL 8.4.11 与 MariaDB 10.11.18(Docker)、Node 22.23.2、JDK 25(release 21)。
- 范围。 11 个场景、24 个变异体(16 个 TS + 8 个 Java)、定向 TS/Java 测试套件,以及与 feat(managed-agent): Add durable remote Shell result delivery #12894 的试合并。
✅ 真实链路上确认成立
- 一轮工具调用覆盖三种 transport。 恰好 3 次物理副作用。第二次模型请求时观察到的检查点阶段为
results_ready。 - 资源与提示操作。
- 二进制
resource_read逐字节一致(300 字节)。 - 经 SSE 的
prompt_get两条消息都在。 - 同一
operationId重复请求返回相同响应。 - 用不同 URI 或不同 server 复用该 ID 会被拒绝,额外副作用为 0。
- 二进制
- 目录。 私有与公开目录都不含 token、
Bearer、URL、command 或 env(检查 10 个关键串,0 命中)。MCP 记录的task_kind为 NULL,公开/tasks列表为空。 - detach 与 load。 detach 释放全部 3 个连接和 Broker tool-session。load 生成新 owner
mcp:<sid>:<digest>,重放已提交的 prompt 结果与原结果一致。 - 作者的恢复主张全部复现(triage 机器人把这些列为"未验证"):
- 丢回包。 同时丢弃 configure 回包和 invoke 回包,经原 ID 的
mcp-status恢复,物理读取恰好 1 次。 - 撤销权限。
can_read=FALSE后 close 返回 204、lease 已释放;新工作被拒(Broker acquire 409)。 - Broker 不可用。 reload 之后 Broker 对所有请求返回 503:已提交结果重放返回 202,close 返回 204,Broker 请求 0 次、副作用 0 次。
- stale 目录。 stale 目录从未授权过调用:下文 F2 中所有被拒调用的副作用都是 0。
- 丢回包。 同时丢弃 configure 回包和 invoke 回包,经原 ID 的
- 在途调用期间替换配置。 在 r1 上有 6 s 调用在跑时把 r1 替换为 r2,返回 202。该调用在 r1 上完成,下一轮只公布并使用 r2。
- 限额。 65 个工具的服务端被标为
tools: partial,公布 33 个(16 KiB 上限)。70 KB 结果变为managed_mcp_output_limit,Session 保持正常。
问题
| # | 触发(普通事件) | 实测 | 影响范围 |
|---|---|---|---|
| F1 | MCP Session 在两轮之间保持 attach | 一直持有 Workspace 执行 lease,同 Workspace 其他 Session 无法跑工具 | Workspace |
| F2 | 服务端发 */list_changed,或 SSE 服务端重启 |
之后所有 MCP 调用都是 managed_mcp_catalog_conflict、副作用 0,直到手工重新配置 |
Session 的 MCP |
| F3 | 取消运行中的 resource_read / prompt_get |
永久 outcome_unknown;prompt 409、detach 503、lease 不释放 |
Session + Workspace |
| F4 | MCP 工具运行中取消 prompt,或工具超过 25 s | Broker execution 为 UNKNOWN;迟到的成功也不能恢复;detach 503 |
Session + Workspace |
| F5 | close 时 Streamable HTTP 服务端不可达 | detach 503,服务端恢复后仍 503 | Session + Workspace |
| F6 | 每次重新配置 | 旧连接不关闭;第 17 次触发限额 | Session 的 MCP |
| F7 | 首次安装被拒(Workspace 忙) | 在再发一次 prompt 之前 detach 一直 503 | Session |
F1 —— 已 attach 的 MCP Session 在两轮之间一直持有 Workspace lease。 HostedWorkspaceToolTurn 对 MCP profile 跳过 broker.release(),HostedMcpSession 只在 close() 里释放。我在同一 Workspace 放了三个 Session,两种数据库结果一致:
- B(MCP profile),A 空闲但仍 attach 时:
POST /prompt→ 503hosted_prompt_admission_failed(Broker acquire 返回 409)。 - C(files profile):
turn_error,没有写入任何文件。 - A detach 后 B 可以工作,随后 B 以同样方式挡住 C。
也就是说,一个已 attach 的 MCP Session 会让同 Workspace 其他所有 Session 无法使用工具。叠加 F3–F5,一个卡死的 Session 会无限期锁住整个 Workspace。
F3/F4 —— 取消、或超过 25 s 的调用,让 Session 及其 Workspace 卡死。 按 MCP 规范,SDK 服务端不会回复已被取消的请求(我查了 shared/protocol.js 的成功与错误两条分支)。所以发出 notifications/cancelled 之后,永远等不到结算。
| 情形 | 服务端台账 | 结果 |
|---|---|---|
取消 resource_read(MySQL 上走 HTTP,MariaDB 上走 SSE) |
handler 在 1.5 s 内被中止 | 12 s 后仍为 outcome_unknown。状态 recoveryBlocked,下一次 prompt 409,detach 503,lease 不释放,同 Workspace 的新 Session 503。 |
| MCP 工具运行中取消 prompt | 取消后 0.5 s 中止 | 该轮没有终态事件。同样是 409 / 503 / lease 不释放 / 503。 |
| 32 s 工具,不取消 | 服务端在 41 s 回复 | Runtime 在 25 s 超时(TIMEOUT_MS 写死,manifest 不接受 timeout 字段)。 |
32 s 这一例里,迟到回复确实结算了 Runtime 自己的视图。但 qwen_tool_execution.execution_state 一直是 UNKNOWN,也从不重查,所以 Session 永远恢复不了。close 时 mcp-release 200,随后 tool-sessions/mcp:<sid>:release 返回 409,detach 返回 503。对比:普通 MCP 用 MCP_DEFAULT_TIMEOUT_MSEC = 10 分钟,可按服务端配置。
F2 / F5 / F6 / F7 —— 首轮之后的连接生命周期。
-
F2。
resources或tools的list_changed会让所有调用类型共用的那一个catalogRevision加一,而 Harness 从不发起mcp-discover。此后:- 同一轮内下一次工具调用得到
catalog_conflict; - 之后每一轮也是;
prompt_get也是;- Harness 的目录仍显示 revision 1、
complete。
SSE 服务端重启效果相同,因为
onerror会让连接退役。SDK 的McpServer在 connect 之后注册、启用或禁用工具时都会发list_changed,所以这个触发条件并不罕见。 - 同一轮内下一次工具调用得到
-
F5。 detach 在 1 s 内返回 503(MySQL 678 ms,MariaDB 228 ms)。服务端恢复后再次 detach 仍返回 503。配置一直停在
releasing,lease 不释放。DELETE 发不出去时terminateSession()会抛错,而当时该连接的在途操作数是 0。 -
F6。 每次
POST /mcp/configurations都让上一个 stdio 进程继续活着:同一 Session、同一服务端,存活进程从 1 涨到 16。第 17 次重新配置返回 409。只有 detach 才回收它们,耗时 6.9 s。这就是 triage 机器人 stage-2 第 4 点,此处给出实测。结合 F2,文档里的恢复手段每个 Session 大约能用 15 次。 -
F7。 首次安装被拒后,
status报告recoveryBlocked:true,但/prompt仍然放行。即使 Workspace 已经空出来,detach 也一直 503。只有之后某次 prompt 把遗留的配置 intent 派发出去,detach 才能成功。
可选方向(决定权在维护者):
- F1: 要么 MCP owner 在两轮之间不再持有 storage lease,要么把"每个 Workspace 同时只能 attach 一个 MCP Session"写成 H1 限制。
- F2: 遇到
catalog_conflict/stale时调用已有的mcp-discover,传输出错后用新的 generation 重连。 - F3: 对
resource_read/prompt_get,把被取消的操作结算为不阻塞的cancelled,因为其结果只是数据。 - F4:
- 把超时做成 manifest 字段;
- 工具调用要么不发
notifications/cancelled,要么限时等待后结算; - close 或 load 时重查 execution 状态,以对账迟到的成功。
- F5: 在途操作为 0 时,DELETE 失败或返回
405应视为已释放。 - F6: 退役连接的在途集合清空后即关闭它。
- F7: 让
close()在本地把仍停在execution: intent的配置 intent 结算为not_started_proven,与操作记录的处理方式一致。
门禁
- CI。
- 在
a390b5d5,Test (ubuntu)红:process-env-guard.test.ts标出QWEN_MANAGED_MCP_CONFIG、PATH和SystemRoot。 - 作者已在
e3ecba66修复,现为绿:22 通过,2 个 pending。 - 本地该 guard 5 次中 4 次通过,唯一一次失败发生在负载均值接近 100 时。
- 在
- main。 feat(serve): add gated Hosted foreground Shell turns #12848 合入后本 PR 为
CONFLICTING:8 个文件,其中 6 个生产文件共 17 个冲突块(其中 8 个在hosted-workspace-tool-turn.ts)。F1 和 F4 正好在这个文件里,合并后的 head 需要重跑相关场景。 - 与 feat(managed-agent): Add durable remote Shell result delivery #12894(
a7c0a391)的合并顺序。- 两个 PR 都新增
V19(V19__managed_mcp_records.sql和V19__managed_tool_publication.sql),git 对这两个文件不报冲突。 - 试合并后的 jar 在 MySQL 8.4 上启动即失败:
Found more than one migration with version 19。 - 试合并还有 8 处文本冲突。
- 后合入的一方必须给自己的 migration 重新编号。
- 两个 PR 都新增
a390b5d5的本地测试套件。- cli 定向 79/79,core 定向 230/230,managed-agent-server 180/180,runtime-broker 395/396。
- Broker 那 1 个失败(
HttpRuntimeTransportTest#closesTheConnectionOnTheDeadlineAndOnCallerCancel)在本机高负载下对mainbd45b95f同样失败,单独跑 5/5 通过。
- 变异测试。 Java 8/8 被杀,TS 13/16 被杀。三个存活者指向没有测试覆盖的行为:
- M7: tools-only 服务端返回的
-32601应表示空的complete列表。 - M11: prompt 路由里的
hasPendingOperations()检查。F3 里的 409 正是这个守卫返回的,但没有单测固定它。 - M13: 由两个 Runtime Session 拥有的配置。
- M7: tools-only 服务端返回的
小问题
/mcp/operations和/mcp/configurations上的身份冲突返回 503,看起来像可重试。用 409 更合适。- 模型看到的工具名是
mcp_<32 位十六进制>,其中含 catalog revision 和 connection generation,所以每次 load 名字都会全变。reload 之后,对话记录里引用的是已不再公布的工具。 - Runtime 的错误回执以裸错误码形式交给模型,例如
managed_mcp_catalog_conflict。 - stdio 服务端运行时
HOME被设为 Workspace 目录(实测HOME = cwd = <ws>/child)。服务端缓存在~下的任何内容都会落进用户的 Workspace。
未验证
- Windows 与 Linux。
- 真实模型;与作者一样,我用的是假模型。
- 重启 Harness 后能否恢复被 F3–F5 卡住的 Session。
- 32 个在途请求的限额。
- 与 feat(serve): implement generic Broker provider controls #12868 的交互(triage 机器人已提出)。
- 解决 feat(serve): add gated Hosted foreground Shell turns #12848 冲突之后的 head。
装置、场景脚本、变异体与原始结果:wenshao/qwen-code@a5ef4914/pr12946
Merge validation — 7a32435Merged main at Local checks passed: full build/typecheck/bundle, formatting/lint, 951 CLI + 104 Core + 160 Broker + 48 Java server tests. All 81 Broker classes embedded in the clean-packaged server match the current compilation. Both independent E2E paths passed against the final artifacts:
The model was deterministic in both paths; Runtime/MCP/Shell processes were real. All test process groups were confirmed exited. Local execution was on macOS. GitHub reports the PR as mergeable without conflicts; fresh CI is running. |
Real-stack verification, round 2 —
|
Review follow-up and merge verification —
|
| Finding | Change / retained boundary |
|---|---|
| F1: idle attached owner blocks other Sessions | Retained and documented as a private H1 constraint: the owner holds the tenant/storage lease until detach, including Workspaces sharing storage. Releasing it while a stdio server can still write would permit concurrent writers. Connection/storage lifetime separation is follow-up work before enablement. |
| F2: stale discovery or retired connection never recovers | Discover before each model request and each new raw command; unchanged usable catalogs retain their revision. Changed catalogs and retired connections publish a new immutable configuration/generation. Already advertised calls keep their original pins. |
| F3: native cancellation suppresses settlement | Cancellation after dispatch is deferred. Runtime keeps observing the original request without sending the native cancellation notification; an actual response supplies the settlement proof. Cancellation before dispatch still records not_started_proven. |
| F4: 25-second timeout / cancellation leaves a tool permanently UNKNOWN | Invocation timeout defaults to ten minutes and is deployment-configurable up to that limit. Hosted observes the original execution for 630 seconds, including temporary UNKNOWN/network outcomes; Java reconciles both tool protocol versions and durably accepts conclusive late results. There is no second start. Results after that window and restart during an unfinished turn still require follow-up checkpoint recovery. |
| F5: failing HTTP DELETE prevents local closure | Remote session deletion is best effort with a one-second bound, followed by bounded local transport closure. Release proves local closure after all requests settle; it does not claim the remote server session was deleted. |
| F6: idle replacements exhaust 16 connections | Close idle retired siblings; close busy retired siblings only after their original requests settle. Closed generations retain their receipts and can acknowledge their original release. |
| F7: a refused first configuration cannot detach | Close cancels an unsent configuration, completes the original Broker acquisition when the Workspace becomes available, then releases it without installing MCP. No retry prompt is required. |
Two further audit findings were fixed before commit: an idle Broker restart must reacquire and validate the original owner before new work/close; an idle Harness restart must advance durable configuration record revisions before issuing new-writer grants. The Runtime gate continues to reject owner replacement at the same revision. Explicit replacement also works after the first definition fails, without retrying that failed definition first.
Other review points
- Status polling now reuses the validated owner and backs off from 100 ms to one second. A failed status query reacquires only the original owner and retries status, never an effect. New work and close still validate ownership at their boundary.
- Idempotent shared Broker acquire and the additional binding/generation fields are intentional. Existing file/Shell reservation identities and Shell input digest remain covered by explicit assertions and the shared Shell integration regression.
- Record-domain recognition is global; execution admission remains private-profile and scoped-manifest gated. This distinction is now documented.
- Acquire failures preserve uncertainty on lost replies; definite Broker rejections restore the previous local acquired flag. Grants require a successfully validated Runtime binding.
- Raw operation ID conflicts now return HTTP 409 without another effect. M7 (method-not-found means authoritative empty discovery) and M11 (pending raw work blocks model prompt admission) have direct regression assertions; the existing M13 two-owner guard remains covered.
- Deferred explicitly: process-lifetime operation receipts / closed-generation tombstones and the release waiter for permanently unknown work remain unbounded history. The concurrent connection/request quotas do not claim to bound this memory. Durable-acknowledged pruning and bounded waiter cleanup are follow-ups before production enablement, recorded here and in both design languages. Tool names intentionally include immutable catalog/connection identity, and stdio HOME remains the Workspace unless a deployment overrides it.
- feat(serve): implement generic Broker provider controls #12868 is still an open semantic integration dependency. The later merge must preserve original-owner status/cancel/release after revocation and drain-before-storage-release ordering while adopting the generic control mechanism; this PR does not claim that cross-branch combination has been verified.
Verification on the final tree
- Build/static checks: repository build, final CLI rebuild/bundle, repository typecheck, changed-file ESLint/Prettier, and both Java modules’ Checkstyle passed. The pre-commit hook also passed without changing the audited tree.
- Focused tests: 961 CLI (121 directly affected + 840 shared consumers), 59 Core, 160 Broker and 54 Server tests passed: 1,234 tests, zero failures. The new grant test also failed under a deliberately disabled-renewal mutation, then passed after restoring the implementation.
- Real stack: macOS, Node 22.22.2, JDK 21.0.11, isolated MariaDB 10.11.19, packaged Hosted Harness → Java Broker/SQL store → provisioned Runtime → local stdio/Streamable HTTP wire fixtures; only the model is deterministic. All 11 formal lifecycle cases passed, with F1 checked as a retained limit:
| Scenario | Observed result |
|---|---|
| F1/F7 | While A remains attached, B is refused. After A detaches, B directly detaches with 204, original MCP owner RELEASED and no lease; a following file Session writes successfully. |
| F2 | Two list-change events recover before raw/model calls; killing an idle stdio server also recovers. Four unique effect IDs, no replay. |
| F3 resource + prompt | Cancel leaves the original request running; only the actual reply settles it. No native cancel notification, one effect per case, subsequent work and detach succeed. |
| F4 tool cancel | Prompt cancel returns 204, execution remains unproven until the actual reply; then SQL SETTLED, next prompt and detach succeed. One start/effect. |
| F4 long tool | Real response at 32,002 ms succeeds under the default timeout; SQL SETTLED/success and detach 204. |
| F4 configured timeout | The same SQL execution is observed UNKNOWN/NULL, then SETTLED/success after the original physical reply. Exactly one Broker start and one native effect. |
| F5 | After a successful HTTP read, shutting the listener still permits detach 204, owner RELEASED and lease NULL; a later Session reads again. Conflicting reuse returns 409 with zero added effects. |
| F6 | All 18 replacements succeed through revision 19; only the current native PID survives each replacement. All 19 PIDs exit by detach. |
| A1 idle Harness restart | SIGKILL the idle Harness and let its 5-second writer lease expire; reload under a new boot ID. Runtime binding/generation and native PID stay the same. New read and detach succeed without replay. |
| A2 initial definition failure | Bad r1 fails explicitly; expected-revision-1 replacement by good r2 succeeds. Next read has one effect and detach returns 204. |
- The original five-effect MCP happy/recovery scenario also passed: byte-preserving resource/prompt results, lost configure/invoke ACK recovery by original ID, revoked-access close, and offline durable replay with 0 Broker requests / 0 new effects.
- The upstream six-Workspace, 100 MiB Shell packaged integration passed against H2, preserving the shared Broker/Shell path. This is separate from the MariaDB MCP scenarios; no MySQL run is claimed for this revision.
- Every recorded lifecycle test process group and native child was confirmed gone. One preliminary test-driver assertion incorrectly included an unrelated file-profile control owner; it was narrowed to the original MCP owner and the entire case rerun successfully.
Artifacts exercised: CLI SHA-256 7db44d0d703e3c45f7d04148834ecdbbe6e9f5b529b4249b6e819d4355a92746; packaged Server SHA-256 7586d6c4a502f4fb8ee573ae6f139e4ad17fe6252da71107ef38216bb1b6e188. The 81 embedded Broker classes match the compiled Broker byte-for-byte. Windows/Linux, a real model, full restart matrix, and cross-branch #12868 integration were not verified locally; SSE retains focused transport coverage. New GitHub CI is tracked separately from these local results.
The prior finding-bearing audit rounds reset the clean count. After the last fixes, two consecutive open-ended and targeted passes were clean, including independent CLI and Core/Java review. These are verification results for this private H1 scope, not an approval or production-enablement claim.
中文摘要:已同时修复冲突和评论确认的问题,并补上审计发现的 Broker/Harness 空闲重启及首次失败配置替换问题。F1 存储租约独占、观察窗口外/turn 中途重启恢复、历史回执与永久等待者的回收仍为明确的 H1 后续项;未宣称这些限制已修复,也未启用生产 profile。
Real-stack verification, round 3 —
|
| # | Round 3 observed | Verdict |
|---|---|---|
| F1 | The MCP owner still holds the Workspace lease: a second MCP Session gets 503, a file Session gets turn_error |
Retained, as documented |
| F2 | After resources or tools list_changed: the same-turn call, the next turn and prompt_get all succeed. After an SSE server restart the next call succeeds (1 effect) |
Fixed |
| F3 | Cancel of a running resource_read answers running at 2.2 s; the operation settles with the full result at 9.9 s. The server ran to done (+8.0 s); detach 204, lease released |
Fixed |
| F4 | Prompt cancel at 2.5 s: the turn ends cancelled at 11.6 s (the tool ran to done +10.0 s); next prompt and detach work. A 32 s tool completes at 33.7 s (default timeout) and at 33.8 s (timeoutMs: 5000, UNKNOWN in between) |
Fixed |
| F5 | HTTP server down at close: detach 204 in 291 ms, lease released | Fixed |
| F6 | 17 reconfigurations: one live stdio process each time, revision 18 accepted, detach reaps it | Fixed |
| F7 | Refused first install: detach 204 once the other Session leaves, no extra prompt (still 503 while it holds the Workspace) | Fixed |
Still holding from earlier rounds:
- three transports with
results_readyordering, byte-exact blobs and 0 catalog leaks; - lost configure and invoke replies → exactly 1 read;
- revoke then close → 204;
- Broker down → replay 202, close 204, 0 Broker requests, 0 effects;
- replacement during an in-flight call;
- the 16 KiB / 60 KiB limits.
B1 — the CI red is caused by one Broker line
Failing job. CI Hosted process fault gates / MySQL 8.4 fails in HostedProcessCrashIT.processCrashesNeverReplayHostedToolsOnMySql (FG6C). CI stops at the first failure, so I ran each case alone on the PR tree:
| FG6C case | At 9b488679 |
Same tree with only that line restored |
|---|---|---|
worker-kill |
status 409 runtime_execution_evidence_unavailable; the gate expects runtime_broker_execution_unknown |
1/1 |
worker-stop |
status 503 managed_runtime_unavailable (retryable); the gate expects 200 |
1/1 |
harness-*, spring-kill |
pass | pass |
Cause. The PR removes && Integer.valueOf(3).equals(record.getReference().get("runtimeProtocol")) from observe(). As a result, status reads of UNKNOWN v2 (file) and MCP executions now go through reconcileExecution, which has to reach the original worker. That changes the Broker's status contract for every Hosted tool profile, not only MCP.
Options (the maintainers' call). Either:
- keep
observe()protocol-3-only and reconcile MCP executions on their own path; or - update the FG6C expectations for v2, with the reason stated in the PR.
Why the author missed it. The author's note says the Shell/file IT ran on H2 only and no MySQL run was claimed for this revision. HostedProcessCrashIT is the gate that catches this change.
Behaviour notes the fixes leave open
- N1 — a reply that can never arrive still ends in a stuck Session, now after 10.5 minutes.
- Connection lost: I restarted the SSE server 2 s into a 20 s call. The execution was
UNKNOWNat 6.8 s, the turn stayed active until 633 s, then the Session wasrecoveryBlockedwith no terminal event. - Tool never answers: on a
timeoutMs: 5000server, the execution wasUNKNOWNat 7.8 s and the turn stayed active until 634 s. - Both cases end the same way: next prompt 409, detach 503, lease held (so, by F1, the Workspace is locked),
qwen_tool_executionUNKNOWN. - The configured
timeoutMsdoes not shorten the Hosted wait, and a lost connection is known at once but still waits out the 630 s window. The author documents "results after that window" as follow-up work; these two cases show the window itself is spent even when the outcome is already decided.
- Connection lost: I restarted the SSE server 2 s into a 20 s call. The execution was
- N2 — one pinned server down makes every turn fail. Because the Harness now refreshes before each model request, a single unavailable pinned server (SSE or Streamable HTTP) makes every turn in the Session end in
turn_error, with 0 effects. That includes a text-only answer and a turn that only calls the healthy stdio server. Each failed turn records a failedmcp_configurationrevision, and the Session recovers once the server returns. One option is to drop that server's tools for the round instead of failing the turn. - N3 — cancellation is deferred. The prompt cancel above took 9.1 s to end the turn because the tool ran to completion; with a 10-minute tool, the user waits up to 10 minutes and the effect still happens. This is documented by the author, and I'm listing it here only with the measured number.
Migrations, mutation, cost, CI
- Flyway ordering. Tested with a trial merge with feat(managed-agent): Add durable remote Shell result delivery #12894 at
033c74d3, which carries V19 and V20.- On a fresh database, the merged jar applies 19 → 20 → 21 and starts.
- If this PR runs first (database at 21), the merged jar then refuses to start with
Detected resolved migration not applied to database: 19 … set -outOfOrder=true.outOfOrderis not enabled. - So feat(managed-agent): Add durable remote Shell result delivery #12894 should land first, or whichever lands second should renumber above the other.
- Mutation (6 cli test files, 122 tests):
- M7 and M11 are now killed, and so are the inverses of the fixes: native cancel, idle sibling close, bounded DELETE, wait-through-unknown, intent cancel on close, unchanged-catalog skip, and grant renewal.
- Survivors:
- X7, no refresh before each model request. This is the F2 fix itself; it is proven end to end above but not by a unit test.
- X3, a busy retired connection is not closed after its last settle.
- X5, the 600 s default timeout falls back to 25 s.
- X9, identity conflicts return 503 again.
- Refresh cost. With text-only turns at a load average of about 25–30: 1
mcp-discoverper server per model request. Time to the first model request was 103–185 ms with 1 server and 154–232 ms with 6 servers. - CI at
9b488679: 21 pass and 1 fail (B1).HostedWorkspaceToolTurnITis 5/5 on MySQL in that same job (115.5 s, including the six-Workspace 100 MiB Shell path). Test, Lint, MariaDB and the Java matrix are green.
Not verified
- Windows/Linux.
- A real model.
- MariaDB this round (CI's MariaDB job is green).
- A Harness restart during an unfinished turn.
- The feat(serve): implement generic Broker provider controls #12868 combination.
Evidence: wenshao/qwen-code@9fd4f16f/pr12946/r3 (per-scenario raw results, S12 timelines, FG6C A/B extracts, Flyway logs, remerge diff, mutants, scripts).
中文版
真实环境验证第三轮 —— 9b488679(F1–F7 处置)
合并参考(一段话): 处置表里的每一项都在 MySQL 8.4.11 真实链路上复现成立,包括作者只有单测覆盖的 SSE:F2–F7 已修复,F1 的行为与现在的文档描述完全一致,身份冲突返回 409,第一轮存活的变异体 M7/M11 现在已被测试固定。改名为 V21 消除了与 #12894 同号导致的启动失败。但目前还不能合入,有一个阻断项: CI 的 Hosted process fault gates / MySQL 8.4 是红的。我在本地复现了它,并用单行 A/B 定位到本 PR 对 RuntimeBrokerHttpServer.observe() 的一行改动。此外还有三条扩大了已记录 H1 限制的行为说明(N1–N3)、一条与 #12894 相关的 Flyway 顺序规则,以及四个存活的变异体。
F1–F7 复验
用 9b488679 重新构建 bundle 和 jar;装置、manifest 和 MCP 服务端与前两轮相同,另加一个 timeoutMs: 5000 的定义。
| # | 第三轮实测 | 结论 |
|---|---|---|
| F1 | MCP owner 仍持有 Workspace lease:第二个 MCP Session 得到 503,文件 Session 得到 turn_error |
按文档保留 |
| F2 | resources 或 tools 的 list_changed 之后,同一轮的下一次调用、下一轮和 prompt_get 都成功;SSE 服务端重启后下一次调用成功(1 次副作用) |
已修复 |
| F3 | 取消运行中的 resource_read:2.2 s 返回 running,9.9 s 带完整结果结算;服务端跑到 done(+8.0 s),detach 204,lease 已释放 |
已修复 |
| F4 | 2.5 s 取消 prompt:turn 在 11.6 s 以 cancelled 结束(工具跑到 done +10.0 s),下一次 prompt 与 detach 正常。32 s 工具在默认超时下 33.7 s 完成,在 timeoutMs: 5000 下 33.8 s 完成(中间为 UNKNOWN) |
已修复 |
| F5 | close 时 HTTP 服务端宕机:detach 在 291 ms 内返回 204,lease 已释放 | 已修复 |
| F6 | 连续 17 次重新配置:每次只有 1 个 stdio 进程存活,修订 18 被接受,detach 会回收 | 已修复 |
| F7 | 首次安装被拒:另一个 Session 离开后直接 detach 204,不需要再发 prompt(对方仍占着 Workspace 时仍是 503) | 已修复 |
前两轮验证过、这一轮仍然成立的:
- 三种 transport 下的
results_ready顺序、字节精确的 blob、目录 0 泄漏; - 丢失 configure 与 invoke 回包 → 恰好 1 次读取;
- 撤销权限后关闭 → 204;
- Broker 不可用 → 重放 202、关闭 204、0 次 Broker 请求、0 次副作用;
- 在途调用期间替换配置;
- 16 KiB / 60 KiB 限额。
B1 —— CI 红由 Broker 的一行改动造成
失败的 job。 CI 的 Hosted process fault gates / MySQL 8.4 在 HostedProcessCrashIT.processCrashesNeverReplayHostedToolsOnMySql(FG6C)处失败。CI 遇到第一个失败就停,所以我在 PR 树上把每个用例单独跑了一遍:
| FG6C 用例 | 9b488679 |
同一棵树只恢复那一行 |
|---|---|---|
worker-kill |
状态查询返回 409 runtime_execution_evidence_unavailable;门禁期望 runtime_broker_execution_unknown |
1/1 |
worker-stop |
状态查询返回 503 managed_runtime_unavailable(可重试);门禁期望 200 |
1/1 |
harness-*、spring-kill |
通过 | 通过 |
原因。 本 PR 从 observe() 里删除了 && Integer.valueOf(3).equals(record.getReference().get("runtimeProtocol"))。于是 UNKNOWN 的 v2(文件)和 MCP execution 的状态查询都会走 reconcileExecution,而它必须联系原来的 worker。这改变的是所有 Hosted 工具 profile 的 Broker 状态契约,不只是 MCP。
可选做法(由维护者决定)。 二选一:
- 保持
observe()只对协议 3 对账,MCP execution 走单独的对账路径;或 - 更新 FG6C 对 v2 的期望,并在 PR 中写明理由。
为什么作者没发现。 作者的说明写明,这一版的 Shell/文件 IT 只在 H2 上跑过,没有声称跑过 MySQL。而 HostedProcessCrashIT 正是能拦住这次改动的门禁。
修复之后仍未解决的行为问题
- N1 —— 回复永远不会来时,Session 仍会卡死,只是变成等 10.5 分钟之后。
- 连接丢失: 在一个 20 s 调用开始 2 s 后重启 SSE 服务端。execution 在 6.8 s 变为
UNKNOWN,turn 一直活动到 633 s,然后 Session 进入recoveryBlocked,没有终态事件。 - 工具永不返回: 在
timeoutMs: 5000的服务端上,execution 在 7.8 s 变为UNKNOWN,turn 一直活动到 634 s。 - 两种情况的结局相同:下一次 prompt 409,detach 503,lease 不释放(按 F1,整个 Workspace 被锁),
qwen_tool_execution为 UNKNOWN。 - 配置的
timeoutMs并不会缩短 Hosted 的等待;连接丢失在一开始就已确知,却仍要等满 630 s 窗口。作者把"窗口之后的结果"记为后续工作;这两个例子说明,即使结局早已确定,这个窗口本身也会被耗尽。
- 连接丢失: 在一个 20 s 调用开始 2 s 后重启 SSE 服务端。execution 在 6.8 s 变为
- N2 —— 绑定的某一个服务端宕机,会让每个 turn 都失败。 由于 Harness 现在每次模型请求前都会刷新,只要有一个绑定的服务端不可用(SSE 或 Streamable HTTP),这个 Session 的每个 turn 都以
turn_error结束、副作用为 0。纯文本回答和只调用健康 stdio 服务端的 turn 也不例外。每个失败的 turn 都会记录一条失败的mcp_configuration修订;服务端恢复后 Session 会自动恢复。一种做法是:这一轮先去掉该服务端的工具,而不是让整个 turn 失败。 - N3 —— 取消被延迟。 上面的 prompt 取消花了 9.1 s 才结束 turn,因为工具要跑完;如果是 10 分钟的工具,用户最多要等 10 分钟,而且副作用照样发生。作者已写明这一点,这里只是补上实测数字。
迁移、变异、开销与 CI
- Flyway 顺序。 与 feat(managed-agent): Add durable remote Shell result delivery #12894(
033c74d3,带 V19 和 V20)试合并后测试:- 全新库上,合并后的 jar 依次应用 19 → 20 → 21 并正常启动;
- 如果本 PR 先运行过(库已到 21),合并后的 jar 启动失败:
Detected resolved migration not applied to database: 19 … set -outOfOrder=true。outOfOrder没有开启。 - 所以 feat(managed-agent): Add durable remote Shell result delivery #12894 应该先合,否则后合的一方要把编号改到对方之上。
- 变异测试(6 个 cli 测试文件,122 个测试):
- M7 和 M11 现在被杀;各项修复的反向变异体也都被杀,包括原生取消、关闭空闲的 sibling、DELETE 限时、等待 unknown、关闭时取消 intent、目录未变时跳过、grant 续期。
- 存活者:
- X7,不在每次模型请求前刷新。这就是 F2 修复本身;上面的端到端测试证明了它,但没有单测覆盖。
- X3,忙碌的已退役连接在最后一次结算后没有被关闭。
- X5,600 s 的默认超时退回 25 s。
- X9,身份冲突又变回 503。
- 刷新开销。 纯文本 turn、负载均值约 25–30 时:每个服务端、每次模型请求 1 次
mcp-discover。到第一次模型请求的时间,1 个服务端为 103–185 ms,6 个服务端为 154–232 ms。 9b488679的 CI: 21 通过、1 失败(B1)。同一个 job 里HostedWorkspaceToolTurnIT在 MySQL 上 5/5 通过(115.5 s,含 6 个 Workspace 的 100 MiB Shell 路径)。Test、Lint、MariaDB 和 Java 矩阵均为绿。
未验证
- Windows/Linux。
- 真实模型。
- 本轮的 MariaDB(CI 的 MariaDB job 为绿)。
- 未完成的 turn 进行中重启 Harness。
- 与 feat(serve): implement generic Broker provider controls #12868 的组合。
证据:wenshao/qwen-code@9fd4f16f/pr12946/r3(各场景原始结果、S12 时间线、FG6C A/B 摘录、Flyway 日志、remerge diff、变异体与脚本)。
|
B1 is fixed in 7444dc9. The shared v2 status contract is preserved: default reads no longer contact an UNKNOWN execution's worker. MCP polling explicitly opts into original-execution reconciliation with Verification before pushing:
These local process tests use macOS/MariaDB and a deterministic model fixture. The separate Linux/MySQL 8.4 CI gate now passes on 7444dc9: all six original FG6C success markers are present, Hosted Workspace tool turns pass 5/5, Hosted Harness MySQL checks pass 2/2, and 29 additional Runtime fault-gate tests pass. The Linux unit suite, lint/static checks, Serve A/B, no-AK integration, Java matrix, MariaDB integration, real daemon E2E and TUI checks also pass. The downstream Web Shell browser smoke gate also passes. Final check snapshot: 23 passed, 8 skipped by workflow policy, no failures; only the separate automatic code review remains in progress. Round-3 follow-ups: N1's full 630-second wait and N2's all-pinned-server availability requirement are now explicit in both design languages and the PR body. N3's deferred cancellation remains documented. Flyway guidance now states that if V21 has already run, pending publication migrations must be renumbered above it, without enabling out-of-order execution. N1–N3 behavior changes and the non-blocking X7/X3/X5/X9 unit-test gaps are deferred to follow-up work; this CI fix does not expand the established H1 scope. 中文说明7444dc9 修复 B1:普通 v2 UNKNOWN 查询恢复被动读取;MCP 通过显式 GET 参数对原 execution 对账,v3 自动对账不变。没有修改 FG6C 期望、重发执行或更换 owner。 旧产物已复现 worker-kill/worker-stop 失败,新产物在真实 MariaDB 和进程信号下通过原版六个 FG6C 场景。另行验证 MCP 配置超时后,同一 SQL execution 接受迟到结果,start 与物理调用各一次,后续 prompt 和 detach 成功。52 项 CLI、161 项 Broker 测试及构建、类型、格式检查通过;提交前连续两轮开放式和定向审计无新问题。新提交的 Linux/MySQL 原失败关卡现已通过,日志确认原版六个 FG6C 场景、5 项 Hosted 工具 turn、2 项 Harness MySQL 检查及另外 29 项 Runtime 故障门禁均通过。Linux 全量单测、Web Shell 浏览器冒烟测试及上述其他检查也已通过。最终状态为 23 通过、8 按工作流策略跳过、0 失败;仅独立的自动代码评审仍在运行。 N1–N3、迁移顺序与变异测试建议的处置见上文;中英文设计及 PR 描述已同步。 |
Real-stack verification, round 4 —
|
|
@qwen-code /triage |
qqqys
left a comment
There was a problem hiding this comment.
Independent Critical-only review — head 7444dc9726e1bc8ce8fdf8d15093b2506e2f2abd
Verdict: COMMENT. I found no Critical I can prove, but I cannot approve: the only review round on this PR left eight Critical-grade candidates explicitly unruled, that round ran against an earlier head, and no round has run against this one. At +8857/−128 across 52 files I could not rule on those eight myself inside this channel's per-PR budget.
Why this is not an approval
The prior round did not verify its own findings. The bot's round at 9b488679 states that both verification waves were killed by harness stalls, that 44 of 51 candidates were never ruled on, and that 44 findings still carried the [unverified] tag when the loop ended. Only the shards covering R1-11/R1-12 and R1-41/R1-42/R1-43/R1-44/R1-51/R1-52 delivered verdicts, and the five findings it did post are all sev:S. Its reverse audit never reached two consecutive dry rounds, and its test-efficacy probe returned inconclusive on all 21 probes with the positive control never running — so mutant coverage there is an absence of measurement, not clean coverage.
The eight unruled Critical-grade candidates, in the round's own words:
- R1-1 — grant lease (300 s) shorter than the runtime timeout (600 s) and the Hosted observation window (630 s).
- R1-2 — a transient
mcp-discoverfailure permanently recovery-blocks the Session, with no reset site forsession.blocked. - R1-3 —
unknown()never removes an operation fromconnection.pending, so zombies block the release drain and consume the worker-wide inflight and connection quotas. - R1-4 —
install()commits executiondispatch_startedfrom a durableoutcome_unknownrun, which the extension state line rejects. - R1-5 — an unmatched string-id response is forwarded to the SDK client.
- R1-6 — the catalog blob is projected without shape validation.
- R1-7 — a Windows
SYSTEMROOT/SystemRootcase collision in the stdio child environment. - R1-8 — the status and cancel routes sit outside the
mcpBusyfence.
An unruled candidate is not a confirmed defect, and I am not asserting any of these. But under this channel's rules a blocking-grade issue I can neither confirm fixed nor refute prevents an approval, and eight of them sit on a private MCP runtime that spawns child processes from catalog-supplied definitions.
The head moved after that round and nothing has re-reviewed it. The round ran at 9b488679; this head is 7444dc97. There are no inline Criticals and no author replies on this PR resolving any of the eight, so I have no evidence either way about whether the pushes since that round addressed them.
What I did verify myself at this head
The new public route is scoped correctly. ManagedMcpController takes TenantContext, forwards both tenant.tenantId() and tenant.actorId() into ManagedMcpCatalogService.get, and answers Cache-Control: no-store — consistent with the sibling public routes, so the catalog is not cached or served cross-tenant by construction.
The migration only widens. V21__managed_mcp_records.sql issues two MODIFY COLUMN … VARCHAR(32) NULL statements against qwen_managed_session_extension_record, relaxing task_kind and task_state to nullable so MCP records can carry their own lifecycle without becoming Session tasks. It drops no column, no index and no data, and it does not rewrite existing rows. The risk it creates is on the read side — every consumer of those two columns must now tolerate null — which is what ManagedExtensionRecordStore.java (+111/−7) and ManagedExtensionProjection.java (+11/−1) would need to handle, and which I did not read.
R1-7 traced part of the way, not ruled on. The stdio child environment at managed-mcp-runtime.ts:603-614 is built in this order: every key of the imported DEFAULT_INHERITED_ENV_VARS blanked to '', then PATH/HOME/USERPROFILE, then SystemRoot spelled that way and only when process.env['SystemRoot'] is truthy, then ...definition.env last. Whether a case collision is reachable therefore turns entirely on whether that imported list contains an uppercase SYSTEMROOT spelling — if it does, the object carries two keys differing only in case, one of them empty, and which one the child receives is not defined by the object literal. I did not open the defining module inside budget, so I am recording this as unresolved rather than filing it. It is a one-line check and the cheapest of the eight to close.
CI. Green at this head: Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), Serve A/B, web-shell E2E Smoke, the full Java matrix, Runtime Broker and Managed Agent MariaDB / Java 21, Hosted process fault gates / MySQL 8.4 / Java 21, Real daemon E2E / Java 11, the TUI parity and no-flicker gates and both Desktop Shell lanes all pass. Only review-pr is pending, which is not a gating check. The prior round also noted that Integration Tests (CLI, No Sandbox) was skipped, so that suite has not run for this change. No failure is attributable to this PR.
Coverage note
I read the new public controller, the migration, and the stdio environment block, and confirmed the review history and CI state. I did not read managed-mcp-runtime.ts (1043 lines) beyond that block, hosted-mcp-session.ts (904), managed-mcp-record.ts (322), managed-mcp-protocol.ts, the changes to managed-session-authority.ts (+90/−9), hosted-harness-session.ts (+184/−9), managed-runtime-tool-executor.ts (+136/−6), hosted-workspace-tool-turn.ts (+44/−21) or managed-context-worker.ts (+53/−28), nor the Java ManagedMcpRecords.java (185), ManagedMcpCatalogService.java (84) or the extension-store changes, nor the 3536 lines of new tests. A Critical-only scan of that surface is not reachable in this channel's budget.
Next step: re-run the review round on this head so the eight candidates get ruled on, and settle R1-7 first since it is one grep. R1-2 and R1-3 are the two I would want answered before merge on durability grounds — a Session that stays recovery-blocked after a transient discover failure, and pending operations that are never reaped, are both fail-closed in the wrong direction for a long-lived worker.
|
Addressed Finding 1 in commit New Hosted Sessions now reject more than 16 server pins with HTTP 400 Runtime connection quota failures now return HTTP 409 Verification:
The real-stack checks use macOS, MariaDB 10.11.19, real Java/Runtime/MCP processes and a deterministic model fixture. No new Java implementation or migration was introduced. CI for the new commit is reported separately. One evidence correction: two prompts against the original 17-pin Session reused the same 16 native PIDs and produced only 16 initialization events total. We reproduced repeated quota failure, but not the report's inferred repeated 16-process spawn cascade. Finding 2's diagnostic-code distinction remains deferred: both malformed pins and wrong-profile input are already rejected, and this patch stays focused on the confirmed capacity defect. 中文说明已处理 Finding 1:新建 Session 超过16个 pin 时立即400,准入与Runtime共用同一上限。旧版本已保存的17–32个 pin 记录仍按原定义加载,保留detach清理入口。prompt准入、配置替换和原始操作初始化遇到连接配额不足时返回409及明确错误码;已202准入的turn仍沿用turn错误语义。 真实旧链路复现了17个pin创建成功、后续prompt通用503;修复后创建直接400,没有Broker请求、Runtime Session、子进程或模型调用。130项定向测试、构建、类型、格式检查及修正后的连续两轮审计通过。 补充校正:重复prompt复用原有16个进程,没有实测到每次重启16个进程。Finding 2作为非阻断诊断改进留待后续。 |
|
R5 Suggestion disposition: this PR has exceeded five review rounds, so the repository’s late-review scope rule applies. The confirmed correctness fixes above are included; the following 26 Suggestions are explicitly deferred for follow-up, not claimed fixed or silently resolved. Their original threads remain open and retain the detailed acceptance criteria. R5-17 is covered by the approval integration regression, and the previously deferred R4-8 is now fixed together with R5-5.
|
|
@qwen-code /triage |
doudouOUC
left a comment
There was a problem hiding this comment.
已完成对当前 head 40b05a74f280e37ae5e0471ec9219934dd971e7b 的全面独立审查,覆盖相对 main 的全部 58 个文件,并核对跨包调用、Runtime/Session/Workspace 归属、授权与 fencing、TypeScript/Java 记录及资源契约、目录发布、取消、原身份恢复和释放路径,以及双语设计和全部现有 review threads。未发现新增的可确认合并阻塞问题,Approve 本 PR 明确限定的私有、显式启用 H1 范围。
此前我提交的两个取消问题已在当前提交修复,并通过独立复验:首轮 MCP 初始化可响应 cancel/deadline;目录刷新进入 replacement 后也保留取消信号,尚未派发的配置不会在中止后继续发送。已经持久提交 dispatch 的工作保持原 operation 身份、最多发送一次;迟到结果按原 ID 恢复,不通过重发制造第二次物理效果。能够证明未发送或已经结算时,后续 prompt 和 detach 正常。
历史 Critical 按当前代码重新判定,不以 thread 的 resolved 标志代替验证。R3-11 仍成立并明确保留:无法结算的远端副作用使 detach/delete 返回 503,attachment、activation 和原恢复 owner 会继续保留。原 Runtime 可能仍在写入,强制释放不能证明已经排空,也会丢失迟到结果的恢复入口;这是本私有切片公开接受的限制,不是已修复项。生产启用前仍需解决客户端/SSE 生命周期与持久恢复所有权的分离及有界 orphan 清理;已有延期 Suggestions 不重复提交。
本地验证:
npm run build、npm run bundle、npm run typecheck、改动 TS 文件 ESLint、适用文件 Prettier 和 diff whitespace 检查通过。- CLI 11 个文件 1047 项、Core 5 个文件 304 项通过。Hosted 的 198 项最终关闭 coverage、串行文件运行全部通过;早先并行运行出现共享 coverage 目录竞争及已记录的等待超时 flake,未把那些运行作为全绿证据。
- Runtime Broker 全套 530 项无 failure/error,其中 1 项最初因未配置 worker bundle 跳过;构建 bundle 后,该真实 worker 用例另行运行通过、无跳过。Managed Agent 定向 40 项通过,包含 H2 MySQL 模式的 SQL、资源闭包和 V23 迁移检查。
- 16 个独立竞争窗口探针及 2 个实际 Hosted HTTP route 探针通过,覆盖 warmup/refresh 的 intent、acquire、durable dispatch acknowledgement、configure response 等待及 cancel/deadline。实际 HTTP 路径验证取消时不准入 input/模型、不初始化第二台 server;未知效果保留 owner,迟到结果恢复后 prompt 与 detach 正常。
独立探针使用当前提交的真实 Hosted 状态机、HTTP routes、本地持久化 stores 和 Broker HTTP 客户端;Broker 为 loopback 协议替身,模型和远程 store connector 也有明确替身。全局 qwen 0.24.6 无法启动此私有 profile,因此采用 test-script fallback;不将这些结果宣称为已部署 Java Broker/SQL/外部 MCP 的全栈 E2E。原生 stdio/HTTP/SSE transport 由 Runtime 单元测试覆盖。本地未重跑物理 fault-gate、MySQL/MariaDB profiles 或完整生产故障恢复矩阵。当前 head 的 Node、Java、MariaDB、MySQL fault-gate、daemon 和 web-shell 测试 CI 已通过;自动 review job 仍在运行。V23 与待合并 V20–V22 的发布顺序仍须按双语设计约束协调。
qqqys
left a comment
There was a problem hiding this comment.
Independent Critical-only review — head 40b05a74f280e37ae5e0471ec9219934dd971e7b
Verdict: COMMENT. One historical Critical is confirmed still standing at this head by reading the code, six more were filed two commits ago and have never been adjudicated at this head, and an 11,184-line / 58-file diff cannot be scanned to the depth this channel's Approve path requires inside one review budget. Under this channel's rules any historical blocking issue that still exists, or cannot be confirmed, forbids an Approve — so this is not a judgement that the PR is unsound, and I am not filing a new Critical of my own.
R3-11 — a recovery-blocked MCP session can never be detached or deleted: STILL STANDS
I re-read this at 40b05a74 rather than taking the thread's word for it. packages/cli/src/serve/hosted-harness-session.ts lines 1071–1091, the single close handler that both /session/:id/detach and DELETE /session/:id route through:
try {
await session.mcp?.close();
await session.managed.close();
for (const stop of session.streams) stop();
sessions.delete(req.params['id']);
res.sendStatus(204);
} catch {
error(res, 503, 'managed_session_close_failed');
} finally {
session.mcpClosing = false;
session.mcpBusy = false;
}The chain holds end to end at this head. HostedMcpSession.close() throws HostedMcpRecoveryRequiredError while hasPendingOperations(), which is precisely the state the design doc expects ("Physical connection loss with an unresolved request remains recovery-blocked") — status() → accept() commits recovery_blocked/outcome_unknown, and that is not in ['settled','failed','cancelled']. The throw lands in a bare catch that answers 503 managed_session_close_failed and returns, so the three statements after await session.mcp?.close() never run:
session.managed.close()is skipped, which is what callsrenewal.stop()andauthority.releaseActivation(). The activation lease is therefore not merely left to expire — the renewal interval atleaseDurationMs/3keeps actively holding it.for (const stop of session.streams) stop()is skipped, so the session's open SSE responses leak.sessions.delete(req.params['id'])is skipped, so theHostedSessionentry leaks for the daemon's lifetime.
The finally resets mcpBusy and mcpClosing, so a retry re-enters the same handler and throws again; the pending record can only be cleared by a conclusive runtime reply that, by hypothesis, will never arrive. There is no force-close escape, because both exit routes share this handler.
I want to be accurate about the status of this finding rather than merely repeat it: @doudouOUC's approval of this head states that R3-11 still stands and is deliberately retained as an accepted limitation of this private, explicitly-enabled H1 slice, with production enablement conditioned on separating the client/SSE lifecycle from durable recovery ownership and adding bounded orphan cleanup. That is a scoping decision the author and maintainer are entitled to make, and it is documented. It does not change how this channel must vote: a blocking issue that still exists at the head is a bar to Approve here, so the choice is to land it with that bar recorded, or to close the leak. The narrow fix is to make the MCP close non-fatal to the rest of the teardown — settle managed.close(), the stream shutdown and the map deletion even when session.mcp.close() throws, then report the MCP failure — so a recovery-blocked MCP session can still release its activation and stop leaking, without pretending the unresolved remote effect was settled.
Six Criticals filed two commits ago have never been ruled on at this head
Round 5 reviewed 06c1df7f and published at 2026-09-30T09:15:31Z with six findings at severity C. The head has moved twice since, and review-pr is still pending on this head, so no automated round has re-adjudicated them here and I had no budget to verify six call chains across two languages myself. They are unconfirmed, not cleared:
- R5-1
hosted-harness-session.ts:705—POST /session/:id/mcp/configurationsadmits a definition replacement while another operation is in flight. - R5-2
hosted-mcp-session.ts:764—close()commitsreleaseState: 'releasing'before dispatchingmcp-release. - R5-3
managed-mcp-runtime.ts:698— the list-changed lookup indexes a plain object literal with a remote-controlled key. - R5-4
managed-mcp-runtime.ts:725— retiring the predecessor connection runs inside thetrywhosecatchtears down. - R5-5
HttpRuntimeTransport.java:618— an MCP operation whose encoding exceedsTOOL_REQUEST_LIMIT_BYTES(256 KiB). - R5-6
RuntimeBrokerService.java:843— the synthesizedSessionContextadopted byreleasedSessionis never evicted.
R5-3 and R5-6 are the two I would read first: a remote-controlled property key and an unbounded cache are both shapes that stay invisible until they are expensive.
Round 3's other eleven Criticals (R3-1 to R3-10, R3-12, R3-13) were filed at 4134aaf2 and are likewise unadjudicated at this head by anything I can cite. The two cancellation regressions @doudouOUC raised at 06c1df7f and 67c908ea are, by their own current-head re-verification, fixed — I did not re-derive those and am not carrying them as my own confirmation.
Gate not completed: the current-head scan
I read the close handler above and the review record; I did not scan the production surface of this diff. Unread, and therefore unconfirmed: hosted-mcp-session.ts (+1039, new), managed-mcp-runtime.ts (+1072, new), managed-mcp-routes.ts (+78, new), hosted-turn-wait.ts (+27, new), managed-mcp-record.ts (+322, new), managed-mcp-protocol.ts (+134, new), ManagedMcpRecords.java (+185, new), ManagedMcpCatalogService.java (+84, new), ManagedMcpProtocol.java (+73, new), ManagedMcpController.java (+24, new), the migration V23__managed_mcp_records.sql, and the changes to hosted-harness-session.ts (+241/−9), hosted-workspace-broker.ts (+96/−9), hosted-workspace-tool-turn.ts (+85/−50), managed-runtime-tool-executor.ts (+133/−1), managed-session-authority.ts (+90/−9), RuntimeBrokerHttpServer.java (+33/−20), RuntimeBrokerService.java (+34/−5), ManagedExtensionRecordStore.java (+111/−7) and the +178/−4 OpenAPI change.
Two of those deserve naming because they carry risk that a green suite does not bound. The V23 migration lands behind a feature that is private and explicitly enabled, and @doudouOUC's own approval records that its release ordering against the pending V20–V22 migrations still has to be coordinated — a migration applied out of order is not something any test in this diff can catch. And the OpenAPI change should be checked the way the last contract PR was: read x-qwen-implementation-status on each touched operation at this head, so a route that is not served cannot be certified as one.
CI
Green at this head: Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), Serve A/B, web-shell E2E Smoke, TUI parity snapshots, OpenTUI no-flicker gate, Hosted process fault gates / MySQL 8.4 / Java 21, Runtime Broker and Managed Agent MariaDB / Java 21, Real daemon E2E / Java 11 and the full Java matrix all pass. review-pr is pending, which this channel does not treat as a gate — though here its completion is what would rule on R5-1 to R5-6 at the head that actually merges.
Next step: the shortest path to an Approve is to let the pending review round rule on R5-1 to R5-6 at this head, and to decide R3-11 explicitly — either close the teardown leak so a recovery-blocked MCP session still releases its activation, or record the accepted limitation somewhere a gate reads, since a still-standing Critical blocks here whatever its justification.
|
Merged latest origin/main The main Shell publication/recovery flow and the existing MCP flow both remain wired through the refactored Hosted turn entry point. MCP retains its shared owner, per-prompt preparation identity, cancellation/reconciliation behavior and post-approval grant renewal; remote Shell uses its publication/v3 path. Both sides' tests are retained, and the three direct MCP constructor fixtures were adapted to the combined signature. The migration dependency is now in this branch: main includes publication V20–V22, followed by MCP V23. Both design languages and the PR description reflect that #12894 has landed. Migrations remain ordered; this does not authorize applying V20–V22 retroactively to a database where standalone V23 was already deployed. Validation for the merge:
Previously deferred Suggestions and the documented R3-11 H1 limitation remain unchanged. New CI will validate the pushed merge commit. |
|
@qwen-code /triage |
qqqys
left a comment
There was a problem hiding this comment.
Independent Critical-only review — head f2618e3027d41921bf6a798cb061c50bb6d6e357
Verdict: COMMENT. The historical Critical I confirmed by reading code at the previous head is unchanged at this one, the six Criticals filed two heads ago sit in files this head does not touch, no reviewer of any kind has ruled on this head, and a 58-file / +11,187−185 diff cannot be scanned to the depth this channel's Approve path requires in one budget. Any historical blocking issue that still stands or cannot be confirmed forbids an Approve here, so this is a COMMENT. I file no new Critical of my own.
This head is a merge of main, not a fix
The PR's own content did not change. Totals moved from +11,184/−184 at 40b05a74 to +11,187/−185 here, and the delta is reconciliation with a newer main rather than new work: hosted-workspace-tool-turn.test.ts gains imports of assertManagedSessionDurableRef and HttpToolPublicationOwner, broker mocks for prepareV3/executeV3/acknowledgeV3, and an extra constructor argument — all of which arrive from main's tool-publication surface, alongside expectWritesStopped from the approval-expiry work. A three-dot compare against 40b05a74 lists 76 files for that reason; the PR's file list is still the same 58. So nothing in this head answers any finding, and my review of 40b05a74 carries over in substance — re-verified below rather than inherited.
Also worth stating plainly, because it changed with the head: @doudouOUC's approval of 40b05a74 is now dismissed, and the review list at f2618e30 is empty. No maintainer approval and no automated round stands at the head that would merge.
R3-11 — a recovery-blocked MCP session can never be detached or deleted: CONFIRMED STILL STANDS
I re-read this at f2618e30 rather than relying on the earlier pass. packages/cli/src/serve/hosted-harness-session.ts lines 1441–1457 are unchanged:
await session.mcp?.close();
await session.managed.close();
for (const stop of session.streams) stop();
sessions.delete(req.params['id']);
res.sendStatus(204);
} catch {
error(res, 503, 'managed_session_close_failed');
} finally {
session.mcpClosing = false;
session.mcpBusy = false;
}HostedMcpSession.close() throws HostedMcpRecoveryRequiredError while hasPendingOperations(), which is the state the design doc itself expects. The throw lands in the bare catch, so the three statements after it never run: session.managed.close() — which is what calls renewal.stop() and authority.releaseActivation(), leaving the renewal interval at leaseDurationMs/3 actively holding the activation lease rather than letting it lapse; the stream shutdown, leaking the session's open SSE responses; and sessions.delete(...), leaking the HostedSession entry for the daemon's lifetime. The finally resets mcpBusy and mcpClosing, so a retry re-enters and throws again, and the pending record can only be cleared by a conclusive runtime reply that by hypothesis never arrives. Both /session/:id/detach (line 1453) and DELETE /session/:id (line 1456) route through this one handler, so there is no force-close escape.
@doudouOUC's now-dismissed approval recorded this as deliberately retained — an accepted limitation of the private, explicitly-enabled H1 slice, with production enablement conditioned on separating the client/SSE lifecycle from durable recovery ownership and adding bounded orphan cleanup. That is a scoping decision the author and maintainer may make, and it is documented. It does not change how this channel votes: a blocking issue still standing at the head bars an Approve here. The narrow fix remains to make the MCP close non-fatal to the rest of the teardown — settle managed.close(), the stream shutdown and the map deletion even when session.mcp.close() throws, then report the MCP failure — so a recovery-blocked session still releases its activation and stops leaking, without claiming the unresolved remote effect was settled.
R5-1 to R5-6: unaddressed at this head, and unadjudicated
Round 5 reviewed 06c1df7f and published six findings at severity C. Every file they name has an identical diff footprint at this head as at 40b05a74, so none was touched by the merge:
- R5-1
WorkspaceRuntimeTransport.java:145(+19/−1) — the grant-mode entry point lets a provably pre-dispatch ownership refusal take a path it should not. - R5-2
hosted-mcp-session.ts:764(+1039, new) —close()commitsreleaseState: 'releasing'before dispatchingmcp-release. - R5-3
managed-mcp-runtime.ts:698(+1072, new) — the list-changed lookup indexes a plain object literal with a remote-controlled key. - R5-4
managed-mcp-runtime.ts:725— retiring the predecessor connection runs inside thetrywhosecatchtears down. - R5-5
HttpRuntimeTransport.java:618(+16/−1) — an MCP operation whose encoding exceedsTOOL_REQUEST_LIMIT_BYTES(256 KiB). - R5-6
RuntimeBrokerService.java:843(+34/−5) — the synthesizedSessionContextadopted byreleasedSessionis never evicted.
I did not re-derive these six myself and am not asserting them as my own findings; I am recording that nothing has ruled on them at the head that would merge, which under this channel's rules is the same as unconfirmed. R5-3 and R5-6 are the two I would read first — a remote-controlled property key and an unbounded cache both stay invisible until they are expensive. Round 3's other eleven Criticals are in the same position.
Gate not completed: the current-head scan
I read the close handler above, the merge delta, and the review record. I did not scan this diff: hosted-mcp-session.ts (+1039, new), managed-mcp-runtime.ts (+1072, new), managed-mcp-routes.ts (+78, new), hosted-turn-wait.ts (+27, new), managed-mcp-record.ts (+322, new), managed-mcp-protocol.ts (+134, new), ManagedMcpRecords.java (+185, new), ManagedMcpCatalogService.java (+84, new), ManagedMcpProtocol.java (+73, new), ManagedMcpController.java (+24, new), the migration V23__managed_mcp_records.sql, and the changes to hosted-harness-session.ts (+241/−9), hosted-workspace-broker.ts (+96/−9), hosted-workspace-tool-turn.ts (+85/−51), managed-runtime-tool-executor.ts (+133/−1), managed-session-authority.ts (+90/−9), RuntimeBrokerHttpServer.java (+33/−20), RuntimeBrokerService.java, ManagedExtensionRecordStore.java (+111/−7) and the +178/−4 OpenAPI change.
Two deserve naming because a green suite does not bound them. V23's release ordering against the migrations now on main — this head merged a main that carries V20, V21 and V22 for the tool-publication work, so V23 is no longer the only pending migration and the ordering constraint the design doc records is now live rather than hypothetical; an out-of-order migration is invisible to every test in this diff. And the OpenAPI change should be checked the way the last contract PR was: read x-qwen-implementation-status on each touched operation at this head, so a route the server does not serve cannot be certified as one.
CI
Two checks fail at this head and several are still pending, so the picture is incomplete rather than green. web-shell E2E Smoke and Desktop Shell (${{ matrix.os }}) both report failure with a duration of 0, and the latter carries its matrix expression unexpanded in its own name — a workflow-definition shape, not a product result. This PR touches no web-shell file, both jobs passed at 40b05a74, and the head between them is a merge of main, so on the available evidence neither failure is attributable to this diff and I am not treating either as a Critical. Lint & Static, Test (ubuntu-latest, Node 22.x), Serve A/B, Hosted process fault gates / MySQL 8.4 / Java 21 and review-pr are pending; pending state is not a gate here, though review-pr completing is what would rule on R5-1 to R5-6 without another manual pass. The Java matrix, Runtime Broker and Managed Agent MariaDB / Java 21, Real daemon E2E / Java 11, Integration Tests (no-AK, No Sandbox), TUI parity snapshots and OpenTUI no-flicker gate pass.
Next step: two things unblock this, and neither needs a full re-scan first. Decide R3-11 explicitly — either make the teardown survive an MCP close failure so a recovery-blocked session still releases its activation, or record the accepted limitation somewhere a gate reads, since a still-standing Critical blocks here whatever its justification. Then get a ruling on R5-1 to R5-6 at the head that will merge, from the pending review round or a targeted pass over those six call chains. Given this PR has been through five maintainer rounds, the repository's own guidance about not letting review rounds balloon it applies: beyond those Criticals, remaining Suggestions are better deferred to a follow-up issue than folded into another round.
doudouOUC
left a comment
There was a problem hiding this comment.
复审冲突解决后的 head f2618e3027d41921bf6a798cb061c50bb6d6e357,重新 Approve。本次以上一轮 40b05a74 的完整审查为基础,核对合入 main 3b18cfe5e 的 merge/remerge diff、全部人工冲突解决,以及 MCP 与远程 Shell 交付的衔接,未发现新增的可确认合并阻塞问题。
合并后的统一 turn 执行路径保留 MCP Session、审批参数与取消信号;MCP 使用原 Broker owner、实际 prompt/call 身份和原 operation 恢复,Shell 的 v3 publication 路径仍独立选择。Runtime 的 publisher 接线保留 MCP executor,Java control 转发、撤权后的原 owner 恢复、失败释放缓存清理及 drain-before-release 顺序均保留。MCP 状态机、记录契约和 authority 文件与已审查版本一致。双语迁移说明已更新为 V20/V21/V22 先于 V23,前置迁移现在已随 main 合入。
当前提交的本地验证:
- build、typecheck、合并涉及 TS 文件 lint/format 和 diff whitespace 检查通过。
- CLI 10 个文件 1113 项、Core 3 个文件 85 项通过;覆盖 Hosted、MCP、Runtime worker、远程 Shell publication 及 Session store。
- Runtime Broker 定向 213 项、Managed Agent 定向 41 项通过,均无 failure/error/skip;Managed Agent 测试前安装了当前提交的 Broker 制品。
- 16 个取消/deadline 竞争窗口探针和 2 个实际 Hosted HTTP route 探针重新运行通过。独立探针仍使用 loopback Broker 协议替身和明确的模型/store connector 替身,不将其作为外部 MCP/已部署 Java Broker/SQL 的全栈 E2E;本轮没有重跑完整物理故障恢复矩阵。
批准范围仍是私有、显式启用 H1。R3-11 的未知远端副作用保留 attachment/activation/原恢复 owner 限制没有改变,也不声明已修复;客户端/SSE 生命周期分离与有界 orphan 清理仍是生产启用前的后续工作。新 head 的 CI 正在重跑,提交本 review 前没有已失败项,部分测试与自动 review 尚未结束。
Real-stack verification, round 9 —
|
| R5 finding | Real stack at 40b05a74 and f2618e30 |
67c908ea |
|---|---|---|
| R5-1 replacement during a turn | turn ends turn_error, not latched; the replacement settles at 6.4 s; the next turn runs on the new server |
turn ends with no terminal event and stays blocked; next prompt 409; detach 503 |
| R5-2 undelivered release, one server | stuck: detach 503 on every retry; lease held | stuck (same) |
| R5-2 undelivered release, two servers | recovers: the 2nd detach re-sends both releases → 204, lease 0 | stuck |
| R5-3 prototype-named notifications | ignored: 0 configurations, catalog unchanged | every one triggers a replacement (config revision 1 → 2 → 3 → 4) |
| R5-4 slow retirement of the old connection | the replacement settles; the new server is used | replacement 503, outcome_unknown |
| R5-5 oversized MCP operation → 413 | not driven here (in round 8 a 300 KiB argument was refused by the Harness before the Broker); mutant J14 killed | — |
R5-7 adopted SessionContext eviction¹ |
not driven on the real stack; mutant J15 killed | — |
¹ Labels follow the linked inline threads. The Critical-only review numbers some items differently: its "R5-6" is the inline R5-7, and its "R5-1" points at WorkspaceRuntimeTransport.java:145, which is not the inline R5-1 above and which I did not drive.
- Build. Fresh bundles and Spring jars from
40b05a74and thenf2618e30, on MySQL 8.4.11. main's new@a2a-js/sdkdependency needed apnpm installfirst. The old arm is my67c908eabuild, with its Broker jar kept aside. The stdio server gained three test hooks:addtoolchanges its tool list at runtime;rawnotifysends a raw notification;--stickymakes a grandchild keep stdout open for about 45 s after the server exits.
- R4-9 (fixed). A lone surrogate in a resource URI, a prompt name, an argument key or an argument value:
- all four return 400
invalid_mcp_operationwith 0 Broker requests and 0 server reads; - the next read settles, detach returns 204 and the lease goes to 0;
- the well-formed emoji control passes through unchanged.
- all four return 400
- Refresh-cancellation P1 (fixed). The model calls
addtool, so the next refresh in the same turn starts a real replacement, and the Broker proxy holds that replacement'sacquire:67c908ea: the turn is cancelled at 2.0 s, but the replacementmcp-configureis still sent at 4.9 s. A detach during the hold returns 204, releasing before the acquisition finished.40b05a74: no configure is sent after the cancel. A detach during the hold returns 503, so nothing is released. The next turn is OK and detach returns 204.- If the replacement's
mcp-configureis held instead (afterdispatch_started), it is sent once and not resent, on both heads.
- R5-1 and R5-4. A turn is running a 3 s tool on
remotewhen an explicit replacement of thestickyserver (rev 1 → 2) is posted. Retiring the old sticky connection takes more than 5 s. The table above has both outcomes. - Approval merged from main. With
approvalMode: default, the MCP tool call raises an approval request. I allowed it at 310 s, after the 300 s grant lease: the tool ran once with the original arguments, the turn endedend_turnwith no Broker errors, and detach returned 204. A 2 s control behaved the same.
-
R5-2, one pinned server (not fixed). The Broker proxy answers the first
mcp-releasewith 503 without forwarding it.- Detach 1 returns 503.
- Detach 2 goes straight to the Broker tool-session release, which returns 409
managed_runtime_identity_conflict("still owns unfinished work"). After that,mcp-statusandacquirereturn 409runtime_session_not_ready, so the original release is never re-sent. Every later detach repeats this, and the lease stays held. - Cause: when every configuration is
releasing,close()triesbroker.release()before looking up or re-sending the release. The Runtime still holds the open connection, so that call fails, and afterwards the runtime Session is not ready for the re-send path. - With two pinned servers, the second configuration is still
active, so this shortcut is skipped. Detach 2 re-sends both releases and returns 204; on67c908eait stayed stuck. - The new unit test retries an undelivered release with the original immutable identity kills the inverse mutant, yet the one-server shape above still fails on the real stack.
-
A release that timed out (not new, not covered). Here the stdio server's stdout outlives the process by about 45 s, so its release exceeds the 5 s drain bound.
- One server: detach returns 503 while the pipe is open and 204 at 52 s, on both heads.
- Two servers, or a server replaced on
40b05a74: detach stayed 503 for as long as I watched (98 s and 134 s), and the lease stayed held. - On
40b05a74, each retry re-sendsmcp-releasewith the same operation ID and gets back the storedoutcome_unknown/managed_mcp_drain_unknown. The retry never drains again, so the later configurations are never released and their server process stays alive.67c908eaonly looks the release up and is equally stuck. - For example, a stdio server launched through a wrapper whose child keeps stdout open can take this path.
-
Flyway. On
f2618e30the jar carries V20–V22 and V23, and an empty database migrates to 1..23 and starts. On40b05a74, applying this PR first and feat(managed-agent): Add durable remote Shell result delivery #12894's V20–V22 afterwards failed validation (Detected resolved migration not applied to database: 20). Now that feat(managed-agent): Add durable remote Shell result delivery #12894 is on main, that order can no longer happen. -
Regression. The full round 1–8 set holds on
40b05a74, except that I did not rerun the worker SIGKILL scenario. The key subset (S1,list_changed, prompt cancel, busy owner, R1-2, hung-read status/cancel, lost replies, orphaned Session) reran onf2618e30with the same results:Area What I saw S1, three transports exact blob, 409 conflicts, 0 leaks, detach and load list_changedtools/list_changedandresources/list_changedrecover in the same turnRaw-op cancel 202 runningPrompt cancel the turn ends cancelledLate settlement 33.6 s F6 17 reconfigures keep 1 live stdio process R1-2 / R1-5 both hold Busy owner B detaches 204 and the holder is unchanged Hung-read status/cancel 200 / 202 Lost replies 1 effect; revoked-access close 204 SSE restart / HTTP down at close served / detach 204 Orphaned Session once the old writer's store lease expires: load 200, prompt 202, detach 204, lease 0 N1 unchanged: UNKNOWNfrom 37.6 s, turn ends at 638.5 s, then blocked -
Mutants. 9 of 10 are killed. Baselines were cli 226/227 and broker 530 with 1 skipped; the one cli failure is the known 404 load flake in serves MCP cancel during dispatch…. K2 survives because it removes the first
configuringcheck inrefresh(), and the same check is repeated right after the discover dispatch.Mutant Change it inverts Killing test K1 R4-9 route check rejects invalid MCP strings before committing or dispatching (4 cases) K3 R5-2 re-send retries an undelivered release with the original immutable identity K4 R5-3 Object.hasOwnignores inherited notification names without changing the catalog K5 R5-4 tolerated sibling drain keeps a healthy replacement while retaining the undrained predecessor hold K6 / K7 refresh signal / close fence stops a refresh replacement cancelled during intent / acquire / dispatch_started and fences close until it drains K8 late-approval grant renewal honors allow approval for an MCP call while retaining its shared owner J14 R5-5 413 HttpRuntimeTransportTest.oversizedMcpRequestHasDefinitiveClassificationand one moreJ15 R5-7 eviction RuntimeBrokerServiceTest.failedAdoptedReleaseLeavesReclamationReachable -
CI.
40b05a74: 24 checks pass.f2618e30: 22 pass, andTest (ubuntu)andreview-prwere still running when I posted.web-shell E2E SmokeandDesktop Shellare listed as failures, but both jobs were cancelled rather than failed. -
Not verified. Windows/Linux native runs, a real model, and MariaDB this round. I also did not rerun the worker SIGKILL scenario, and did not drive R5-5 or R5-7 end to end.
Evidence: wenshao/qwen-code@78144486/pr12946/r9. It contains per-scenario results for both arms and for f2618e30, the regression summaries, mutants by test name, the Flyway boots, and the harness, including the three new stdio server hooks.
中文版
真实环境验证第九轮 —— 40b05a74 与 f2618e30
接第八轮与作者的 40b05a74 更新说明。本轮也回答了最新 Critical-only 评审提出的问题:R5 各条在这个 head 上是否成立。全程以上一个 head 67c908ea 作为 A/B 基线。
我写报告期间 head 变成了 f2618e30,它合入了已包含 #12894 的 main。这次合并没有改动 hosted-mcp-session.ts 和 managed-mcp-runtime.ts,但 Hosted 的 turn 路径和 Broker 都有变化。所以我重新构建了 f2618e30,在它上面重跑了关键场景。下面每一项结果在两个 head 上都相同。
合并参考(一段话):
-
真实环境中已修复。 下面各项都与
67c908ea做了 A/B:- 我第八轮报告的 R4-9 陷阱;
- 刷新替换时的取消 P1;
- R5-1 和 R5-3;
- R5-4 的配置那一半。
从 main 合入的审批路径,在超过 300 s 的 MCP grant 租期后才批准时也能正常工作。
-
未修复:只 pin 一台服务端的 Session 上的 R5-2。 这是最常见的形状。如果它的第一个
mcp-release没送达 Runtime,detach 会一直 503,Workspace lease 一直被持有,表现与67c908ea完全相同。只有两台及以上服务端的 Session 现在能恢复。 -
不是新问题,但 R5-2 的重试也没覆盖到。 排空超时(
managed_mcp_drain_unknown)的 release,之后每次都会得到同样的存储结果。pin 两台服务端的 Session 永远 detach 不了,两个 head 都一样。 -
Flyway。 feat(managed-agent): Add durable remote Shell result delivery #12894 已经合入。
f2618e30带有 V20–V22 和 V23,空库会迁移到 1..23,合并顺序的问题已经不存在。
建议合并前修掉单服务端 Session 的 R5-2,或者把它与 R3-11 一起明确列为接受的限制。
| R5 发现 | 40b05a74 与 f2618e30 真实环境 |
67c908ea |
|---|---|---|
| R5-1 turn 进行中的替换 | 这一轮以 turn_error 结束,不被锁住;替换在 6.4 s 结算;下一轮使用新服务端 |
这一轮没有终态事件,并被永久锁住;下一个 prompt 409;detach 503 |
| R5-2 release 未送达,单服务端 | 卡死:每次重试 detach 都是 503;lease 被持有 | 卡死(相同) |
| R5-2 release 未送达,两台服务端 | 能恢复:第 2 次 detach 重发两个 release → 204,lease 归零 | 卡死 |
| R5-3 原型名通知 | 被忽略:0 次配置,目录不变 | 每一种都触发一次替换(配置修订 1 → 2 → 3 → 4) |
| R5-4 旧连接退役慢 | 替换结算,使用新服务端 | 替换 503,outcome_unknown |
| R5-5 超大 MCP 操作 → 413 | 本轮未驱动(第八轮中 300 KiB 的参数在到达 Broker 前就被 Harness 拒绝);变异体 J14 被杀 | — |
R5-7 被收养的 SessionContext 移除¹ |
未在真实环境驱动;变异体 J15 被杀 | — |
¹ 编号以文中链接的行内评论为准。Critical-only 评审的部分编号不同:它的"R5-6"即行内的 R5-7;它的"R5-1"指向 WorkspaceRuntimeTransport.java:145,与上表的行内 R5-1 不是同一条,那一项我没有驱动。
-
构建。 先后用
40b05a74和f2618e30全新构建 bundle 和 Spring jar,跑在 MySQL 8.4.11 上。main 新增了@a2a-js/sdk依赖,所以先执行了pnpm install。旧臂是我的67c908ea构建,其 Broker jar 单独保留。stdio 服务端新增三个测试钩子:addtool:运行时修改自己的工具列表;rawnotify:发送原始通知;--sticky:服务端退出后,由孙进程再占住 stdout 约 45 s。
-
R4-9(已修复)。 资源 URI、prompt 名称、参数键或参数值里带孤立代理项时:
- 四种情况都返回 400
invalid_mcp_operation,Broker 请求为 0,服务端读取为 0; - 下一次读取 settled,detach 返回 204,lease 归零;
- 合法 emoji 的对照原样通过。
- 四种情况都返回 400
-
刷新替换时的取消 P1(已修复)。 模型调用
addtool,同一轮里的下一次 refresh 于是发起真正的替换;Broker 代理扣住这次替换的acquire:67c908ea: 2.0 s 时这一轮被取消,但 4.9 s 时仍派发了替换用的mcp-configure。扣住期间 detach 返回 204,在获取完成前就释放了。40b05a74: 取消后没有派发任何 configure。扣住期间 detach 返回 503,没有释放任何东西。下一轮正常,detach 返回 204。- 如果改为扣住替换的
mcp-configure(已进入dispatch_started),两个 head 上都只发送一次,不会重发。
-
R5-1 与 R5-4。 一轮对话正在
remote上执行一个 3 s 的工具时,提交对sticky服务端的显式替换(rev 1 → 2)。旧的 sticky 连接退役要超过 5 s。两个 head 的结果见上表。 -
从 main 合入的审批路径。 在
approvalMode: default下,MCP 工具调用会发起审批请求。我在 310 s 时批准(已超过 300 s 的 grant 租期):工具用原参数执行了 1 次,这一轮以end_turn结束,Broker 没有报错,detach 返回 204。2 s 就批准的对照组表现相同。 -
R5-2,单服务端(未修复)。 Broker 代理对第一个
mcp-release直接返回 503,不转发。- 第 1 次 detach 返回 503。
- 第 2 次 detach 直接走 Broker 的 tool-session release,返回 409
managed_runtime_identity_conflict("still owns unfinished work")。此后mcp-status和acquire都返回 409runtime_session_not_ready,原来的 release 再也没有重发。之后每次 detach 都重复这个过程,lease 一直被持有。 - 原因:所有配置都处于
releasing时,close()会先尝试broker.release(),再去查询或重发 release。这时 Runtime 里的连接仍然开着,这一步必然失败,之后 runtime Session 也就不再就绪,重发路径走不通。 - pin 两台服务端时,第二个配置仍是
active,不会走这个快速路径。第 2 次 detach 重发了两个 release,返回 204;在67c908ea上则一直卡死。 - 新增的单测 retries an undelivered release with the original immutable identity 能杀掉对应的反向变异体,但上面单服务端的形状在真实环境中仍然失败。
-
排空超时的 release(不是新问题,也没被覆盖)。 这里 stdio 服务端的 stdout 比进程多存活约 45 s,所以它的 release 超过了 5 s 的排空上限。
- 单服务端: 管道还开着时 detach 返回 503,52 s 时返回 204,两个 head 相同。
- 两台服务端,或在
40b05a74上被替换过的服务端: 在我观察的整个时间里(98 s 和 134 s)detach 一直是 503,lease 一直被持有。 - 在
40b05a74上,每次重试都用同一个 operation ID 重发mcp-release,得到的都是存下来的outcome_unknown/managed_mcp_drain_unknown。重试不会重新排空,所以后面的配置永远得不到释放,它们的服务端进程也一直活着。67c908ea只查询不重发,同样卡死。 - 例如,通过包装器启动、其子进程一直占着 stdout 的 stdio 服务端,就可能走到这条路径。
-
Flyway。
f2618e30的 jar 带有 V20–V22 和 V23,空库迁移到 1..23 并正常启动。在40b05a74上,如果本 PR 先应用、之后才加入 feat(managed-agent): Add durable remote Shell result delivery #12894 的 V20–V22,校验会失败(Detected resolved migration not applied to database: 20)。现在 feat(managed-agent): Add durable remote Shell result delivery #12894 已在 main 上,这种顺序不会再出现。 -
回归。 第一到第八轮的完整集合在
40b05a74上都仍然成立,只是这次没有重跑 worker SIGKILL 场景。关键子集(S1、list_changed、prompt 取消、busy owner、R1-2、挂起读取的查询/取消、回复丢失、孤儿 Session)在f2618e30上重跑,结果相同:项目 结果 S1,三种 transport blob 逐字节一致,409 冲突,零泄漏,detach 与 load list_changedtools/list_changed和resources/list_changed在同一轮内恢复原始操作取消 202 runningprompt 取消 这一轮以 cancelled结束迟到结算 33.6 s F6 17 次 reconfigure 只有 1 个存活的 stdio 进程 R1-2 / R1-5 都成立 busy owner B 以 204 detach,持有者不变 挂起读取的查询/取消 200 / 202 回复丢失 1 次副作用;撤销访问后 close 204 SSE 重启 / close 时 HTTP 宕机 可继续服务 / detach 204 孤儿 Session 旧写者的存储租约过期后:load 200,prompt 202,detach 204,lease 归零 N1 不变:37.6 s 起为 UNKNOWN,这一轮在 638.5 s 结束,随后被阻塞 -
变异体。 10 个中 9 个被杀。基线为 cli 226/227、broker 530(跳过 1 个);cli 那 1 个失败是已知的 404 负载抖动,出现在 serves MCP cancel during dispatch…。K2 存活,因为它删掉的是
refresh()里的第一次configuring检查,而 discover 派发后紧接着还有一次同样的检查。变异体 反转的改动 杀掉它的测试 K1 R4-9 路由校验 rejects invalid MCP strings before committing or dispatching(4 个用例) K3 R5-2 重发 retries an undelivered release with the original immutable identity K4 R5-3 Object.hasOwnignores inherited notification names without changing the catalog K5 R5-4 容忍旧连接排空失败 keeps a healthy replacement while retaining the undrained predecessor hold K6 / K7 refresh 传递 signal / close 守卫 stops a refresh replacement cancelled during intent / acquire / dispatch_started and fences close until it drains K8 延迟审批时续 grant honors allow approval for an MCP call while retaining its shared owner J14 R5-5 返回 413 HttpRuntimeTransportTest.oversizedMcpRequestHasDefinitiveClassification等J15 R5-7 移除 RuntimeBrokerServiceTest.failedAdoptedReleaseLeavesReclamationReachable -
CI。
40b05a74:24 项通过。f2618e30:22 项通过,发帖时Test (ubuntu)和review-pr仍在运行。web-shell E2E Smoke和Desktop Shell被列为失败,但这两个任务其实是被取消,并非测试失败。 -
未验证。 Windows/Linux 原生运行、真实模型、本轮的 MariaDB。另外没有重跑 worker SIGKILL 场景,R5-5 和 R5-7 也没有做端到端驱动。
证据:wenshao/qwen-code@78144486/pr12946/r9。其中包括两个臂以及 f2618e30 的各场景结果、回归汇总、按测试名归因的变异体、Flyway 启动记录,以及装置脚本(含三个新的 stdio 服务端钩子)。
…e contract to 1.25 Merging main brought in #12946's V23__managed_mcp_records.sql next to this PR's V23__managed_actions.sql; Flyway refuses two migrations with the same version, so the server would not start. The Actions migration becomes V24 with its content unchanged. The contract takes 1.25.0 on top of main's 1.24.0 (#12946), with a v1.25 note for the Actions routes this PR serves.
|
The round 9 release-recovery findings are addressed in follow-up PR #13109, since this PR has merged. The earlier R5-2 fix was incomplete for a single pinned server: its test double did not preserve the Broker state transition after an unsuccessful owner release. The follow-up persists connection drain proof before owner release, settles late physical closure on the original release receipt, and retains old-writer lost-ack recovery without treating network errors as closure proof. Final real-stack verification passed all five scenarios, including a baseline Harness writing the legacy records and a new Harness recovering them. Verification report. Unknown invocations still retain their holds. 第九轮报告的两项释放恢复问题已在后续 PR #13109 修复。此前单 server 的 R5-2 判为完全修复不准确:测试替身遗漏了失败 owner 释放造成的 Broker 状态变化。新 PR 已通过五项最终真实链路验证,包含真实旧写者记录的升级接管;未知调用仍保留占用。 |
A Workspace Session whose connector was built without the Managed Action store silently fell back to the deployment approval mode and skipped the Harness mode confirmation -- the case the design declares fail-closed. createOrLoad now refuses it, so the two remaining null checks in the connector are gone and an un-wired store fails loudly at admission. requireOwner let an IllegalArgumentException from actorKey escape as an unmapped 500. It now answers 403 actor_scope_mismatch, matching the neighbouring workspace routes. endedCode maps cancelled to action_cancelled but no test asserted it: a cancelled Action now fails its response operation with that code and returns no resolution. The design docs still named the Actions migration V23 after it moved to V24; V23 belongs to the Session MCP catalog (#12946).
…3101) * feat(managed-agent): Serve durable permission Actions (D6b) * test(managed-agent): Avoid concurrent Action mock stubbing * fix(managed-agent): renumber the Actions migration to V24 and bump the contract to 1.25 Merging main brought in QwenLM#12946's V23__managed_mcp_records.sql next to this PR's V23__managed_actions.sql; Flyway refuses two migrations with the same version, so the server would not start. The Actions migration becomes V24 with its content unchanged. The contract takes 1.25.0 on top of main's 1.24.0 (QwenLM#12946), with a v1.25 note for the Actions routes this PR serves. * fix(managed-agent): address the D6b review follow-ups A Workspace Session whose connector was built without the Managed Action store silently fell back to the deployment approval mode and skipped the Harness mode confirmation -- the case the design declares fail-closed. createOrLoad now refuses it, so the two remaining null checks in the connector are gone and an un-wired store fails loudly at admission. requireOwner let an IllegalArgumentException from actorKey escape as an unmapped 500. It now answers 403 actor_scope_mismatch, matching the neighbouring workspace routes. endedCode maps cancelled to action_cancelled but no test asserted it: a cancelled Action now fails its response operation with that code and returns no resolution. The design docs still named the Actions migration V23 after it moved to V24; V23 belongs to the Session MCP catalog (QwenLM#12946). --------- Co-authored-by: 易良 <[email protected]> Co-authored-by: yiliang114 <[email protected]>
… panel (QwenLM#13107) * feat(managed-agent): Serve durable permission Actions (D6b) * test(managed-agent): Avoid concurrent Action mock stubbing * feat(web-shell): show and answer Hosted tool approvals in the Managed panel The Managed panel passed pendingApproval={null}, so a Hosted Session waiting on a D6a approval showed nothing to answer. Build on the D6b Actions contract (QwenLM#13101): - The provider gains an optional Actions reader. The Java provider lists requested permission Actions through actions/query and answers through actions/respond with the Action's inputRevision and policyRevision and a per-Action, per-option idempotency key. The Session summary reports the actions capability only when the service does. - action.updated projects to action_updated. It carries no Turn, so the transcript skips it instead of settling the Turn being streamed. - useManagedActions re-reads pending approvals on action_updated, on a stream gap, after an expiry and after an answer. It hides an answered approval and brings it back if the answer fails. - The page maps the pending Action onto the shared approval card, with allow/deny as allow_once/reject_once so labels are localized, and attaches it to the tool row `${turnId}:${functionCallId}` so the row expands and shows its arguments. Not built, type-checked, linted or tested locally. * feat(web-shell): export the pending Action type for custom Managed providers * fix(managed-agent): renumber the Actions migration to V24 and bump the contract to 1.25 Merging main brought in QwenLM#12946's V23__managed_mcp_records.sql next to this PR's V23__managed_actions.sql; Flyway refuses two migrations with the same version, so the server would not start. The Actions migration becomes V24 with its content unchanged. The contract takes 1.25.0 on top of main's 1.24.0 (QwenLM#12946), with a v1.25 note for the Actions routes this PR serves. * fix(web-shell): memoize pending actions to keep the respond callback stable The conditional selected a fresh empty array on every render, so the useCallback dependency changed identity each time and ESLint's react-hooks/exhaustive-deps warning failed the lint gate. * style(web-shell): format the Managed approvals page test with Prettier * fix(web-shell): retry failed approval reads and name them as such A transient failure of actions/query on the first read left a pending approval invisible until a manual Refresh: summary polls do not re-read Actions, and without an Action there is no expiry timer. The page then showed "The approval answer could not be sent" although nothing was sent. - useManagedActions reports loadError and answerError separately, and retries a failed read after 2, 5 and 10 seconds before giving up. retry() reads again on demand and restarts the bound. - The page shows "Pending approvals could not be loaded." with a Retry button for a read failure, and keeps the answer message for a failed answer. - Tests cover the automatic recovery, the retry bound and the manual retry. * fix(web-shell): recover failed approval answers and bound expiry reads * fix(web-shell): Show available managed approval arguments * fix(web-shell): keep the approval card when an answer was not applied * fix(web-shell): keep Managed approvals stable across reloads and ended Actions Treat an unknown session summary as unknown rather than as no Actions capability, so a reload keeps the shown approval answerable. Drop an approval the service reports as ended instead of offering a retry, scope answer failures to the approval that failed, stop retrying client errors, restart the retry ladder whenever reads resume, and carry the arguments-unavailable notice inside the approval card. * fix(web-shell): land only the Critical R1-1 fix for Managed approvals The previous commit also carried the review's Suggestions. The PR is past its review-round budget, so only R1-1 stays here: an unknown session summary keeps the shown approval instead of reading as no Actions capability. The Suggestions move to the follow-up recorded in QwenLM#12867. * fix(web-shell): match Managed approvals to itemId-keyed tool rows Since QwenLM#13037, Managed tool rows are keyed ${turnId}:${itemId} and Java tool Items always carry an itemId, so findManagedApprovalTool no longer found the row for a pending Action: the card could not show that row's arguments or expand it. Keep the producer's call ID on the row and match on it, still within the Action's Turn. Candidate patch from wenshao's round-4 real-stack verification (R4-1). * test(web-shell): pin the page half of the Managed approval reload fix Reverting the page gate back to detail.summary?.capabilities.actions === true kept every test green (G9). Hold a reload's Session summary and assert the shown approval stays and no extra Actions read happens. Candidate test from wenshao's round-3 real-stack verification. * test(web-shell): spell canCancel in the Managed approvals fixtures The three approvals fixtures handed to `summary()` omitted `capabilities.canCancel`, which is a required member, so each one was a TS2741 that no gate in this package reports (both tsconfigs exclude `client/**/*.test.tsx` and vitest does not typecheck). They also modelled a summary the real mapper cannot emit, which always sets both flags. Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmupe138l16 * fix(web-shell): keep the Harness call title on Managed approval cards `toManagedPermissionRequest` reads `tool.title` for the approval `title`, but no producer ever set it: `managedEventsToMessages` only assigned args/status/times/rawOutput, so the left branch was dead and the card description degenerated into a restatement of the heading while the Harness`s own per-call title was dropped. Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmupe138l16 * fix(web-shell): stop Managed approval state outliving its Action Two lifecycle leaks in `useManagedActions`: - a successful re-read cleared `loadError` but never `answerError`, so "the answer could not be confirmed" stayed on screen after the Harness ended the Action, and then labelled the next, unrelated card. The warning now carries the Action it belongs to and is dropped once that Action is no longer pending. - a withdrawn reader cleared `pending` without resetting the retry budget, so a reader restored after the ladder was exhausted got a single attempt with no retry scheduled, contradicting the hook`s own comment. Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmupe138l16 * fix(web-shell): stop retrying approval reads the service already answered R1-5: a 4xx other than 408/429 is the service's answer, not a hiccup, so a deleted Session no longer costs four guaranteed-failing /actions/query requests and a Retry that can never succeed. The classification the submit path already used moves to managed-request-error.ts and is reused, so both failure paths of the feature agree. R1-4: the hook reports whether the pending approvals were read for this Session, so a failed background re-read is named as a refresh instead of claiming nothing could be loaded next to the loaded, actionable card. R1-8: the "arguments are unavailable" caveat renders as a sibling of the approval dialog, so ToolApproval takes an extra description id and the panel describes the caveat too — otherwise a screen-reader user confirms a Hosted tool call hearing only the tool name. Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmupe138l16 * test(web-shell): pin four Managed approval paths that no test observed R1-10, four sites the review mutation-tested as unpinned: 1. the `stream_gap` arm of the re-read trigger — the durable transcript is the only route that reports a gap here, since use-managed-session breaks the live loop on a gap before merging it; the hook doc now says so. 2. the session-switch reset effect and the cross-session mask — the new test switches Session while the previous card is displayed and its failed answer is still flagged, so deleting the effect leaves the warning behind and reverting the mask leaves the stale card answerable. 3. the `pendingApproval` wiring into MessageList — the page test's mock now captures the prop, which is what keeps the turn owning the pending call from being folded away. 4. the `setAnswerError(undefined)` clear on a successful answer — the retry test now ends on the alert being gone, not just on the card. Every assertion was checked to fail with the production line it covers reverted (8/8 mutations RED). Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-closeout/jmupe138l16 * fix(web-shell): isolate late Managed approval replies Ignore answer settlements once their Session or pending Action has changed, and clear or display warnings only for the Action they belong to. Pin late success and failure after Session switches and Action replacement. * fix(web-shell): drop a Managed approval the service reports as ended Answering an Action that already expired, was cancelled or was answered elsewhere returns 409 with a code the contract defines as ended, and no retry can succeed. Keep the card hidden, clear the warning and read the list again instead of offering a retry. A read that still lists the Action shows it again, so a wrong report cannot hide an approvable Action. The 403 actor_scope_mismatch is not an ended Action and keeps its handling. --------- Co-authored-by: Shaojin Wen <[email protected]> Co-authored-by: yiliang114 <[email protected]> Co-authored-by: Shaojin Wen <[email protected]> Co-authored-by: qwen-code-dev-bot <[email protected]> Co-authored-by: yiliang114 <[email protected]> Co-authored-by: Qwen-Coder <[email protected]>













What this PR does
Adds the private
hosted-workspace-mcp/1profile for Stage H1, including the prerequisite Hosted → Broker → Runtime wiring. The Runtime owns stdio, Streamable HTTP and SSE connections and credentials. Models receive pinned tool schemas and use the ordinary durable tool lifecycle; resource reads and prompt gets retain their complete MCP responses in separate operation records. An authenticated public endpoint projects the Session's committed catalog without exposing connection configuration.Configuration changes pin definition, catalog and connection revisions. Recovery retains original operation identities, replays committed results without reconnecting, and resumes only intents proven never dispatched. Dispatched requests use deferred cancellation so native MCP cancellation cannot suppress their settlement replies. Transient discovery failures terminate the current turn without permanently blocking the Session. Reconciliation keeps unknown execution evidence until the original configuration settles; duplicate completed MCP replies do not disrupt other calls. Discovery is checked before new work, idle replaced connections are closed, and HTTP session deletion failure does not prevent proven local closure. Empty configuration intents can be cancelled and their original Broker ownership released. Idle Harness reload renews durable grant revisions before the new writer uses the original connection; an explicitly selected replacement can recover an initially failed definition.
Shared Broker acquisition is intentionally idempotent for an existing owner and exposes the Workspace generation and Runtime binding/generation. File and Shell reservations retain the actual prompt and call identities. Runtime invocation timeout is deployment-configurable up to ten minutes; Hosted observes the original execution for 630 seconds. MCP polling explicitly requests reconciliation and accepts conclusive late results from the original execution; ordinary v2 status reads retain their passive behavior and v3 retains automatic reconciliation.
Turn cancellation also stops refresh replacements before dispatch. Explicit replacement can overlap a turn without permanently blocking it, while already admitted calls retain their pins. Undelivered release controls retry the same idempotent identity; unknown invocations remain query-only. Malformed Unicode is rejected before raw-operation admission, and tool approval renews its grant before execution preparation.
Why it's needed
H0c provides extension records, while the existing private Hosted tool profile supports only file tools. H1 needs a working invocation chain, revision ownership, durable resource/prompt results and failure recovery before MCP can run through the Managed path. This PR includes the necessary narrow Broker control forwarding; unrelated control verbs remain unsupported.
Reviewer Test Plan
How to verify
Evidence (Before & After)
Before: the globally installed CLI rejects Hosted startup; the repository's existing private Hosted profile advertises only file tools. After: a deterministic model drives real stdio and Streamable HTTP tools through the Java Broker, provisioned Runtime and isolated SQL store. Binary resource bytes and both prompt messages survive, duplicates do not add effects, dropped responses reconcile by their original IDs, and normal close succeeds. Detailed E2E observations are posted separately.
Validation includes the repository build, typecheck and bundle, focused TypeScript and Java tests, lint/format checks and independent real-stack failure reproduction and verification. Detailed results and audit evidence are posted in the accompanying verification comment.
Tested on
Environment (optional)
Node.js 22.22.2, JDK 21.0.11, MariaDB 10.11.19, locally built CLI and Java server, deterministic fake model, real local MCP transports and synthetic credentials.
Risk & Scope
Design: English · 简体中文.
Linked Issues
Implements the H1 slice of #12827 and the necessary narrow control prerequisite discussed in #12765. This does not close the remaining Stage H work.
中文说明
本 PR 做了什么
为 H1 增加私有
hosted-workspace-mcp/1配置,并包含 Hosted → Broker → Runtime 调用所需的前置接线。Runtime 持有 stdio、Streamable HTTP、SSE 连接和凭据。模型获得固定修订的工具 schema,调用沿用普通工具的持久化生命周期;资源读取和提示获取在独立操作记录中保留完整 MCP 响应。经过认证的公开接口展示 Session 已提交的目录,不暴露连接配置。配置变更固定定义、目录和连接修订。恢复保留原操作身份,重放已提交结果无需重连,只对能够证明从未派发的 intent 恢复首次发送。已派发请求采用延迟取消,避免原生 MCP 取消抑制结算响应。临时发现失败以错误结束当轮,不会永久阻塞 Session。配置恢复保留未知执行证据,直到原操作明确结算;重复的已完成 MCP 回包不会干扰其他调用。新工作前检查发现状态,空闲旧连接会关闭;HTTP Session 删除失败不阻断已证明的本地关闭。未派发配置 intent 可取消并释放原 Broker 所有权。Harness 空闲重载后先推进持久授权修订,再由新 writer 使用原连接;显式选择的新定义可恢复首次配置失败。
共享 Broker acquire 有意对原 owner 幂等,并暴露 Workspace generation 和 Runtime binding/generation。文件与 Shell 预留保留实际 prompt/call 身份。Runtime 调用超时可由部署配置,最长十分钟;Hosted 按原 execution 观察 630 秒。MCP 轮询显式请求对账并接受原 execution 的明确迟到结果;普通 v2 状态查询保留被动读取语义,v3 保留自动对账。
取消 turn 也会阻止刷新替换在派发前继续执行。显式替换可与 turn 重叠而不永久锁住 Session,已准入调用保留原 pin。未送达的 release 使用同一幂等身份重试,未知调用仍只查询。原始操作准入前拒绝非法 Unicode,工具审批通过后在准备执行前更新 grant。
为什么需要
H0c 提供扩展记录,但现有私有 Hosted 工具配置仅支持文件工具。H1 需要实际可运行的调用链、修订归属、持久化资源/提示结果以及失败恢复,才能经 Managed 路径执行 MCP。本 PR 包含必要且范围有限的 Broker control 转发,其他 control 动词仍不受支持。
评审测试计划
如何验证
证据(变更前后)
变更前:全局 CLI 拒绝 Hosted 启动;仓库已有私有 Hosted 配置只发布文件工具。变更后:确定性模型经 Java Broker、实际启动的 Runtime 和隔离 SQL 存储调用真实 stdio、Streamable HTTP 工具。二进制资源字节及两条提示消息完整保留,重复请求不增加副作用,丢响应后按原 ID 对账,正常关闭成功。详细 E2E 观察结果单独评论。
验证包含仓库 build、typecheck、bundle,定向 TypeScript 和 Java 测试、lint/格式检查,以及独立实栈故障复现和复验。具体结果及审计证据另附验证评论。
测试平台
环境
Node.js 22.22.2、JDK 21.0.11、MariaDB 10.11.19、本地构建 CLI 与 Java 服务、确定性假模型、真实本地 MCP transport 和合成凭据。
风险与范围
设计:English · 简体中文。
关联 Issue
实现 #12827 的 H1 切片,以及 #12765 中相关的必要有限 control 前置接线。本 PR 不关闭 Stage H 的其余工作。