Skip to content

feat(managed-agent): Implement private Hosted MCP runtime (H1) - #12946

Merged
wenshao merged 17 commits into
mainfrom
codex/managed-mcp-h1
Sep 30, 2026
Merged

wenshao merged 17 commits into
mainfrom
codex/managed-mcp-h1

Conversation

@wenshao

@wenshao wenshao commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Adds the private hosted-workspace-mcp/1 profile for Stage H1, including the prerequisite Hosted → Broker → Runtime wiring. The Runtime owns stdio, Streamable HTTP and SSE connections and credentials. Models receive pinned tool schemas and use the ordinary durable tool lifecycle; resource reads and prompt gets retain their complete MCP responses in separate operation records. An authenticated public endpoint projects the Session's committed catalog without exposing connection configuration.

Configuration changes pin definition, catalog and connection revisions. Recovery retains original operation identities, replays committed results without reconnecting, and resumes only intents proven never dispatched. Dispatched requests use deferred cancellation so native MCP cancellation cannot suppress their settlement replies. Transient discovery failures terminate the current turn without permanently blocking the Session. Reconciliation keeps unknown execution evidence until the original configuration settles; duplicate completed MCP replies do not disrupt other calls. Discovery is checked before new work, idle replaced connections are closed, and HTTP session deletion failure does not prevent proven local closure. Empty configuration intents can be cancelled and their original Broker ownership released. Idle Harness reload renews durable grant revisions before the new writer uses the original connection; an explicitly selected replacement can recover an initially failed definition.

Shared Broker acquisition is intentionally idempotent for an existing owner and exposes the Workspace generation and Runtime binding/generation. File and Shell reservations retain the actual prompt and call identities. Runtime invocation timeout is deployment-configurable up to ten minutes; Hosted observes the original execution for 630 seconds. MCP polling explicitly requests reconciliation and accepts conclusive late results from the original execution; ordinary v2 status reads retain their passive behavior and v3 retains automatic reconciliation.

Turn cancellation also stops refresh replacements before dispatch. Explicit replacement can overlap a turn without permanently blocking it, while already admitted calls retain their pins. Undelivered release controls retry the same idempotent identity; unknown invocations remain query-only. Malformed Unicode is rejected before raw-operation admission, and tool approval renews its grant before execution preparation.

Why it's needed

H0c provides extension records, while the existing private Hosted tool profile supports only file tools. H1 needs a working invocation chain, revision ownership, durable resource/prompt results and failure recovery before MCP can run through the Managed path. This PR includes the necessary narrow Broker control forwarding; unrelated control verbs remain unsupported.

Reviewer Test Plan

How to verify

  • Configure deployment-owned stdio and HTTP MCP definitions, create an explicitly pinned private Hosted Session, and verify that its advertised tools execute through the Runtime and their results are committed before the next model request.
  • Read a binary resource and a prompt with multiple messages. Repeat the same operation ID and then reuse it with different arguments: the first should replay, the second should reject, and neither should repeat the physical effect.
  • Drop configuration and invocation responses, query the original IDs, and verify recovery without another invocation. Revoke Workspace read access and close the existing Session; it should release its original owner. Replay a completed operation after Session reload while the Broker is unavailable; its committed result should remain available without reconnecting.
  • Exercise cancellation before dispatch and deferred cancellation after dispatch, a tool lasting over 25 seconds, late settlement after a configured timeout, discovery invalidation, more than 16 sequential replacements, HTTP outage at close, and a refused first install. Confirm that the original effects are neither resent nor falsely settled, idle retired processes exit, and close works after the Workspace becomes available. Restart an idle Harness while retaining the Runtime and confirm new work and detach can use renewed grants.

Evidence (Before & After)

Before: the globally installed CLI rejects Hosted startup; the repository's existing private Hosted profile advertises only file tools. After: a deterministic model drives real stdio and Streamable HTTP tools through the Java Broker, provisioned Runtime and isolated SQL store. Binary resource bytes and both prompt messages survive, duplicates do not add effects, dropped responses reconcile by their original IDs, and normal close succeeds. Detailed E2E observations are posted separately.

Validation includes the repository build, typecheck and bundle, focused TypeScript and Java tests, lint/format checks and independent real-stack failure reproduction and verification. Detailed results and audit evidence are posted in the accompanying verification comment.

Tested on

OS Status
🍏 macOS ✅ tested
🪟 Windows ⚠️ not tested
🐧 Linux ⚠️ not tested

Environment (optional)

Node.js 22.22.2, JDK 21.0.11, MariaDB 10.11.19, locally built CLI and Java server, deterministic fake model, real local MCP transports and synthetic credentials.

Risk & Scope

  • Main risk or tradeoff: this crosses Session persistence, Broker ownership and Runtime lifecycle boundaries. Unknown effects remain blocked when their original Runtime cannot provide evidence. Limits are per Runtime instance: 16 connections, 32 in-flight calls, 16 KiB per catalog category and 60 KiB per raw result. New Hosted Sessions accept at most 16 server pins; saved legacy Sessions retain their original pins so they can load and detach. Live replacement needs one free connection slot: with all 16 slots occupied by pinned servers, a catalog change or replacement keeps failing until detach/load; deployments needing live refresh must reserve that slot. Capacity failures during prompt admission, configuration replacement and raw-operation initialization expose a distinct quota error; failures after prompt admission remain turn errors.
  • Retained H1 limits: one attached MCP owner holds the tenant/storage lease until detach, including between turns. A lost connection or a server that never replies can consume the full 630-second Hosted observation window despite a shorter Runtime timeout or cancellation. Every pinned server must refresh before a model request, so one unavailable server also blocks text-only turns and healthy-server calls until it recovers. Tool settlement after the 630-second observation window and Harness restart during an unfinished turn still need checkpoint recovery. Runtime receipt history and closed-connection tombstones grow until process exit; concurrency limits do not bound that memory. These remain follow-up work before production enablement.
  • Not validated / out of scope: production profile enablement, SDK reverse clients, aggregate budgets across processes, object-storage results, Windows/Linux and the full production restart matrix. SSE and concurrent quotas have focused transport/contract coverage.
  • Breaking changes / migration notes: the additive MCP profile is opt-in. Migration V23 permits extension records without task projections and follows the V20/V21/V22 publication migrations now merged from feat(managed-agent): Add durable remote Shell result delivery #12894. The branch includes these prerequisites; deploy the combined migrations in increasing order. If MCP V23 has already been applied, any later pending publication migrations must be renumbered above it; do not enable out-of-order migration. MCP operations do not become Session tasks. The generic controls from feat(serve): implement generic Broker provider controls #12868 are integrated, preserving original-owner recovery after access revocation and drain-before-release ordering.

Design: English · 简体中文.

Linked Issues

Implements the H1 slice of #12827 and the necessary narrow control prerequisite discussed in #12765. This does not close the remaining Stage H work.

中文说明

本 PR 做了什么

为 H1 增加私有 hosted-workspace-mcp/1 配置,并包含 Hosted → Broker → Runtime 调用所需的前置接线。Runtime 持有 stdio、Streamable HTTP、SSE 连接和凭据。模型获得固定修订的工具 schema,调用沿用普通工具的持久化生命周期;资源读取和提示获取在独立操作记录中保留完整 MCP 响应。经过认证的公开接口展示 Session 已提交的目录,不暴露连接配置。

配置变更固定定义、目录和连接修订。恢复保留原操作身份,重放已提交结果无需重连,只对能够证明从未派发的 intent 恢复首次发送。已派发请求采用延迟取消,避免原生 MCP 取消抑制结算响应。临时发现失败以错误结束当轮,不会永久阻塞 Session。配置恢复保留未知执行证据,直到原操作明确结算;重复的已完成 MCP 回包不会干扰其他调用。新工作前检查发现状态,空闲旧连接会关闭;HTTP Session 删除失败不阻断已证明的本地关闭。未派发配置 intent 可取消并释放原 Broker 所有权。Harness 空闲重载后先推进持久授权修订,再由新 writer 使用原连接;显式选择的新定义可恢复首次配置失败。

共享 Broker acquire 有意对原 owner 幂等,并暴露 Workspace generation 和 Runtime binding/generation。文件与 Shell 预留保留实际 prompt/call 身份。Runtime 调用超时可由部署配置,最长十分钟;Hosted 按原 execution 观察 630 秒。MCP 轮询显式请求对账并接受原 execution 的明确迟到结果;普通 v2 状态查询保留被动读取语义,v3 保留自动对账。

取消 turn 也会阻止刷新替换在派发前继续执行。显式替换可与 turn 重叠而不永久锁住 Session,已准入调用保留原 pin。未送达的 release 使用同一幂等身份重试,未知调用仍只查询。原始操作准入前拒绝非法 Unicode,工具审批通过后在准备执行前更新 grant。

为什么需要

H0c 提供扩展记录,但现有私有 Hosted 工具配置仅支持文件工具。H1 需要实际可运行的调用链、修订归属、持久化资源/提示结果以及失败恢复,才能经 Managed 路径执行 MCP。本 PR 包含必要且范围有限的 Broker control 转发,其他 control 动词仍不受支持。

评审测试计划

如何验证

  • 配置由部署管理的 stdio 和 HTTP MCP 定义,创建显式固定修订的私有 Hosted Session,确认模型看到的工具在 Runtime 执行,并且结果在下一次模型请求前提交。
  • 读取二进制资源和包含多条消息的提示。重复相同操作 ID,再以不同参数复用该 ID:前者应重放,后者应拒绝,两者都不能重复物理副作用。
  • 丢弃配置与调用响应,再按原 ID 查询,确认恢复不会重新调用。撤销 Workspace 读取权限后关闭现有 Session,确认原 owner 被释放。Session 重载后使 Broker 不可用,重放已完成操作,确认可以直接取得已提交结果且没有重连。
  • 覆盖派发前取消与派发后延迟取消、超过 25 秒的工具、配置超时后的迟到结算、目录失效、连续超过 16 次替换、关闭时 HTTP 故障,以及首次配置被拒。确认原副作用不被重发或错误结算、空闲旧进程退出,Workspace 可用后可以直接关闭。保持 Runtime 存活,仅重启空闲 Harness,确认新工作和 detach 使用更新后的 grant。

证据(变更前后)

变更前:全局 CLI 拒绝 Hosted 启动;仓库已有私有 Hosted 配置只发布文件工具。变更后:确定性模型经 Java Broker、实际启动的 Runtime 和隔离 SQL 存储调用真实 stdio、Streamable HTTP 工具。二进制资源字节及两条提示消息完整保留,重复请求不增加副作用,丢响应后按原 ID 对账,正常关闭成功。详细 E2E 观察结果单独评论。

验证包含仓库 build、typecheck、bundle,定向 TypeScript 和 Java 测试、lint/格式检查,以及独立实栈故障复现和复验。具体结果及审计证据另附验证评论。

测试平台

OS 状态
🍏 macOS ✅ 已测试
🪟 Windows ⚠️ 未测试
🐧 Linux ⚠️ 未测试

环境

Node.js 22.22.2、JDK 21.0.11、MariaDB 10.11.19、本地构建 CLI 与 Java 服务、确定性假模型、真实本地 MCP transport 和合成凭据。

风险与范围

  • 主要风险或取舍:跨越 Session 持久化、Broker 归属和 Runtime 生命周期边界。原 Runtime 无法提供证据时,未知副作用保持阻塞。限制按每个 Runtime 实例计算:16 个连接、32 个在途调用、每类目录 16 KiB、每份原始结果 60 KiB。新建 Hosted Session 最多接受 16 个 server pin;旧 Session 保留原 pin,仍可加载并 detach。实时替换需要一个空闲连接槽位:16 个槽位均被固定 server 占用时,目录变更或替换持续失败,需 detach/load 恢复;需要实时刷新的部署应预留空位。prompt 准入、配置替换及原始操作初始化中的容量不足返回明确配额错误;prompt 已准入后的失败仍以 turn 错误报告。
  • 保留的 H1 限制:attached MCP owner 在 detach 前持续占有 tenant/storage 租约,包括 turn 间隙。连接丢失或 server 永不回复可能耗尽 630 秒 Hosted 观察窗口,更短的 Runtime 超时或取消不会缩短该窗口。模型请求前必须刷新所有固定 server,一个 server 不可用也会阻塞纯文本 turn 和健康 server 的调用,直到它恢复。工具在 630 秒观察窗口后结算或 Harness 在未完成 turn 中途重启,仍需后续 checkpoint 恢复。Runtime 历史回执和已关闭连接 tombstone 随历史增长,直到进程退出;并发配额不限制这部分内存。这些是生产启用前的后续工作。
  • 未验证或不在范围内:生产配置启用、SDK 反向客户端、跨进程总预算、对象存储结果、Windows/Linux 和完整生产重启矩阵。SSE 与并发配额有定向 transport/契约测试覆盖。
  • 兼容性与迁移:新增 MCP 配置需要显式启用。V23 迁移允许扩展记录不产生任务投影,位于 feat(managed-agent): Add durable remote Shell result delivery #12894 已合入的 V20/V21/V22 publication migration 之后。本分支已包含这些前置,部署时按递增顺序执行全部迁移。如果 MCP V23 已执行,后到且尚未应用的 publication migration 必须改到更高编号,不能启用 out-of-order migration 绕过。MCP 操作不成为 Session 任务。已合入 feat(serve): implement generic Broker provider controls #12868 的通用 control,保留撤权后的原 owner 恢复及先 drain 再 release 的顺序。

设计:English · 简体中文。

关联 Issue

实现 #12827 的 H1 切片,以及 #12765 中相关的必要有限 control 前置接线。本 PR 不关闭 Stage H 的其余工作。

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

E2E and pre-commit verification

Verified on macOS with Node.js 22.22.2, JDK 21.0.11 and an isolated MariaDB 10.11.19 instance. The locally built Hosted CLI used the Java Broker and provisioned Runtime with real stdio and Streamable HTTP MCP servers; only the model was a deterministic fixture. Credentials and test records were synthetic.

  • Three model requests completed with exactly two physical tool calls, one per transport. The SQL results_ready checkpoint was committed before the next model request.
  • Binary resource responses retained their exact base64 bytes; the prompt retained both messages. There were exactly five physical MCP operations total: two tools, two resource reads and one prompt get.
  • Duplicate operation IDs replayed; conflicting content was rejected without another effect. Dropped configuration and invocation responses recovered through the original IDs without resending the effect.
  • After setting Workspace can_read=FALSE in SQL, the existing owner's acquire, MCP release/status and final Broker release all returned 200. The two connections closed on their original owner.
  • After close and reload, the Broker proxy returned 503. Replaying the committed prompt result and closing again added zero Broker requests and zero MCP effects.
  • All test process groups were confirmed exited after completion.

The first expanded run used a stale Broker dependency embedded in the Spring JAR and failed the revoked-access close case. After reinstalling the Broker dependency and a clean server package, all 81 embedded Broker classes matched the current compiled classes; the complete E2E above passed. Final CLI SHA-256: 604ef0a1f586b0aae9b6a96dc07e650e06093e452ebf3ea91da7797aefa17ca1. Final server JAR SHA-256: 3e967c89df9f2ca9e8a82029a4d62524fb435246f840d932c8d498ba94c75064.

Additional checks passed: root build/typecheck/bundle; 904 CLI + 346 Core + 157 Broker + 43 Java server tests; changed-file ESLint/Prettier; staged diff checks. Independent failure reproductions verified both dispatch/cancel race orderings, durable replay, real 25-second timeout followed by native cancellation and late settlement, discovery invalidation during listing, original-owner recovery and grant/effect identity isolation.

Repeated open-ended and targeted audits found defects that were fixed before commit. The final two consecutive passes found no remaining actionable issues in snapshot c35e3302c8cf685f99945f9a757b4c5dfe4b3a3be9cefe7cc0591a13b13d6341, based on origin/main at bd45b95f826b3873e25ae84d78ae35f1599c141a.

Windows/Linux and the full production restart matrix were not exercised locally. Full SQL E2E covers stdio and Streamable HTTP; SSE, replacement, quotas and cancellation have focused transport/contract coverage. Production profile enablement and SDK reverse clients remain outside this private H1 slice.

@wenshao
wenshao marked this pull request as ready for review September 28, 2026 14:05
@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification — H1 Hosted MCP runtime @ e3ecba66

Merge reference in one paragraph. On a locally built real stack, everything this PR claims holds on both MySQL 8.4.11 and MariaDB 10.11.18. That includes SSE, which the PR covered only in unit tests, and every recovery claim in the author's E2E note, each re-run independently. The CI red leg I found at a390b5d5 (process-env guard) was fixed by the author during my run. The PR is not ready to land as it stands, for three reasons:

  1. It is now CONFLICTING with main. feat(serve): add gated Hosted foreground Shell turns #12848 landed and collides in 8 serve/ files; the 6 production files among them hold 17 hunks.
  2. It collides with open feat(managed-agent): Add durable remote Shell result delivery #12894 on Flyway V19. Git does not flag this; the merged server fails to start.
  3. Five behaviour findings (F1–F5). Ordinary events leave a Session stuck or its MCP unusable, and through F1 they also lock the whole Workspace. The events are: a user cancel, a tool that runs longer than 25 s, an MCP server outage or restart, and a list_changed notification.

Because the profile is private and opt-in, maintainers could accept F1–F5 as documented H1 limits with follow-up issues. I would not let F3–F5 reach any enablement.

What was run

  • Build under test. The a390b5d5 bundle (dist/cli.js) plus the Spring jar with the embedded Broker. During the run the head moved to e3ecba66, which adds only 16 lines to process-env-guard.test.ts. The diff of all other paths is empty, so every runtime result below applies to e3ecba66.
  • Stack. Packaged Hosted Harness (qwen serve --profile hosted-harness) → embedded Java Broker (session store on) → provisioned Runtime worker with a deployment manifest in QWEN_MANAGED_MCP_CONFIG → real MCP servers built on the repo's @modelcontextprotocol/sdk 1.30.0:
    • stdio, as a child process of the worker;
    • Streamable HTTP, with a Bearer header;
    • SSE, with a Bearer header.
  • Instrumentation. Only the model is a fixture. Broker traffic goes through a counting and fault-injecting proxy. Every physical MCP request is appended to a JSONL ledger. The test-only piece is a small Java main that seeds Workspaces and supplies the catalog actor.
  • Environment. MySQL 8.4.11 and MariaDB 10.11.18 (Docker), Node 22.23.2, JDK 25 (release 21).
  • Scope. 11 scenarios, 24 mutants (16 TS + 8 Java), the focused TS/Java suites, and a trial merge with feat(managed-agent): Add durable remote Shell result delivery #12894.

✅ Confirmed on the real stack

happy path

  • One tool round across all three transports. It produced exactly 3 physical effects. The checkpoint phase observed at the second model request was results_ready.
  • Resource and prompt operations.
    • A binary resource_read came back byte-exact (300 bytes).
    • prompt_get over SSE kept both messages.
    • Repeating the same operationId returned the identical response.
    • Reusing the ID with a different URI or server was refused, with 0 extra effects.
  • Catalogs. Neither the private nor the public catalog contains the token, Bearer, URL, command or env; I checked 10 needles and got 0 hits. MCP records have task_kind NULL, and the public /tasks list stays empty.
  • Detach and load. Detach releases all 3 connections and the Broker tool-session. Load creates a new owner mcp:<sid>:<digest>, and replaying a committed prompt result is identical.
  • Author's recovery claims, all reproduced (the triage bot listed these as "not verified"):
    • Dropped replies. Dropping both the configure reply and the invoke reply recovers through mcp-status on the original IDs, with exactly 1 physical read.
    • Revoked access. With can_read=FALSE, close returns 204 and the lease is released; new work is refused (Broker acquire 409).
    • Broker down. After a reload, with the Broker returning 503 for everything, the committed replay returns 202, close returns 204, with 0 Broker requests and 0 effects.
    • Stale catalogs. A stale catalog never authorised a call: every refused call in F2 below has 0 effects.
  • Replacement during an in-flight call. Replacing r1 → r2 while a 6 s call ran on r1 returned 202. The call finished on r1, and the next turn advertised and used only r2.
  • Limits. A 65-tool server gets tools: partial with 33 tools advertised (16 KiB cap). A 70 KB result becomes managed_mcp_output_limit, and the Session stays healthy.

Findings

# Trigger (ordinary event) Observed Scope
F1 An MCP Session stays attached between turns It keeps the Workspace execution lease, so other Sessions in the Workspace can't run tools Workspace
F2 */list_changed from the server, or an SSE server restart Every later MCP call gets managed_mcp_catalog_conflict, 0 effects, until someone reconfigures by hand Session MCP
F3 Cancel of a running resource_read or prompt_get outcome_unknown for good; prompt 409, detach 503, lease held Session + Workspace
F4 Prompt cancel during an MCP tool, or a tool taking more than 25 s Broker execution UNKNOWN; a late success doesn't heal it; detach 503 Session + Workspace
F5 Streamable HTTP server unreachable at close detach 503, and still 503 after the server comes back Session + Workspace
F6 Each reconfiguration The previous connection stays open; the 17th reconfigure hits the quota Session MCP
F7 First install refused (Workspace busy) Detach 503 until another prompt is sent Session

F1 — An attached MCP Session holds the Workspace lease between turns. HostedWorkspaceToolTurn skips broker.release() for the MCP profile, and HostedMcpSession releases only in close(). I put three Sessions in one Workspace; the outcome is the same on both databases:

  • B (MCP profile), while A sits idle but attached: POST /prompt → 503 hosted_prompt_admission_failed (the Broker acquire returned 409).
  • C (files profile): turn_error, and nothing is written.
  • After A detaches, B works, and then B blocks C in the same way.

So one attached MCP Session makes its Workspace unusable for tools in every other Session. Combined with F3–F5, one stuck Session locks the Workspace indefinitely.

F1

F3/F4 — Cancel, or a call longer than 25 s, leaves the Session and its Workspace stuck. Per the MCP spec, the SDK server drops the response to a request it was told to cancel. I checked the success and error paths in shared/protocol.js. So after notifications/cancelled no settlement ever comes back.

Case Server ledger Result
Cancel a resource_read (HTTP on MySQL, SSE on MariaDB) Handler aborted within 1.5 s outcome_unknown 12 s later. Status recoveryBlocked, next prompt 409, detach 503, lease held, a fresh Session in the same Workspace 503.
Prompt cancel while an MCP tool runs Aborted 0.5 s after the cancel The turn gets no terminal event. Same 409 / 503 / held / 503.
A 32 s tool, no cancel Server answered at 41 s The Runtime timed out at 25 s (TIMEOUT_MS, hard-coded; the manifest rejects a timeout field).

The late reply in the 32 s case does settle the Runtime's own view. But qwen_tool_execution.execution_state stays UNKNOWN and is never re-queried, so the Session never heals. Close runs mcp-release 200, then tool-sessions/mcp:<sid>:release returns 409, and detach returns 503. For comparison, ordinary MCP uses MCP_DEFAULT_TIMEOUT_MSEC = 10 min, configurable per server.

F3/F4

F2 / F5 / F6 / F7 — connection lifecycle after the first turn.

F2 F5 F6 F7

  • F2. A list_changed notification for resources or tools bumps the single catalogRevision that every call kind is checked against, and the Harness never dispatches mcp-discover. From then on:

    • the next tool call in the same turn gets catalog_conflict;
    • so does every later turn;
    • so does prompt_get;
    • and the Harness catalog still says revision 1, complete.

    An SSE server restart has the same effect, because onerror retires the connection. The SDK's McpServer emits list_changed whenever a tool is registered, enabled or disabled after connect, so this trigger is not exotic.

  • F5. Detach returns 503 within 1 s: 678 ms on MySQL, 228 ms on MariaDB. After the server comes back, detach returns 503 again. The configuration stays releasing and the lease stays held. terminateSession() throws when the DELETE cannot be sent, although that connection had 0 pending operations.

  • F6. Each POST /mcp/configurations left the previous stdio process running: 1 → 16 live processes for one server in one Session. The 17th reconfiguration got 409. Only detach reaped them, and that took 6.9 s. This is stage-2 point 4 of the triage bot's review, now measured. Together with F2, the documented recovery path works about 15 times per Session.

  • F7. After a refused first install, status reports recoveryBlocked:true while /prompt still admits. Detach stays 503 even after the Workspace frees up. It succeeds only after a later prompt dispatches the leftover configuration intent.

Possible directions (the choice is the maintainers'):

  • F1: either stop holding the storage lease between turns for MCP owners, or document "one attached MCP Session per Workspace" as an H1 limit.
  • F2: dispatch the existing mcp-discover on catalog_conflict / stale, and reconnect with a new generation after a transport error.
  • F3: for resource_read / prompt_get, settle a cancelled operation as a non-blocking cancelled, since its result is data only.
  • F4:
    • make the timeout a manifest field;
    • either don't send notifications/cancelled for tool calls, or bound the wait and settle them;
    • re-query the execution status on close or load, so a late success is reconciled.
  • F5: with 0 pending operations, treat a failed or 405 DELETE as released.
  • F6: close a retired connection once its pending set drains.
  • F7: have close() locally settle configuration intents that are still at execution: intent, as not_started_proven, the same way operations are handled.

Gates

gates

  • CI.
    • At a390b5d5, Test (ubuntu) was red: process-env-guard.test.ts flagged QWEN_MANAGED_MCP_CONFIG, PATH and SystemRoot.
    • The author fixed it in e3ecba66, and it is now green: 22 pass, 2 pending.
    • Locally the guard passes 4 of 5 runs; the one failure was at a load average near 100.
  • Main. The PR is CONFLICTING after feat(serve): add gated Hosted foreground Shell turns #12848 landed: 8 files, 17 hunks in the 6 production files (8 of them in hosted-workspace-tool-turn.ts). That file is where F1 and F4 live, so the post-merge head needs a re-run of the relevant scenarios.
  • Merge order with feat(managed-agent): Add durable remote Shell result delivery #12894 (a7c0a391).
    • Both PRs add V19 (V19__managed_mcp_records.sql and V19__managed_tool_publication.sql), and git reports no conflict on those files.
    • The trial-merged jar on MySQL 8.4 fails at startup with Found more than one migration with version 19.
    • The trial merge also has 8 textual conflicts.
    • Whichever PR lands second has to renumber its migration.
  • Local suites at a390b5d5.
    • cli focused: 79/79. core focused: 230/230. managed-agent-server: 180/180. runtime-broker: 395/396.
    • The one Broker failure, HttpRuntimeTransportTest#closesTheConnectionOnTheDeadlineAndOnCallerCancel, also fails on main bd45b95f on this host under load, and passes 5/5 alone.
  • Mutation. Java: 8/8 killed. TS: 13/16 killed. Three survivors point at untested behaviour:
    • M7: a -32601 from a tools-only server should mean an empty complete list.
    • M11: the prompt route's hasPendingOperations() check. This guard returned the 409 in F3, yet no unit test pins it.
    • M13: configurations owned by two Runtime Sessions.

Smaller notes

  • Identity conflicts on /mcp/operations and /mcp/configurations return 503, which looks retryable. A 409 would fit better.
  • Model-facing tool names are mcp_<32 hex>. They include the catalog revision and connection generation, so every name changes on each load. After a reload, the transcript refers to tools that are no longer advertised.
  • The Runtime's error receipts reach the model as bare codes, e.g. managed_mcp_catalog_conflict.
  • stdio servers run with HOME set to the Workspace directory (observed: HOME = cwd = <ws>/child). Anything a server caches under ~ therefore lands in the user's Workspace.

Not verified

Harness, scenarios, mutants and raw results: wenshao/qwen-code@a5ef4914/pr12946

中文版

真实环境验证 —— H1 Hosted MCP runtime @ e3ecba66

合并参考(一段话): 在本地构建的真实环境中,本 PR 声称的内容在 MySQL 8.4.11 与 MariaDB 10.11.18 上均成立。其中包括 PR 只有单测覆盖的 SSE,也包括作者 E2E 说明里的每一条恢复主张(均独立重跑)。我在 a390b5d5 发现的 CI 红腿(process-env guard)已由作者在我验证过程中修复。目前仍不建议直接合入,原因有三:

  1. 现在与 main 冲突。 feat(serve): add gated Hosted foreground Shell turns #12848 已合入,与本 PR 在 8 个 serve/ 文件上冲突,其中 6 个生产文件共 17 个冲突块。
  2. 与在飞的 feat(managed-agent): Add durable remote Shell result delivery #12894 撞 Flyway V19。 git 不会提示,合并后的服务端无法启动。
  3. 五个行为问题(F1–F5)。 普通事件就会让 Session 卡死或 MCP 不可用,经由 F1 还会锁住整个 Workspace。这些事件包括:用户取消、超过 25 s 的工具、MCP 服务端宕机或重启、list_changed 通知。

该 profile 是私有且需显式启用的,维护者可以选择把 F1–F5 记为 H1 已知限制并另开后续 issue。但我认为 F3–F5 不应带入任何生产启用。

跑了什么

  • 被测构建。 a390b5d5 的 bundle(dist/cli.js)加内嵌 Broker 的 Spring jar。验证途中 head 前进到 e3ecba66,只在 process-env-guard.test.ts 增加了 16 行,其余路径 diff 为空,因此以下运行时结果全部适用于 e3ecba66。
  • 链路。 打包后的 Hosted Harness(qwen serve --profile hosted-harness)→ 内嵌 Java Broker(开启 session store)→ 由部署 manifest(QWEN_MANAGED_MCP_CONFIG)拉起的 Runtime worker → 基于仓库自带 @modelcontextprotocol/sdk 1.30.0 的真实 MCP 服务端:
    • stdio,作为 worker 的子进程;
    • Streamable HTTP,带 Bearer 头;
    • SSE,带 Bearer 头。
  • 观测手段。 只有模型是假的。Broker 流量经计数、可注入故障的代理;每个物理 MCP 请求都写入 JSONL 台账。测试专用部分只有一个 Java main,用于预置 Workspace 并提供目录接口所需的 actor。
  • 环境。 MySQL 8.4.11 与 MariaDB 10.11.18(Docker)、Node 22.23.2、JDK 25(release 21)。
  • 范围。 11 个场景、24 个变异体(16 个 TS + 8 个 Java)、定向 TS/Java 测试套件,以及与 feat(managed-agent): Add durable remote Shell result delivery #12894 的试合并。

✅ 真实链路上确认成立

  • 一轮工具调用覆盖三种 transport。 恰好 3 次物理副作用。第二次模型请求时观察到的检查点阶段为 results_ready。
  • 资源与提示操作。
    • 二进制 resource_read 逐字节一致(300 字节)。
    • 经 SSE 的 prompt_get 两条消息都在。
    • 同一 operationId 重复请求返回相同响应。
    • 用不同 URI 或不同 server 复用该 ID 会被拒绝,额外副作用为 0。
  • 目录。 私有与公开目录都不含 token、Bearer、URL、command 或 env(检查 10 个关键串,0 命中)。MCP 记录的 task_kind 为 NULL,公开 /tasks 列表为空。
  • detach 与 load。 detach 释放全部 3 个连接和 Broker tool-session。load 生成新 owner mcp:<sid>:<digest>,重放已提交的 prompt 结果与原结果一致。
  • 作者的恢复主张全部复现(triage 机器人把这些列为"未验证"):
    • 丢回包。 同时丢弃 configure 回包和 invoke 回包,经原 ID 的 mcp-status 恢复,物理读取恰好 1 次。
    • 撤销权限。 can_read=FALSE 后 close 返回 204、lease 已释放;新工作被拒(Broker acquire 409)。
    • Broker 不可用。 reload 之后 Broker 对所有请求返回 503:已提交结果重放返回 202,close 返回 204,Broker 请求 0 次、副作用 0 次。
    • stale 目录。 stale 目录从未授权过调用:下文 F2 中所有被拒调用的副作用都是 0。
  • 在途调用期间替换配置。 在 r1 上有 6 s 调用在跑时把 r1 替换为 r2,返回 202。该调用在 r1 上完成,下一轮只公布并使用 r2。
  • 限额。 65 个工具的服务端被标为 tools: partial,公布 33 个(16 KiB 上限)。70 KB 结果变为 managed_mcp_output_limit,Session 保持正常。

问题

# 触发(普通事件) 实测 影响范围
F1 MCP Session 在两轮之间保持 attach 一直持有 Workspace 执行 lease,同 Workspace 其他 Session 无法跑工具 Workspace
F2 服务端发 */list_changed,或 SSE 服务端重启 之后所有 MCP 调用都是 managed_mcp_catalog_conflict、副作用 0,直到手工重新配置 Session 的 MCP
F3 取消运行中的 resource_read / prompt_get 永久 outcome_unknown;prompt 409、detach 503、lease 不释放 Session + Workspace
F4 MCP 工具运行中取消 prompt,或工具超过 25 s Broker execution 为 UNKNOWN;迟到的成功也不能恢复;detach 503 Session + Workspace
F5 close 时 Streamable HTTP 服务端不可达 detach 503,服务端恢复后仍 503 Session + Workspace
F6 每次重新配置 旧连接不关闭;第 17 次触发限额 Session 的 MCP
F7 首次安装被拒(Workspace 忙) 在再发一次 prompt 之前 detach 一直 503 Session

F1 —— 已 attach 的 MCP Session 在两轮之间一直持有 Workspace lease。 HostedWorkspaceToolTurn 对 MCP profile 跳过 broker.release(),HostedMcpSession 只在 close() 里释放。我在同一 Workspace 放了三个 Session,两种数据库结果一致:

  • B(MCP profile),A 空闲但仍 attach 时: POST /prompt → 503 hosted_prompt_admission_failed(Broker acquire 返回 409)。
  • C(files profile): turn_error,没有写入任何文件。
  • A detach 后 B 可以工作,随后 B 以同样方式挡住 C。

也就是说,一个已 attach 的 MCP Session 会让同 Workspace 其他所有 Session 无法使用工具。叠加 F3–F5,一个卡死的 Session 会无限期锁住整个 Workspace。

F3/F4 —— 取消、或超过 25 s 的调用,让 Session 及其 Workspace 卡死。 按 MCP 规范,SDK 服务端不会回复已被取消的请求(我查了 shared/protocol.js 的成功与错误两条分支)。所以发出 notifications/cancelled 之后,永远等不到结算。

情形 服务端台账 结果
取消 resource_read(MySQL 上走 HTTP,MariaDB 上走 SSE) handler 在 1.5 s 内被中止 12 s 后仍为 outcome_unknown。状态 recoveryBlocked,下一次 prompt 409,detach 503,lease 不释放,同 Workspace 的新 Session 503。
MCP 工具运行中取消 prompt 取消后 0.5 s 中止 该轮没有终态事件。同样是 409 / 503 / lease 不释放 / 503。
32 s 工具,不取消 服务端在 41 s 回复 Runtime 在 25 s 超时(TIMEOUT_MS 写死,manifest 不接受 timeout 字段)。

32 s 这一例里,迟到回复确实结算了 Runtime 自己的视图。但 qwen_tool_execution.execution_state 一直是 UNKNOWN,也从不重查,所以 Session 永远恢复不了。close 时 mcp-release 200,随后 tool-sessions/mcp:<sid>:release 返回 409,detach 返回 503。对比:普通 MCP 用 MCP_DEFAULT_TIMEOUT_MSEC = 10 分钟,可按服务端配置。

F2 / F5 / F6 / F7 —— 首轮之后的连接生命周期。

  • F2。 resources 或 tools 的 list_changed 会让所有调用类型共用的那一个 catalogRevision 加一,而 Harness 从不发起 mcp-discover。此后:

    • 同一轮内下一次工具调用得到 catalog_conflict;
    • 之后每一轮也是;
    • prompt_get 也是;
    • Harness 的目录仍显示 revision 1、complete。

    SSE 服务端重启效果相同,因为 onerror 会让连接退役。SDK 的 McpServer 在 connect 之后注册、启用或禁用工具时都会发 list_changed,所以这个触发条件并不罕见。

  • F5。 detach 在 1 s 内返回 503(MySQL 678 ms,MariaDB 228 ms)。服务端恢复后再次 detach 仍返回 503。配置一直停在 releasing,lease 不释放。DELETE 发不出去时 terminateSession() 会抛错,而当时该连接的在途操作数是 0。

  • F6。 每次 POST /mcp/configurations 都让上一个 stdio 进程继续活着:同一 Session、同一服务端,存活进程从 1 涨到 16。第 17 次重新配置返回 409。只有 detach 才回收它们,耗时 6.9 s。这就是 triage 机器人 stage-2 第 4 点,此处给出实测。结合 F2,文档里的恢复手段每个 Session 大约能用 15 次。

  • F7。 首次安装被拒后,status 报告 recoveryBlocked:true,但 /prompt 仍然放行。即使 Workspace 已经空出来,detach 也一直 503。只有之后某次 prompt 把遗留的配置 intent 派发出去,detach 才能成功。

可选方向(决定权在维护者):

  • F1: 要么 MCP owner 在两轮之间不再持有 storage lease,要么把"每个 Workspace 同时只能 attach 一个 MCP Session"写成 H1 限制。
  • F2: 遇到 catalog_conflict / stale 时调用已有的 mcp-discover,传输出错后用新的 generation 重连。
  • F3: 对 resource_read / prompt_get,把被取消的操作结算为不阻塞的 cancelled,因为其结果只是数据。
  • F4:
    • 把超时做成 manifest 字段;
    • 工具调用要么不发 notifications/cancelled,要么限时等待后结算;
    • close 或 load 时重查 execution 状态,以对账迟到的成功。
  • F5: 在途操作为 0 时,DELETE 失败或返回 405 应视为已释放。
  • F6: 退役连接的在途集合清空后即关闭它。
  • F7: 让 close() 在本地把仍停在 execution: intent 的配置 intent 结算为 not_started_proven,与操作记录的处理方式一致。

门禁

  • CI。
    • 在 a390b5d5,Test (ubuntu) 红:process-env-guard.test.ts 标出 QWEN_MANAGED_MCP_CONFIG、PATH 和 SystemRoot。
    • 作者已在 e3ecba66 修复,现为绿:22 通过,2 个 pending。
    • 本地该 guard 5 次中 4 次通过,唯一一次失败发生在负载均值接近 100 时。
  • main。 feat(serve): add gated Hosted foreground Shell turns #12848 合入后本 PR 为 CONFLICTING:8 个文件,其中 6 个生产文件共 17 个冲突块(其中 8 个在 hosted-workspace-tool-turn.ts)。F1 和 F4 正好在这个文件里,合并后的 head 需要重跑相关场景。
  • 与 feat(managed-agent): Add durable remote Shell result delivery #12894(a7c0a391)的合并顺序。
    • 两个 PR 都新增 V19(V19__managed_mcp_records.sql 和 V19__managed_tool_publication.sql),git 对这两个文件不报冲突。
    • 试合并后的 jar 在 MySQL 8.4 上启动即失败:Found more than one migration with version 19。
    • 试合并还有 8 处文本冲突。
    • 后合入的一方必须给自己的 migration 重新编号。
  • a390b5d5 的本地测试套件。
    • cli 定向 79/79,core 定向 230/230,managed-agent-server 180/180,runtime-broker 395/396。
    • Broker 那 1 个失败(HttpRuntimeTransportTest#closesTheConnectionOnTheDeadlineAndOnCallerCancel)在本机高负载下对 main bd45b95f 同样失败,单独跑 5/5 通过。
  • 变异测试。 Java 8/8 被杀,TS 13/16 被杀。三个存活者指向没有测试覆盖的行为:
    • M7: tools-only 服务端返回的 -32601 应表示空的 complete 列表。
    • M11: prompt 路由里的 hasPendingOperations() 检查。F3 里的 409 正是这个守卫返回的,但没有单测固定它。
    • M13: 由两个 Runtime Session 拥有的配置。

小问题

  • /mcp/operations 和 /mcp/configurations 上的身份冲突返回 503,看起来像可重试。用 409 更合适。
  • 模型看到的工具名是 mcp_<32 位十六进制>,其中含 catalog revision 和 connection generation,所以每次 load 名字都会全变。reload 之后,对话记录里引用的是已不再公布的工具。
  • Runtime 的错误回执以裸错误码形式交给模型,例如 managed_mcp_catalog_conflict。
  • stdio 服务端运行时 HOME 被设为 Workspace 目录(实测 HOME = cwd = <ws>/child)。服务端缓存在 ~ 下的任何内容都会落进用户的 Workspace。

未验证

装置、场景脚本、变异体与原始结果:wenshao/qwen-code@a5ef4914/pr12946

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Merge validation — 7a32435

Merged main at a765229c0a and resolved the overlapping Hosted Shell/MCP changes. The shared tool turn retains both profiles, asynchronous pinned declarations, the Shell input digest and the MCP prompt/owner distinction. Two independent open-ended and targeted audit passes found no remaining issue in the resolved paths.

Local checks passed: full build/typecheck/bundle, formatting/lint, 951 CLI + 104 Core + 160 Broker + 48 Java server tests. All 81 Broker classes embedded in the clean-packaged server match the current compilation.

Both independent E2E paths passed against the final artifacts:

  • MCP: real Hosted → Java Broker → Runtime → stdio/Streamable HTTP with isolated MariaDB. Exactly five physical operations and three model requests; committed result ordering, original-ID lost-ACK recovery, revoked-access close and offline durable replay all passed. Offline replay added zero Broker requests or MCP effects.
  • Foreground Shell: the upstream packaged-Harness integration test passed with six workspaces, complete 100 MiB output, reload, lost start reply, publication faults, cancellation and at-most-once checks. All six producers exited before a fresh reader verified the retained output. This test uses H2 in MySQL mode.

The model was deterministic in both paths; Runtime/MCP/Shell processes were real. All test process groups were confirmed exited. Local execution was on macOS. GitHub reports the PR as mergeable without conflicts; fresh CI is running.

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification, round 2 — 7a324350 (main merged)

Follow-up to round 1. 7a324350 merges main a765229c into e3ecba66, bringing in #12848 (Hosted Shell turns), #12942 and #12918. It contains no other change.

Merge reference in one paragraph. The conflict resolution is faithful to both sides, and GitHub CI is green at this head. I rebuilt the bundle and jar from 7a324350 and re-ran every round-1 scenario on the real stack. Each outcome is identical: the author's claims still hold, and F1–F7 all reproduce unchanged, because the merge touches none of that logic. The V19 clash with open #12894 (a7c0a391) is also unchanged. My round-1 recommendation therefore stands. The conflict blocker is gone. What remains is the #12894 merge order, plus a maintainer decision on F1–F5: fix them here, or record them as H1 limits with follow-ups.

Same scenarios, new head

round 2 matrix

  • Rebuild and scope. Fresh dist/cli.js and Spring jar from 7a324350; MySQL 8.4.11; the same rig, manifest and MCP servers (stdio, Streamable HTTP, SSE) as round 1. I compared all 11 scenarios one by one.
  • Timings, from the MCP server ledger.
    • The cancelled resource_read was aborted 1.4 s after it started.
    • The tool was aborted 1.1 s after the prompt cancel.
    • The 32 s tool finished at +32.0 s, 7 s after the Runtime's 25 s timeout.
  • New measurement. After the prompt cancel (F4), detach made 190 Broker calls before its 503: 95 acquire, 94 mcp-status and 1 mcp-release. This puts a number on triage stage-2 point 3: lookup() re-acquires the owner on every 100 ms poll.
  • The stuck Workspaces stay stuck. The lease snapshot taken minutes later in S2 still shows the three Sessions stuck by F3/F4 holding their Workspaces.

Merge-resolution audit and gates

round 2 audit

  • Hand-resolved parts (git show --remerge-diff: 9 files, +76/−133):
    • toolProfile accepts files, shell or mcp, one per Session, so Shell and MCP never meet in one turn.
    • Declarations are (shell ? Shell tools : file tools) plus the MCP tools.
    • prepare(runtimeCallId, digest, shellInputDigest, promptId) keeps both sides' arguments.
    • Shell's input, digest and timeout stay on the Shell branch.
    • The worker keeps both the remote Shell publishers and the MCP runtime.
    • if (!this.mcp) await this.broker.release() is unchanged, so the F1 root cause is too.
  • GitHub CI at 7a324350: 22 pass. That includes Hosted process fault gates / MySQL 8.4 (HostedWorkspaceToolTurnIT 5/5 in 96.8 s, with the six-Workspace 100 MiB Shell path and FG6a/b/d/e) and the MariaDB job. The author's note says the Shell IT ran on H2 locally; CI runs it on MySQL.
  • Local, same command as CI (-Phosted-harness-mysql verify checkstyle:check), at a load average of 50–100:
    • unit 182/182, HostedHarnessMySqlIT 2/2, 0 Checkstyle violations;
    • HostedWorkspaceToolTurnIT 3/5 in 439 s. The failures were FG6B writers:renew → 409 (writer lease lapsed) and a 130 s driver timeout on the six-Workspace path. Both are timing failures in drivers this PR does not change, and the class passes 5/5 on CI in 96.8 s. Re-running just those two alone at a load of 31–54 passed 2/2 in 203.7 s.
  • Mutation on the merged tree (6 cli test files, 112 tests):
    • Three new resolution mutants are killed:
      • Shell profile advertises only file tools (15 tests fail);
      • prepare drops promptId (5 fail);
      • prepare drops the Shell input digest (1 fails).
    • The promptId mutant survives on the a390b5d5 tree, so triage stage-2 point 7 is now pinned.
    • M13 (two-owner configurations) is now killed; it survived in round 1.
    • M7 (-32601 → empty complete list) and M11 (the prompt route's hasPendingOperations() guard, which is what returns the 409 in F3) still survive.

Unchanged from round 1

Evidence: wenshao/qwen-code@9e3c79fe/pr12946/r2 (raw results per scenario, remerge diff, CI and local gate extracts, mutant scripts).

中文版

真实环境验证第二轮 —— 7a324350(已合入 main)

接第一轮。7a324350 把 main a765229c 合入 e3ecba66,带进 #12848(Hosted Shell turn)、#12942 和 #12918,没有其他改动。

合并参考(一段话): 冲突解决对两边都忠实,GitHub CI 在该 head 上全绿。我用 7a324350 重新构建 bundle 和 jar,在真实链路上重跑了第一轮的全部场景,结果逐项一致:作者的主张依旧成立,F1–F7 全部原样复现,因为本次合并没有触及这些逻辑。与在飞的 #12894(a7c0a391)撞 V19 的问题也没变。所以第一轮的建议不变:冲突这个阻断已消除,剩下的是 #12894 的合并顺序,以及 F1–F5 由维护者决定——在本 PR 修复,还是记为 H1 限制另开后续。

同一组场景,新 head

  • 重建与范围。 用 7a324350 重新构建 dist/cli.js 和 Spring jar;MySQL 8.4.11;装置、manifest 和 MCP 服务端(stdio、Streamable HTTP、SSE)与第一轮相同。11 个场景逐项对比。
  • 时长(来自 MCP 服务端台账)。
    • 被取消的 resource_read 在开始后 1.4 s 被中止。
    • prompt 取消后 1.1 s,工具被中止。
    • 32 s 的工具在 +32.0 s 完成,比 Runtime 的 25 s 超时晚 7 s。
  • 新测量。 prompt 取消(F4)之后,detach 在返回 503 前一共发了 190 次 Broker 请求:95 次 acquire、94 次 mcp-status、1 次 mcp-release。这给 triage stage-2 第 3 点(lookup() 每 100 ms 轮询都重新 acquire owner)提供了实测数字。
  • 卡住的 Workspace 一直卡着。 几分钟后 S2 的 lease 快照里,被 F3/F4 卡住的三个 Session 仍各自占着自己的 Workspace。

冲突解决审计与门禁

  • 手工解决部分(git show --remerge-diff:9 个文件,+76/−133):
    • toolProfile 接受 files、shell 或 mcp,每个 Session 只能选一个,所以 Shell 与 MCP 不会出现在同一轮。
    • 工具声明为 (shell ? Shell 工具 : 文件工具) 再加上 MCP 工具。
    • prepare(runtimeCallId, digest, shellInputDigest, promptId) 保留了两边的参数。
    • Shell 的 input、摘要和超时只走 Shell 分支。
    • worker 同时保留远程 Shell publisher 和 MCP runtime。
    • if (!this.mcp) await this.broker.release() 没变,所以 F1 的根因也没变。
  • GitHub CI(7a324350): 22 项通过。其中包括 Hosted process fault gates / MySQL 8.4(HostedWorkspaceToolTurnIT 5/5、96.8 s,含 6 个 Workspace 的 100 MiB Shell 路径和 FG6a/b/d/e)以及 MariaDB job。作者说明里的 Shell IT 是本地用 H2 跑的,CI 上用的是真 MySQL。
  • 本地、与 CI 相同的命令(-Phosted-harness-mysql verify checkstyle:check),负载均值 50–100:
    • 单测 182/182,HostedHarnessMySqlIT 2/2,Checkstyle 0 违规;
    • HostedWorkspaceToolTurnIT 3/5,耗时 439 s。失败的是 FG6B 的 writers:renew → 409(writer 租约过期),以及 6-Workspace 路径 driver 130 s 超时。两者都是本 PR 未改动的 driver 的时序失败,同一类在 CI 上 96.8 s 内 5/5 通过。在负载 31–54 时单独重跑这两个用例:2/2 通过,203.7 s。
  • 合并后代码的变异测试(6 个 cli 测试文件,112 个测试):
    • 三个针对冲突解决的新变异体全部被杀:
      • Shell profile 只公布文件工具(15 个测试失败);
      • prepare 丢掉 promptId(5 个失败);
      • prepare 丢掉 Shell 输入摘要(1 个失败)。
    • promptId 这个变异体在 a390b5d5 的树上是存活的,所以 triage stage-2 第 7 点现在已被测试固定。
    • M13(两个 owner 的配置)现在被杀;第一轮它存活。
    • M7(-32601 → 空的 complete 列表)和 M11(prompt 路由的 hasPendingOperations() 守卫,F3 里的 409 就是它返回的)仍然存活。

与第一轮相同的部分

证据:wenshao/qwen-code@9e3c79fe/pr12946/r2(各场景原始结果、remerge diff、CI 与本地门禁摘录、变异脚本)。

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Review follow-up and merge verification — 9b4886793a

Addresses the triage review and both real-stack reports, round 2. Merged main at 693528b004; the D5 Turn list/detail contract and MCP catalog contract are both retained, with OpenAPI version 1.21. The unreleased MCP migration is now V21, avoiding the V19/V20 files on #12894. Deployment ordering must still be checked when either branch lands.

Finding disposition

Finding Change / retained boundary
F1: idle attached owner blocks other Sessions Retained and documented as a private H1 constraint: the owner holds the tenant/storage lease until detach, including Workspaces sharing storage. Releasing it while a stdio server can still write would permit concurrent writers. Connection/storage lifetime separation is follow-up work before enablement.
F2: stale discovery or retired connection never recovers Discover before each model request and each new raw command; unchanged usable catalogs retain their revision. Changed catalogs and retired connections publish a new immutable configuration/generation. Already advertised calls keep their original pins.
F3: native cancellation suppresses settlement Cancellation after dispatch is deferred. Runtime keeps observing the original request without sending the native cancellation notification; an actual response supplies the settlement proof. Cancellation before dispatch still records not_started_proven.
F4: 25-second timeout / cancellation leaves a tool permanently UNKNOWN Invocation timeout defaults to ten minutes and is deployment-configurable up to that limit. Hosted observes the original execution for 630 seconds, including temporary UNKNOWN/network outcomes; Java reconciles both tool protocol versions and durably accepts conclusive late results. There is no second start. Results after that window and restart during an unfinished turn still require follow-up checkpoint recovery.
F5: failing HTTP DELETE prevents local closure Remote session deletion is best effort with a one-second bound, followed by bounded local transport closure. Release proves local closure after all requests settle; it does not claim the remote server session was deleted.
F6: idle replacements exhaust 16 connections Close idle retired siblings; close busy retired siblings only after their original requests settle. Closed generations retain their receipts and can acknowledge their original release.
F7: a refused first configuration cannot detach Close cancels an unsent configuration, completes the original Broker acquisition when the Workspace becomes available, then releases it without installing MCP. No retry prompt is required.

Two further audit findings were fixed before commit: an idle Broker restart must reacquire and validate the original owner before new work/close; an idle Harness restart must advance durable configuration record revisions before issuing new-writer grants. The Runtime gate continues to reject owner replacement at the same revision. Explicit replacement also works after the first definition fails, without retrying that failed definition first.

Other review points

  • Status polling now reuses the validated owner and backs off from 100 ms to one second. A failed status query reacquires only the original owner and retries status, never an effect. New work and close still validate ownership at their boundary.
  • Idempotent shared Broker acquire and the additional binding/generation fields are intentional. Existing file/Shell reservation identities and Shell input digest remain covered by explicit assertions and the shared Shell integration regression.
  • Record-domain recognition is global; execution admission remains private-profile and scoped-manifest gated. This distinction is now documented.
  • Acquire failures preserve uncertainty on lost replies; definite Broker rejections restore the previous local acquired flag. Grants require a successfully validated Runtime binding.
  • Raw operation ID conflicts now return HTTP 409 without another effect. M7 (method-not-found means authoritative empty discovery) and M11 (pending raw work blocks model prompt admission) have direct regression assertions; the existing M13 two-owner guard remains covered.
  • Deferred explicitly: process-lifetime operation receipts / closed-generation tombstones and the release waiter for permanently unknown work remain unbounded history. The concurrent connection/request quotas do not claim to bound this memory. Durable-acknowledged pruning and bounded waiter cleanup are follow-ups before production enablement, recorded here and in both design languages. Tool names intentionally include immutable catalog/connection identity, and stdio HOME remains the Workspace unless a deployment overrides it.
  • feat(serve): implement generic Broker provider controls #12868 is still an open semantic integration dependency. The later merge must preserve original-owner status/cancel/release after revocation and drain-before-storage-release ordering while adopting the generic control mechanism; this PR does not claim that cross-branch combination has been verified.

Verification on the final tree

  • Build/static checks: repository build, final CLI rebuild/bundle, repository typecheck, changed-file ESLint/Prettier, and both Java modules’ Checkstyle passed. The pre-commit hook also passed without changing the audited tree.
  • Focused tests: 961 CLI (121 directly affected + 840 shared consumers), 59 Core, 160 Broker and 54 Server tests passed: 1,234 tests, zero failures. The new grant test also failed under a deliberately disabled-renewal mutation, then passed after restoring the implementation.
  • Real stack: macOS, Node 22.22.2, JDK 21.0.11, isolated MariaDB 10.11.19, packaged Hosted Harness → Java Broker/SQL store → provisioned Runtime → local stdio/Streamable HTTP wire fixtures; only the model is deterministic. All 11 formal lifecycle cases passed, with F1 checked as a retained limit:
Scenario Observed result
F1/F7 While A remains attached, B is refused. After A detaches, B directly detaches with 204, original MCP owner RELEASED and no lease; a following file Session writes successfully.
F2 Two list-change events recover before raw/model calls; killing an idle stdio server also recovers. Four unique effect IDs, no replay.
F3 resource + prompt Cancel leaves the original request running; only the actual reply settles it. No native cancel notification, one effect per case, subsequent work and detach succeed.
F4 tool cancel Prompt cancel returns 204, execution remains unproven until the actual reply; then SQL SETTLED, next prompt and detach succeed. One start/effect.
F4 long tool Real response at 32,002 ms succeeds under the default timeout; SQL SETTLED/success and detach 204.
F4 configured timeout The same SQL execution is observed UNKNOWN/NULL, then SETTLED/success after the original physical reply. Exactly one Broker start and one native effect.
F5 After a successful HTTP read, shutting the listener still permits detach 204, owner RELEASED and lease NULL; a later Session reads again. Conflicting reuse returns 409 with zero added effects.
F6 All 18 replacements succeed through revision 19; only the current native PID survives each replacement. All 19 PIDs exit by detach.
A1 idle Harness restart SIGKILL the idle Harness and let its 5-second writer lease expire; reload under a new boot ID. Runtime binding/generation and native PID stay the same. New read and detach succeed without replay.
A2 initial definition failure Bad r1 fails explicitly; expected-revision-1 replacement by good r2 succeeds. Next read has one effect and detach returns 204.
  • The original five-effect MCP happy/recovery scenario also passed: byte-preserving resource/prompt results, lost configure/invoke ACK recovery by original ID, revoked-access close, and offline durable replay with 0 Broker requests / 0 new effects.
  • The upstream six-Workspace, 100 MiB Shell packaged integration passed against H2, preserving the shared Broker/Shell path. This is separate from the MariaDB MCP scenarios; no MySQL run is claimed for this revision.
  • Every recorded lifecycle test process group and native child was confirmed gone. One preliminary test-driver assertion incorrectly included an unrelated file-profile control owner; it was narrowed to the original MCP owner and the entire case rerun successfully.

Artifacts exercised: CLI SHA-256 7db44d0d703e3c45f7d04148834ecdbbe6e9f5b529b4249b6e819d4355a92746; packaged Server SHA-256 7586d6c4a502f4fb8ee573ae6f139e4ad17fe6252da71107ef38216bb1b6e188. The 81 embedded Broker classes match the compiled Broker byte-for-byte. Windows/Linux, a real model, full restart matrix, and cross-branch #12868 integration were not verified locally; SSE retains focused transport coverage. New GitHub CI is tracked separately from these local results.

The prior finding-bearing audit rounds reset the clean count. After the last fixes, two consecutive open-ended and targeted passes were clean, including independent CLI and Core/Java review. These are verification results for this private H1 scope, not an approval or production-enablement claim.

中文摘要:已同时修复冲突和评论确认的问题,并补上审计发现的 Broker/Harness 空闲重启及首次失败配置替换问题。F1 存储租约独占、观察窗口外/turn 中途重启恢复、历史回执与永久等待者的回收仍为明确的 H1 后续项;未宣称这些限制已修复,也未启用生产 profile。

@wenshao

wenshao commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification, round 3 — 9b488679 (F1–F7 dispositions)

Follow-up to round 1 and round 2, checking the author's disposition table.

Merge reference in one paragraph. Every disposition in that table reproduces on the real stack on MySQL 8.4.11, including SSE, which the author covered only in unit tests: F2–F7 are fixed, F1 behaves exactly as now documented, identity conflicts return 409, and the round-1 mutant survivors M7/M11 are now pinned. The V21 rename removes the duplicate-version crash with #12894. It is not mergeable yet, because of one blocker: CI's Hosted process fault gates / MySQL 8.4 job is red. I reproduced it locally and traced it, by a one-line A/B test, to this PR's change in RuntimeBrokerHttpServer.observe(). There are also three behaviour notes (N1–N3) that widen the documented H1 limits, a Flyway ordering rule for #12894, and four surviving mutants.

F1–F7, re-run

fix verification

Fresh bundle and jar from 9b488679; same rig, manifest and MCP servers as before, plus a timeoutMs: 5000 definition.

# Round 3 observed Verdict
F1 The MCP owner still holds the Workspace lease: a second MCP Session gets 503, a file Session gets turn_error Retained, as documented
F2 After resources or tools list_changed: the same-turn call, the next turn and prompt_get all succeed. After an SSE server restart the next call succeeds (1 effect) Fixed
F3 Cancel of a running resource_read answers running at 2.2 s; the operation settles with the full result at 9.9 s. The server ran to done (+8.0 s); detach 204, lease released Fixed
F4 Prompt cancel at 2.5 s: the turn ends cancelled at 11.6 s (the tool ran to done +10.0 s); next prompt and detach work. A 32 s tool completes at 33.7 s (default timeout) and at 33.8 s (timeoutMs: 5000, UNKNOWN in between) Fixed
F5 HTTP server down at close: detach 204 in 291 ms, lease released Fixed
F6 17 reconfigurations: one live stdio process each time, revision 18 accepted, detach reaps it Fixed
F7 Refused first install: detach 204 once the other Session leaves, no extra prompt (still 503 while it holds the Workspace) Fixed

Still holding from earlier rounds:

  • three transports with results_ready ordering, byte-exact blobs and 0 catalog leaks;
  • lost configure and invoke replies → exactly 1 read;
  • revoke then close → 204;
  • Broker down → replay 202, close 204, 0 Broker requests, 0 effects;
  • replacement during an in-flight call;
  • the 16 KiB / 60 KiB limits.

B1 — the CI red is caused by one Broker line

new findings

Failing job. CI Hosted process fault gates / MySQL 8.4 fails in HostedProcessCrashIT.processCrashesNeverReplayHostedToolsOnMySql (FG6C). CI stops at the first failure, so I ran each case alone on the PR tree:

FG6C case At 9b488679 Same tree with only that line restored
worker-kill status 409 runtime_execution_evidence_unavailable; the gate expects runtime_broker_execution_unknown 1/1
worker-stop status 503 managed_runtime_unavailable (retryable); the gate expects 200 1/1
harness-*, spring-kill pass pass

Cause. The PR removes && Integer.valueOf(3).equals(record.getReference().get("runtimeProtocol")) from observe(). As a result, status reads of UNKNOWN v2 (file) and MCP executions now go through reconcileExecution, which has to reach the original worker. That changes the Broker's status contract for every Hosted tool profile, not only MCP.

Options (the maintainers' call). Either:

  • keep observe() protocol-3-only and reconcile MCP executions on their own path; or
  • update the FG6C expectations for v2, with the reason stated in the PR.

Why the author missed it. The author's note says the Shell/file IT ran on H2 only and no MySQL run was claimed for this revision. HostedProcessCrashIT is the gate that catches this change.

Behaviour notes the fixes leave open

  • N1 — a reply that can never arrive still ends in a stuck Session, now after 10.5 minutes.
    • Connection lost: I restarted the SSE server 2 s into a 20 s call. The execution was UNKNOWN at 6.8 s, the turn stayed active until 633 s, then the Session was recoveryBlocked with no terminal event.
    • Tool never answers: on a timeoutMs: 5000 server, the execution was UNKNOWN at 7.8 s and the turn stayed active until 634 s.
    • Both cases end the same way: next prompt 409, detach 503, lease held (so, by F1, the Workspace is locked), qwen_tool_execution UNKNOWN.
    • The configured timeoutMs does not shorten the Hosted wait, and a lost connection is known at once but still waits out the 630 s window. The author documents "results after that window" as follow-up work; these two cases show the window itself is spent even when the outcome is already decided.
  • N2 — one pinned server down makes every turn fail. Because the Harness now refreshes before each model request, a single unavailable pinned server (SSE or Streamable HTTP) makes every turn in the Session end in turn_error, with 0 effects. That includes a text-only answer and a turn that only calls the healthy stdio server. Each failed turn records a failed mcp_configuration revision, and the Session recovers once the server returns. One option is to drop that server's tools for the round instead of failing the turn.
  • N3 — cancellation is deferred. The prompt cancel above took 9.1 s to end the turn because the tool ran to completion; with a 10-minute tool, the user waits up to 10 minutes and the effect still happens. This is documented by the author, and I'm listing it here only with the measured number.

Migrations, mutation, cost, CI

gates

  • Flyway ordering. Tested with a trial merge with feat(managed-agent): Add durable remote Shell result delivery #12894 at 033c74d3, which carries V19 and V20.
    • On a fresh database, the merged jar applies 19 → 20 → 21 and starts.
    • If this PR runs first (database at 21), the merged jar then refuses to start with Detected resolved migration not applied to database: 19 … set -outOfOrder=true. outOfOrder is not enabled.
    • So feat(managed-agent): Add durable remote Shell result delivery #12894 should land first, or whichever lands second should renumber above the other.
  • Mutation (6 cli test files, 122 tests):
    • M7 and M11 are now killed, and so are the inverses of the fixes: native cancel, idle sibling close, bounded DELETE, wait-through-unknown, intent cancel on close, unchanged-catalog skip, and grant renewal.
    • Survivors:
      • X7, no refresh before each model request. This is the F2 fix itself; it is proven end to end above but not by a unit test.
      • X3, a busy retired connection is not closed after its last settle.
      • X5, the 600 s default timeout falls back to 25 s.
      • X9, identity conflicts return 503 again.
  • Refresh cost. With text-only turns at a load average of about 25–30: 1 mcp-discover per server per model request. Time to the first model request was 103–185 ms with 1 server and 154–232 ms with 6 servers.
  • CI at 9b488679: 21 pass and 1 fail (B1). HostedWorkspaceToolTurnIT is 5/5 on MySQL in that same job (115.5 s, including the six-Workspace 100 MiB Shell path). Test, Lint, MariaDB and the Java matrix are green.

Not verified

Evidence: wenshao/qwen-code@9fd4f16f/pr12946/r3 (per-scenario raw results, S12 timelines, FG6C A/B extracts, Flyway logs, remerge diff, mutants, scripts).

中文版

真实环境验证第三轮 —— 9b488679(F1–F7 处置)

接第一轮与第二轮,核对作者的处置表。

合并参考(一段话): 处置表里的每一项都在 MySQL 8.4.11 真实链路上复现成立,包括作者只有单测覆盖的 SSE:F2–F7 已修复,F1 的行为与现在的文档描述完全一致,身份冲突返回 409,第一轮存活的变异体 M7/M11 现在已被测试固定。改名为 V21 消除了与 #12894 同号导致的启动失败。但目前还不能合入,有一个阻断项: CI 的 Hosted process fault gates / MySQL 8.4 是红的。我在本地复现了它,并用单行 A/B 定位到本 PR 对 RuntimeBrokerHttpServer.observe() 的一行改动。此外还有三条扩大了已记录 H1 限制的行为说明(N1–N3)、一条与 #12894 相关的 Flyway 顺序规则,以及四个存活的变异体。

F1–F7 复验

用 9b488679 重新构建 bundle 和 jar;装置、manifest 和 MCP 服务端与前两轮相同,另加一个 timeoutMs: 5000 的定义。

# 第三轮实测 结论
F1 MCP owner 仍持有 Workspace lease:第二个 MCP Session 得到 503,文件 Session 得到 turn_error 按文档保留
F2 resources 或 tools 的 list_changed 之后,同一轮的下一次调用、下一轮和 prompt_get 都成功;SSE 服务端重启后下一次调用成功(1 次副作用) 已修复
F3 取消运行中的 resource_read:2.2 s 返回 running,9.9 s 带完整结果结算;服务端跑到 done(+8.0 s),detach 204,lease 已释放 已修复
F4 2.5 s 取消 prompt:turn 在 11.6 s 以 cancelled 结束(工具跑到 done +10.0 s),下一次 prompt 与 detach 正常。32 s 工具在默认超时下 33.7 s 完成,在 timeoutMs: 5000 下 33.8 s 完成(中间为 UNKNOWN) 已修复
F5 close 时 HTTP 服务端宕机:detach 在 291 ms 内返回 204,lease 已释放 已修复
F6 连续 17 次重新配置:每次只有 1 个 stdio 进程存活,修订 18 被接受,detach 会回收 已修复
F7 首次安装被拒:另一个 Session 离开后直接 detach 204,不需要再发 prompt(对方仍占着 Workspace 时仍是 503) 已修复

前两轮验证过、这一轮仍然成立的:

  • 三种 transport 下的 results_ready 顺序、字节精确的 blob、目录 0 泄漏;
  • 丢失 configure 与 invoke 回包 → 恰好 1 次读取;
  • 撤销权限后关闭 → 204;
  • Broker 不可用 → 重放 202、关闭 204、0 次 Broker 请求、0 次副作用;
  • 在途调用期间替换配置;
  • 16 KiB / 60 KiB 限额。

B1 —— CI 红由 Broker 的一行改动造成

失败的 job。 CI 的 Hosted process fault gates / MySQL 8.4 在 HostedProcessCrashIT.processCrashesNeverReplayHostedToolsOnMySql(FG6C)处失败。CI 遇到第一个失败就停,所以我在 PR 树上把每个用例单独跑了一遍:

FG6C 用例 9b488679 同一棵树只恢复那一行
worker-kill 状态查询返回 409 runtime_execution_evidence_unavailable;门禁期望 runtime_broker_execution_unknown 1/1
worker-stop 状态查询返回 503 managed_runtime_unavailable(可重试);门禁期望 200 1/1
harness-*、spring-kill 通过 通过

原因。 本 PR 从 observe() 里删除了 && Integer.valueOf(3).equals(record.getReference().get("runtimeProtocol"))。于是 UNKNOWN 的 v2(文件)和 MCP execution 的状态查询都会走 reconcileExecution,而它必须联系原来的 worker。这改变的是所有 Hosted 工具 profile 的 Broker 状态契约,不只是 MCP。

可选做法(由维护者决定)。 二选一:

  • 保持 observe() 只对协议 3 对账,MCP execution 走单独的对账路径;或
  • 更新 FG6C 对 v2 的期望,并在 PR 中写明理由。

为什么作者没发现。 作者的说明写明,这一版的 Shell/文件 IT 只在 H2 上跑过,没有声称跑过 MySQL。而 HostedProcessCrashIT 正是能拦住这次改动的门禁。

修复之后仍未解决的行为问题

  • N1 —— 回复永远不会来时,Session 仍会卡死,只是变成等 10.5 分钟之后。
    • 连接丢失: 在一个 20 s 调用开始 2 s 后重启 SSE 服务端。execution 在 6.8 s 变为 UNKNOWN,turn 一直活动到 633 s,然后 Session 进入 recoveryBlocked,没有终态事件。
    • 工具永不返回: 在 timeoutMs: 5000 的服务端上,execution 在 7.8 s 变为 UNKNOWN,turn 一直活动到 634 s。
    • 两种情况的结局相同:下一次 prompt 409,detach 503,lease 不释放(按 F1,整个 Workspace 被锁),qwen_tool_execution 为 UNKNOWN。
    • 配置的 timeoutMs 并不会缩短 Hosted 的等待;连接丢失在一开始就已确知,却仍要等满 630 s 窗口。作者把"窗口之后的结果"记为后续工作;这两个例子说明,即使结局早已确定,这个窗口本身也会被耗尽。
  • N2 —— 绑定的某一个服务端宕机,会让每个 turn 都失败。 由于 Harness 现在每次模型请求前都会刷新,只要有一个绑定的服务端不可用(SSE 或 Streamable HTTP),这个 Session 的每个 turn 都以 turn_error 结束、副作用为 0。纯文本回答和只调用健康 stdio 服务端的 turn 也不例外。每个失败的 turn 都会记录一条失败的 mcp_configuration 修订;服务端恢复后 Session 会自动恢复。一种做法是:这一轮先去掉该服务端的工具,而不是让整个 turn 失败。
  • N3 —— 取消被延迟。 上面的 prompt 取消花了 9.1 s 才结束 turn,因为工具要跑完;如果是 10 分钟的工具,用户最多要等 10 分钟,而且副作用照样发生。作者已写明这一点,这里只是补上实测数字。

迁移、变异、开销与 CI

  • Flyway 顺序。 与 feat(managed-agent): Add durable remote Shell result delivery #12894(033c74d3,带 V19 和 V20)试合并后测试:
    • 全新库上,合并后的 jar 依次应用 19 → 20 → 21 并正常启动;
    • 如果本 PR 先运行过(库已到 21),合并后的 jar 启动失败:Detected resolved migration not applied to database: 19 … set -outOfOrder=true。outOfOrder 没有开启。
    • 所以 feat(managed-agent): Add durable remote Shell result delivery #12894 应该先合,否则后合的一方要把编号改到对方之上。
  • 变异测试(6 个 cli 测试文件,122 个测试):
    • M7 和 M11 现在被杀;各项修复的反向变异体也都被杀,包括原生取消、关闭空闲的 sibling、DELETE 限时、等待 unknown、关闭时取消 intent、目录未变时跳过、grant 续期。
    • 存活者:
      • X7,不在每次模型请求前刷新。这就是 F2 修复本身;上面的端到端测试证明了它,但没有单测覆盖。
      • X3,忙碌的已退役连接在最后一次结算后没有被关闭。
      • X5,600 s 的默认超时退回 25 s。
      • X9,身份冲突又变回 503。
  • 刷新开销。 纯文本 turn、负载均值约 25–30 时:每个服务端、每次模型请求 1 次 mcp-discover。到第一次模型请求的时间,1 个服务端为 103–185 ms,6 个服务端为 154–232 ms。
  • 9b488679 的 CI: 21 通过、1 失败(B1)。同一个 job 里 HostedWorkspaceToolTurnIT 在 MySQL 上 5/5 通过(115.5 s,含 6 个 Workspace 的 100 MiB Shell 路径)。Test、Lint、MariaDB 和 Java 矩阵均为绿。

未验证

证据:wenshao/qwen-code@9fd4f16f/pr12946/r3(各场景原始结果、S12 时间线、FG6C A/B 摘录、Flyway 日志、remerge diff、变异体与脚本)。

@wenshao

wenshao commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator Author

B1 is fixed in 7444dc9. The shared v2 status contract is preserved: default reads no longer contact an UNKNOWN execution's worker. MCP polling explicitly opts into original-execution reconciliation with GET ...?reconcile=true; v3 automatic reconciliation remains unchanged. Start/cancel paths, ownership/generation checks and the original FG6C assertions are unchanged. No dispatch retry or replacement owner was added.

Verification before pushing:

  • Reproduced both worker-kill and worker-stop on the previous artifacts. On the rebuilt artifacts, all six original FG6C process scenarios pass against isolated MariaDB 10.11.19: harness-prepare, harness-start, harness-result, spring-kill, worker-kill and worker-stop. Both Java and TypeScript gate-source hashes match the failing baseline.
  • The separate real-MCP late-result case passes: the same SQL execution transitions UNKNOWN → SETTLED/success through explicit reconciliation, with one Broker start and one native invocation. The next prompt and detach succeed; all test process groups and injected workers exited.
  • 52 focused CLI tests and 161 Broker tests pass, along with build, bundle, typecheck, lint/format and Java Checkstyle. New regression tests were red before the fix.
  • Two consecutive open-ended and targeted audit passes completed without new findings, including independent reviewers. The committed tree matches the audited tree.

These local process tests use macOS/MariaDB and a deterministic model fixture. The separate Linux/MySQL 8.4 CI gate now passes on 7444dc9: all six original FG6C success markers are present, Hosted Workspace tool turns pass 5/5, Hosted Harness MySQL checks pass 2/2, and 29 additional Runtime fault-gate tests pass. The Linux unit suite, lint/static checks, Serve A/B, no-AK integration, Java matrix, MariaDB integration, real daemon E2E and TUI checks also pass. The downstream Web Shell browser smoke gate also passes. Final check snapshot: 23 passed, 8 skipped by workflow policy, no failures; only the separate automatic code review remains in progress.

Round-3 follow-ups: N1's full 630-second wait and N2's all-pinned-server availability requirement are now explicit in both design languages and the PR body. N3's deferred cancellation remains documented. Flyway guidance now states that if V21 has already run, pending publication migrations must be renumbered above it, without enabling out-of-order execution. N1–N3 behavior changes and the non-blocking X7/X3/X5/X9 unit-test gaps are deferred to follow-up work; this CI fix does not expand the established H1 scope.

中文说明

7444dc9 修复 B1:普通 v2 UNKNOWN 查询恢复被动读取;MCP 通过显式 GET 参数对原 execution 对账,v3 自动对账不变。没有修改 FG6C 期望、重发执行或更换 owner。

旧产物已复现 worker-kill/worker-stop 失败,新产物在真实 MariaDB 和进程信号下通过原版六个 FG6C 场景。另行验证 MCP 配置超时后,同一 SQL execution 接受迟到结果,start 与物理调用各一次,后续 prompt 和 detach 成功。52 项 CLI、161 项 Broker 测试及构建、类型、格式检查通过;提交前连续两轮开放式和定向审计无新问题。新提交的 Linux/MySQL 原失败关卡现已通过,日志确认原版六个 FG6C 场景、5 项 Hosted 工具 turn、2 项 Harness MySQL 检查及另外 29 项 Runtime 故障门禁均通过。Linux 全量单测、Web Shell 浏览器冒烟测试及上述其他检查也已通过。最终状态为 23 通过、8 按工作流策略跳过、0 失败;仅独立的自动代码评审仍在运行。

N1–N3、迁移顺序与变异测试建议的处置见上文;中英文设计及 PR 描述已同步。

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification, round 4 — 7444dc97 (B1 fix)

Follow-up to round 3 and the author's B1 note.

Merge reference in one paragraph. B1 is fixed. The fix leaves the FG6C gate sources untouched (no diff since 9b488679), and the full FG6C class passes locally on MySQL 8.4. The MCP late-result path still works through the new ?reconcile=true opt-in, and all four mutants of the fix are killed. CI is green: 23 pass, only review-pr pending. The round-3 behaviour notes N1–N3 and the Flyway rule are now stated in both design languages; the behaviour itself is unchanged, as the author said. From my side there is no remaining correctness blocker. What is left is a maintainer decision: accept the documented H1 limits (F1, N1–N3), leave the four unit-test gaps (X3/X5/X7/X9) for follow-up, and settle the merge order with #12894 (still open with V19/V20; main is at V18).

round 4

  • Build. Fresh bundle and Spring jar from 7444dc97. The Broker class embedded in the server jar is byte-identical to the compiled one.
  • B1 / FG6C. Full HostedProcessCrashIT on local MySQL 8.4: harness-prepare, harness-start, harness-result, spring-kill, worker-kill and worker-stop give 6 × HOSTED_PROCESS_CRASH_OK, 1/1 in 119.4 s. The two cases that failed in round 3 now pass without any change to their expectations.
  • MCP paths after the change.
    • A 32 s tool on a timeoutMs: 5000 server goes UNKNOWN, then the late result is accepted through ?reconcile=true: end_turn at 37.0 s, and the next prompt and detach succeed. This is the path that needs the opt-in.
    • A 32 s tool on the default timeout ends at 37.0 s.
    • A prompt cancel at 5.5 s ends the turn at 15.0 s (deferred cancellation, as documented).
    • S1 regression passes: three transports, results_ready ordering, exact blob, both prompt messages, 409 conflicts, 0 catalog leaks, detach and load.
  • New probe — Runtime worker SIGKILLed during an MCP tool call. The reconcile status read answers 409 once, and the turn ends 2 s after the kill: recoveryBlocked, next prompt 409, detach 503, lease held, qwen_tool_execution UNKNOWN. So worker loss fails fast into the documented blocked state, the same outcome FG6C expects for file tools. That contrasts with N1 (a lost MCP server connection), which still waits out the 630 s window.
  • Mutants of the fix. All four are killed:
    • Y1, the MCP opt-in is ignored: 2 Broker tests fail.
    • Y2, every UNKNOWN read reconciles again (B1 comes back): 2 fail.
    • Y3, Hosted polls without ?reconcile=true: hosted-workspace-broker.test.ts fails.
    • Y4, the reconcile value is not validated: 1 fails.
  • Still open, all documented by the author.
    • F1: one attached MCP owner per storage lease.
    • N1: a lost connection or a server that never replies means a 630 s wait, then a blocked Session.
    • N2: every pinned server must refresh before a model request.
    • N3: cancellation is deferred.
    • Flyway: if V21 has already run, feat(managed-agent): Add durable remote Shell result delivery #12894's migrations must be renumbered above it.
    • Unit-test gaps: X3, X5, X7 (the F2 refresh wiring) and X9.
  • Not verified. Windows/Linux, a real model, MariaDB this round (CI's MariaDB job is green), a Harness restart mid-turn, and feat(serve): implement generic Broker provider controls #12868.

Evidence: wenshao/qwen-code@17e44727/pr12946/r4.

中文版

真实环境验证第四轮 —— 7444dc97(B1 修复)

接第三轮与作者的 B1 说明。

合并参考(一段话): B1 已修复。 这次修复没有动 FG6C 门禁的源码(自 9b488679 以来无 diff),完整的 FG6C 在本地 MySQL 8.4 上通过。MCP 的迟到结果路径通过新的 ?reconcile=true 显式参数仍然可用,针对这次修复的四个变异体全部被杀。CI 全绿:23 通过,只剩 review-pr 在跑。第三轮提出的 N1–N3 行为说明和 Flyway 规则已写入中英文设计文档;行为本身没有变,与作者所说一致。我这边已没有剩余的正确性阻断项。 剩下的是维护者的决定:是否接受已记录的 H1 限制(F1、N1–N3),是否把四个单测缺口(X3/X5/X7/X9)留作后续,以及与 #12894 的合并顺序(它仍带着 V19/V20 在飞;main 在 V18)。

  • 构建。 用 7444dc97 重新构建 bundle 和 Spring jar。服务端 jar 内嵌的 Broker class 与编译产物逐字节一致。
  • B1 / FG6C。 在本地 MySQL 8.4 上跑完整的 HostedProcessCrashIT:harness-prepare、harness-start、harness-result、spring-kill、worker-kill、worker-stop 共 6 个 HOSTED_PROCESS_CRASH_OK,1/1,119.4 s。第三轮失败的两个用例现在通过,而且没有修改它们的期望。
  • 改动后的 MCP 路径。
    • 在 timeoutMs: 5000 的服务端上跑 32 s 工具:先变为 UNKNOWN,随后通过 ?reconcile=true 接受迟到结果,37.0 s end_turn,之后的 prompt 和 detach 都成功。这条路径依赖这个显式参数。
    • 默认超时下的 32 s 工具在 37.0 s 结束。
    • 5.5 s 发起 prompt 取消,turn 在 15.0 s 结束(延迟取消,与文档一致)。
    • S1 回归通过:三种 transport、results_ready 顺序、blob 精确、两条 prompt 消息、409 冲突、目录 0 泄漏、detach 与 load。
  • 新探针:MCP 工具调用期间 SIGKILL Runtime worker。 带对账的状态查询返回一次 409,turn 在 kill 之后 2 s 结束:recoveryBlocked,下一次 prompt 409,detach 503,lease 未释放,qwen_tool_execution 为 UNKNOWN。也就是说,worker 丢失会快速进入文档所述的阻塞状态,与 FG6C 对文件工具的期望一致。对比 N1:MCP 服务端连接丢失时仍要等满 630 s 窗口。
  • 针对修复的变异体。 四个全部被杀:
    • Y1,忽略 MCP 的显式参数:2 个 Broker 测试失败。
    • Y2,每次 UNKNOWN 读取都重新对账(B1 回来了):2 个失败。
    • Y3,Hosted 轮询不带 ?reconcile=true:hosted-workspace-broker.test.ts 失败。
    • Y4,不校验 reconcile 的取值:1 个失败。
  • 仍未解决的部分(作者均已写入文档)。
    • F1:每个存储 lease 只允许一个已 attach 的 MCP owner。
    • N1:连接丢失或服务端始终不回复时,要等 630 s,然后 Session 被阻塞。
    • N2:每次模型请求前,所有绑定的服务端都必须刷新成功。
    • N3:取消被延迟执行。
    • Flyway:如果 V21 已经执行过,feat(managed-agent): Add durable remote Shell result delivery #12894 的 migration 必须重新编号到它之上。
    • 单测缺口:X3、X5、X7(F2 的刷新接线)和 X9。
  • 未验证。 Windows/Linux、真实模型、本轮的 MariaDB(CI 的 MariaDB job 为绿)、turn 进行中重启 Harness、feat(serve): implement generic Broker provider controls #12868。

证据:wenshao/qwen-code@17e44727/pr12946/r4。

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head 7444dc9726e1bc8ce8fdf8d15093b2506e2f2abd

Verdict: COMMENT. I found no Critical I can prove, but I cannot approve: the only review round on this PR left eight Critical-grade candidates explicitly unruled, that round ran against an earlier head, and no round has run against this one. At +8857/−128 across 52 files I could not rule on those eight myself inside this channel's per-PR budget.

Why this is not an approval

The prior round did not verify its own findings. The bot's round at 9b488679 states that both verification waves were killed by harness stalls, that 44 of 51 candidates were never ruled on, and that 44 findings still carried the [unverified] tag when the loop ended. Only the shards covering R1-11/R1-12 and R1-41/R1-42/R1-43/R1-44/R1-51/R1-52 delivered verdicts, and the five findings it did post are all sev:S. Its reverse audit never reached two consecutive dry rounds, and its test-efficacy probe returned inconclusive on all 21 probes with the positive control never running — so mutant coverage there is an absence of measurement, not clean coverage.

The eight unruled Critical-grade candidates, in the round's own words:

  1. R1-1 — grant lease (300 s) shorter than the runtime timeout (600 s) and the Hosted observation window (630 s).
  2. R1-2 — a transient mcp-discover failure permanently recovery-blocks the Session, with no reset site for session.blocked.
  3. R1-3 — unknown() never removes an operation from connection.pending, so zombies block the release drain and consume the worker-wide inflight and connection quotas.
  4. R1-4 — install() commits execution dispatch_started from a durable outcome_unknown run, which the extension state line rejects.
  5. R1-5 — an unmatched string-id response is forwarded to the SDK client.
  6. R1-6 — the catalog blob is projected without shape validation.
  7. R1-7 — a Windows SYSTEMROOT/SystemRoot case collision in the stdio child environment.
  8. R1-8 — the status and cancel routes sit outside the mcpBusy fence.

An unruled candidate is not a confirmed defect, and I am not asserting any of these. But under this channel's rules a blocking-grade issue I can neither confirm fixed nor refute prevents an approval, and eight of them sit on a private MCP runtime that spawns child processes from catalog-supplied definitions.

The head moved after that round and nothing has re-reviewed it. The round ran at 9b488679; this head is 7444dc97. There are no inline Criticals and no author replies on this PR resolving any of the eight, so I have no evidence either way about whether the pushes since that round addressed them.

What I did verify myself at this head

The new public route is scoped correctly. ManagedMcpController takes TenantContext, forwards both tenant.tenantId() and tenant.actorId() into ManagedMcpCatalogService.get, and answers Cache-Control: no-store — consistent with the sibling public routes, so the catalog is not cached or served cross-tenant by construction.

The migration only widens. V21__managed_mcp_records.sql issues two MODIFY COLUMN … VARCHAR(32) NULL statements against qwen_managed_session_extension_record, relaxing task_kind and task_state to nullable so MCP records can carry their own lifecycle without becoming Session tasks. It drops no column, no index and no data, and it does not rewrite existing rows. The risk it creates is on the read side — every consumer of those two columns must now tolerate null — which is what ManagedExtensionRecordStore.java (+111/−7) and ManagedExtensionProjection.java (+11/−1) would need to handle, and which I did not read.

R1-7 traced part of the way, not ruled on. The stdio child environment at managed-mcp-runtime.ts:603-614 is built in this order: every key of the imported DEFAULT_INHERITED_ENV_VARS blanked to '', then PATH/HOME/USERPROFILE, then SystemRoot spelled that way and only when process.env['SystemRoot'] is truthy, then ...definition.env last. Whether a case collision is reachable therefore turns entirely on whether that imported list contains an uppercase SYSTEMROOT spelling — if it does, the object carries two keys differing only in case, one of them empty, and which one the child receives is not defined by the object literal. I did not open the defining module inside budget, so I am recording this as unresolved rather than filing it. It is a one-line check and the cheapest of the eight to close.

CI. Green at this head: Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), Serve A/B, web-shell E2E Smoke, the full Java matrix, Runtime Broker and Managed Agent MariaDB / Java 21, Hosted process fault gates / MySQL 8.4 / Java 21, Real daemon E2E / Java 11, the TUI parity and no-flicker gates and both Desktop Shell lanes all pass. Only review-pr is pending, which is not a gating check. The prior round also noted that Integration Tests (CLI, No Sandbox) was skipped, so that suite has not run for this change. No failure is attributable to this PR.

Coverage note

I read the new public controller, the migration, and the stdio environment block, and confirmed the review history and CI state. I did not read managed-mcp-runtime.ts (1043 lines) beyond that block, hosted-mcp-session.ts (904), managed-mcp-record.ts (322), managed-mcp-protocol.ts, the changes to managed-session-authority.ts (+90/−9), hosted-harness-session.ts (+184/−9), managed-runtime-tool-executor.ts (+136/−6), hosted-workspace-tool-turn.ts (+44/−21) or managed-context-worker.ts (+53/−28), nor the Java ManagedMcpRecords.java (185), ManagedMcpCatalogService.java (84) or the extension-store changes, nor the 3536 lines of new tests. A Critical-only scan of that surface is not reachable in this channel's budget.

Next step: re-run the review round on this head so the eight candidates get ruled on, and settle R1-7 first since it is one grep. R1-2 and R1-3 are the two I would want answered before merge on durability grounds — a Session that stays recovery-blocked after a transient discover failure, and pending operations that are never reaped, are both fail-closed in the wrong direction for a long-lived worker.

@wenshao

wenshao commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed Finding 1 in commit 2159a735ee, from the sandbox verification.

New Hosted Sessions now reject more than 16 server pins with HTTP 400 invalid_hosted_mcp_servers, before creating the Session or contacting the Broker. Admission and Runtime allocation use the same cap. The original 32-pin format remains accepted when loading an already saved Session, with exact definition matching, so older records can still attach and detach for cleanup.

Runtime connection quota failures now return HTTP 409 managed_mcp_connection_quota from prompt admission, configuration replacement and raw-operation initialization. Other errors retain their existing handling. A failure discovered inside an already accepted prompt remains a turn error; no attempt is made to change an HTTP 202 response after admission.

Verification:

  • Before the fix, the real Hosted → Java Broker → Runtime → stdio chain accepted 17 pins and failed both subsequent prompts with the generic 503. After the fix, the same creation returns 400 with zero Broker requests, Runtime Sessions, native processes or model calls.
  • The real 16-pin happy path still completes a model-driven native tool call. At full capacity, explicit configuration and resource requests now both return 409 managed_mcp_connection_quota. A warmed prompt still returns 202 and reports the later quota failure within its turn, without another model/native call. Detach returns 204, the original Runtime Session is RELEASED and the storage lease is empty. All test processes exited.
  • Eight new regression cases cover 16/17/32 creation, loading and detaching saved 17/32-pin records, all three error entry points and prompt retry after capacity becomes available. The initial regression run failed in the expected five places; all 130 focused tests now pass.
  • Repository build, bundle and typecheck, focused ESLint and formatting pass. Two consecutive clean open-ended and targeted audits followed a correction for legacy Session cleanup.

The real-stack checks use macOS, MariaDB 10.11.19, real Java/Runtime/MCP processes and a deterministic model fixture. No new Java implementation or migration was introduced. CI for the new commit is reported separately.

One evidence correction: two prompts against the original 17-pin Session reused the same 16 native PIDs and produced only 16 initialization events total. We reproduced repeated quota failure, but not the report's inferred repeated 16-process spawn cascade. Finding 2's diagnostic-code distinction remains deferred: both malformed pins and wrong-profile input are already rejected, and this patch stays focused on the confirmed capacity defect.

中文说明

已处理 Finding 1:新建 Session 超过16个 pin 时立即400,准入与Runtime共用同一上限。旧版本已保存的17–32个 pin 记录仍按原定义加载,保留detach清理入口。prompt准入、配置替换和原始操作初始化遇到连接配额不足时返回409及明确错误码;已202准入的turn仍沿用turn错误语义。

真实旧链路复现了17个pin创建成功、后续prompt通用503;修复后创建直接400,没有Broker请求、Runtime Session、子进程或模型调用。130项定向测试、构建、类型、格式检查及修正后的连续两轮审计通过。

补充校正:重复prompt复用原有16个进程,没有实测到每次重启16个进程。Finding 2作为非阻断诊断改进留待后续。

@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

R5 Suggestion disposition: this PR has exceeded five review rounds, so the repository’s late-review scope rule applies. The confirmed correctness fixes above are included; the following 26 Suggestions are explicitly deferred for follow-up, not claimed fixed or silently resolved. Their original threads remain open and retain the detailed acceptance criteria. R5-17 is covered by the approval integration regression, and the previously deferred R4-8 is now fixed together with R5-5.

Finding Deferred follow-up
R5-8 Clarify quota documentation.
R5-10 Pin the idempotent prompt replay fence.
R5-11 Add the private catalog route regression.
R5-12 Refine route error classification.
R5-13 Move mock assertions out of potentially swallowed paths.
R5-14 Test acquire response binding validation.
R5-15 Add the int64 upper-bound validation.
R5-16 Extend Broker control coverage.
R5-18 Extend resolver refusal coverage.
R5-43 Remove waitFor timing flakiness.
R5-19 Strengthen successful result payload assertions.
R5-20 Cover route-level 400 responses.
R5-21 Refine deep structuredClone/RangeError classification.
R5-22 Extend manifest tests.
R5-23 Extend result-mapping tests.
R5-24 Cover the JSON media envelope.
R5-25 Remove or exercise the unused fixture contractVersion.
R5-34 Evaluate task projection index performance.
R5-35 Reduce quadratic definition validation.
R5-36 Consolidate duplicate Java helpers.
R5-37 Strengthen Java test assertions.
R5-38 Consolidate duplicate fixture loaders.
R5-39 Disambiguate projection fixture expectations.
R5-6 Restore stronger cancellation sequencing evidence in the fault gate.
R5-40 Revisit acquire envelope placement.
R5-41 Add a valid prompt_get Java case.

@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

doudouOUC
doudouOUC previously approved these changes Sep 30, 2026

@doudouOUC doudouOUC left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已完成对当前 head 40b05a74f280e37ae5e0471ec9219934dd971e7b 的全面独立审查,覆盖相对 main 的全部 58 个文件,并核对跨包调用、Runtime/Session/Workspace 归属、授权与 fencing、TypeScript/Java 记录及资源契约、目录发布、取消、原身份恢复和释放路径,以及双语设计和全部现有 review threads。未发现新增的可确认合并阻塞问题,Approve 本 PR 明确限定的私有、显式启用 H1 范围。

此前我提交的两个取消问题已在当前提交修复,并通过独立复验:首轮 MCP 初始化可响应 cancel/deadline;目录刷新进入 replacement 后也保留取消信号,尚未派发的配置不会在中止后继续发送。已经持久提交 dispatch 的工作保持原 operation 身份、最多发送一次;迟到结果按原 ID 恢复,不通过重发制造第二次物理效果。能够证明未发送或已经结算时,后续 prompt 和 detach 正常。

历史 Critical 按当前代码重新判定,不以 thread 的 resolved 标志代替验证。R3-11 仍成立并明确保留:无法结算的远端副作用使 detach/delete 返回 503,attachment、activation 和原恢复 owner 会继续保留。原 Runtime 可能仍在写入,强制释放不能证明已经排空,也会丢失迟到结果的恢复入口;这是本私有切片公开接受的限制,不是已修复项。生产启用前仍需解决客户端/SSE 生命周期与持久恢复所有权的分离及有界 orphan 清理;已有延期 Suggestions 不重复提交。

本地验证:

  • npm run build、npm run bundle、npm run typecheck、改动 TS 文件 ESLint、适用文件 Prettier 和 diff whitespace 检查通过。
  • CLI 11 个文件 1047 项、Core 5 个文件 304 项通过。Hosted 的 198 项最终关闭 coverage、串行文件运行全部通过;早先并行运行出现共享 coverage 目录竞争及已记录的等待超时 flake,未把那些运行作为全绿证据。
  • Runtime Broker 全套 530 项无 failure/error,其中 1 项最初因未配置 worker bundle 跳过;构建 bundle 后,该真实 worker 用例另行运行通过、无跳过。Managed Agent 定向 40 项通过,包含 H2 MySQL 模式的 SQL、资源闭包和 V23 迁移检查。
  • 16 个独立竞争窗口探针及 2 个实际 Hosted HTTP route 探针通过,覆盖 warmup/refresh 的 intent、acquire、durable dispatch acknowledgement、configure response 等待及 cancel/deadline。实际 HTTP 路径验证取消时不准入 input/模型、不初始化第二台 server;未知效果保留 owner,迟到结果恢复后 prompt 与 detach 正常。

独立探针使用当前提交的真实 Hosted 状态机、HTTP routes、本地持久化 stores 和 Broker HTTP 客户端;Broker 为 loopback 协议替身,模型和远程 store connector 也有明确替身。全局 qwen 0.24.6 无法启动此私有 profile,因此采用 test-script fallback;不将这些结果宣称为已部署 Java Broker/SQL/外部 MCP 的全栈 E2E。原生 stdio/HTTP/SSE transport 由 Runtime 单元测试覆盖。本地未重跑物理 fault-gate、MySQL/MariaDB profiles 或完整生产故障恢复矩阵。当前 head 的 Node、Java、MariaDB、MySQL fault-gate、daemon 和 web-shell 测试 CI 已通过;自动 review job 仍在运行。V23 与待合并 V20–V22 的发布顺序仍须按双语设计约束协调。

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head 40b05a74f280e37ae5e0471ec9219934dd971e7b

Verdict: COMMENT. One historical Critical is confirmed still standing at this head by reading the code, six more were filed two commits ago and have never been adjudicated at this head, and an 11,184-line / 58-file diff cannot be scanned to the depth this channel's Approve path requires inside one review budget. Under this channel's rules any historical blocking issue that still exists, or cannot be confirmed, forbids an Approve — so this is not a judgement that the PR is unsound, and I am not filing a new Critical of my own.

R3-11 — a recovery-blocked MCP session can never be detached or deleted: STILL STANDS

I re-read this at 40b05a74 rather than taking the thread's word for it. packages/cli/src/serve/hosted-harness-session.ts lines 1071–1091, the single close handler that both /session/:id/detach and DELETE /session/:id route through:

try {
  await session.mcp?.close();
  await session.managed.close();
  for (const stop of session.streams) stop();
  sessions.delete(req.params['id']);
  res.sendStatus(204);
} catch {
  error(res, 503, 'managed_session_close_failed');
} finally {
  session.mcpClosing = false;
  session.mcpBusy = false;
}

The chain holds end to end at this head. HostedMcpSession.close() throws HostedMcpRecoveryRequiredError while hasPendingOperations(), which is precisely the state the design doc expects ("Physical connection loss with an unresolved request remains recovery-blocked") — status() → accept() commits recovery_blocked/outcome_unknown, and that is not in ['settled','failed','cancelled']. The throw lands in a bare catch that answers 503 managed_session_close_failed and returns, so the three statements after await session.mcp?.close() never run:

  • session.managed.close() is skipped, which is what calls renewal.stop() and authority.releaseActivation(). The activation lease is therefore not merely left to expire — the renewal interval at leaseDurationMs/3 keeps actively holding it.
  • for (const stop of session.streams) stop() is skipped, so the session's open SSE responses leak.
  • sessions.delete(req.params['id']) is skipped, so the HostedSession entry leaks for the daemon's lifetime.

The finally resets mcpBusy and mcpClosing, so a retry re-enters the same handler and throws again; the pending record can only be cleared by a conclusive runtime reply that, by hypothesis, will never arrive. There is no force-close escape, because both exit routes share this handler.

I want to be accurate about the status of this finding rather than merely repeat it: @doudouOUC's approval of this head states that R3-11 still stands and is deliberately retained as an accepted limitation of this private, explicitly-enabled H1 slice, with production enablement conditioned on separating the client/SSE lifecycle from durable recovery ownership and adding bounded orphan cleanup. That is a scoping decision the author and maintainer are entitled to make, and it is documented. It does not change how this channel must vote: a blocking issue that still exists at the head is a bar to Approve here, so the choice is to land it with that bar recorded, or to close the leak. The narrow fix is to make the MCP close non-fatal to the rest of the teardown — settle managed.close(), the stream shutdown and the map deletion even when session.mcp.close() throws, then report the MCP failure — so a recovery-blocked MCP session can still release its activation and stop leaking, without pretending the unresolved remote effect was settled.

Six Criticals filed two commits ago have never been ruled on at this head

Round 5 reviewed 06c1df7f and published at 2026-09-30T09:15:31Z with six findings at severity C. The head has moved twice since, and review-pr is still pending on this head, so no automated round has re-adjudicated them here and I had no budget to verify six call chains across two languages myself. They are unconfirmed, not cleared:

  • R5-1 hosted-harness-session.ts:705 — POST /session/:id/mcp/configurations admits a definition replacement while another operation is in flight.
  • R5-2 hosted-mcp-session.ts:764 — close() commits releaseState: 'releasing' before dispatching mcp-release.
  • R5-3 managed-mcp-runtime.ts:698 — the list-changed lookup indexes a plain object literal with a remote-controlled key.
  • R5-4 managed-mcp-runtime.ts:725 — retiring the predecessor connection runs inside the try whose catch tears down.
  • R5-5 HttpRuntimeTransport.java:618 — an MCP operation whose encoding exceeds TOOL_REQUEST_LIMIT_BYTES (256 KiB).
  • R5-6 RuntimeBrokerService.java:843 — the synthesized SessionContext adopted by releasedSession is never evicted.

R5-3 and R5-6 are the two I would read first: a remote-controlled property key and an unbounded cache are both shapes that stay invisible until they are expensive.

Round 3's other eleven Criticals (R3-1 to R3-10, R3-12, R3-13) were filed at 4134aaf2 and are likewise unadjudicated at this head by anything I can cite. The two cancellation regressions @doudouOUC raised at 06c1df7f and 67c908ea are, by their own current-head re-verification, fixed — I did not re-derive those and am not carrying them as my own confirmation.

Gate not completed: the current-head scan

I read the close handler above and the review record; I did not scan the production surface of this diff. Unread, and therefore unconfirmed: hosted-mcp-session.ts (+1039, new), managed-mcp-runtime.ts (+1072, new), managed-mcp-routes.ts (+78, new), hosted-turn-wait.ts (+27, new), managed-mcp-record.ts (+322, new), managed-mcp-protocol.ts (+134, new), ManagedMcpRecords.java (+185, new), ManagedMcpCatalogService.java (+84, new), ManagedMcpProtocol.java (+73, new), ManagedMcpController.java (+24, new), the migration V23__managed_mcp_records.sql, and the changes to hosted-harness-session.ts (+241/−9), hosted-workspace-broker.ts (+96/−9), hosted-workspace-tool-turn.ts (+85/−50), managed-runtime-tool-executor.ts (+133/−1), managed-session-authority.ts (+90/−9), RuntimeBrokerHttpServer.java (+33/−20), RuntimeBrokerService.java (+34/−5), ManagedExtensionRecordStore.java (+111/−7) and the +178/−4 OpenAPI change.

Two of those deserve naming because they carry risk that a green suite does not bound. The V23 migration lands behind a feature that is private and explicitly enabled, and @doudouOUC's own approval records that its release ordering against the pending V20–V22 migrations still has to be coordinated — a migration applied out of order is not something any test in this diff can catch. And the OpenAPI change should be checked the way the last contract PR was: read x-qwen-implementation-status on each touched operation at this head, so a route that is not served cannot be certified as one.

CI

Green at this head: Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), Serve A/B, web-shell E2E Smoke, TUI parity snapshots, OpenTUI no-flicker gate, Hosted process fault gates / MySQL 8.4 / Java 21, Runtime Broker and Managed Agent MariaDB / Java 21, Real daemon E2E / Java 11 and the full Java matrix all pass. review-pr is pending, which this channel does not treat as a gate — though here its completion is what would rule on R5-1 to R5-6 at the head that actually merges.

Next step: the shortest path to an Approve is to let the pending review round rule on R5-1 to R5-6 at this head, and to decide R3-11 explicitly — either close the teardown leak so a recovery-blocked MCP session still releases its activation, or record the accepted limitation somewhere a gate reads, since a still-standing Critical blocks here whatever its justification.

@wenshao
wenshao dismissed stale reviews from ghost and doudouOUC via f2618e3 September 30, 2026 11:50
@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Merged latest origin/main 3b18cfe5e4ab7ea72f1a92736186dacf753bf727 in f2618e3, resolving the conflicts introduced by #12894 and retaining #13098's queued approval expiry guard.

The main Shell publication/recovery flow and the existing MCP flow both remain wired through the refactored Hosted turn entry point. MCP retains its shared owner, per-prompt preparation identity, cancellation/reconciliation behavior and post-approval grant renewal; remote Shell uses its publication/v3 path. Both sides' tests are retained, and the three direct MCP constructor fixtures were adapted to the combined signature.

The migration dependency is now in this branch: main includes publication V20–V22, followed by MCP V23. Both design languages and the PR description reflect that #12894 has landed. Migrations remain ordered; this does not authorize applying V20–V22 retroactively to a database where standalone V23 was already deployed.

Validation for the merge:

  • Build, typecheck, bundle, lint, formatting and diff checks passed.
  • CLI: 1077 tests across 10 files; Core: 85 tests across 3 files; Managed Agent: 56 tests including H2/Flyway migrations through V23 and publication/MCP/Workspace authorization contracts.
  • Java Broker: 538 cases, zero failures/errors. The two initially skipped interop cases were rerun after supplying the built artifacts and both passed without skips.
  • Independent source verification: 9/9 probes passed, covering refresh cancellation windows, Unicode admission, delayed-approval grant renewal, and two actual Hosted HTTP prompt turns through the refactored execution helper. The two turns retain the same MCP owner, use distinct turn IDs, make no v3 calls, and release the owner only at detach. Model/Broker transport and approval time are controlled; this is not an OSS or deployed Java/SQL E2E claim.
  • Two consecutive clean audits, open-ended and directed, completed before commit. The committed tree cea90f9c809f3a53773d282e81501fb80e835a78 matches the audited tree; hooks made no additional changes.

Previously deferred Suggestions and the documented R3-11 H1 limitation remain unchanged. New CI will validate the pushed merge commit.

@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head f2618e3027d41921bf6a798cb061c50bb6d6e357

Verdict: COMMENT. The historical Critical I confirmed by reading code at the previous head is unchanged at this one, the six Criticals filed two heads ago sit in files this head does not touch, no reviewer of any kind has ruled on this head, and a 58-file / +11,187−185 diff cannot be scanned to the depth this channel's Approve path requires in one budget. Any historical blocking issue that still stands or cannot be confirmed forbids an Approve here, so this is a COMMENT. I file no new Critical of my own.

This head is a merge of main, not a fix

The PR's own content did not change. Totals moved from +11,184/−184 at 40b05a74 to +11,187/−185 here, and the delta is reconciliation with a newer main rather than new work: hosted-workspace-tool-turn.test.ts gains imports of assertManagedSessionDurableRef and HttpToolPublicationOwner, broker mocks for prepareV3/executeV3/acknowledgeV3, and an extra constructor argument — all of which arrive from main's tool-publication surface, alongside expectWritesStopped from the approval-expiry work. A three-dot compare against 40b05a74 lists 76 files for that reason; the PR's file list is still the same 58. So nothing in this head answers any finding, and my review of 40b05a74 carries over in substance — re-verified below rather than inherited.

Also worth stating plainly, because it changed with the head: @doudouOUC's approval of 40b05a74 is now dismissed, and the review list at f2618e30 is empty. No maintainer approval and no automated round stands at the head that would merge.

R3-11 — a recovery-blocked MCP session can never be detached or deleted: CONFIRMED STILL STANDS

I re-read this at f2618e30 rather than relying on the earlier pass. packages/cli/src/serve/hosted-harness-session.ts lines 1441–1457 are unchanged:

await session.mcp?.close();
await session.managed.close();
for (const stop of session.streams) stop();
sessions.delete(req.params['id']);
res.sendStatus(204);
} catch {
  error(res, 503, 'managed_session_close_failed');
} finally {
  session.mcpClosing = false;
  session.mcpBusy = false;
}

HostedMcpSession.close() throws HostedMcpRecoveryRequiredError while hasPendingOperations(), which is the state the design doc itself expects. The throw lands in the bare catch, so the three statements after it never run: session.managed.close() — which is what calls renewal.stop() and authority.releaseActivation(), leaving the renewal interval at leaseDurationMs/3 actively holding the activation lease rather than letting it lapse; the stream shutdown, leaking the session's open SSE responses; and sessions.delete(...), leaking the HostedSession entry for the daemon's lifetime. The finally resets mcpBusy and mcpClosing, so a retry re-enters and throws again, and the pending record can only be cleared by a conclusive runtime reply that by hypothesis never arrives. Both /session/:id/detach (line 1453) and DELETE /session/:id (line 1456) route through this one handler, so there is no force-close escape.

@doudouOUC's now-dismissed approval recorded this as deliberately retained — an accepted limitation of the private, explicitly-enabled H1 slice, with production enablement conditioned on separating the client/SSE lifecycle from durable recovery ownership and adding bounded orphan cleanup. That is a scoping decision the author and maintainer may make, and it is documented. It does not change how this channel votes: a blocking issue still standing at the head bars an Approve here. The narrow fix remains to make the MCP close non-fatal to the rest of the teardown — settle managed.close(), the stream shutdown and the map deletion even when session.mcp.close() throws, then report the MCP failure — so a recovery-blocked session still releases its activation and stops leaking, without claiming the unresolved remote effect was settled.

R5-1 to R5-6: unaddressed at this head, and unadjudicated

Round 5 reviewed 06c1df7f and published six findings at severity C. Every file they name has an identical diff footprint at this head as at 40b05a74, so none was touched by the merge:

  • R5-1 WorkspaceRuntimeTransport.java:145 (+19/−1) — the grant-mode entry point lets a provably pre-dispatch ownership refusal take a path it should not.
  • R5-2 hosted-mcp-session.ts:764 (+1039, new) — close() commits releaseState: 'releasing' before dispatching mcp-release.
  • R5-3 managed-mcp-runtime.ts:698 (+1072, new) — the list-changed lookup indexes a plain object literal with a remote-controlled key.
  • R5-4 managed-mcp-runtime.ts:725 — retiring the predecessor connection runs inside the try whose catch tears down.
  • R5-5 HttpRuntimeTransport.java:618 (+16/−1) — an MCP operation whose encoding exceeds TOOL_REQUEST_LIMIT_BYTES (256 KiB).
  • R5-6 RuntimeBrokerService.java:843 (+34/−5) — the synthesized SessionContext adopted by releasedSession is never evicted.

I did not re-derive these six myself and am not asserting them as my own findings; I am recording that nothing has ruled on them at the head that would merge, which under this channel's rules is the same as unconfirmed. R5-3 and R5-6 are the two I would read first — a remote-controlled property key and an unbounded cache both stay invisible until they are expensive. Round 3's other eleven Criticals are in the same position.

Gate not completed: the current-head scan

I read the close handler above, the merge delta, and the review record. I did not scan this diff: hosted-mcp-session.ts (+1039, new), managed-mcp-runtime.ts (+1072, new), managed-mcp-routes.ts (+78, new), hosted-turn-wait.ts (+27, new), managed-mcp-record.ts (+322, new), managed-mcp-protocol.ts (+134, new), ManagedMcpRecords.java (+185, new), ManagedMcpCatalogService.java (+84, new), ManagedMcpProtocol.java (+73, new), ManagedMcpController.java (+24, new), the migration V23__managed_mcp_records.sql, and the changes to hosted-harness-session.ts (+241/−9), hosted-workspace-broker.ts (+96/−9), hosted-workspace-tool-turn.ts (+85/−51), managed-runtime-tool-executor.ts (+133/−1), managed-session-authority.ts (+90/−9), RuntimeBrokerHttpServer.java (+33/−20), RuntimeBrokerService.java, ManagedExtensionRecordStore.java (+111/−7) and the +178/−4 OpenAPI change.

Two deserve naming because a green suite does not bound them. V23's release ordering against the migrations now on main — this head merged a main that carries V20, V21 and V22 for the tool-publication work, so V23 is no longer the only pending migration and the ordering constraint the design doc records is now live rather than hypothetical; an out-of-order migration is invisible to every test in this diff. And the OpenAPI change should be checked the way the last contract PR was: read x-qwen-implementation-status on each touched operation at this head, so a route the server does not serve cannot be certified as one.

CI

Two checks fail at this head and several are still pending, so the picture is incomplete rather than green. web-shell E2E Smoke and Desktop Shell (${{ matrix.os }}) both report failure with a duration of 0, and the latter carries its matrix expression unexpanded in its own name — a workflow-definition shape, not a product result. This PR touches no web-shell file, both jobs passed at 40b05a74, and the head between them is a merge of main, so on the available evidence neither failure is attributable to this diff and I am not treating either as a Critical. Lint & Static, Test (ubuntu-latest, Node 22.x), Serve A/B, Hosted process fault gates / MySQL 8.4 / Java 21 and review-pr are pending; pending state is not a gate here, though review-pr completing is what would rule on R5-1 to R5-6 without another manual pass. The Java matrix, Runtime Broker and Managed Agent MariaDB / Java 21, Real daemon E2E / Java 11, Integration Tests (no-AK, No Sandbox), TUI parity snapshots and OpenTUI no-flicker gate pass.

Next step: two things unblock this, and neither needs a full re-scan first. Decide R3-11 explicitly — either make the teardown survive an MCP close failure so a recovery-blocked session still releases its activation, or record the accepted limitation somewhere a gate reads, since a still-standing Critical blocks here whatever its justification. Then get a ruling on R5-1 to R5-6 at the head that will merge, from the pending review round or a targeted pass over those six call chains. Given this PR has been through five maintainer rounds, the repository's own guidance about not letting review rounds balloon it applies: beyond those Criticals, remaining Suggestions are better deferred to a follow-up issue than folded into another round.

@doudouOUC doudouOUC left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复审冲突解决后的 head f2618e3027d41921bf6a798cb061c50bb6d6e357,重新 Approve。本次以上一轮 40b05a74 的完整审查为基础,核对合入 main 3b18cfe5e 的 merge/remerge diff、全部人工冲突解决,以及 MCP 与远程 Shell 交付的衔接,未发现新增的可确认合并阻塞问题。

合并后的统一 turn 执行路径保留 MCP Session、审批参数与取消信号;MCP 使用原 Broker owner、实际 prompt/call 身份和原 operation 恢复,Shell 的 v3 publication 路径仍独立选择。Runtime 的 publisher 接线保留 MCP executor,Java control 转发、撤权后的原 owner 恢复、失败释放缓存清理及 drain-before-release 顺序均保留。MCP 状态机、记录契约和 authority 文件与已审查版本一致。双语迁移说明已更新为 V20/V21/V22 先于 V23,前置迁移现在已随 main 合入。

当前提交的本地验证:

  • build、typecheck、合并涉及 TS 文件 lint/format 和 diff whitespace 检查通过。
  • CLI 10 个文件 1113 项、Core 3 个文件 85 项通过;覆盖 Hosted、MCP、Runtime worker、远程 Shell publication 及 Session store。
  • Runtime Broker 定向 213 项、Managed Agent 定向 41 项通过,均无 failure/error/skip;Managed Agent 测试前安装了当前提交的 Broker 制品。
  • 16 个取消/deadline 竞争窗口探针和 2 个实际 Hosted HTTP route 探针重新运行通过。独立探针仍使用 loopback Broker 协议替身和明确的模型/store connector 替身,不将其作为外部 MCP/已部署 Java Broker/SQL 的全栈 E2E;本轮没有重跑完整物理故障恢复矩阵。

批准范围仍是私有、显式启用 H1。R3-11 的未知远端副作用保留 attachment/activation/原恢复 owner 限制没有改变,也不声明已修复;客户端/SSE 生命周期分离与有界 orphan 清理仍是生产启用前的后续工作。新 head 的 CI 正在重跑,提交本 review 前没有已失败项,部分测试与自动 review 尚未结束。

@wenshao
wenshao enabled auto-merge September 30, 2026 12:08
@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification, round 9 — 40b05a74 and f2618e30

Follow-up to round 8 and the author's 40b05a74 update. This round also answers the question in the latest Critical-only review: whether the R5 Criticals hold at this head. The previous head, 67c908ea, is the A/B baseline throughout.

The head moved to f2618e30 while I was writing; it merges main, which now includes #12894. That merge leaves hosted-mcp-session.ts and managed-mcp-runtime.ts unchanged, but the Hosted turn path and the Broker changed. So I rebuilt f2618e30 and reran the key scenarios on it. Every result below is the same on both heads.

Merge reference in one paragraph.

  • Fixed on the real stack. Each item below was A/B'd against 67c908ea:

    The approval path merged from main also works when the answer comes after the 300 s MCP grant lease.

  • Not fixed: R5-2 for a Session with one pinned server. This is the most common shape. If its first mcp-release never reaches the Runtime, detach stays 503 and the Workspace lease stays held, exactly as on 67c908ea. Only Sessions with two or more servers now recover.

  • Not new, and not covered by the R5-2 retry. A release that timed out (managed_mcp_drain_unknown) is answered with the same stored result forever. A Session with two servers never detaches, on either head.

  • Flyway. feat(managed-agent): Add durable remote Shell result delivery #12894 has landed. f2618e30 carries V20–V22 and V23, and an empty database migrates to 1..23, so the ordering concern is gone.

I recommend fixing R5-2 for the one-server Session before merge, or explicitly accepting it next to R3-11.

round 9 fixes

R5 finding Real stack at 40b05a74 and f2618e30 67c908ea
R5-1 replacement during a turn turn ends turn_error, not latched; the replacement settles at 6.4 s; the next turn runs on the new server turn ends with no terminal event and stays blocked; next prompt 409; detach 503
R5-2 undelivered release, one server stuck: detach 503 on every retry; lease held stuck (same)
R5-2 undelivered release, two servers recovers: the 2nd detach re-sends both releases → 204, lease 0 stuck
R5-3 prototype-named notifications ignored: 0 configurations, catalog unchanged every one triggers a replacement (config revision 1 → 2 → 3 → 4)
R5-4 slow retirement of the old connection the replacement settles; the new server is used replacement 503, outcome_unknown
R5-5 oversized MCP operation → 413 not driven here (in round 8 a 300 KiB argument was refused by the Harness before the Broker); mutant J14 killed —
R5-7 adopted SessionContext eviction¹ not driven on the real stack; mutant J15 killed —

¹ Labels follow the linked inline threads. The Critical-only review numbers some items differently: its "R5-6" is the inline R5-7, and its "R5-1" points at WorkspaceRuntimeTransport.java:145, which is not the inline R5-1 above and which I did not drive.

  • Build. Fresh bundles and Spring jars from 40b05a74 and then f2618e30, on MySQL 8.4.11. main's new @a2a-js/sdk dependency needed a pnpm install first. The old arm is my 67c908ea build, with its Broker jar kept aside. The stdio server gained three test hooks:
    • addtool changes its tool list at runtime;
    • rawnotify sends a raw notification;
    • --sticky makes a grandchild keep stdout open for about 45 s after the server exits.
  • R4-9 (fixed). A lone surrogate in a resource URI, a prompt name, an argument key or an argument value:
    • all four return 400 invalid_mcp_operation with 0 Broker requests and 0 server reads;
    • the next read settles, detach returns 204 and the lease goes to 0;
    • the well-formed emoji control passes through unchanged.
  • Refresh-cancellation P1 (fixed). The model calls addtool, so the next refresh in the same turn starts a real replacement, and the Broker proxy holds that replacement's acquire:
    • 67c908ea: the turn is cancelled at 2.0 s, but the replacement mcp-configure is still sent at 4.9 s. A detach during the hold returns 204, releasing before the acquisition finished.
    • 40b05a74: no configure is sent after the cancel. A detach during the hold returns 503, so nothing is released. The next turn is OK and detach returns 204.
    • If the replacement's mcp-configure is held instead (after dispatch_started), it is sent once and not resent, on both heads.
  • R5-1 and R5-4. A turn is running a 3 s tool on remote when an explicit replacement of the sticky server (rev 1 → 2) is posted. Retiring the old sticky connection takes more than 5 s. The table above has both outcomes.
  • Approval merged from main. With approvalMode: default, the MCP tool call raises an approval request. I allowed it at 310 s, after the 300 s grant lease: the tool ran once with the original arguments, the turn ended end_turn with no Broker errors, and detach returned 204. A 2 s control behaved the same.

round 9 release paths, Flyway and regression

  • R5-2, one pinned server (not fixed). The Broker proxy answers the first mcp-release with 503 without forwarding it.

    • Detach 1 returns 503.
    • Detach 2 goes straight to the Broker tool-session release, which returns 409 managed_runtime_identity_conflict ("still owns unfinished work"). After that, mcp-status and acquire return 409 runtime_session_not_ready, so the original release is never re-sent. Every later detach repeats this, and the lease stays held.
    • Cause: when every configuration is releasing, close() tries broker.release() before looking up or re-sending the release. The Runtime still holds the open connection, so that call fails, and afterwards the runtime Session is not ready for the re-send path.
    • With two pinned servers, the second configuration is still active, so this shortcut is skipped. Detach 2 re-sends both releases and returns 204; on 67c908ea it stayed stuck.
    • The new unit test retries an undelivered release with the original immutable identity kills the inverse mutant, yet the one-server shape above still fails on the real stack.
  • A release that timed out (not new, not covered). Here the stdio server's stdout outlives the process by about 45 s, so its release exceeds the 5 s drain bound.

    • One server: detach returns 503 while the pipe is open and 204 at 52 s, on both heads.
    • Two servers, or a server replaced on 40b05a74: detach stayed 503 for as long as I watched (98 s and 134 s), and the lease stayed held.
    • On 40b05a74, each retry re-sends mcp-release with the same operation ID and gets back the stored outcome_unknown/managed_mcp_drain_unknown. The retry never drains again, so the later configurations are never released and their server process stays alive. 67c908ea only looks the release up and is equally stuck.
    • For example, a stdio server launched through a wrapper whose child keeps stdout open can take this path.
  • Flyway. On f2618e30 the jar carries V20–V22 and V23, and an empty database migrates to 1..23 and starts. On 40b05a74, applying this PR first and feat(managed-agent): Add durable remote Shell result delivery #12894's V20–V22 afterwards failed validation (Detected resolved migration not applied to database: 20). Now that feat(managed-agent): Add durable remote Shell result delivery #12894 is on main, that order can no longer happen.

  • Regression. The full round 1–8 set holds on 40b05a74, except that I did not rerun the worker SIGKILL scenario. The key subset (S1, list_changed, prompt cancel, busy owner, R1-2, hung-read status/cancel, lost replies, orphaned Session) reran on f2618e30 with the same results:

    Area What I saw
    S1, three transports exact blob, 409 conflicts, 0 leaks, detach and load
    list_changed tools/list_changed and resources/list_changed recover in the same turn
    Raw-op cancel 202 running
    Prompt cancel the turn ends cancelled
    Late settlement 33.6 s
    F6 17 reconfigures keep 1 live stdio process
    R1-2 / R1-5 both hold
    Busy owner B detaches 204 and the holder is unchanged
    Hung-read status/cancel 200 / 202
    Lost replies 1 effect; revoked-access close 204
    SSE restart / HTTP down at close served / detach 204
    Orphaned Session once the old writer's store lease expires: load 200, prompt 202, detach 204, lease 0
    N1 unchanged: UNKNOWN from 37.6 s, turn ends at 638.5 s, then blocked
  • Mutants. 9 of 10 are killed. Baselines were cli 226/227 and broker 530 with 1 skipped; the one cli failure is the known 404 load flake in serves MCP cancel during dispatch…. K2 survives because it removes the first configuring check in refresh(), and the same check is repeated right after the discover dispatch.

    Mutant Change it inverts Killing test
    K1 R4-9 route check rejects invalid MCP strings before committing or dispatching (4 cases)
    K3 R5-2 re-send retries an undelivered release with the original immutable identity
    K4 R5-3 Object.hasOwn ignores inherited notification names without changing the catalog
    K5 R5-4 tolerated sibling drain keeps a healthy replacement while retaining the undrained predecessor hold
    K6 / K7 refresh signal / close fence stops a refresh replacement cancelled during intent / acquire / dispatch_started and fences close until it drains
    K8 late-approval grant renewal honors allow approval for an MCP call while retaining its shared owner
    J14 R5-5 413 HttpRuntimeTransportTest.oversizedMcpRequestHasDefinitiveClassification and one more
    J15 R5-7 eviction RuntimeBrokerServiceTest.failedAdoptedReleaseLeavesReclamationReachable
  • CI. 40b05a74: 24 checks pass. f2618e30: 22 pass, and Test (ubuntu) and review-pr were still running when I posted. web-shell E2E Smoke and Desktop Shell are listed as failures, but both jobs were cancelled rather than failed.

  • Not verified. Windows/Linux native runs, a real model, and MariaDB this round. I also did not rerun the worker SIGKILL scenario, and did not drive R5-5 or R5-7 end to end.

Evidence: wenshao/qwen-code@78144486/pr12946/r9. It contains per-scenario results for both arms and for f2618e30, the regression summaries, mutants by test name, the Flyway boots, and the harness, including the three new stdio server hooks.

中文版

真实环境验证第九轮 —— 40b05a74 与 f2618e30

接第八轮与作者的 40b05a74 更新说明。本轮也回答了最新 Critical-only 评审提出的问题:R5 各条在这个 head 上是否成立。全程以上一个 head 67c908ea 作为 A/B 基线。

我写报告期间 head 变成了 f2618e30,它合入了已包含 #12894 的 main。这次合并没有改动 hosted-mcp-session.ts 和 managed-mcp-runtime.ts,但 Hosted 的 turn 路径和 Broker 都有变化。所以我重新构建了 f2618e30,在它上面重跑了关键场景。下面每一项结果在两个 head 上都相同。

合并参考(一段话):

  • 真实环境中已修复。 下面各项都与 67c908ea 做了 A/B:

    从 main 合入的审批路径,在超过 300 s 的 MCP grant 租期后才批准时也能正常工作。

  • 未修复:只 pin 一台服务端的 Session 上的 R5-2。 这是最常见的形状。如果它的第一个 mcp-release 没送达 Runtime,detach 会一直 503,Workspace lease 一直被持有,表现与 67c908ea 完全相同。只有两台及以上服务端的 Session 现在能恢复。

  • 不是新问题,但 R5-2 的重试也没覆盖到。 排空超时(managed_mcp_drain_unknown)的 release,之后每次都会得到同样的存储结果。pin 两台服务端的 Session 永远 detach 不了,两个 head 都一样。

  • Flyway。 feat(managed-agent): Add durable remote Shell result delivery #12894 已经合入。f2618e30 带有 V20–V22 和 V23,空库会迁移到 1..23,合并顺序的问题已经不存在。

建议合并前修掉单服务端 Session 的 R5-2,或者把它与 R3-11 一起明确列为接受的限制。

R5 发现 40b05a74 与 f2618e30 真实环境 67c908ea
R5-1 turn 进行中的替换 这一轮以 turn_error 结束,不被锁住;替换在 6.4 s 结算;下一轮使用新服务端 这一轮没有终态事件,并被永久锁住;下一个 prompt 409;detach 503
R5-2 release 未送达,单服务端 卡死:每次重试 detach 都是 503;lease 被持有 卡死(相同)
R5-2 release 未送达,两台服务端 能恢复:第 2 次 detach 重发两个 release → 204,lease 归零 卡死
R5-3 原型名通知 被忽略:0 次配置,目录不变 每一种都触发一次替换(配置修订 1 → 2 → 3 → 4)
R5-4 旧连接退役慢 替换结算,使用新服务端 替换 503,outcome_unknown
R5-5 超大 MCP 操作 → 413 本轮未驱动(第八轮中 300 KiB 的参数在到达 Broker 前就被 Harness 拒绝);变异体 J14 被杀 —
R5-7 被收养的 SessionContext 移除¹ 未在真实环境驱动;变异体 J15 被杀 —

¹ 编号以文中链接的行内评论为准。Critical-only 评审的部分编号不同:它的"R5-6"即行内的 R5-7;它的"R5-1"指向 WorkspaceRuntimeTransport.java:145,与上表的行内 R5-1 不是同一条,那一项我没有驱动。

  • 构建。 先后用 40b05a74 和 f2618e30 全新构建 bundle 和 Spring jar,跑在 MySQL 8.4.11 上。main 新增了 @a2a-js/sdk 依赖,所以先执行了 pnpm install。旧臂是我的 67c908ea 构建,其 Broker jar 单独保留。stdio 服务端新增三个测试钩子:

    • addtool:运行时修改自己的工具列表;
    • rawnotify:发送原始通知;
    • --sticky:服务端退出后,由孙进程再占住 stdout 约 45 s。
  • R4-9(已修复)。 资源 URI、prompt 名称、参数键或参数值里带孤立代理项时:

    • 四种情况都返回 400 invalid_mcp_operation,Broker 请求为 0,服务端读取为 0;
    • 下一次读取 settled,detach 返回 204,lease 归零;
    • 合法 emoji 的对照原样通过。
  • 刷新替换时的取消 P1(已修复)。 模型调用 addtool,同一轮里的下一次 refresh 于是发起真正的替换;Broker 代理扣住这次替换的 acquire:

    • 67c908ea: 2.0 s 时这一轮被取消,但 4.9 s 时仍派发了替换用的 mcp-configure。扣住期间 detach 返回 204,在获取完成前就释放了。
    • 40b05a74: 取消后没有派发任何 configure。扣住期间 detach 返回 503,没有释放任何东西。下一轮正常,detach 返回 204。
    • 如果改为扣住替换的 mcp-configure(已进入 dispatch_started),两个 head 上都只发送一次,不会重发。
  • R5-1 与 R5-4。 一轮对话正在 remote 上执行一个 3 s 的工具时,提交对 sticky 服务端的显式替换(rev 1 → 2)。旧的 sticky 连接退役要超过 5 s。两个 head 的结果见上表。

  • 从 main 合入的审批路径。 在 approvalMode: default 下,MCP 工具调用会发起审批请求。我在 310 s 时批准(已超过 300 s 的 grant 租期):工具用原参数执行了 1 次,这一轮以 end_turn 结束,Broker 没有报错,detach 返回 204。2 s 就批准的对照组表现相同。

  • R5-2,单服务端(未修复)。 Broker 代理对第一个 mcp-release 直接返回 503,不转发。

    • 第 1 次 detach 返回 503。
    • 第 2 次 detach 直接走 Broker 的 tool-session release,返回 409 managed_runtime_identity_conflict("still owns unfinished work")。此后 mcp-status 和 acquire 都返回 409 runtime_session_not_ready,原来的 release 再也没有重发。之后每次 detach 都重复这个过程,lease 一直被持有。
    • 原因:所有配置都处于 releasing 时,close() 会先尝试 broker.release(),再去查询或重发 release。这时 Runtime 里的连接仍然开着,这一步必然失败,之后 runtime Session 也就不再就绪,重发路径走不通。
    • pin 两台服务端时,第二个配置仍是 active,不会走这个快速路径。第 2 次 detach 重发了两个 release,返回 204;在 67c908ea 上则一直卡死。
    • 新增的单测 retries an undelivered release with the original immutable identity 能杀掉对应的反向变异体,但上面单服务端的形状在真实环境中仍然失败。
  • 排空超时的 release(不是新问题,也没被覆盖)。 这里 stdio 服务端的 stdout 比进程多存活约 45 s,所以它的 release 超过了 5 s 的排空上限。

    • 单服务端: 管道还开着时 detach 返回 503,52 s 时返回 204,两个 head 相同。
    • 两台服务端,或在 40b05a74 上被替换过的服务端: 在我观察的整个时间里(98 s 和 134 s)detach 一直是 503,lease 一直被持有。
    • 在 40b05a74 上,每次重试都用同一个 operation ID 重发 mcp-release,得到的都是存下来的 outcome_unknown/managed_mcp_drain_unknown。重试不会重新排空,所以后面的配置永远得不到释放,它们的服务端进程也一直活着。67c908ea 只查询不重发,同样卡死。
    • 例如,通过包装器启动、其子进程一直占着 stdout 的 stdio 服务端,就可能走到这条路径。
  • Flyway。 f2618e30 的 jar 带有 V20–V22 和 V23,空库迁移到 1..23 并正常启动。在 40b05a74 上,如果本 PR 先应用、之后才加入 feat(managed-agent): Add durable remote Shell result delivery #12894 的 V20–V22,校验会失败(Detected resolved migration not applied to database: 20)。现在 feat(managed-agent): Add durable remote Shell result delivery #12894 已在 main 上,这种顺序不会再出现。

  • 回归。 第一到第八轮的完整集合在 40b05a74 上都仍然成立,只是这次没有重跑 worker SIGKILL 场景。关键子集(S1、list_changed、prompt 取消、busy owner、R1-2、挂起读取的查询/取消、回复丢失、孤儿 Session)在 f2618e30 上重跑,结果相同:

    项目 结果
    S1,三种 transport blob 逐字节一致,409 冲突,零泄漏,detach 与 load
    list_changed tools/list_changed 和 resources/list_changed 在同一轮内恢复
    原始操作取消 202 running
    prompt 取消 这一轮以 cancelled 结束
    迟到结算 33.6 s
    F6 17 次 reconfigure 只有 1 个存活的 stdio 进程
    R1-2 / R1-5 都成立
    busy owner B 以 204 detach,持有者不变
    挂起读取的查询/取消 200 / 202
    回复丢失 1 次副作用;撤销访问后 close 204
    SSE 重启 / close 时 HTTP 宕机 可继续服务 / detach 204
    孤儿 Session 旧写者的存储租约过期后:load 200,prompt 202,detach 204,lease 归零
    N1 不变:37.6 s 起为 UNKNOWN,这一轮在 638.5 s 结束,随后被阻塞
  • 变异体。 10 个中 9 个被杀。基线为 cli 226/227、broker 530(跳过 1 个);cli 那 1 个失败是已知的 404 负载抖动,出现在 serves MCP cancel during dispatch…。K2 存活,因为它删掉的是 refresh() 里的第一次 configuring 检查,而 discover 派发后紧接着还有一次同样的检查。

    变异体 反转的改动 杀掉它的测试
    K1 R4-9 路由校验 rejects invalid MCP strings before committing or dispatching(4 个用例)
    K3 R5-2 重发 retries an undelivered release with the original immutable identity
    K4 R5-3 Object.hasOwn ignores inherited notification names without changing the catalog
    K5 R5-4 容忍旧连接排空失败 keeps a healthy replacement while retaining the undrained predecessor hold
    K6 / K7 refresh 传递 signal / close 守卫 stops a refresh replacement cancelled during intent / acquire / dispatch_started and fences close until it drains
    K8 延迟审批时续 grant honors allow approval for an MCP call while retaining its shared owner
    J14 R5-5 返回 413 HttpRuntimeTransportTest.oversizedMcpRequestHasDefinitiveClassification 等
    J15 R5-7 移除 RuntimeBrokerServiceTest.failedAdoptedReleaseLeavesReclamationReachable
  • CI。 40b05a74:24 项通过。f2618e30:22 项通过,发帖时 Test (ubuntu) 和 review-pr 仍在运行。web-shell E2E Smoke 和 Desktop Shell 被列为失败,但这两个任务其实是被取消,并非测试失败。

  • 未验证。 Windows/Linux 原生运行、真实模型、本轮的 MariaDB。另外没有重跑 worker SIGKILL 场景,R5-5 和 R5-7 也没有做端到端驱动。

证据:wenshao/qwen-code@78144486/pr12946/r9。其中包括两个臂以及 f2618e30 的各场景结果、回归汇总、按测试名归因的变异体、Flyway 启动记录,以及装置脚本(含三个新的 stdio 服务端钩子)。

@wenshao
wenshao added this pull request to the merge queue Sep 30, 2026
Merged via the queue into main with commit 7827a3f Sep 30, 2026
94 of 115 checks passed
yiliang114 pushed a commit that referenced this pull request Sep 30, 2026
…e contract to 1.25

Merging main brought in #12946's V23__managed_mcp_records.sql next to this
PR's V23__managed_actions.sql; Flyway refuses two migrations with the same
version, so the server would not start. The Actions migration becomes V24
with its content unchanged. The contract takes 1.25.0 on top of main's
1.24.0 (#12946), with a v1.25 note for the Actions routes this PR serves.
@wenshao

wenshao commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

The round 9 release-recovery findings are addressed in follow-up PR #13109, since this PR has merged. The earlier R5-2 fix was incomplete for a single pinned server: its test double did not preserve the Broker state transition after an unsuccessful owner release.

The follow-up persists connection drain proof before owner release, settles late physical closure on the original release receipt, and retains old-writer lost-ack recovery without treating network errors as closure proof. Final real-stack verification passed all five scenarios, including a baseline Harness writing the legacy records and a new Harness recovering them. Verification report. Unknown invocations still retain their holds.

第九轮报告的两项释放恢复问题已在后续 PR #13109 修复。此前单 server 的 R5-2 判为完全修复不准确:测试替身遗漏了失败 owner 释放造成的 Broker 状态变化。新 PR 已通过五项最终真实链路验证,包含真实旧写者记录的升级接管;未知调用仍保留占用。

yiliang114 pushed a commit that referenced this pull request Sep 30, 2026
A Workspace Session whose connector was built without the Managed Action
store silently fell back to the deployment approval mode and skipped the
Harness mode confirmation -- the case the design declares fail-closed.
createOrLoad now refuses it, so the two remaining null checks in the
connector are gone and an un-wired store fails loudly at admission.

requireOwner let an IllegalArgumentException from actorKey escape as an
unmapped 500. It now answers 403 actor_scope_mismatch, matching the
neighbouring workspace routes.

endedCode maps cancelled to action_cancelled but no test asserted it: a
cancelled Action now fails its response operation with that code and
returns no resolution.

The design docs still named the Actions migration V23 after it moved to
V24; V23 belongs to the Session MCP catalog (#12946).
pull Bot pushed a commit to mcx/qwen-code that referenced this pull request Sep 30, 2026
…3101)

* feat(managed-agent): Serve durable permission Actions (D6b)

* test(managed-agent): Avoid concurrent Action mock stubbing

* fix(managed-agent): renumber the Actions migration to V24 and bump the contract to 1.25

Merging main brought in QwenLM#12946's V23__managed_mcp_records.sql next to this
PR's V23__managed_actions.sql; Flyway refuses two migrations with the same
version, so the server would not start. The Actions migration becomes V24
with its content unchanged. The contract takes 1.25.0 on top of main's
1.24.0 (QwenLM#12946), with a v1.25 note for the Actions routes this PR serves.

* fix(managed-agent): address the D6b review follow-ups

A Workspace Session whose connector was built without the Managed Action
store silently fell back to the deployment approval mode and skipped the
Harness mode confirmation -- the case the design declares fail-closed.
createOrLoad now refuses it, so the two remaining null checks in the
connector are gone and an un-wired store fails loudly at admission.

requireOwner let an IllegalArgumentException from actorKey escape as an
unmapped 500. It now answers 403 actor_scope_mismatch, matching the
neighbouring workspace routes.

endedCode maps cancelled to action_cancelled but no test asserted it: a
cancelled Action now fails its response operation with that code and
returns no resolution.

The design docs still named the Actions migration V23 after it moved to
V24; V23 belongs to the Session MCP catalog (QwenLM#12946).

---------

Co-authored-by: 易良 <[email protected]>
Co-authored-by: yiliang114 <[email protected]>
neoneye pushed a commit to agent-memory-atlas-archive/QwenLM--qwen-code that referenced this pull request Oct 1, 2026
… panel (QwenLM#13107)

* feat(managed-agent): Serve durable permission Actions (D6b)

* test(managed-agent): Avoid concurrent Action mock stubbing

* feat(web-shell): show and answer Hosted tool approvals in the Managed panel

The Managed panel passed pendingApproval={null}, so a Hosted Session
waiting on a D6a approval showed nothing to answer. Build on the D6b
Actions contract (QwenLM#13101):

- The provider gains an optional Actions reader. The Java provider lists
  requested permission Actions through actions/query and answers through
  actions/respond with the Action's inputRevision and policyRevision and
  a per-Action, per-option idempotency key. The Session summary reports
  the actions capability only when the service does.
- action.updated projects to action_updated. It carries no Turn, so the
  transcript skips it instead of settling the Turn being streamed.
- useManagedActions re-reads pending approvals on action_updated, on a
  stream gap, after an expiry and after an answer. It hides an answered
  approval and brings it back if the answer fails.
- The page maps the pending Action onto the shared approval card, with
  allow/deny as allow_once/reject_once so labels are localized, and
  attaches it to the tool row `${turnId}:${functionCallId}` so the row
  expands and shows its arguments.

Not built, type-checked, linted or tested locally.

* feat(web-shell): export the pending Action type for custom Managed providers

* fix(managed-agent): renumber the Actions migration to V24 and bump the contract to 1.25

Merging main brought in QwenLM#12946's V23__managed_mcp_records.sql next to this
PR's V23__managed_actions.sql; Flyway refuses two migrations with the same
version, so the server would not start. The Actions migration becomes V24
with its content unchanged. The contract takes 1.25.0 on top of main's
1.24.0 (QwenLM#12946), with a v1.25 note for the Actions routes this PR serves.

* fix(web-shell): memoize pending actions to keep the respond callback stable

The conditional selected a fresh empty array on every render, so the
useCallback dependency changed identity each time and ESLint's
react-hooks/exhaustive-deps warning failed the lint gate.

* style(web-shell): format the Managed approvals page test with Prettier

* fix(web-shell): retry failed approval reads and name them as such

A transient failure of actions/query on the first read left a pending
approval invisible until a manual Refresh: summary polls do not re-read
Actions, and without an Action there is no expiry timer. The page then
showed "The approval answer could not be sent" although nothing was
sent.

- useManagedActions reports loadError and answerError separately, and
  retries a failed read after 2, 5 and 10 seconds before giving up.
  retry() reads again on demand and restarts the bound.
- The page shows "Pending approvals could not be loaded." with a Retry
  button for a read failure, and keeps the answer message for a failed
  answer.
- Tests cover the automatic recovery, the retry bound and the manual
  retry.

* fix(web-shell): recover failed approval answers and bound expiry reads

* fix(web-shell): Show available managed approval arguments

* fix(web-shell): keep the approval card when an answer was not applied

* fix(web-shell): keep Managed approvals stable across reloads and ended Actions

Treat an unknown session summary as unknown rather than as no Actions
capability, so a reload keeps the shown approval answerable. Drop an
approval the service reports as ended instead of offering a retry, scope
answer failures to the approval that failed, stop retrying client errors,
restart the retry ladder whenever reads resume, and carry the
arguments-unavailable notice inside the approval card.

* fix(web-shell): land only the Critical R1-1 fix for Managed approvals

The previous commit also carried the review's Suggestions. The PR is past
its review-round budget, so only R1-1 stays here: an unknown session
summary keeps the shown approval instead of reading as no Actions
capability. The Suggestions move to the follow-up recorded in QwenLM#12867.

* fix(web-shell): match Managed approvals to itemId-keyed tool rows

Since QwenLM#13037, Managed tool rows are keyed ${turnId}:${itemId} and Java tool
Items always carry an itemId, so findManagedApprovalTool no longer found the
row for a pending Action: the card could not show that row's arguments or
expand it. Keep the producer's call ID on the row and match on it, still
within the Action's Turn.

Candidate patch from wenshao's round-4 real-stack verification (R4-1).

* test(web-shell): pin the page half of the Managed approval reload fix

Reverting the page gate back to detail.summary?.capabilities.actions === true
kept every test green (G9). Hold a reload's Session summary and assert the
shown approval stays and no extra Actions read happens.

Candidate test from wenshao's round-3 real-stack verification.

* test(web-shell): spell canCancel in the Managed approvals fixtures

The three approvals fixtures handed to `summary()` omitted
`capabilities.canCancel`, which is a required member, so each one was a
TS2741 that no gate in this package reports (both tsconfigs exclude
`client/**/*.test.tsx` and vitest does not typecheck). They also modelled a
summary the real mapper cannot emit, which always sets both flags.

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-closeout/jmupe138l16

* fix(web-shell): keep the Harness call title on Managed approval cards

`toManagedPermissionRequest` reads `tool.title` for the approval
`title`, but no producer ever set it: `managedEventsToMessages` only
assigned args/status/times/rawOutput, so the left branch was dead and the
card description degenerated into a restatement of the heading while the
Harness`s own per-call title was dropped.

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-closeout/jmupe138l16

* fix(web-shell): stop Managed approval state outliving its Action

Two lifecycle leaks in `useManagedActions`:

- a successful re-read cleared `loadError` but never `answerError`, so
  "the answer could not be confirmed" stayed on screen after the Harness
  ended the Action, and then labelled the next, unrelated card. The
  warning now carries the Action it belongs to and is dropped once that
  Action is no longer pending.
- a withdrawn reader cleared `pending` without resetting the retry budget,
  so a reader restored after the ladder was exhausted got a single attempt
  with no retry scheduled, contradicting the hook`s own comment.

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-closeout/jmupe138l16

* fix(web-shell): stop retrying approval reads the service already answered

R1-5: a 4xx other than 408/429 is the service's answer, not a hiccup, so a
deleted Session no longer costs four guaranteed-failing /actions/query
requests and a Retry that can never succeed. The classification the submit
path already used moves to managed-request-error.ts and is reused, so both
failure paths of the feature agree.

R1-4: the hook reports whether the pending approvals were read for this
Session, so a failed background re-read is named as a refresh instead of
claiming nothing could be loaded next to the loaded, actionable card.

R1-8: the "arguments are unavailable" caveat renders as a sibling of the
approval dialog, so ToolApproval takes an extra description id and the panel
describes the caveat too — otherwise a screen-reader user confirms a Hosted
tool call hearing only the tool name.

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-closeout/jmupe138l16

* test(web-shell): pin four Managed approval paths that no test observed

R1-10, four sites the review mutation-tested as unpinned:

1. the `stream_gap` arm of the re-read trigger — the durable transcript is
   the only route that reports a gap here, since use-managed-session breaks
   the live loop on a gap before merging it; the hook doc now says so.
2. the session-switch reset effect and the cross-session mask — the new test
   switches Session while the previous card is displayed and its failed
   answer is still flagged, so deleting the effect leaves the warning behind
   and reverting the mask leaves the stale card answerable.
3. the `pendingApproval` wiring into MessageList — the page test's mock now
   captures the prop, which is what keeps the turn owning the pending call
   from being folded away.
4. the `setAnswerError(undefined)` clear on a successful answer — the retry
   test now ends on the alert being gone, not just on the card.

Every assertion was checked to fail with the production line it covers
reverted (8/8 mutations RED).

Co-authored-by: Qwen-Coder <[email protected]>
Patrol-Run: qwen-pr-closeout/jmupe138l16

* fix(web-shell): isolate late Managed approval replies

Ignore answer settlements once their Session or pending Action has changed,
and clear or display warnings only for the Action they belong to.

Pin late success and failure after Session switches and Action replacement.

* fix(web-shell): drop a Managed approval the service reports as ended

Answering an Action that already expired, was cancelled or was answered
elsewhere returns 409 with a code the contract defines as ended, and no
retry can succeed. Keep the card hidden, clear the warning and read the
list again instead of offering a retry. A read that still lists the Action
shows it again, so a wrong report cannot hide an approvable Action. The
403 actor_scope_mismatch is not an ended Action and keeps its handling.

---------

Co-authored-by: Shaojin Wen <[email protected]>
Co-authored-by: yiliang114 <[email protected]>
Co-authored-by: Shaojin Wen <[email protected]>
Co-authored-by: qwen-code-dev-bot <[email protected]>
Co-authored-by: yiliang114 <[email protected]>
Co-authored-by: Qwen-Coder <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants