Repository navigation
fix(mcp): Support larger Apps, scoped tool calls and isolated origins - #12258
Conversation
|
Local verification on
The real daemon-backed WebShell rendered the configured larger resource inside its MCP App sandbox; clicking the fixture button changed it to “Interaction verified”. After stopping and restarting the daemon, the saved transcript retained exactly 1,048,577 HTML bytes, rendered the App, and supported the same interaction. Three unedited screenshots are embedded in the PR body; the fixture and screenshots contain only synthetic test data. Full build + bundle, repository-wide typecheck, four focused suites (347 tests), ESLint, Prettier and diff checks passed. Tests also cover UTF-8/base64 limits, configured boundary/clamping, cancellation, pool isolation, metadata/session tool projections, and worst-case JSON escaping of the final 4 MiB ceiling. Independent review caught a replay-budget risk in the initial larger ceiling; the final ceiling leaves approximately 8 MiB of envelope headroom under the existing 32 MiB replay limit. Limit of this evidence: the real Amplitude endpoint was contacted but returned 中文最终提交已在 macOS arm64 / Node.js 24.19.0 上验证。6 个本地 CLI E2E 场景全部通过:默认小资源成功、默认超大资源拒绝、配置后超大资源成功、默认 11 秒读取超时、配置 15 秒后慢读取成功、配置 500 ms 后及时超时。 真实守护进程/WebShell 沙箱成功渲染同一份 1,048,577 字节 HTML,按钮交互成功。停止并重启守护进程后,保存的记录仍保留完整 HTML,回放展示和交互再次通过。PR 正文附三张未修改的实际浏览器截图,均为合成测试数据。 完整构建及 bundle、全仓类型检查、347 项相关测试、ESLint、Prettier 和 diff 检查均通过。独立审查发现的序列化回放预算问题已修正:最终 HTML 上限为 4 MiB,即使最坏 JSON 转义也为现有 32 MiB 回放限制预留约 8 MiB 空间。 真实 Amplitude 端点返回 |
|
Addressed the review's documentation finding: the pool-key comment now explicitly covers shared tool-snapshot settings and explains why App resource limits must remain in the fingerprint—they are stored during discovery and are not re-projected per session. This follow-up changes comments only. The pool-fragmentation tradeoff is now recorded in the PR's English and Chinese risk sections. Different configured policies require separate entries for correct isolation; redesigning the per-name budget is outside this resource-loading fix. The request and abort signal intentionally share the same resolved deadline, as already documented and covered by the short-timeout case. Validation: full repository build and typecheck, 27 pool-key tests, ESLint, Prettier, and a clean diff audit. Existing rendering/replay screenshots and E2E evidence remain applicable because executable code is unchanged. 已修复 review 指出的注释问题,说明 App 限制属于共享工具快照配置,因此必须参与连接池指纹。本次仅修改注释。连接池条目增加的权衡已补入 PR 中英文风险说明;超时信号与 SDK 使用同一截止时间的行为保持不变。全仓构建、类型检查、27 项连接池测试、Lint、格式检查和 diff 自查通过。 |
|
Verdict: merge-ready — 828/828 scripted assertions passed, 0 unexpected failures. Verified head: 中文摘要结论:可以合并(merge-ready)——828/828 条脚本断言全部通过,无意外失败。
Central claim and A/B proofClaim: each MCP server can opt into a larger App HTML limit and a longer App resource-read deadline via Method: a real stdio MCP fixture server (
Every rejection preserved the successful tool result ( Secondary claims
Reviewer Test Plan walk-through
Vacuity check (mutation A/B)Mutation:
Harness artifact worth noting for future rounds, not a product defect: two earlier mutation attempts appeared to "hang" — in fact vitest was rendering megabyte-scale assertion diffs on failed 1 MiB-string comparisons ( Gates
FindingsNone. No blocking or non-blocking findings from the exercised surface. Not covered
MethodologyLocal maintainer-driven round on Linux aarch64 (12-core, Node v24.14.0). Head and control were checked out as sibling worktrees ( Maintainer-driven local verification round (/verify-pr protocol). Harnesses, raw logs, and cell records preserved locally under |
|
@qwen-code /takeover from 2 |
|
🤝 Takeover engaged: the autofix loop now manages this PR — it will address new review feedback and resolve base conflicts until the label is removed or the round cap is reached. This window's round counter starts at 2 (the rounds this PR spent in review before takeover), so the Critical-only brake engages after 3 more change-producing round(s) instead of a full fresh 5. Remove the 中文说明🤝 已接管:autofix 循环现在管理此 PR —— 将持续处理新的评审反馈与 base 冲突,直到移除标签或达到轮次上限。本窗口轮次计数从 2 起算(即本 PR 托管前已进行的评审轮数),因此再经过 3 个产生改动的轮次即进入 Critical-only,而非重新计满 5 轮。移除 |
|
🤖 AutoFix updated a stale base — the fix did not pass verification, but this PR was behind What I found before stopping: See the Qwen Autofix agent step logs for model/tool output. 中文说明🤖 AutoFix 更新了一个过期的 base —— 修复未通过验证,但本 PR 落后于 Run log: https://github.com/QwenLM/qwen-code/actions/runs/35470296850 🧠 Handled by Qwen Code · model/模型 |
…2258) Address review on the App resource limits change: - WebShell historical-page admission now drops each MCP App display's html whole (never truncated) when a page exceeds its byte budget, so two ceiling-size documents degrade to fallbackText instead of failing the whole page as non-retryable unavailable; the design doc ceiling rationale now names all five budgets a retained document crosses. - Pool fingerprint hashes App resource limits at their enforced (clamped/floored/defaulted) values via a shared module, so configs with byte-identical enforced policy share one pool entry. - Limit warnings name the setting in the source that declares the server (extension manifest or .mcp.json) instead of always pointing at settings.json, whose same-named entry would shadow rather than merge. - Timeout warnings credit the key that produced the deadline (timeout when no explicit App timeout is configured). - A non-finite configured limit now logs a debug warning naming the server, key, rejected value, and fallback instead of being silently discarded. - Add a reconnect-replay test pinning that App resource limits survive rediscovery, and sync the daemon transport-pool doc with the fingerprint contract.
|
🤖 Addressed the latest review feedback (round 4/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 4/100 轮)。改动内容与我反驳保留之处如下: Autofix round — PR #12258 review feedbackCommit: Eight of the nine findings from the CHANGES_REQUESTED review are implemented and verified; the ninth (R1-2, carrying the two keys through the hand-enumerated config surfaces) is deferred to the next round with a reply on its thread — it is the widest change and this round's batch was already at its bound. Dispositions
Mutation probes
Verification
中文说明Autofix 本轮处理 — PR #12258 评审意见提交: CHANGES_REQUESTED 评审中的九条意见,本轮实现并验证了八条;第九条(R1-2,把两个配置键打通到各手工枚举的配置面)推迟到下一轮,并已在该意见的讨论串中回复说明——它是九条中涉及面最广的一条,而本轮批次已达到上限。 处理结论
变异探针
验证
🧭 Gate advisory — this round modified areas outside the PR footprint (machine-measured, not agent-authored):
Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧵 Resolved all 8 selected review thread(s). · 已关闭全部选中的 8 条评审线程。 Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。 🧠 Handled by Qwen Code · model/模型 |
|
Local follow-up completed at
Validation: full repository build, bundle and typecheck; 64 bridge + 354 core + 105 WebShell + 796 CLI + 105 SDK tests = 1,424 passed; SDK public-type compiler fence; ESLint and Prettier. New disconnect and configuration tests failed before the fixes and passed afterward. Independent verification exercised seven live-queue scenarios, five historical-page/navigation scenarios, SDK initialization and actual The screenshot uses synthetic data and shows a 4 MiB App with subsequent conversation text. The existing pinned historical-window limit can still require navigating to an adjacent turn and back; the real browser verified recovery, so that separate memory-protection behavior was preserved. Authenticated Amplitude rendering remains unverified because OAuth was not completed. The non-blocking timeout-warning attribution issue is deferred: when the default 10-second cap supplies the deadline, the current warning names the general timeout although raising it cannot override that cap; an explicit App timeout is required. This does not change execution or acceptance behavior. Following the repository's review-scope rule after roughly five rounds, this wording refinement is left out of the two correctness fixes. 本轮已修复剩余实时断连及配置丢字段问题,推送两个独立提交。全量构建、类型检查、1,424 项相关测试和独立本地验证通过;截图为真实 WebShell 上的合成数据。原有历史缓存窗口限制可通过邻近对话导航恢复,保持不变。真实 Amplitude 仍待 OAuth 验证。默认超时警告的归因措辞作为非阻塞改进留待后续,避免本轮继续扩大修复范围。 |
|
Authenticated Amplitude follow-up on
Only the screenshot is published. OAuth tokens stayed in memory; raw HTML and tool metadata stayed in a private local temporary directory. Temporary test servers were stopped. Documentation-only change: inspected diff and 已完成真实 OAuth 实测:2.23 MiB,79–89 秒。4 MiB 足以容纳本次资源,文档超时示例修正为 120 秒。真实 HTML 原样通过配置后的 core/CLI 并在 WebShell 运行,但图表仍因宿主缺少数据获取能力而失败;截图明确记录该剩余限制,未声称 Amplitude 完整兼容已修复。 |
|
Qwen Code resolved the merge conflicts and pushed the branch update. Merge resolution — PR #12258 ← origin/mainOne conflict, one file: Root causeMain picked up #12539 ( Textual, not semanticMain changed prose; the PR changed the code that prose described. Delegation is behaviour-preserving because async discoverToolsForServer(serverName: string, reconnect = false) {
this.removeMcpToolsByServer(serverName);
// prompts + resources removal unchanged
await this.mcpClientManager.discoverMcpToolsForServer(
serverName, this.config, reconnect);
}That sentence being main's whole edit, I carried it to where the deletion now lives — appended to the existing comment in Load-bearing
Not verifiedNo build, typecheck, lint or tests run here. 中文说明唯一冲突文件 根因:main 合入 #12539( 属于文本冲突:委托调用等价—— 关键约束:两个移除方法必须继续不清理 未验证:未跑构建、类型检查、lint 或测试。 |
|
@qwen-code /takeover from 10 |
|
🤝 Takeover engaged: the autofix loop now manages this PR — it will address new review feedback and resolve base conflicts until the label is removed or the round cap is reached. This window's round counter starts at 10 (the rounds this PR spent in review before takeover), so the Critical-only brake engages after 0 more change-producing round(s) instead of a full fresh 5. Remove the 中文说明🤝 已接管:autofix 循环现在管理此 PR —— 将持续处理新的评审反馈与 base 冲突,直到移除标签或达到轮次上限。本窗口轮次计数从 10 起算(即本 PR 托管前已进行的评审轮数),因此再经过 0 个产生改动的轮次即进入 Critical-only,而非重新计满 5 轮。移除 |
|
🔀 Base updated: red check(s) [TUI parity snapshots (ink vs opentui), Test (ubuntu-latest, Node 22.x), OpenTUI no-flicker gate, Lint & Static (ubuntu-latest, Node 22.x), Integration Tests (no-AK, No Sandbox), web-shell E2E Smoke (ubuntu-latest, Node 22.x)] pass on current main — merged current main via update-branch; CI will re-run. 中文说明🔀 已更新 base:红色检查 [TUI parity snapshots (ink vs opentui), Test (ubuntu-latest, Node 22.x), OpenTUI no-flicker gate, Lint & Static (ubuntu-latest, Node 22.x), Integration Tests (no-AK, No Sandbox), web-shell E2E Smoke (ubuntu-latest, Node 22.x)] 在当前 main 上通过 —— 已通过 update-branch 合入当前 main,CI 将重新运行。 |
The main merge left sourceSessionId twice on TranscriptViewport, which fails the web-shell typecheck and stops every job at install. Co-authored-by: Cursor <[email protected]>
|
R11 follow-up on cab293b: 0d6eb9e fixes the two standing R8 findings, with no new feature scope.
Validation: frozen pnpm installation, full build and bundle, full workspace/integration typechecks, focused lint/format checks, 172 core tests and 38 App component tests passed. Three independent DOM probes passed, including a genuinely pending request cancelled at the component/host-adapter boundary. Minimal temporary mutations proved the new checks fail when the timeout attribution branch or input/result teardown is removed, and when the timeout comparison is inverted. These are core/component checks, not a new vendor or remote browser acceptance run. R1-1 was rechecked against the exact tree. Replacing a very short HTML payload with a fallback can increase bytes, but the page accounts for the signed byte difference and rejects an over-budget result. Both replay-envelope callers revalidate the serialized envelope after their bounded degradation pass. I found no confirmed budget bypass or unbounded retry; this does not claim every oversized page can be recovered. The truncated unresolved review note provides no additional trigger to reproduce. The maintainer-reproduced cross-tab HTTP/1.1 wedge remains open, as previously recorded; this commit does not solve it. The event-stream delivery proposal still needs the stated maintainer decision about pre-merge work versus a follow-up. The existing non-critical suggestions remain deferred under the convergence rule, and earlier Amplitude/Tableau evidence retains its original scope. The core change in this round is limited to diagnostic attribution; no timeout or tool-execution policy changed. |
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: Autofix review round — no change (verification-only)Head verified: Critical-only mode is active for this window (seeded at round 10 by takeover). This round re-verified every actionable Critical finding against the exact tree and found all of them fixed at HEAD. The failed-checks list is empty. The one remaining functional item is an explicit maintainer decision, which this round deliberately does not settle. Standing Critical findings — verified fixed at HEAD
R1-1 (aggregate App-HTML budget) — traced at HEAD, no confirmed bypassThe standing "unresolved, please confirm" item was re-traced leg by leg against the exact tree:
No silent over-budget admission, no unbounded loop, and no empty-content frame was found in any leg; a document that cannot be degraded enough fails with an explicit, visible error. This matches the author's R11 recheck. The two related fails-closed hardening notes from rounds 6/8 (newest-segment eviction order; non-shrinking loop exit) sit in this window's deferred audit record and were not touched. Maintainer items
Deferred non-Critical feedbackThe bot Suggestions listed in this window's "Deferred non-Critical feedback" section are an audit record; per the Critical-only brake they were not worked, resolved, or replied to. Verification
Total: 1,169 focused tests passed at the unchanged head. No source or test file was modified; 中文说明Autofix 审查轮次 —— 无改动(仅验证)验证的 head: 本窗口处于仅 Critical 模式(由 takeover 从第 10 轮起算)。本轮对当前代码逐项复核了所有可执行的 Critical 发现,确认它们在 HEAD 均已修复。失败检查列表为空。唯一剩余的功能性事项是维护者明确需要自己决定的问题,本轮有意不代为定夺。 未决 Critical 发现 —— 已在 HEAD 验证修复
R1-1(App HTML 聚合预算)—— 已在 HEAD 逐环节追踪,未确认存在绕过对仍标记「未解决,请确认」的条目按当前代码逐环节复核:
各环节均未发现静默放行超预算、无界循环或空内容帧;无法充分降级的文档会以显式可见的错误失败。这与作者第 11 轮的复核结论一致。第 6/8 轮留下的两条相关 fails-closed 加固记录(最新段的驱逐顺序、不收缩时的循环退出)属于本窗口的延后审计记录,本轮未触碰。 维护者事项
延后的非 Critical 反馈本窗口「Deferred non-Critical feedback」一节列出的机器人 Suggestion 属于审计记录;按仅 Critical 的收敛规则,本轮未处理、未解决、未回复这些条目。 验证
合计:在未改动的 head 上 1,169 项聚焦测试全部通过。未修改任何源码或测试文件; Deferred non-Critical feedbackCritical-only mode is active: the round counter reached 5 (this window was seeded at round 10 by
中文说明已进入仅处理 Critical 的模式:轮次计数已达 5(本窗口由 Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧵 Resolved all 17 selected review thread(s). · 已关闭全部选中的 17 条评审线程。 🧠 Handled by Qwen Code · model/模型 |
Maintainer verification, round 7 (head
|
| Server config | Actual deadline | before (warning names) | after (warning names) |
|---|---|---|---|
timeout unset (600 000 ms default) |
10 000 ms | mcpServers.fxdefault.timeout ✗ |
mcpServers.fxdefault.appResourceTimeoutMs ✓ |
timeout: 120000 (operator already raised it) |
10 000 ms | mcpServers.fxgen.timeout ✗ (same advice again) |
mcpServers.fxgen.appResourceTimeoutMs ✓ |
timeout: 5000 (control) |
5 000 ms | …fxshort.timeout |
…fxshort.timeout (same) |
appResourceTimeoutMs: 5000 (control) |
5 000 ms | …fxappshort.appResourceTimeoutMs |
same |
appResourceTimeoutMs: 20000 (the new advice) |
20 000 ms | App renders, read done at 12.0 s | App renders, read done at 12.0 s |
The fixture's ledger shows the server-side abort in each case. Reads aborted at 10 000–10 003 ms or 5 001 ms, and both builds give the same numbers. So only the text of the warning changed.
2. R8-3: a failed handoff now tears the App down (fault injection)
Can this happen naturally? The host sends sandbox-resource-ready, tool-input and tool-result as one-way postMessage notifications, and their payloads are JSON. A handler that throws inside the App cannot reject them on the host side. So I found no way to make them fail in a real run.
Injection: I patched the served WebShell bundle so that the host's PostMessageTransport.send throws for ui/notifications/tool-result. This is the "transport errors" trigger named in R8-3. Both builds got exactly the same patch.
| control (after, no injection) | before + injection | after + injection | |
|---|---|---|---|
| Card | App renders | text fallback, iframe display:none, src kept |
text fallback, iframe src removed |
| App document | alive | still alive (__ready=true, vendor frame loaded) |
about:blank, gone |
App's own stable_app_tool call after 5 s |
approval shown | approval shown for an App the user was told had failed | no approval within 15 s |
| Server executions after approval | 1 | 1 | 0 (0 /mcp-app/tools/call requests) |
Not checked in a browser: the "stale failure after a remount" case. The unit tests cover it, and mutation M3 below confirms the test catches it.
3. R8-2 and the historical page
After the App turn, I added 12 turns of 900 KB each and reloaded. The App turn fell outside the live window (0 live cards). Opening turn 2 from the session timeline gave the historical viewport. There:
- the App rendered;
- the host advertised
serverTools; stable_app_toolasked for approval and ran once;- the historical page showed 0 feedback buttons.
Round 4 could only check this path with unit tests.
4. Round 4–6 cases on the new head
| Case | Round 6 (e6c7a795d0) |
Round 7 (45c260aad1) |
|---|---|---|
| 5 concurrent App calls in one tab, one approval every ~8 s | 5/5 | 5/5. First result at 3 s, last at 36 s. |
| Kill the MCP server, then make App calls | 1st fails, restart once, then ok | Same. The 1st fails (MCP App tool call failed.). The server restarts once. The next stable and token calls succeed. stable_app_tool ran 3× for 3 successes, so nothing was replayed. |
| App-only tools in the next model turn | excluded | Excluded. |
Canary SECRET-FIXTURE outside the App |
none | None. 0 hits in model requests, session files, runtime and daemon logs. The only hit is the fixture's own --secret argument in settings.json. |
| Isolation probes and attacks | blocked | Blocked. The App gets a unique <uuid>.localhost origin. It cannot read the top document or storage. Daemon API and WebSocket requests fail. A consumed sandbox URL returns 404. Top-level navigation is blocked, and popups return null. |
Forwarded single port (--block-isolated) |
not re-run | Chromium and WebKit: switch to mode=data after about 10 s. The App renders, the token call works, and the vendor frame keeps its own origin. |
| 3 tabs × 1 App call | 0/3, 300 s | 0/3, 300 s. 183 approval clicks, 0 executions. At 6 s, 3 session event streams and 3 POST …/mcp-app/tools/call are still open. |
5. Tests and mutation
- Focused suites on head: 861/861 passing.
- core
mcp-tool,tool-registry: 251 - web-shell
McpApp: 38 - web-shell
TranscriptViewport.mcp-app,ChatPane: 175 - CLI
mcp-app-sandbox,history-replay-page: 48 - acp-bridge
eventBus,compactionEngine,bridgeClient: 349
- core
- CI on
45c260aad1is green. - Mutations on
0d6eb9e2db's code: 6 of 10 were killed.
| Mutation | Result |
|---|---|
| M1: drop the ceiling branch | killed (4 tests) |
| M2: invert the ceiling check | killed (6) |
M3: remove the generation guard in failInitialization |
killed (1) |
M4: keep oncalltool after a failure |
killed (3) |
M6: tool-input/result catch goes back to setError |
killed (3) |
M8: sandbox-resource catch goes back to setError |
killed (1) |
M5: no clearTimeout in failInitialization |
survived; practically equivalent, because the timer callbacks already check active |
M7: no active check before sendToolResult |
survived; practically equivalent, because a late send only reaches the guarded failInitialization |
M9: connect() catch goes back to setError |
survived; test gap |
M10: normal cleanup keeps oncalltool |
survived; test gap |
M9 and M10 are cheap follow-up tests. Neither blocks the merge: a connect() failure is rare, and cleanup also aborts and closes the bridge.
Not covered this round
- No real vendor run (Tableau Cloud or Amplitude) and no remote HTTPS run.
- WebKit was only re-run for the data-mode case.
- Triage F-2/F-3 and the live-vs-replay App-call count mismatch were not re-run. Their code is unchanged since round 6, and they remain follow-ups.
Merge reference
- Nothing new blocks the merge. R8-1, R8-2 and R8-3 are fixed as described. The merges of main changed none of the PR's behavior.
- The open decision is the same as rounds 5–6: the cross-tab wedge. One option is to merge and track event-stream delivery (round-4 option (a)) as a follow-up. The other is to require it before merge.
- Cheap follow-ups: tests for M9/M10, the F-3 replacer function and the F-2 refuse-when-full rule.
- The review decision is still
CHANGES_REQUESTED, from the bot's earlier review.
Earlier rounds: round 5, round 6.
中文版
维护者验证,第 7 轮(head 45c260aad1)
结论:两个新修复在真实浏览器里都按描述生效,也没有回归。没有新的阻塞项。唯一的开放决策仍是第 5–6 轮的跨标签页卡死。
- R8-1 已修复。 超时警告原来让人去改
mcpServers.<srv>.timeout,但这个设置抬不高 App 的期限。现在改为指向appResourceTimeoutMs,照着新提示设置后 App 能渲染出来。期限本身没有变化。 - R8-3 已修复。 我注入了一次
tool-result交接失败。在没有修复的构建里,卡片显示了回退文本,但隐藏的 App 仍然存活:它自己发起的工具调用弹出了审批,并在服务器上执行了。有修复时,iframe 被卸载,不出审批,服务器执行 0 次。我没有找到自然触发这个失败的方式,所以请把它看作防御性修复。 - R8-2 已修复。 构建通过。另外首次在真实浏览器里验证了从较早轮次发起的 App 调用,在
historical页面上正常工作。 - 第 4–6 轮结论仍成立。 5 个并发 App 调用 5/5;MCP 崩溃后重启 1 次、不重放;金丝雀密钥没有泄露;data 模式回退在 Chromium 和 WebKit 上都正常。
- 仍开放,没有变化: 跨标签页卡死(3 个标签页各 1 个 App 调用:执行 0/3,300 秒超时)。triage F-2/F-3 也没变,因为
mcp-app-sandbox.ts自第 6 轮以来没有改动。
自第 6 轮以来的变化
第 6 轮测的是 e6c7a795d0。此后有两个修复提交,以及三次合并 main:0dcbf8b12a、90b0bbd41b、45c260aad1。
- 只由本 PR 改动的文件有 52 个。 其中的改动只有
0d6eb9e2db的 4 个文件:mcp-tool.ts、McpApp.tsx以及各自的测试。 - main 也改过的文件有 21 个。 我在忽略空白的前提下,对比了这些文件里 PR 自己的 ± 行在第 6 轮和现在的差别。只剩两处:
tool-registry.ts(/resolve解决的冲突):只是一段注释挪了位置。discoverToolsForServer现在调用removeMcpToolsByServer。这个函数做的删除和 main 的内联循环相同,另外还会清理mcpAppTools。main fix(core): keep the deferred-tool bridge halves on the same tool #12539 的规则保持不变:清掉 reveal 状态,不动 reviewed-declaration 记录。ChatPane.tsx:main 的 feat(web-shell): support selective host artifact integration #12588 加了和 PR 相同的sourceSessionId={connection.sessionId}属性。合并90b0bbd41b留下了两份,导致 typecheck 失败(TS17001)。cab293b3dd删掉了其中一份。两份的值相同,行为没有变化。
环境
- 同一棵树的两个构建(
pnpm install --frozen-lockfile、pnpm run build、pnpm run bundle):- 修复后: head
45c260aad1。 - 修复前: 同一个 head,只回退
0d6eb9e2db的两个源文件,所以两棵源码树只在这两个文件上不同。bundle 里 core 的上限分支出现 1 次对 0 次,web-shell 的oncalltool=void 0出现 2 次对 0 次。
- 修复后: head
- 浏览器和链路: macOS arm64 上的 Chromium 149(Playwright 1.61.1)。浏览器 → 打包的 WebShell → 127.0.0.1 上的
dist/cli.js serve→ ACP 子进程 → 基于官方@modelcontextprotocol/ext-appsApp SDK 的真实 stdio MCP 服务器。 - 脚本化的部分: 只有模型选哪个工具。每次审批都是在审批面板里真实点击。
- 装置: 沿用第 5 轮的装置,新增两个 App 变体:
- 慢资源:
resources/read耗时 12 秒; - 自调用 App:连接 5 秒后,自己调用 App 可见工具
stable_app_tool。
- 慢资源:
- 脚本、每次运行的
result.json和变异结果都在pr-12258-round7/。
1. R8-1:超时警告现在指向能改变期限的设置
每个服务器都指向同一个慢资源,只有配置不同。
| 服务器配置 | 实际期限 | 修复前(警告指向) | 修复后(警告指向) |
|---|---|---|---|
未设 timeout(默认 600 000 ms) |
10 000 ms | mcpServers.fxdefault.timeout ✗ |
mcpServers.fxdefault.appResourceTimeoutMs ✓ |
timeout: 120000(操作者已经调高过) |
10 000 ms | mcpServers.fxgen.timeout ✗(重复同一条建议) |
mcpServers.fxgen.appResourceTimeoutMs ✓ |
timeout: 5000(对照) |
5 000 ms | …fxshort.timeout |
…fxshort.timeout(相同) |
appResourceTimeoutMs: 5000(对照) |
5 000 ms | …fxappshort.appResourceTimeoutMs |
相同 |
appResourceTimeoutMs: 20000(按新建议) |
20 000 ms | App 渲染,12.0 秒读完 | App 渲染,12.0 秒读完 |
fixture 的台账记下了每种情况下服务器端的中止时间:10 000–10 003 ms 或 5 001 ms,两个构建的数字相同。所以变化的只有警告的文字。
2. R8-3:交接失败后 App 现在会被拆除(故障注入)
这种失败会自然发生吗? 宿主发送的 sandbox-resource-ready、tool-input、tool-result 都是单向 postMessage 通知,载荷是 JSON。App 内部的处理函数抛异常,也不会让宿主这边的发送失败。所以我在真实运行里没有找到让它们失败的方法。
注入方式: 我修改了服务端下发的 WebShell bundle,让宿主的 PostMessageTransport.send 在发送 ui/notifications/tool-result 时抛错。这正是 R8-3 里说的「传输出错」触发条件。两个构建打的是完全相同的补丁。
| 对照(修复后,不注入) | 修复前 + 注入 | 修复后 + 注入 | |
|---|---|---|---|
| 卡片 | App 渲染 | 文本回退,iframe display:none,src 保留 |
文本回退,iframe src 被移除 |
| App 文档 | 存活 | 仍然存活(__ready=true,vendor 帧已加载) |
about:blank,已卸载 |
App 在 5 秒后自己调用 stable_app_tool |
弹出审批 | 为一个已告知用户「失败」的 App 弹出审批 | 15 秒内没有审批 |
| 批准后服务器执行次数 | 1 | 1 | 0(0 次 /mcp-app/tools/call 请求) |
没有在浏览器里验证的: 「重新挂载后旧 mount 迟到的失败」这个场景。单测覆盖了它,下面的变异 M3 也确认测试能抓住。
3. R8-2 与历史页
App 轮次之后,我又加了 12 轮、每轮 900 KB 的回复,然后刷新页面。App 轮次落到了 live 窗口之外(live 卡片 0 张)。从会话时间线打开第 2 轮,进入了 historical 视图。在这个视图里:
- App 正常渲染;
- 宿主通告了
serverTools; stable_app_tool弹出审批,并执行了 1 次;- 历史页上的反馈按钮为 0。
第 4 轮时这条路径只能靠单测验证。
4. 在新 head 上重跑第 4–6 轮的用例
| 用例 | 第 6 轮(e6c7a795d0) |
第 7 轮(45c260aad1) |
|---|---|---|
| 单标签页 5 个并发 App 调用,约每 8 秒批准一次 | 5/5 | 5/5。 第一个结果在 3 秒,最后一个在 36 秒。 |
| 杀掉 MCP 服务器后再发 App 调用 | 第 1 次失败,重启一次,之后正常 | 相同。 第 1 次失败(MCP App tool call failed.),服务器重启一次,之后的 stable 和 token 调用都成功。stable_app_tool 成功 3 次、执行 3 次,没有重放。 |
| 下一个模型轮次里的 App 专用工具 | 排除 | 排除。 |
App 之外出现金丝雀 SECRET-FIXTURE |
无 | 无。 模型请求、会话文件、runtime 和 daemon 日志里都是 0 次。唯一命中的是 settings.json 里 fixture 自己的 --secret 参数。 |
| 隔离探针与攻击 | 被拦截 | 被拦截。 App 拿到唯一的 <uuid>.localhost origin,读不到顶层文档和存储;daemon API 和 WebSocket 请求失败;已用过的沙箱 URL 返回 404;顶层导航被拦截,弹窗返回 null。 |
单端口转发(--block-isolated) |
未重跑 | Chromium 和 WebKit: 约 10 秒后切换到 mode=data。App 渲染,token 调用正常,vendor 帧保留自己的 origin。 |
| 3 个标签页 × 1 个 App 调用 | 0/3,300 秒 | 0/3,300 秒。 点了 183 次审批,执行 0 次。第 6 秒时仍有 3 条会话事件流和 3 个 POST …/mcp-app/tools/call 没有结束。 |
5. 测试与变异
- head 上的定向测试套件:861/861 通过。
- core
mcp-tool、tool-registry:251 - web-shell
McpApp:38 - web-shell
TranscriptViewport.mcp-app、ChatPane:175 - CLI
mcp-app-sandbox、history-replay-page:48 - acp-bridge
eventBus、compactionEngine、bridgeClient:349
- core
45c260aad1上的 CI 全绿。- 针对
0d6eb9e2db代码的变异:10 个杀掉 6 个。
| 变异 | 结果 |
|---|---|
| M1:删掉上限分支 | 被杀(4 个测试) |
| M2:把上限判断取反 | 被杀(6) |
M3:去掉 failInitialization 里的 generation 守卫 |
被杀(1) |
M4:失败后保留 oncalltool |
被杀(3) |
M6:tool-input/result 的 catch 改回 setError |
被杀(3) |
M8:sandbox-resource 的 catch 改回 setError |
被杀(1) |
M5:failInitialization 里不 clearTimeout |
存活;实际等价,因为定时器回调本来就检查 active |
M7:sendToolResult 前不检查 active |
存活;实际等价,因为迟到的发送只会进入带守卫的 failInitialization |
M9:connect() 的 catch 改回 setError |
存活;测试盲区 |
M10:正常清理时保留 oncalltool |
存活;测试盲区 |
M9 和 M10 可以作为低成本的后续测试补上。两者都不阻塞合并:connect() 失败很少见,而且清理时也会中止并关闭桥接。
本轮未覆盖
- 没有跑真实 vendor(Tableau Cloud 或 Amplitude),也没有跑远端 HTTPS。
- WebKit 只重跑了 data 模式这一项。
- triage F-2/F-3,以及 App 调用计数在实时视图和回放里不一致的问题,本轮没有重跑。它们的代码自第 6 轮以来没有改动,仍作为后续项。
合并参考
- 没有新的阻塞项。 R8-1、R8-2、R8-3 都按描述修好了。几次合并 main 没有改变 PR 的任何行为。
- 开放决策与第 5–6 轮相同: 跨标签页卡死。一种选择是先合入,把事件流交付(第 4 轮方案 (a))作为后续跟踪;另一种是要求合入前完成。
- 低成本后续项: M9/M10 的测试、F-3 改用替换函数、F-2 在表满时拒绝新注册。
- 评审决定仍是
CHANGES_REQUESTED,来自机器人较早的一次评审。
|
@qwen-code /triage |
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: Autofix round: no code changeHead reviewed: What the actionable feedback containsThe only item in this round's actionable sections is @wenshao's round-7 maintainer verification (issue-level comment
There are no inline comments, no review bodies, and no failed or still-red checks in this round's actionable sections. Why no code change was made
Open decision that needs a maintainer (unchanged from rounds 5–6)The cross-tab wedge: 3 tabs × 1 App call → 0/3 executed, 300 s timeout; at 6 s, 3 session event streams and 3
This choice is a scope/product judgment, not an autonomous one: option (a) is a daemon event-stream delivery change in core infrastructure, and implementing it unilaterally in a critical-only round would contradict both the brake and the maintainer's framing. The recommendation from the evidence at hand is Option A — every verified behavior on the head passes, the wedge has a narrow multi-tab trigger, and the maintainer's own round-7 verdict records no new blocker — but the call belongs to the maintainer. It is left open here so it reads as an explicit question, not a verdict this round already reached. State checks performed
中文说明Autofix 本轮:无代码改动评审的 head: 本轮可执行反馈的内容本轮可执行区域里唯一的条目是 @wenshao 的第 7 轮维护者验证(issue 级评论
本轮可执行区域中没有内联评论、没有评审正文,也没有失败或持续失败的检查。 本轮不做代码改动的原因
仍需维护者决策的开放问题(与第 5–6 轮相同)跨标签页卡死: 3 个标签页各 1 个 App 调用 → 执行 0/3,300 秒超时;第 6 秒时仍有 3 条会话事件流和 3 个
这个选择属于范围/产品判断,不应由自动化流程自行决定:方案 (a) 是核心基础设施中 daemon 事件流交付的改动,在仅 Critical 的轮次里单方面实施它,既违背刹车的目的,也违背维护者的表述。就现有证据而言的倾向性建议是方案 A——当前 head 上所有已验证行为均通过,卡死的触发条件局限于多标签页场景,且维护者第 7 轮的结论记录了无新阻塞项——但最终决定权在维护者。此处将其保持开放,是为了让维护者读到一个明确的问题,而不是本轮已经擅自得出的结论。 本轮执行的状态检查
Deferred non-Critical feedbackCritical-only mode is active: the round counter reached 5 (this window was seeded at round 10 by
中文说明已进入仅处理 Critical 的模式:轮次计数已达 5(本窗口由 Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧠 Handled by Qwen Code · model/模型 |
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: Autofix round — no action takenNo changes were made in this round. There was nothing actionable to address:
To resume automated handling of the deferred non-Critical feedback, a maintainer can start a fresh counting window by commenting VerificationNo code changes were made, so no build, typecheck, lint, or test runs were required. The round's inputs were verified as follows:
中文说明Autofix 本轮 —— 未做任何改动本轮未进行任何修改,因为没有需要处理的可执行事项:
如需恢复对延后的非 Critical 反馈的自动处理,维护者可在 PR 上评论 验证由于未做任何代码改动,无需运行构建、类型检查、lint 或测试。本轮输入的核实方式如下:
Deferred non-Critical feedbackCritical-only mode is active: the round counter reached 5 (this window was seeded at round 10 by
中文说明已进入仅处理 Critical 的模式:轮次计数已达 5(本窗口由 Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧠 Handled by Qwen Code · model/模型 |
|
@qwen-code /triage |
qqqys
left a comment
There was a problem hiding this comment.
Independent Critical-only review — head 45c260aad
Not approving, on budget rather than on a finding I can prove. This is a 73-file, +5839/-428 change whose production surface the repository's own Stage 0 measured at 1833 lines across five packages, 394 of them in packages/core/src. An independent Critical-only pass over that surface, to the confidence bar this channel holds itself to, does not fit in the window I have; publishing an approval I cannot back with evidence would be worse than publishing this.
What I did confirm: R1-1's terminal leg is closed at this head
The last automated round at this commit carried R1-1 forward as "Unresolved, please confirm" — but it says so only because its verifier wave exceeded its wall clock (22 of 54 agents returned, so none of that round's 30 findings carries a verdict). The finding itself is worth separating from the round's inability to rule on it, because R1-1 named the restore-envelope leg as terminal: a recorded session whose replay page held eight ceiling-size App documents could not be loaded at all, history-replay-page.ts rethrew instead of degrading, acpAgent.ts mapped that to RequestError(-32012, transcript_page_too_large), and the message named neither the App documents nor appResourceMaxBytes.
That leg is closed in the code at this head. history-replay-page.ts now carries blankReplayedMcpAppHtml (copies a replayed update with its MCP App html blanked, keeping fallbackText so the client renders text in place of the App) and degradeReplayEnvelopeAppHtml (blanks oldest-first until the serialized envelope fits), and the page-assembly overflow pass does the same before it will throw — lines 293-316 blank App html oldest-first across both the retained and the delivered update, and HistoryReplayLimitError at line 321 is reached only for a page that still does not fit after every App document is degraded. The file's own comment states the invariant: it "cannot fail closed on a page whose App documents degrade cleanly". So the unit of loss is no longer the whole page for the life of the session.
I did not re-trace the other legs of R1-1 (the compacted replay window, which the finding itself already downgraded to a degrade rather than data loss because fullTranscriptAvailable: true and the frozen pagination anchor let a client re-page), nor the 16 historical Critical threads that prior posted rounds ruled fixed, nor the four other blockers the last round re-traced without publishing a verdict. All 21 Critical threads are marked resolved and all 33 still-unresolved threads are [Suggestion]; a resolved flag is not proof of a fix, and I am not treating it as one — I am saying plainly that I verified one leg myself and inherited the rest.
The one current item I verified in code: the App-call limiter is per-tab
McpApp.tsx bounds concurrent App tool calls with module scope:
// Share the limit across cards so App requests leave HTTP/1.1 connections for
// session events and approval requests. Hold the slot until the request settles.
let activeAppToolCalls = 0;
const waitingAppToolCalls = new Set<() => void>();The mechanism is sound for what a module-scoped counter can do — the slot is acquired against an AbortSignal, held until the request settles, and handed to the next waiter on release. But a module-scoped counter bounds one JavaScript realm, which is one browser tab. The stated purpose is to leave HTTP/1.1 connections for session events and approval requests, and the browser's per-origin connection budget is shared across every tab of that origin, not per tab. N tabs each admitting up to the per-tab limit can therefore consume the very budget the limiter exists to protect, and the sessions and approvals it is protecting are the ones in the other tabs.
I am reporting this as an open item rather than as a Critical I have proven, because whether it is a defect depends on a decision the diff cannot make: whether the aggregate belongs on the daemon side (where the POST /session/:id/mcp-app/tools/call route already sits, and where a cross-client bound is enforceable) or whether per-tab is the intended contract and the comment should say so. The last automated round reached the same root cause at the same line and called it a product decision. Either answer is defensible; what is not defensible is a comment promising a guarantee the scope cannot deliver.
CI
Green at this head across everything that ran, which matters here because R6-1 — the route-catalog drift Critical — was originally evidenced by a failing unit suite: Test (ubuntu-latest, Node 22.x) passes in 22m18s, Lint & Static in 12m1s, and so do web-shell E2E Smoke, Capture web-shell visuals, Serve A/B, Integration Tests (no-AK, No Sandbox), the TUI parity and no-flicker gates, Live Host (macos-latest), the full Java matrix including Real daemon E2E and Hosted no-tool processes, and review-pr. No failure is attributable to this PR. Note for whoever merges: the CHANGES_REQUESTED decision still showing on this PR is carried by a review at the superseded cab293b3d and does not reflect this head; a maintainer has approved at 45c260aad.
What would let the next pass close
Two things, both cheap for the author and neither a request for new work:
- A statement of where the cross-tab App-call bound is intended to live — daemon-side, per-tab by contract, or a follow-up issue — so the limiter's comment and its scope agree.
- If the remaining R1-1 legs and the 16 inherited Critical rulings have already been re-traced against
45c260aadsomewhere I have not read, a pointer to that; the last round explicitly declined to re-trace them under its own budget.
Verdict: COMMENT — No Critical proven at this head, and one historical terminal leg (R1-1's restore envelope) positively confirmed fixed in code. What blocks an approval is that the independent Critical-only pass over 1833 production lines across five packages could not be completed in budget, and that the App-call limiter's per-tab scope does not deliver the cross-tab guarantee its own comment states.
…gate Both sides edited the same `if (webShellDir)` block in createServeApp: this branch passes the desktop-relay CSP flag to mountWebShellAssets, while main (QwenLM#12258) rewired mountMcpAppSandbox to take an origin-allow callback and return a disposer that run-qwen-serve stops on shutdown. The two changes are adjacent, not overlapping, so keep both: the 4-arg mountWebShellAssets call and main's sandbox wiring with its app.locals capture.
|
R12 recheck: this PR was merged at 2026-09-27 04:56:49 UTC as 81c260b. The final PR head is 45c260a. I checked the latest reviews and verification reports, the final CI rollup, and the four files from 0d6eb9e against the final head. Those four files are unchanged. CI has 33 successful checks, 165 skipped checks and two cancelled superseded route jobs, with no failed or pending check. R8-1 and R8-3 have explicit fixed rulings; the maintainer's round-7 report independently confirms both fixes in a real browser, including that failed Apps no longer execute tool calls. The latest sandbox round reports 105/105 executed assertions and no new blocking finding. No code change or new test run was needed in this recheck. This does not close the known cross-tab HTTP/1.1 wedge: the current limiter is per page, and cross-tab-safe delivery remains follow-up work. The event-stream proposal should preserve client-targeted private App results and cancellation. The previously deferred Low findings F-1/F-2/F-3/F-5 also remain open; merge and green CI do not mean those have been fixed. Earlier vendor/remote evidence retains its original scope. |
When no `appResourceTimeoutMs` is configured, the App resource read deadline
derives from the server `timeout` through
`boundedAppLimit(mcpTimeout, DEFAULT, 1, DEFAULT)`, so a `timeout` at or above
`MCP_APP_RESOURCE_TIMEOUT_DEFAULT_MS` (10 s) yields exactly 10 s. The timeout
warning named only `appResourceTimeoutMs` in that case, which hides both which
key produced the limit and why raising `timeout` changed nothing, and points at
a key the operator never set. `docs/users/features/mcp.md` already documents the
real model ("the deadline remains the smaller of the general `timeout` and
10,000 ms"), so the diagnostic contradicted the repository's own docs.
The warning now names `timeout` as the source plus the cap that pinned it and
the one key that can go past it. Below the cap nothing changes: the cap is not
binding there, raising `timeout` does lift the deadline, and naming the cap
would send the operator the wrong way. An explicit `appResourceTimeoutMs` still
owns the deadline and is still named alone.
Deferred review finding ic:5746879970 from QwenLM#12258.
Authored by Mycroft, the synthetic co-founder at Anton Dzyatkovsky's lab
(autonomous mode; named responsible person: Anton Dziatkovskii). The test runs
above were independently re-executed before submission.
Assisted-by: Claude Code / claude-opus-5
Machine: MacBook-Anton
Account: tonydzi
Operator: Anton Dziatkovskii
Signed-off-by: tonydzi <[email protected]>







Important
Remote HTTPS renderer verified at
d6d532eb8b:d6d532eb8b's unmodified production App component and sandbox handler were deployed in an isolated fixture on the actual remote instance, reached through its existing authenticated HTTPS gateway and port proxy. The public Tableau Superstore chart rendered, West filtering worked, and page reload followed by filtering passed. Report and screenshots. Authentication/tool results were mocked; the existing remote Qwen service was not upgraded. Full remote daemon/session integration and authenticated Tableau Cloud acceptance remain unverified. Subsequent history-binding, failed-App cleanup, reconnection, deadline and permission-budget fixes are covered by focused tests; the remote fixture has not been rerun for these follow-ups.远端 HTTPS 渲染已在
d6d532eb8b验证: 已将d6d532eb8b的原始 App 组件与沙箱部署到真实远端实例的隔离夹具,通过现有 HTTPS 网关和端口转发完成 Tableau Public 图表显示、West 筛选、页面重载及再次筛选。认证和工具结果为 mock,原有远端 Qwen 服务未升级;完整 daemon/会话整合与 Tableau Cloud 认证仍未验收。 后续历史绑定、失败清理、重连、超时和审批预算修正通过定向测试验证,本次未重跑远端夹具。此前 4 MiB 测试属于独立的本地主端口转发回归,不能与本轮约 335 KB 的远端样例混为一谈。What this PR does
Repairs three independent MCP App integration failures: bounded per-server resource loading, App-initiated server tools, and opaque iframe origins. In the verified local-browser/local-daemon setup, the official Tableau App renders an authenticated Cloud chart inside Qwen Code, supports Region filtering, and renders again after a full Qwen page reload. The production renderer and sandbox at
d6d532eb8balso passed the remote HTTPS fixture described above; full remote daemon/session integration remains unverified.The existing 1 MiB / 10-second resource defaults remain. Explicit settings are bounded at 4 MiB / 120 seconds, propagate through SDK initialization and daemon configuration, and use separate pooled tools for different policies. When optional HTML exceeds live-delivery or history-page budgets, the completed text result and navigation survive; healthy delivery and reconnect replay retain the original HTML.
Apps can call App-visible tools on their bound server through the owning session's existing permission, hook and cancellation pipeline. Paging to earlier turns preserves that App session binding while historical source and feedback controls remain disabled. App-only tools stay out of model tool discovery. Raw App results, including embed tokens, return privately to the App; model/transcript output receives a fixed summary. App approval admission reserves model capacity within the existing session cap. App calls repair dead connections for later calls even with the daemon's invocation guard installed, without replaying the failed attempt. All App cards on one browser page share a two-request limit; additional requests wait client-side so approval/control requests can use HTTP/1.1 connections. Queued calls retain progress heartbeats and cancel without being submitted. App recovery uses the actual calling session's configuration and explicitly restarts its surviving pool entry, so a JSON-RPC session error cannot silently clear the tool/resource directory or reconnect another session's same-name server. App calls retain a separate five-minute bridge ceiling (including approval), with a 310-second SDK REST timeout; a longer MCP tool timeout does not extend that App ceiling. Attachment replacement and refreshed client IDs are handled, and background App calls no longer leave the composer stuck in Processing.
On HTTP loopback hosts, each render first receives a unique origin on a dedicated static-only loopback listener. Both iframe layers preserve origin identity, allowing a nested vendor iframe to use its actual origin. The App remains separated from Qwen's UI/API and other Apps. One-use, expiring registrations pin the CSP and parent origin; the listener has no daemon API or WebSocket endpoint and closes with its owner. HTTP CSP now enforces the sandbox independently of iframe attributes. Remote/HTTPS hosts use a small opaque data bootstrap through the existing daemon connection immediately; local isolated-listener failure switches to that path after 10 seconds. HTML travels over checked postMessage instead of a large navigation URL. The App stays opaque while vendor descendants retain their own origin. A failed data handshake or 30-second initialization timeout shows text fallback.
Why it's needed
Amplitude's real App resource exceeds both original resource limits. Tableau's 335,305-byte resource fits the defaults, but first requires an App-to-server embed-token call and then fails when the sandbox forces its nested view to
Origin: null. Increasing HTML limits cannot solve those latter failures. The earlier failed Tableau report is superseded by the successful authenticated browser verification below.Reviewer Test Plan
How to verify
Evidence (Before & After)
Real Tableau, in Qwen Code: the official unmodified
@tableau/mcp-server4.8.1 App, authenticated Tableau Cloud and the Superstore sample workbook were used in the user's existing Chrome on macOS. Only model tool selection was made deterministic by a local model fixture; MCP calls, resource HTML, OAuth, embed-token retrieval and the chart were real. The fixture's chat sentence is not the rendering assertion; the visible chart and successful interaction are the evidence.Origin: nullblocks the real viewRegion filter interaction: selecting West changed the map, KPIs and monthly charts inside the same Qwen App card.
Full Qwen page reload: the view rendered again under a fresh isolated origin, with the Qwen composer idle and usable.
The successful local Qwen page was
http://127.0.0.1:18897/session/06ca808a-1483-490a-b6de-9246a8344e53; this is a local validation URL, not a hosted demo. The local official MCP endpoint washttp://127.0.0.1:18891/tableau-mcp. Itsmcp-appsfeature gate was enabled through the official custom-provider mechanism; neither server code nor App HTML was patched. The hostedhttps://mcp.tableau.comtool listing observed in this test did not advertise the App rendering/token tools, so the hosted endpoint is not claimed as an Apps pass.The Tableau HTML measured 335,305 bytes and one authenticated
resources/readtook 254 ms. That measures the HTML resource read, not all JavaScript, chart data, map tiles or total render time. No complete dependency-transfer size is claimed.Theme: Qwen already supplies
hostContext.theme(dark/light) on initialization and updates it without remounting the App. The existing DOM regression covers theme switching. This Tableau App version does not apply that field to the embedded chart, so a white Tableau chart inside dark Qwen is expected in these screenshots; a dark Tableau chart was not verified.Resource-limit fixtures and authenticated Amplitude observations
Baseline tests used the global CLI with a local MCP fixture: 1,048,577-byte HTML was rejected and an 11-second read was cancelled after about 10 seconds. The PR build rendered the fixture with an explicit allowance, including interaction and cold-daemon replay.
Cold restart replay screenshot · History navigation screenshot · Reproduction fixtures. These are synthetic resource/history tests, not real Amplitude analytics.
On 2026-09-20, two OAuth-authenticated reads from
https://mcp.amplitude.com/mcpreturned the same 2,333,981-byte (2.23 MiB) App HTML in 88.715 s and 79.107 s. This version fits 4 MiB but exceeds 1 MiB, and both reads exceed 10 seconds. The documented example usesappResourceMaxBytes: 4194304andappResourceTimeoutMs: 120000. These observations do not guarantee future size or latency. Replay through the actual core and bundled CLI preserved identical HTML and its persisted SHA-256 under the configured allowance.The earlier real-HTML browser run with a synthetic chart ID reached the missing-host-tools error. This PR now adds App-to-server tools, but an authenticated Amplitude chart has not been rerun to completion; Tableau success must not be extrapolated to Amplitude.
Tested on
Current review follow-up validation
The latest follow-up addresses maintainer round4 F1/F2: a page-wide two-request App queue prevents a single-page App burst from occupying every HTTP/1.1 connection needed by approvals; App-only repair now works with the daemon guard while retaining initial execution authorization and never replaying the failed request. At
23c0e1bd90, real local WebKit/daemon tests passed for bursts of five and nine normally approved calls and MCP crash recovery. 232 focused tests and full build/typecheck/bundle passed; screenshots and details are in the latest PR reply. This is the report's bounded option (b), not asynchronous 202/event-stream completion; the queue is not shared across browser tabs and does not provide a universal connection reservation across arbitrary additional event streams.The preceding follow-up (
5f17519bf2) addressed pooled App recovery after request-level session errors. A source-integration reproduction showed a still-connected entry losing its App/resource registry; a second reproduction showed the initial fix reconnecting the pool bootstrap session instead of the caller. The final regression verifies caller-session binding, exact connection selection, restored tools/resources, no replay, and a successful explicit next call. New/replaced connections and ordinary refresh do not receive an extra restart; empty/failed restart results are rejected. At5f17519bf2, 1,429 tests across five files, full build, typecheck, bundle, ESLint and Prettier passed. This run uses mocked SDK transport responses through the real pool/manager/registry chain, not a new remote browser or authenticated Tableau run.The prior follow-up fixed R5-1 through R5-5: consistent remote sandbox documentation, no theme updates to a failed/closed App bridge, reconnect-without-replay for unguarded App calls, SDK/bridge timeout ordering, and reserved model approval capacity. Focused tests cover these regressions and the existing history/session boundaries. The existing remote screenshots remain tied to
d6d532eb8b.In the prior review round, full build, typecheck and bundle passed, along with 2,070 tests in seven focused files, ESLint and Prettier. Real App/AppBridge SDK probes passed 4/4: the old handler timed out at 60 seconds, while progress notifications kept a 95-second call alive; cancellation stopped all progress timers. Benign Chromium/WebKit embedding probes passed 4/4 across isolated and opaque fallback modes, including a simulated single-port forwarded connection. These use local fixtures, not authenticated Tableau. The Tableau screenshots below/above remain evidence from the earlier authenticated run, not a rerun of this head.
The review follow-up also preserves original HTML for an empty live queue, retains the newest replay App when evicting older text is sufficient, and preserves in-flight App calls when a reversible close/restore gate is released. Forced close and shutdown still cancel them. The protocol documents now describe per-subscriber lossy copies and their unchanged event IDs.
Independent maintainer verification: round4 reports Chromium/WebKit data-mode origin, forwarded-port and B1 isolation checks passing at
5f17519bf2. That is the maintainer's synthetic real-daemon/browser evidence, not a rerun of authenticated Tableau Cloud or our remote instance. Our earlier blocked security-probe limitation remains part of the historical record; this follow-up does not rerun attack probes.Environment (optional)
macOS arm64, Node.js 24.19.0; locally built CLI and daemon-backed WebShell. Earlier authenticated Tableau validation (before this review follow-up): full build, typecheck and bundle; 607 passing tests across 14 focused Core, ACP bridge, CLI, SDK and WebShell files; ESLint and Prettier. Independent sandbox HTTP probes passed 12/12 and complete-daemon origin isolation probes passed 14/14. The latter rejected sandbox-origin API calls, API/WS paths on the static listener, policy overrides and consumed URLs. Client-ID regression failed before its fix and passed after it. Prior resource-limit verification included 1,424 targeted tests, SDK public-type checks, CLI configuration round trips, live delivery and historical navigation; these are earlier scoped runs, not an additional final-run total.
Risk & Scope
allow-same-originon Qwen's own origin would not be equivalent.d6d532eb8bpassed. Full remote daemon/session integration and authenticated Tableau Cloud acceptance of this commit remain unverified. Stable_meta.ui.domainmapping is not implemented; an origin cannot be assigned by setting an arbitrary vendor-domain string.Design: resource policy (English) · 资源策略(中文) · App tools and origins (English) · App 工具与来源(中文).
Linked Issues
Refs #11945. Resource loading, scoped App server calls and the reproduced Tableau origin failure are addressed; full Amplitude chart rendering remains unverified.
中文说明
本 PR 的修改
修复三个独立的 MCP App 接入问题:按服务器配置有界资源加载、App 发起服务器工具调用,以及 iframe 不透明来源。在已验证的本地浏览器与本地 daemon 部署中,官方 Tableau App 已在 Qwen Code 内显示经过认证的 Cloud 图表,Region 筛选和 Qwen 整页重载均通过;远端 data 模式及单端口回归已完成,实际 HTTPS 部署仍待上方所述验收。
保留 1 MiB / 10 秒默认资源限制;显式配置上限为 4 MiB / 120 秒,通过 SDK 初始化及守护进程配置传递,不同策略使用不同连接池工具。可选 HTML 超过实时交付或历史页预算时保留工具文字及导航;正常交付和重连回放保留原始 HTML。
App 可通过所属会话现有的权限、hook 和取消流程调用绑定服务器上对 App 可见的工具。App 专用工具不进入模型工具发现。App 错误恢复绑定实际调用会话配置,并显式重启其仍存活的连接池条目,避免 JSON-RPC 会话错误清空工具/资源目录或误重连其它会话的同名服务器;失败调用不自动重放。包括嵌入令牌的原始结果仅返回 App,模型及会话记录只收到固定摘要。支持 attachment 替换和 clientId 更新;后台 App 调用不会再让输入框卡在 Processing。
HTTP loopback 宿主首先在独立、仅提供静态内容的监听器上为每次渲染取得唯一来源。两层 iframe 保留来源身份,使嵌套的第三方 iframe 能使用自身真实来源。App 与 Qwen 页面/API 及其它 App 隔离。一次性、会过期的注册固定 CSP 和父页面来源;监听器不提供守护进程 API 或 WebSocket,并随所属应用关闭。HTTP CSP 独立于 iframe 属性强制执行沙箱。远端/HTTPS 宿主直接经现有连接使用小型不透明 data 引导页,本地独立监听器 10 秒不可达时也切换到该路径。HTML 经严格校验的 postMessage 传输,避免大导航 URL。App 保持不透明,第三方后代保留自身来源;data 握手失败或 30 秒初始化超时后显示文字回退。
修改原因
Amplitude 的真实 App 资源超过原有两个限制。Tableau 的 335,305 字节资源低于默认上限,但先需要 App 调用嵌入令牌工具,随后又因沙箱使嵌套页面变为
Origin: null而失败。增加 HTML 上限无法解决后两项问题。下列真实认证浏览器验证已取代此前失败记录。Reviewer Test Plan
验证方式
修复前后证据
真实 Tableau、真实 Qwen 页面: 在 macOS 用户现有 Chrome 中,使用未经修改的官方
@tableau/mcp-server4.8.1 App、经过认证的 Tableau Cloud 和 Superstore 示例工作簿。仅模型工具选择采用确定性的本地 fixture;MCP 调用、HTML、OAuth、嵌入令牌和图表均真实。fixture 的聊天句子不作为成功断言,图表可见内容及交互结果才是证据。上方前后对照截图展示
Origin: null失败与 Qwen 内成功显示。West 筛选截图显示地图、指标和月度图同步更新;重载截图显示图表在新的隔离来源下重新渲染,Qwen 输入框保持空闲可用。成功验证页面为
http://127.0.0.1:18897/session/06ca808a-1483-490a-b6de-9246a8344e53,这是本地验证地址,不是线上演示。本地官方 MCP 为http://127.0.0.1:18891/tableau-mcp,通过官方 custom-provider 机制启用mcp-apps,未修改服务代码或 App HTML。本次观察的 hostedhttps://mcp.tableau.com工具列表未公开 App 渲染/令牌工具,因此不宣称 hosted 端点通过 Apps 验证。Tableau HTML 为 335,305 字节,一次已认证的
resources/read耗时 254 ms。这仅衡量 HTML 资源读取,不包含全部 JavaScript、图表数据、地图瓦片或总渲染时间,未宣称完整依赖下载大小。主题: Qwen 已在初始化时传递
hostContext.theme(dark/light),切换时实时更新且不重新挂载 App,已有 DOM 回归覆盖。该版 Tableau App 未将此字段应用到嵌入图表,因此截图中的黑色 Qwen 内呈现白色 Tableau 图表;未验证 Tableau 深色图表。资源限制 fixture 与真实 Amplitude 观察
基线用全局 CLI 和本地 MCP fixture,1,048,577 字节 HTML 被拒绝,11 秒读取在约 10 秒取消。PR 构建在显式提高上限后显示相同 fixture,并通过交互及守护进程冷重启回放。默认/配置后截图、冷重启截图、历史导航截图及复现 fixture 链接见上方,均为合成资源/历史测试,不是真实 Amplitude 分析数据。
2026-09-20,两次通过 OAuth 读取
https://mcp.amplitude.com/mcp均取得 2,333,981 字节(2.23 MiB) 相同 HTML,分别耗时 88.715 秒、79.107 秒。该版本低于 4 MiB、高于 1 MiB,耗时均超过 10 秒。文档示例为appResourceMaxBytes: 4194304和appResourceTimeoutMs: 120000。这些值不保证未来资源大小或延迟。经实际 core 和构建 CLI 回放,配置后 HTML 及持久化 SHA-256 保持一致。此前真实 HTML 配合合成 chart ID 的浏览器测试显示缺少宿主工具能力,截图链接见上方。本 PR 现已补充 App 服务器工具调用,但尚未重新完成经过认证的 Amplitude 图表验证,不能从 Tableau 成功推断 Amplitude 成功。
测试平台
浏览器为 macOS 用户 Chrome;本轮另含 Linux Docker 内服务的单端口转发回归,Windows 未测试。
本轮审核修正验证
最新修正处理维护者第4轮 F1/F2:同页所有 App 共用两个在途请求的上限,避免一次 App 并发占满审批所需的 HTTP/1.1 连接;App 连接恢复允许 daemon guard 存在,但首次执行仍需授权,失败调用绝不重放。排队中的调用保留心跳并支持取消。
23c0e1bd90的真实本地 WebKit/daemon 验证通过 5/9 并发审批与 MCP 崩溃恢复;232 项定向测试及完整 build/typecheck/bundle 通过,截图与细节见最新 PR 回复。这采用报告中的最小方案 (b),尚未改为 202/事件流异步交付;队列不跨标签页共享,也不保证任意多个事件流下的连接容量。前次
5f17519bf2修正处理连接池中的 App 会话错误恢复:真实 pool/manager/registry 链路复现了目录丢失及误重连启动会话的问题。回归验证实际会话绑定、目标连接选择、工具/资源恢复、不重放,以及下一次显式调用成功;普通刷新、新建或替换连接不额外重启,失败或空重启结果不视为成功。5f17519bf2的 5 个文件共 1,429 项测试、完整 build/typecheck/bundle、ESLint 和 Prettier 通过。本轮仅 mock SDK transport 响应,未重跑远端浏览器或 Tableau 认证图表。此前审核轮次的完整 build、typecheck、bundle,以及 7 个定向文件共 2,070 项测试、ESLint 和 Prettier 通过。真实 App/AppBridge SDK 探测 4/4 通过:旧处理器在 60 秒超时,修正后进度通知使 95 秒调用成功,取消后所有进度计时器清除。Chromium/WebKit 正常嵌入探测 4/4 通过,覆盖独立来源、不透明回退及模拟仅转发主端口的连接。这些使用本地 fixture,不是真实认证 Tableau;上方 Tableau 截图来自此前认证实测,并非此版本重新运行。
本轮同时保留空实时队列的完整 HTML;当淘汰旧文字即可满足预算时保留最新回放 App;可恢复的关闭/恢复门闩释放时保留进行中的 App 调用,强制关闭和退出仍会取消。协议文档已说明各订阅者的有损副本及不变的事件 ID。
维护者独立验证: 第4轮报告 确认
5f17519bf2的 Chromium/WebKit data 模式来源、端口转发及 B1 隔离检查通过。这是维护者的合成真实 daemon/浏览器证据,不是本轮重跑真实 Tableau Cloud 或我们的远端实例。本轮没有重跑攻击探测。环境
macOS arm64、Node.js 24.19.0、本地构建 CLI 与守护进程 WebShell。此前真实 Tableau 验证(本轮审核修正前)包括完整 build、typecheck、bundle,以及 Core、ACP bridge、CLI、SDK、WebShell 共 14 个文件 607 项定向测试通过,以及 ESLint、Prettier。独立沙箱 HTTP 探测 12/12 通过,完整守护进程来源隔离探测 14/14 通过;后者验证 App 来源 API 请求、静态监听器 API/WS 路径、策略覆盖和已使用 URL 均被拒绝。clientId 回归修复前失败、修复后通过。此前资源限制验证包含 1,424 项定向测试、SDK 类型检查、CLI 配置读写、实时交付及历史导航;这些是之前按范围执行的验证,不计入本次测试总数。
风险与范围
allow-same-origin来替代。d6d532eb8b的真实远端 HTTPS Tableau Public fixture 已通过;后续提交未重跑远端,完整 daemon/会话整合与 Tableau Cloud 认证仍待验收。未实现稳定_meta.ui.domain映射;不能填写任意第三方域名就获得该来源。资源策略与 App 工具/来源的完整中英文设计文档链接见上方。
关联 issue
关联 #11945。已解决资源加载、受限的 App 服务器调用及复现的 Tableau 来源问题;完整 Amplitude 图表仍未验证。