Skip to content

feat(web-shell): search within the current conversation - #12234

Merged
samuelhsin merged 17 commits into
QwenLM:mainfrom
samuelhsin:feat/12231-webshell-conversation-search
Sep 23, 2026
Merged

samuelhsin merged 17 commits into
QwenLM:mainfrom
samuelhsin:feat/12231-webshell-conversation-search

Conversation

@samuelhsin

@samuelhsin samuelhsin commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

Adds a compact search icon below the left Web Shell session timeline. The 14px icon uses a narrow transparent button sized for the navigation gutter. It opens a dialog for literal, case-insensitive searches of user and assistant messages, including code. Results show highlighted snippets and jump to the matching message with a temporary highlight, including messages outside the loaded live window. The search icon shares the timeline visibility and is hidden whenever the timeline is hidden.

The host option conversationSearchThreshold defaults to 10: the icon is hidden at 10 messages and appears at 11. Historical messages and newly admitted live messages count toward the threshold. Searches read persisted history page by page and retain at most 200 snippets; navigation reuses the existing historical viewport and its cache budget. The dialog supports result navigation, Escape, focus restoration, error/retry states, and English/Chinese labels.

Why it's needed

In a long conversation, users often remember a keyword but cannot quickly find the original answer or code. The sidebar content search delivered in #10612 finds sessions; this change locates content inside the session already open. It is scoped separately from the CLI/VS Code requests in #6824 and #11111.

Reviewer Test Plan

How to verify

  • Open a conversation with 10 user/assistant messages: the icon is absent. Add an eleventh message: it appears below the left session timeline. Set conversationSearchThreshold={20} in an embedding: the boundary moves to 20/21.
  • Search for text appearing only in an old assistant response outside the live window. A highlighted excerpt appears; selecting it opens the correct historical position and highlights the message. The composer draft remains intact.
  • Search Chinese text or a literal code token. Check the empty state, previous/next results, keyboard navigation, Escape, and focus return to the trigger.
  • Close the dialog or change the query while history is loading. Stale results must not replace the new search, and cancelled navigation must not leave the viewport loading.
  • Verify light/dark themes, English/Chinese, and desktop/mobile widths.

Evidence (Before & After)

Screenshots / 界面截图

Synthetic fixtures only; no private conversation data. / 仅使用合成数据,不含私人会话。

Timeline-aligned search entry / 与时间轴刻度对齐的搜索入口

Timeline search entry

Stable-height search dialog, dark / 固定高度搜索弹框(深色)

Dark search dialog

Chinese, light / 中文浅色弹框

Chinese search dialog

The search entry hides with the timeline. Input, matching results, and empty results keep the dialog height and vertical position stable. Latest verification: 11 browser tests passed, including alignment, hidden narrow-layout entry, and stable dialog geometry.

搜索入口跟随时间轴显隐;输入、有结果和无结果时弹框高度与垂直位置保持稳定。最新 11 项浏览器测试通过,覆盖图标对齐、窄屏隐藏入口及弹框尺寸稳定性。

Before: the conversation footer only offered scroll-to-bottom; finding an old passage required manual scrolling. After: synthetic browser tests find an assistant message outside the live replay, navigate to its persisted identity, observe the highlight, and confirm the unsent draft is unchanged. The 10/11 threshold and all eight theme/language/width combinations are covered by browser tests. All fixtures are synthetic; no private conversations or original user screenshots are included.

Earlier validation at 713d23d095: 97 focused unit tests, 11 browser tests, and 2 visual scenarios (four light/dark captures) passed. Root build, typecheck, bundle, and changed-file lint/format checks passed locally. All executed GitHub CI checks subsequently passed at 9f8c93e302; skipped jobs were not executed.

Tested on

OS Status
macOS ✅
Windows ⚠️
Linux ⚠️

Environment (optional)

Local Web Shell with the repository's synthetic daemon Playwright harness; Chromium at 1440px and 390px. No real model calls are required by these tests.

Risk & Scope

  • Each debounced query scans the current persisted conversation, so very long histories incur read I/O. Result memory is bounded, and closing/changing a query stops further page requests and discards stale responses.
  • Older daemons without turn navigation search loaded messages only, with an explicit notice. Tool output, thinking, cross-session search, and a settings-page threshold editor are outside this change.
  • No daemon API or storage migration; the optional host prop is additive. Windows/Linux were not locally tested.

External message navigation / 对外消息定位能力

Implemented in this PR: shellRef.current.navigateToMessage({ sessionId, recordId, signal? }), with WebShellMessageNavigationRequest and WebShellMessageNavigationResult exported from the package root. Both WebShell embedding components expose it through the existing WebShellApi.

The host first opens the target session/workspace through the existing provider lifecycle, then invokes navigation with the stable persisted user/assistant transcript record ID. WebShell resolves that exact record, loads older history outside the live/rendered window, and activates the existing scroll/highlight behavior. The operation also works when the timeline/search icon is hidden, without opening the search dialog or meeting its message-count threshold.

Results distinguish located, not_found, not_ready, session_mismatch, unsupported, cancelled, and error. located means the target has been activated; React applies the scroll/highlight on the following render, so this is not an animation-completion signal. An accepted new request, AbortSignal, changed session/workspace/transcript owner, or unmount invalidates stale work. Rejected calls from stale API owners, already-aborted signals, mismatched sessions, or empty record IDs do not cancel an active valid navigation. Identical text does not substitute for record identity. Drafts and ongoing responses are untouched.

Resolution currently scans persisted history until the exact record is found, retaining one hit and reusing the bounded history viewport. Very old targets incur linear reads; no server-side record index or new daemon endpoint is introduced.

Integration boundary: the existing cross-session search endpoint still returns one session plus one snippet, without a message record ID. Host search results must supply that ID before they can use exact navigation. This PR does not add multiple-message search results, archive search, or downstream host UI changes. A host must handle not_ready and wait for the selected session/history viewport before trying again; navigation itself never changes the active session.

Earlier validation at 9f8c93e302: root build, typecheck, bundle, and the real pre-commit hook passed locally; 1,008 App tests and 87 history/store tests passed. All executed GitHub CI checks at this commit passed (including Ubuntu unit tests, lint/static, integration, desktop checks, screenshots and E2E smoke).

Verification: 23 hook tests cover cancellation/ownership/readiness/errors; store tests cover exact IDs, duplicate text, stopping at the target, and missing records. Four real embedding browser scenarios call the public API directly with the timeline hidden, assert historical content and visible highlight, preserve the draft, and verify a missing record does not change the displayed content. The original 11 search-browser scenarios plus a new ArrowUp/ArrowDown wrap-around, previous/next button, and Enter-to-locate scenario pass: 16 browser scenarios in total including the four public-API cases.

中文:本 PR 已实现通过 shellRef 对外定位消息。宿主打开目标会话后传入稳定的持久化记录 ID,WebShell 加载未渲染历史并滚动高亮,不依赖搜索弹框、图标显隐或阈值。返回明确状态,取消及切换会话/工作目录后旧请求失效,保留草稿。现有跨会话搜索仍没有消息记录 ID,也不支持归档搜索;宿主搜索接入是另外的改动,不能把本 PR 的公开定位能力等同于已完成整个宿主功能。

Public API browser evidence / 公开接口本地截图

Synthetic embedding harness; public API invoked directly, timeline and search trigger hidden. The highlighted target is outside the initial live window; the composer still contains the unsent synthetic draft. / 合成嵌入场景,直接调用公开接口,旧消息已加载高亮且草稿保留。

Light desktop / 浅色桌面

Dark desktop / 深色桌面

Light narrow Chinese / 浅色中文窄屏

Dark narrow Chinese / 深色中文窄屏

Linked Issues

Closes #12231

中文说明

本 PR 做了什么

在 Web Shell 左侧会话时间轴下方新增 14px 小尺寸搜索图标,按钮采用适配窄栏的透明样式,点击弹框按字面量、不区分大小写搜索用户和助手消息,包括代码。结果显示高亮摘要,点击后定位并短暂高亮对应消息,包含当前 live window 之外的历史内容。搜索入口跟随时间轴显隐,时间轴隐藏时也隐藏。

宿主配置 conversationSearchThreshold 默认为 10:10 条消息时隐藏,11 条时显示。计数包含历史消息和新接纳的实时消息。搜索分页读取持久化历史,最多保留 200 条摘要;定位复用现有历史浏览窗口和缓存预算。弹框支持结果导航、Escape、焦点恢复、错误与重试状态,以及中英文文案。

为什么需要

长会话中,用户经常记得关键词,却难以找到原来的回答或代码。#10612 已实现的侧边栏内容搜索用于查找会话,本次用于定位已经打开的会话内部内容,与 #6824、#11111 的 CLI / VS Code 需求分开处理。

评审验证计划

如何验证

  • 打开有 10 条用户/助手消息的会话:不显示图标。增加第 11 条消息后:图标出现在左侧会话时间轴下方。嵌入时设置 conversationSearchThreshold={20},显示边界应变为 20/21。
  • 搜索仅存在于 live window 之外的旧助手回复中的文本:显示高亮摘要;选择后打开正确历史位置并高亮消息,输入草稿保持不变。
  • 搜索中文或代码字面量,检查无结果状态、上一个/下一个结果、键盘操作、Escape 关闭及焦点回到入口。
  • 历史加载期间关闭弹框或修改查询:过期结果不能覆盖新查询,取消定位后不能残留加载状态。
  • 检查浅/深主题、中/英文、桌面/移动端宽度。

证据(前后对比)

改动前,会话底部只有置底功能,找回旧内容需要手动滚动。改动后,合成数据浏览器测试能找到 live replay 之外的助手消息,按持久化身份定位,观察高亮,并确认未发送的草稿保持不变。浏览器测试覆盖 10/11 条边界和全部八种主题/语言/宽度组合。所有测试数据均为合成数据,不包含私人会话或用户原始截图。

此前提交 713d23d095 验证:97 项针对性单测、11 项浏览器测试、2 项截图场景(浅深主题共四张)通过;根目录 build、typecheck、bundle 和改动文件 lint/format 本地通过。随后 9f8c93e302 的已执行 CI 检查全部通过;跳过项未执行。

测试平台

OS 状态
macOS ✅
Windows ⚠️
Linux ⚠️

环境(可选)

本地 Web Shell 配合仓库的合成 daemon Playwright 测试设施;Chromium 宽度为 1440px 和 390px。这些测试不需要真实模型调用。

风险与范围

  • 每次防抖查询会扫描当前持久化会话,超长历史会产生读取 I/O。结果内存有界,关闭或修改查询会停止后续分页请求并丢弃过期响应。
  • 不支持 turn navigation 的旧 daemon 仅搜索已加载消息,并明确提示。工具输出、思考内容、跨会话搜索及设置页阈值编辑器不在本次范围内。
  • 不改动 daemon API 或存储格式;可选宿主 prop 为增量配置。本地未验证 Windows/Linux。

关联 Issue

Closes #12231

@github-actions github-actions Bot added the review/self-reported The linked issue was opened by the PR author (self-reported) label Sep 19, 2026
@samuelhsin

samuelhsin commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator Author

Historical verification for the initial implementation. The entry placement and narrow-layout behavior were subsequently revised; see the latest follow-up report below.
此记录对应初始实现,入口位置与窄屏显隐已调整,请以最新跟进报告为准。

E2E verification report

Verified on macOS with Chromium using synthetic daemon fixtures; no real conversation data, screenshots, credentials, or model calls were used.

  • 11/11 browser tests passed on the final implementation (23.0s): default 10/11 visibility boundary; search and navigation to persisted assistant content outside the live window; destination highlight and unsent draft preservation; Chinese text and code-token searches; no-result state; Escape/focus restoration; and all 8 light/dark × English/Chinese × 1440px/390px combinations. The search button's position to the right of the bottom arrow is asserted.
  • 475/475 focused unit tests passed across 9 files, including result caps, empty attachment-message counts, deduplicating live/persisted messages, stale request cancellation, partial-history errors, Unicode matching, and navigation across 3 historical pages with a 2-page cache budget.
  • Root npm run build, npm run typecheck, and npm run bundle passed. Changed-file ESLint and Prettier checks and the real pre-commit hook passed.

Browser reproduction command (from packages/web-shell):

PLAYWRIGHT_PORT=15231 npx playwright test --config playwright.config.ts --project=chromium client/e2e/web-shell.conversation-search.spec.ts

Baseline: the upstream paginated-history suite passed 3/3 before implementation. The globally installed CLI (0.24.0-preview.0) could start an isolated daemon, but did not contain Web Shell assets, so browser baseline and feature verification used the repository's Web Shell test harness. Windows/Linux and real-provider end-to-end calls were not tested.

During verification, Escape focus restoration, cancelled navigation cleanup, and mixed persisted/live message counting were corrected and re-tested. The final full build also passed after a message-kind type narrowing fix.

中文

在 macOS Chromium 上使用合成 daemon 数据验证,不使用真实会话、原始截图、凭据或模型调用。

  • 最终 11/11 项浏览器测试通过(23.0 秒):默认 10/11 条显示边界;定位 live window 之外的持久化助手内容;目标高亮、草稿保留;中文与代码关键词;无结果;Escape 与焦点恢复;浅/深 × 中/英文 × 1440px/390px 八种组合。断言搜索按钮位于置底箭头右侧。
  • 9 个文件的 475/475 项针对性单测通过:包括结果数量上限、空附件消息计数、实时/历史去重、取消过期请求、部分历史错误、Unicode 匹配、两页缓存下跨三页历史定位。
  • 根目录 build、typecheck、bundle、改动文件的 ESLint/Prettier 和真实 pre-commit hook 均通过。

基线:实现前历史分页浏览器测试 3/3 通过。全局 CLI 0.24.0-preview.0 能启动隔离 daemon,但不含 Web Shell 静态资源,因此浏览器基线及功能验证使用仓库测试设施。未验证 Windows/Linux 或真实模型调用。

验证过程中修正并复测了 Escape 焦点恢复、取消定位后的状态清理,以及持久化/实时消息混合计数。消息类型收窄修正后,完整构建重新通过。

@samuelhsin
samuelhsin requested review from callmeYe, qwen-code-ci-bot and wenshao and removed request for qwen-code-ci-bot and wenshao September 19, 2026 13:18
@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Review and CI follow-up

Addressed the actionable findings on this PR:

  • Merged current main (42f9d13cda) without conflicts. The failing Lint & Static job stopped at the lint-gate freshness check because the branch lacked c56d60438b61; that commit is now included. No CI gate was disabled.
  • Fixed a real daemon API mismatch: the first history-search request sent start=0 without a snapshot and received HTTP 400 (invalid_transcript_cursor). Search now obtains the head snapshot first and uses it for subsequent indexed reads. Unit and browser mocks now reject the previously invalid request. A real local daemon returned HTTP 200 and a complete scan of 24 synthetic messages with one matching result after the fix.
  • Added dedicated light/dark scenarios to the automatic visual preview suite, covering the timeline entry and search dialog (four captures), addressing the bot's missing-coverage finding.
  • Fixed the merged build's HTML export size failure without increasing its limit. Search-only translations now stay with the interactive component instead of entering the shared read-only export. Export runtime: 1,927,859 bytes, below the unchanged 1,930,000-byte ceiling (previously 1,930,119; independently measured main baseline 1,927,710).

Validation: 97 focused unit tests, 11 browser tests, and 2 visual scenarios passed. English/Chinese search behavior, timeline alignment/visibility, fixed dialog height, historical navigation, and draft preservation are covered. Root npm run build, npm run typecheck, and npm run bundle all passed on the pushed commit 713d23d095. The new GitHub CI run is still pending. Changed-file ESLint/Prettier and the real pre-commit hook passed.

The initial E2E report above is marked historical because the final entry moved to the timeline and now hides with it. No review threads were resolved and no approval was submitted.

中文

已处理:合入最新 main 修复 CI 规则过期;修正历史搜索首次请求缺少快照导致的真实 HTTP 400;补齐自动截图场景;将搜索专用文案移出只读导出共享翻译包,使导出体积回到原有上限内,未放宽检查。

真实 daemon 已验证完整扫描 24 条合成消息并命中 1 条。97 项相关单测、11 项端到端测试、2 项浅深主题截图场景通过。旧验证报告已标注为历史记录;未解决 review thread、未提交批准。

The issue and PR description now also track the requested public embedding API for exact persisted-message navigation. That additional API is explicitly marked as not implemented yet; the current patch implements internal in-conversation navigation.

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Public message navigation implemented and verified

Commit 9f8c93e302 adds shellRef.current.navigateToMessage({ sessionId, recordId, signal? }) and exports its request/result types. This supersedes the earlier comment that marked the public API as pending.

The host opens the session/workspace first; the API finds the exact persisted user/assistant record, loads historical content outside the live window, and activates scrolling/highlighting. It also works with the timeline hidden and leaves the draft unchanged. Cancellation, stale owners, unsupported history, readiness, missing records, and failures return distinct statuses. located activates the target; the following React render performs the visual update.

Local validation on macOS Chromium, with synthetic data only:

  • 19/19 public API hook tests.
  • 87/87 history/search store tests, including duplicate text with distinct record IDs.
  • 1,008/1,008 App regression tests.
  • 15/15 browser tests: the existing 11 search cases plus 4 direct public-API cases covering light/dark, English desktop and Chinese narrow layouts. They assert the actual historical message and highlight, unchanged draft, and a missing record leaving the view unchanged.
  • Root build, typecheck, bundle and real pre-commit hook passed. HTML export remains 1,927,859 bytes, under the unchanged cap.

Four screenshots are embedded in the PR description, showing the highlighted historical target and preserved synthetic draft. Screenshots are on a separate evidence branch and do not add image binaries to this code PR.

The existing cross-session content-search endpoint still returns session+snippet without a stable message record ID. Host search integration, multiple-message results and archived-session search are not claimed as completed here. New-head CI is pending.

中文:已实现公开消息定位接口,并将 issue/PR 的“尚未实现”更新为实际契约。直接调用公开 API 的四种浏览器场景均验证旧消息加载、高亮、草稿保留和缺失记录行为;四张脱敏截图已放入 PR 正文。本地完整构建、类型检查、打包及回归通过,新提交 CI 待完成。

@wenshao

wenshao commented Sep 19, 2026

Copy link
Copy Markdown
Collaborator

Verdict: merge-ready — 1109/1109 scripted assertions passed, 0 failed. Verified head: 9f8c93e302ffe3ff16ff1b5cec80274b603b9a4a (base 42f9d13cdae0a7d462019487646bf7fb4995bf24). Maintainer-driven local round by @wenshao, 2026-09-19.

中文摘要

结论:可以合并(merge-ready)——1109/1109 项脚本断言全部通过,验证 head 为 9f8c93e302。

  • A/B 对照:将 PR 自带的浏览器端 e2e 用例原样移植到 base(42f9d13)运行。中央功能用例在 base 上 10/10 全部失败(搜索按钮不存在、等待超时),在 head 上 15/15 全部通过——证明该功能是本次 diff 带来的(见下方 A/B 表与截图 01/02)。
  • 变异测试:对核心逻辑做 5 处单点变异(正则转义、阈值边界、消息去重、200 条上限、跨块去重),5/5 均被现有测试精确击杀,含一个同文件阳性对照,证明测试非空转(截图 03)。
  • 门禁:web-shell typecheck 干净、ESLint 干净;两处门禁均做了活性探针(植入类型错误/“debugger”均被捕获后移除)。设计文档中英双语互链齐全。
  • 建议级发现(不阻塞):弹框的结果导航交互——ArrowUp/ArrowDown 循环选择、Enter 选中、上一条/下一条按钮——在 PR 描述和 Reviewer Test Plan 中被声明,但没有任何测试覆盖(Escape 与焦点还原已有 e2e 覆盖)。
  • 未覆盖:Windows/Linux 真实桌面环境(作者已声明)、旧 daemon 的 legacy 降级路径的端到端表现、跨会话搜索(PR 明确排除)、超过 200 条命中时的 UI 呈现(store 层截断已测)。

Central claim and A/B proof

Central claim: above a configurable message threshold (default 10), a search icon appears below the session timeline and opens a dialog that performs literal, case-insensitive search over user/assistant messages — including persisted history outside the loaded live window — and navigates to the matching message with a temporary highlight.
Secondary claims: (a) conversationSearchThreshold host option moves the boundary; (b) hosts can call shellRef.current.navigateToMessage({ sessionId, recordId, signal? }) with typed result statuses.

The PR's own browser e2e suites (real Chromium, real HTTP/WebSocket fetches intercepted only at the Playwright route layer — the full client and SDK wire code executes) were run unmodified at head, then copied verbatim into the base worktree and run against the base build. Same spec files, same fake-daemon scenario, both arms.

Spec cell (chromium) Head 9f8c93e Base 42f9d13
threshold 10 → icon absent pass pass (absent)
threshold 11 → icon present pass fail (no button, timeout)
locate persisted history outside live window, draft kept pass fail (no button, timeout)
content/layout light·dark × en·zh-CN × 1440 4× pass 4× fail
content/layout × 390px (icon hidden) 4× pass 4× pass (hidden)
conversation-search total 11/11 6 failed / 5 passed
external navigateToMessage located/not_found × 4 4/4 4/4 failed (api.navigateToMessage undefined)

Every cell whose oracle requires the feature flips from red at base to green at head; every cell that asserts absence/hidden states is green on both arms — the diff is exactly what makes the difference. Witnesses: 01-ab-head-e2e.png (head cells green), 02-ab-base-e2e.png (base cells timing out waiting for a search button that does not exist).

UI evidence from the head runs: 04-search-dialog-history-hit.png (case-insensitive unique-needle → UNIQUE-NEEDLE highlighted snippet, 1/1 result, draft intact), 05-history-hit-located-flash.png (historical viewport opened at the hit with flash highlight, unsent draft preserved), 06-external-nav-located.png (host-driven navigateToMessage landing on the historical record), 07-timeline-search-entry.png (14px entry aligned with the timeline rail).

Mutation matrix (vacuity check)

Five single-point mutants of the new logic, run against the PR's own unit suites. All killed by the assertion that exists to catch them; M1 is the in-file positive control proving the harness can go red.

# Mutation Killer test Result
M1 drop regex escaping in createConversationSearchSnippet (control) escapes patterns and preserves original Unicode offsets killed
M2 threshold gate <= → < (icon at exactly 10) shows only above the default threshold…, honors a custom threshold and zero killed (2 red)
M3 drop !sameMessage message-count dedupe store suite, 1 red killed
M4 hits cap 200 → 100 bounds stored results while continuing to count all matching messages killed
M5 drop recordId !== lastMatchedRecordId cross-block dedupe counts empty attachment messages and finds later fragments without duplicate hits killed

Witness: 03-mutation-cap100.png. No survivors; no combination rows needed (no two hunks defend the same hazard).

Targeted gates

  • Unit tests (head): the 5 new/changed test files — conversation-search.test.ts, ConversationSearch.test.tsx, useMessageNavigation.test.tsx, TranscriptViewport.pending.test.tsx, App.test.tsx — 1066/1066 passed (vitest, jsdom).
  • Boundary probes (astral-plane offsets, cross-script case folding, metacharacter literals, ellipsis clipping; temp vitest file, removed after): 3/3 passed.
  • Typecheck: tsc -p tsconfig.json --noEmit clean; liveness probe: a planted text: number signature error produced 6 error TS lines, removed after.
  • ESLint on all 10 changed production files: clean; liveness probe: a planted debugger; statement reported no-debugger error, removed after. (The config does not enable unused-vars; the probe rule chosen is one it enforces.)
  • Design docs: docs/design/2026-09-19-web-shell-conversation-search.{md,zh-CN.md} both present with reciprocal language links and matching structure.
  • Merge freshness: origin/main has advanced exactly one commit past the PR base (28029ac7, touches only GitModePopover.module.css, no overlap with PR files); merge-base equals the PR base, so the verified tree is what would land.

Reviewer Test Plan walk-through

Step Result
10/11 boundary; conversationSearchThreshold={20} moves it Covered: e2e 10/11 cells; custom threshold in honors a custom threshold and zero
Search text only in an old assistant message outside the live window; navigates with highlight; draft intact Covered: e2e history cell incl. draft assertion and peer-side anchor census (record-2 received by the fake daemon)
Chinese text / code token / empty state / prev-next / keyboard / Escape / focus return Chinese, code, empty state, Escape, focus return covered by e2e; prev/next buttons and ArrowUp/Down/Enter uncovered — see Findings
Close dialog / change query while loading; stale results must not win; cancelled navigation must not wedge the viewport Covered: debounces queries and ignores an earlier search…, invalidates an in-flight search…, clears pending search navigation…, cancels a distant selection…
Light/dark, en/zh, desktop/mobile Covered: 8-combination e2e matrix

Findings

  1. Suggestion (test coverage) — The dialog's result-navigation interactions claimed in the PR body and required by the Reviewer Test Plan ("previous/next results, keyboard navigation") have no test anywhere: the chevron up/down buttons, ArrowUp/ArrowDown selection cycling, and Enter-to-choose in ConversationSearch.tsx (lines ~255–275). Grep over client/ finds zero Arrow/Enter references in ConversationSearch.test.tsx; the e2e specs drive results only by click. The logic reads correct (modulo cycling, isComposing IME guard, intent-token cancellation), so this is a completeness gap, not a defect. Suggest one jsdom test cycling selection with ArrowDown wrap-around and choosing with Enter.

Not covered

  • Real Windows/Linux desktops (author disclosed macOS-only local testing); this round ran on Linux aarch64 with Chromium headless-shell.
  • The legacy/degraded daemon path end-to-end (older daemon without turn-navigation): the store returns mode: 'legacy' and the dialog shows the loaded-only notice; covered by unit tests (reports unsupported for legacy navigation), not driven in a browser.
  • Long-history I/O cost on real daemons (author-declared accepted tradeoff); the store-side 200-hit cap, page-budget, and stale-cancellation behavior are unit-tested.
  • Cross-session search, tool output, thinking content — explicitly out of the PR's scope.
  • Per-commit attribution of the 6 commits; the aggregate base..head diff was verified.
  • Full-repo CI gates (repo-wide lint/format/test); scoped to the changed workspace and files.

Evidence

01-ab-head-e2e — head 9f8c93e: threshold-11 + history-locate cells green

01-ab-head-e2e

02-ab-base-e2e — base 42f9d13: identical cells time out waiting for a search button that does not exist

02-ab-base-e2e

03-mutation-cap100 — mutation M4 (hits cap 200→100) turns the bounds test red

03-mutation-cap100

04-search-dialog-history-hit — case-insensitive unique-needle → UNIQUE-NEEDLE highlighted hit, 1/1 result, composer draft intact

04-search-dialog-history-hit

05-history-hit-located-flash — historical viewport opened at the hit with flash highlight; unsent draft preserved

05-history-hit-located-flash

06-external-nav-located — host-driven navigateToMessage lands on the historical record

06-external-nav-located

07-timeline-search-entry — 14px search entry aligned with the timeline rail

07-timeline-search-entry

Methodology

Local maintainer round on Linux aarch64 (Orange Pi 6 Plus), Node 22, Chromium headless-shell 149 (Playwright 1.61.1). Head worktree at 9f8c93e3 built with npm ci (exit 0, 2000 packages). Base worktree at 42f9d13 reuses the head install via hardlinked node_modules — clean control because the PR touches no package.json/lockfile; internal workspace links verified to resolve inside the base tree (readlink -f node_modules/@qwen-code/qwen-code-core → …/pr12234-base/packages/core), and git diff base..head confirms zero changes outside packages/web-shell + docs/, so the hardlinked dist/ of sibling packages is base-equivalent. E2E: real Chromium against the suite's route-level fake daemon (SSE + REST), spec files copied verbatim into the base tree. Mutations applied by single-line edits, each reverted and the tree confirmed clean (git status empty) after. Harness logs: tmp/pr12234-verify-20260919-223813/ (report.md, verdict.txt, assertions.json, evidence/*.png).

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Follow-up review: navigation cancellation and keyboard coverage

Addressed the non-blocking coverage suggestion in the maintainer verification report: a real browser test now exercises ArrowUp/ArrowDown selection, wrap-around at both ends, previous/next buttons, and Enter-to-select, then asserts that the selected historical message is visibly highlighted.

Also reproduced and fixed an independent cancellation defect: navigateToMessage advanced its generation before rejecting a stale-owner or already-aborted call, so that rejected call could cancel a valid in-flight navigation. Generation now advances only after the request passes all preconditions. Regression tests prove stale owners, aborted signals, mismatched sessions, and empty record IDs leave the active request intact; valid newer requests still supersede older ones.

Commit: 427bcc6d1f.

Validation: 23/23 hook tests, 16/16 browser tests, root typecheck, root build/bundle, changed-file lint/format, and the real pre-commit hook passed. The two original regression cases failed before the fix and passed afterward. No visible layout changed; the existing synthetic screenshots in the PR remain representative.

All executed CI checks passed on the preceding head 9f8c93e302, including Ubuntu tests, lint/static, integration, desktop checks, visuals and Web Shell E2E smoke. The new commit's CI is pending. No review threads were resolved and no approval/merge was submitted.

中文:补齐维护者建议的键盘与前后按钮浏览器验证,并修复“被拒绝的定位请求误取消有效定位”的竞态。先复现两项失败,再修复并验证 23 项接口单测、16 项浏览器测试及完整构建/类型检查/打包;无界面变化,现有脱敏截图仍适用。上一提交 CI 已全绿,新提交等待 CI。

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Rechecked head 427bcc6d1fc3f135251ffadeed7f9c7db98e3a35.

  • Every executed CI check passed, including Ubuntu unit tests, lint/static, no-AK integration, Linux/Windows desktop checks, visual captures and Web Shell E2E smoke. Skipped jobs are not counted as passes.
  • No new review comments or review threads were found. The maintainer's keyboard-coverage suggestion is covered by the previous follow-up.
  • Re-reviewed request-generation ownership, rejected-call cancellation, historical navigation, and keyboard-selection behavior. No additional actionable defect was identified in this pass; no code changes or redundant test reruns were made.
  • Updated the PR description's stale pending-CI status. The branch is mergeable without conflicts; GitHub still reports the overall merge state as blocked. No merge, approval or review-thread resolution was performed.

中文:本轮复查当前提交的已执行 CI 全部通过,无新增评审意见或待处理线程。未发现新的确定缺陷,因此不追加无必要代码改动;已更新正文中过时的 CI 等待状态。代码无合并冲突,GitHub 整体合并状态仍为 BLOCKED,未执行合并或批准。

@samuelhsin
samuelhsin marked this pull request as ready for review September 19, 2026 19:03
@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Reviewed all 63 inline comments on 427bcc6d1fc3f135251ffadeed7f9c7db98e3a35 and kept this follow-up to the five Critical correctness issues. Fixed in c554b17f1876c941189bbb68658cce7a1b55d5ce:

  • R1-1: Search state/dialog now have a stable owner; moving or hiding the timeline trigger preserves the open query. Closing restores focus to the current trigger or composer; switching session still resets the search.
  • R1-2: Exact user records can resolve to their verified live prompt alias. Assistant records load their own historical block (including a later page after a live user anchor), never the user alias. Navigation waits for an in-flight boundary load and retries retryable boundaries.
  • R1-60: IME committing Enter (keyCode=229, even with isComposing=false) does not select a result.
  • R1-62: Incomplete scans cannot erase established visibility counts. Healthy early-stop threshold probes still work; disconnected scans are skipped and reconnection reissues the probe/query.
  • R1-76: Reproduced cancellation losing the old historical reading page after exceeding the retention budget. Search now retains that page until successful navigation. The review's separate claim that successful navigation returns the wrong target was not established by its witness; the confirmed cancellation defect is fixed and regression-tested.

Validation at this commit:

  • 1,207 focused unit/integration tests passed across 9 files, including App, search dialog, historical viewport, navigation store/page table and public navigation hook.
  • 16/16 Chromium browser scenarios passed using synthetic daemon fixtures, including Escape/focus restoration, draft preservation and hidden-timeline public navigation.
  • Root build, typecheck and bundle passed; changed-file ESLint/Prettier and the real pre-commit hook passed.
  • All five Critical fixes have red-before/green-after regression evidence. Independent review found no remaining defect in this follow-up; two final self-audit passes were clean.

The remaining 58 Suggestion comments are deferred from this correctness follow-up: API/product choices (including per-message vs per-occurrence search and visibility rules), other non-blocking behavioral, accessibility/performance/UI suggestions, fixture/test expansion, and documentation/cleanup. They are not being marked fixed. Existing public API shape and timeline-linked entry visibility remain as specified. Maintainer product/API decisions still require maintainer judgment; this report is not a merge approval.

Local browser evidence is macOS Chromium with synthetic transport, not native IME hardware or Windows/Linux E2E. CI status for the new head is reported separately; previous green CI belongs to the previous commit.

中文:本轮仅处理 5 条 Critical,已补复现及回归验证;其余 58 条 Suggestion 明确后置,不声称已修复。保留原公开 API 与时间轴显隐约定,不扩展功能范围。R1-76 确认并修复的是取消搜索后丢失原历史阅读页,而非其证据未证明的成功定位误跳。

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Rechecked head c554b17f1876c941189bbb68658cce7a1b55d5ce: all executed CI checks passed, including unit/static, integration, desktop, visual capture and Web Shell smoke. Automatic re-review was queued; no new substantive review had been published when this pass began.

Continued through the deferred comments and fixed three concrete interaction issues in d6e494bf2149fd1a55f571a97c1bf7212b621172:

  • R1-64: Selection follows message identity when progressive history results prepend rows. Enter still navigates to the message the user selected instead of whatever happens to occupy its old index.
  • R1-63: A fulfilled but incomplete scan now offers the same Retry action as a rejected scan. Retry resubmits the unchanged query; the warning remains suppressed during an active progressive scan.
  • R1-83: The input/results expose combobox/listbox/option roles, controls, active-descendant and selected state. Arrow navigation retains input focus and accessibility state follows the selected message after history updates.

The existing browser keyboard scenario now checks those ARIA relationships, matches message 1 exactly (cannot pass on message 10), and is tagged @smoke so PR CI runs it. This addresses these specific test assertions; it does not claim all broader test-matrix suggestions are resolved.

Verification:

  • 40/40 focused component/viewport tests passed. The three new regressions each failed on the previous head before the fixes.
  • All 16 Chromium browser scenarios verified: 15 passed in the full run; the historical-result scenario initially used the old button role, was updated to option, and passed its targeted rerun. Production source was unchanged between those browser runs.
  • Root build, typecheck, bundle, changed-file lint/format, and the real pre-commit hook passed.
  • Independent review and two self-audit passes found no additional defect in this follow-up.

The previous five Critical fixes remain in place. Other deferred suggestions and maintainer decisions about public API/product scope remain open. This is not a merge approval; new-head CI is distinct from the green checks on c554b17f18. Browser tests use synthetic daemon data on macOS Chromium; no native screen-reader device session was performed.

中文:本轮继续修复 R1-64 选中项漂移、R1-63 不完整结果无法重试、R1-83 读屏语义缺失,并加强已有键盘 smoke 测试。40 项相关测试与 16 个浏览器场景通过;上一提交已执行 CI 全绿,新提交 CI 单独跟进。

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Rechecked the exact previous head d6e494bf2149fd1a55f571a97c1bf7212b621172 and the current review comments. The original five Critical fixes remain intact; no new substantive review was present.

Fixed the current CI blockers in 631d60265a8d52f02f2c94d10ace15f5b1a4ba77:

  • Merged current main fda4abad194882b8176a0bdba5fa77b6d275b07f after a clean merge-tree preflight. The failed lint lane explicitly required the newer pnpm lint gate; smoke and visual capture could not find the new pnpm store-cache action on the old branch. Both now use the current main files.
  • The previous CI install/build also exceeded the old document renderer budget (1,999,691 > 1,930,000 bytes). The merge incorporates main's existing fix(export): raise the document runtime budget for transcript growth #12298 budget correction, with no PR-specific budget increase. The new local build measures 1,916,436 bytes of renderer JS, below the 2,030,000-byte ceiling.
  • Found one remaining consumer of the old search input role in the visual screenshot scenario. Reproduced its 60-second timeout, changed only the locator from searchbox to combobox, and verified both themes.

Validation on the merged tree:

  • Root build, typecheck and bundle passed. The initial local dependency migration needed executable permissions restored on old npm bin targets; typecheck was rerun successfully after build completed.
  • 1,211 unit tests passed across nine App/search/history/navigation files.
  • All 16 Chromium search and external message-navigation scenarios passed; both dark/light visual scenarios passed and screenshots were inspected. Browser transport uses synthetic transcript data.
  • The real pre-commit hook ran Prettier and ESLint successfully. Its final staging step printed an ignored-directory warning; the commit succeeded, and comparison against the precomputed merge tree confirms the only extra change is the one-line visual locator fix. Working tree is clean.
  • Independent review and final self-audits found no additional necessary fix. The existing product/API scope decisions and deferred Suggestions remain open; this is not a merge approval.

New-head CI is pending and is not covered by these local results. Previous-head failed runs: CI, visual capture.

中文:本轮同步主干解决 pnpm CI action、过期 lint gate 和旧导出预算阻塞,并修正截图测试遗漏的输入框角色。1,211 项单测、16 个浏览器场景、暗亮两项截图、构建/类型检查/bundle 均通过;新提交 CI 单独跟进。

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Rechecked review round 2 against its exact head 631d60265a8d52f02f2c94d10ace15f5b1a4ba77. All executed product CI checks on that head passed (Ubuntu unit/static, no-AK integration, Linux/Windows desktop, Web Shell smoke and visual capture); macOS/Windows unit and CLI integration jobs were skipped, not counted as passed.

Addressed the confirmed correctness findings in e162f3b250d3ba7437a31e5706cc07fdc94f6868:

  • R2-1/R2-2: Exact-record history retains its live continuation when there is no remaining forward cursor. Navigation can reopen a trimmed live boundary using the existing bounded recovery path; an unrecoverable boundary still exits without spinning.
  • R1-76: A full or overlapping pinned reading window no longer wedges search navigation. On capacity failure, search reads toward the exact target without admitting intermediate pages, then releases the old pin only for the final synchronous admission. This deliberately avoids releasing the pin before an asynchronous retry: cancellation during the fallback preserves all five original reading pages. The target may be on a later page without the user turn's anchor block.
  • R2-5/R2-6: The internal viewport handle preserves a distinct cancelled result. The public API reports cancelled, while the dialog remains usable without a false failure alert or false successful close. Genuine false results remain errors.
  • R2-4/R2-10: Chat visibility is supplied by App's actual covering-panel/full-page/fullscreen conditions. Hiding chat dismisses and invalidates search. The host API returns not_ready while chat is hidden and cancels in-flight navigation if it becomes hidden; it does not silently change the host's selected view. The README documents that the host should restore chat first.
  • R2-7: Modified arrow/Enter combinations remain text/OS shortcuts. Plain-key navigation and both existing IME guards remain intact.
  • R2-8/R2-9: Disconnected search discloses loaded-only coverage rather than a definitive empty history result. In-flight host resolution/navigation failures caused by disconnect return not_ready; reconnect resumes dialog search.

Verification:

  • 1,235 tests passed across nine App/search/history/navigation files. The final legacy-mode correction additionally passes all 29 search component tests.
  • All 18 Chromium search/external-navigation scenarios passed, including two new covering-panel smoke scenarios; dark/light visual capture passed 2/2. The first browser pass caught an overly broad empty-state guard; it was narrowed and the full 18-scenario rerun passed without weakening assertions.
  • Root build/typecheck/bundle, changed-file ESLint/Prettier and real pre-commit hook passed. The document renderer remains 1,916,436 bytes of JS; no budget change is included.
  • Independent review plus two final self-audit passes found no further necessary fix. Fixtures are synthetic and no real model request or user transcript is involved.

Deferred R2-11/R2-12/R2-13/R2-14/R2-15 (trigger-render optimization and additional test-mutation coverage) with the existing non-blocking Suggestions, to keep repeated review rounds focused on demonstrated correctness defects. These are recorded, not silently treated as resolved. Product/public-API maintainer decisions remain open; this is not a merge approval. New-head CI is tracked separately.

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Reviewed round 3 against its exact head e162f3b250d3ba7437a31e5706cc07fdc94f6868. That head's executed product CI checks passed; skipped jobs were kept separate.

Fixed three demonstrated state-consistency defects in 09e07146dcec68bf45408d2b2ad85596c9c3a35c:

  • R3-1: Cancelling the current locate clears its pending selected turn, so the rail resumes following the actual reading position. Cleanup is gated by the selection generation: cancelling an older request cannot clear a newer loading or ready selection. The existing wheel-cancel test reproduced the stale loading state before the fix.
  • R3-15: If a capacity fallback discovers the exact record in the live window, it now publishes the successful location and clears its old locate error before returning success. The reproducer previously returned live while the store still reported unavailable and a window-full locate error.
  • R3-6: A viewport capacity failure restores its boundary to loadable before falling back, instead of leaving an error boundary after successful recovery. The reproducer confirmed that returning to the previously attempted turn after successful fallback otherwise retained an error edge that could not continue scrolling. A real offline boundary failure remains an error with retry semantics. Non-viewport load behavior is unchanged.

R3-6 and R3-15 were labeled Suggestions, but were included because runtime reproduction confirmed incorrect state and a blocked history continuation, not merely missing test coverage. The rest of the new Suggestions remain deferred under the convergence policy: R3-2 through R3-5, R3-7 through R3-14, and R3-16 through R3-23, plus the previously deferred R2-11/R2-13/R2-15. No new public API or unrelated cleanup is included.

Validation on the final tree:

  • Five focused search/store/page-table/pending-viewport/turn-rail test files: 165/165 passed. The cancellation and fallback state regressions failed before the fixes; newer loading/ready selections and genuine network failures have positive controls.
  • Chromium search and external navigation: 18/18 passed. Dark/light visual capture: 2/2 passed. All browser fixtures are synthetic.
  • Root build, typecheck, bundle, ESLint/Prettier and the real pre-commit hook passed. Hook output did not change the verified source tree. Document renderer JS remains 1,916,436 bytes.
  • Independent review and two self-audit passes found no additional demonstrated defect. Working tree is clean; this commit changes one production file and three regression-test files.

New-head CI is tracked separately from the preceding head's green results. Existing maintainer/product/API review decisions remain open; this is not a merge approval.

@samuelhsin

samuelhsin commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator Author

Rechecked exact head 09e07146dcec68bf45408d2b2ad85596c9c3a35c.

No new published actionable CR findings or additional confirmed correctness defects were found in this pass. Re-read the latest cancellation/recovery changes and their callers; R3-1 remains fixed despite its thread still being marked unresolved. R3-6 and R3-15 recovery fixes also remain intact. Previously recorded Suggestion deferrals remain unchanged.

Fresh local verification: all 165 tests in the five focused navigation/search/viewport/page-table suites passed. Both the CI run and visual capture run completed successfully on this exact SHA: Ubuntu unit tests, lint/static checks, no-AK integration, Linux/Windows desktop builds, Web Shell E2E smoke, and visual capture. macOS/Windows unit tests and CLI integration were skipped, not counted as passes. On the subsequent recheck, the separate automated review-pr job has moved from queued to in progress (run 35553822329, same SHA), with no new published CR findings; this update does not claim a new automated review result or maintainer approval.

Local HEAD, tracking branch and remote branch match, divergence is 0 0, the working tree is clean, and GitHub reports MERGEABLE. No source changes or new commit were needed this round.

@wenshao

wenshao commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

Round 2 (delta, real daemon) — the feature works end to end against a real qwen serve, but I found one real defect that none of the mock-daemon suites can see. I'd like it fixed before merge; it is small and local. Verified head 09e07146dcec68bf45408d2b2ad85596c9c3a35c, tested as it would land: merged onto main df3f9732a6 (tree 5a91871a). Round 1 is here; this comment only covers what that round did not.

中文说明

第 2 轮(增量,真实 daemon)—— 功能在真实 qwen serve 上端到端可用,但发现了一个 mock daemon 测试看不到的真实缺陷。建议合并前修掉,改动小且局部。 验证 head 09e07146dc,按"实际落地的样子"测试:合并到 main df3f9732a6 之上(tree 5a91871a)。第 1 轮见这里,本评论只写上一轮没覆盖的部分。

本轮新增了什么

第 1 轮(head 9f8c93e)以及 PR 自带的全部浏览器用例,用的都是 Playwright 路由层的假 daemon。本轮去掉了所有 mock:Chromium → daemon 自带的生产版 Web Shell → 真实 qwen serve(从合并树打包)→ 真实 chats/<id>.jsonl。会话由 600 个真实轮次播种(daemon 后面接一个确定性的 OpenAI 兼容假模型),共 1200 条消息、2.8 MB JSONL;浏览前重启 daemon,保证走的是冷加载路径。每个断言的期望值都从磁盘上的 JSONL 直接算出。

结果

场景组 通过
长会话:搜索 + 定位(600 真实轮次,冷加载) 25/25
回放路径下的阈值 + 窄屏 4/4
公开 navigateToMessage + conversationSearchThreshold(真实 record id) 18/18
旧 daemon 降级路径(第 1 轮未覆盖) 6/6
本页实时输入 / 工具调用 / 第二客户端 探针 8/10
合计 61/63 —— 两个失败是同一个缺陷

真实 daemon 上确认可用的:小写查询命中第 3 轮里的大写文本(在初始渲染窗口之外 570 轮);中文;代码块内的 token;a+|b(c)*[d]?\.$^{2} 按字面量匹配、.* 返回 0 条;两轮里完全相同的文本得到两条独立命中,ArrowDown+Enter 落到第 200 轮而不是第 10 轮,且 daemon 收到的 atRecordId 正是第 200 轮的 uuid;600 条匹配 → 恰好 200 行 + 上限提示;草稿保留;Escape 后焦点回到入口;1200 条消息的会话上阈值在 1199/1200 处精确翻转。公开 API 用磁盘上的真实 uuid:located / not_found(未知 id、空白 id、system 记录 id)/ session_mismatch / cancelled(已 abort 的 signal;被更新请求取代)全部符合文档,相同文本按身份而非文本解析,12 次调用后宿主草稿不变。

发现 1(缺陷)—— 在本标签页输入的消息会被重复计数、重复列出

现象。 在全新会话里,所有轮次都在当前标签页输入:

已输入轮次 1 2 3 4 5 6
真实消息数 2 4 6 8 10 12
入口是否可见 否 否 否 是 是 是
PR 声明(> 10) 否 否 否 否 否 是

入口在 8 条消息时就出现,而不是第 11 条。同一会话、同样 8 条消息,刷新之后入口又消失(图 1)。同一原因还导致搜索结果重复:9 条用户消息的会话全部从磁盘回放时是 9 行;在本页再发 2 轮后变成 13 行(应为 11),状态栏显示 3 / 13,两条本页输入的消息各出现两次(图 2);刷新后恢复 11 行。带工具调用的轮次里,工具调用之前的助手文本同样出现两次(图 3)—— 对一个编码 agent 来说这是最常见的轮次形态。

根因(已在真实数据上插桩确认,插桩已移除)。 本页输入的 live 用户块是 { kind: 'user', text },既没有 sourceRecordIds,也没有 promptId;而它的持久化孪生记录两者都有(promptId 来自 daemonPromptId)。从 DaemonSessionProvider.tsx 看,live 轮次里只有最后一条助手记录会被打标(通过 assistant.done 的 branchPoint.assistantRecordUuid),这与实测一致:工具调用后的文本出现一次,调用前的文本出现两次。于是:

  • scanConversation 里的 unstampedByPrompt 回声匹配永远命中不了 —— live 侧 promptKey(block) 是 undefined,根本没被登记 —— 这个块留在 remainingLive 里,又被 messageCount + remainingLive.size 加了一次。k 轮实时输入后计数是 3k 而不是 2k,3·4 = 12 > 10。
  • ConversationSearch.tsx 的 results 只按 record id(liveRecords)去重,持久化命中和 live 命中都保留。

这正是 bot 第 1 轮 reverse-audit 标注为"未探索"的那个前提(daemon 是否给持久化用户块写 promptId)—— 写了,但缺的是 live 那一侧。conversation-search.test.ts 里的相关用例(counts the unpersisted live tail without double counting persisted blocks or local echoes)是手工给 local-user 块设了 promptId: 'prompt-1' 才通过的。这不是合并 main 引入的:构造 live 块的 packages/sdk-typescript/src/daemon/ui/transcript.ts 在 PR 的 merge-base 与 main 之间没有变化。

影响面。 没有错误导航 —— 两个孪生行点击后都落到同一条 live 消息并高亮,无多余请求,也不涉及数据。但它与 PR 的首要声明("10 条隐藏、11 条显示")在最常见的使用路径上不符,且重复行在每个实时会话里都可见。

修复方向。 我试了最直观的 7 行补丁(把 store 的回声匹配结果带到 hit 上,组件里过滤)—— 不起作用,因为回声匹配本身就没命中。需要让 live 用户块拿到身份:要么在 POST /session/:id/prompt 的 202 返回 promptId 时回填到本地回声块上(之后 PR 已有的回声层对计数就生效了,再加上组件侧过滤即可消除重复行),要么让回声层对无 key 的 live 块退回到"角色 + 文本 + 顺序"匹配。无论哪种,请加一个 live 块不带 promptId 的用例,这才是真实形态。工具调用前的助手片段需要同样处理。

观察(不阻塞)

  1. 200 行上限保留的是最旧的匹配,丢掉最新的。 [...older, ...live].slice(0, 200) 加上从 ordinal 0 开始扫描:600 条匹配时显示的是第 1–200 轮,live 窗口里的命中全部被截掉。长会话里搜常见词,最相关(最近)的结果反而不可达。建议保留最新 200 条或倒序扫描。
  2. 每次查询都重读整段历史。 600 轮时每个防抖后的查询是 21 次 /transcript + 4 次 /turn-index,连续 12 个查询每次都完整重发,无复用。本机回环约 0.7 秒出结果;高延迟的远程 daemon 上会线性变慢。作者已声明这是接受的取舍;对话框打开期间按 snapshot 缓存物化文本可以很便宜地解决。
  3. 变异测试中 1 个存活:阈值探针上的 result.complete || result.messageCount > limit 守卫(R1-62 的探针侧)去掉后 219 个测试全绿。搜索侧的同类守卫是被杀死的。这个守卫在真实 store 上基本不可达(只有 legacy/无 client 才会返回不完整且计数偏小的结果,而此时探针已被 mode !== 'ready' 挡住),属于防御性代码缺测试,不是缺陷。
  4. 摘要显示的是原始 Markdown(能看到 ```ts 围栏)。纯外观问题。

第 1 轮遗留项

  • 键盘/上一条/下一条交互当时无测试 → 现在有单测 + @smoke 浏览器用例,且我在真实 daemon 上用 ArrowDown+Enter 复验通过。已关闭。
  • 旧 daemon 降级路径未在浏览器里驱动过 → 本轮用一个透明代理从 /capabilities 里剥掉 session_turn_navigation:页面正常加载,入口出现在 legacy 时间轴下方,对话框显示 "Only loaded messages are available. Update Qwen Code to search all history.",只搜已加载消息,深层历史的 needle 如实返回无结果(图 6)。已关闭。
  • 超过 200 条命中时的 UI → 已观察,见观察 1。

针对第 1 轮之后修复提交的变异矩阵

对 427bcc6、c554b17、d6e494b、e162f3b、09e0714 的修复逐一做单点回退,跑 PR 自带的 7 个套件(每次 219 个测试)。对照组全绿;10 个变异杀死 9 个,每个都被为它而写的那条测试精确杀死(R3-1、R3-15、R3-6、427bcc6 的 generation 顺序、R1-60 IME、R1-62 搜索侧、R1-63、R1-64、R2-5)。唯一存活者见观察 3。这些修复是被测试钉住的,不是"恰好通过"。详见图 7。

合并树上的门禁

  • 合并新鲜度:main 自 merge-base fda4abad 起前进了 45 个提交,其中 21 个涉及 packages/web-shell;与 PR 重叠的 5 个文件(App.tsx、App.module.css、App.test.tsx、README.md、visuals/screenshots.spec.ts)全部自动合并无冲突。
  • pnpm install --frozen-lockfile 32 秒,根目录 npm run build 3 分 52 秒,npm run bundle 通过。
  • PR 自带的 Playwright 用例:18/18(Linux x64,Chromium 149)。相邻回归用例(history-viewport、turn-navigation-scroll-follow、smoke、session-overview):50/50。
  • packages/web-shell 全量 vitest:9490/9491。唯一失败是 BranchPickerPopover.test.tsx 里一个焦点断言,PR 未触碰该文件,隔离重跑 3/3 通过 —— 是我并行跑 Playwright 造成的负载抖动。
  • tsc -p tsconfig.json --noEmit 退出 0;24 个改动的 TS/TSX 文件 eslint --max-warnings 0 退出 0;全部改动文件 Prettier 检查通过。

未覆盖

  • 真实 Windows/macOS 桌面;WebKit/Firefox。
  • 真实模型(假模型是确定性的;本缺陷发生在客户端,与模型无关)。
  • 会话压缩(compaction)之后的搜索、子代理轮次、图片/附件消息。
  • 第二客户端发来的用户消息是否也会重复(我只断言了它的助手回复恰好出现一次)。
  • 真实 daemon 下的 base A/B —— 第 1 轮已证明 base 上没有该功能,未重复。
  • bot 对当前 head 的自动 review 在我发布时仍在运行中。

方法

Linux x86_64,Node 22.22.2。合并树工作区:真实安装 + 根构建 + 打包;daemon 使用隔离的 HOME/QWEN_HOME。浏览器用例跑的是 daemon 自带的生产构建(dist/web-shell),只有公开 API/阈值两组用 vite dev 代理到同一个 daemon 并挂一个仅用于验证的嵌入宿主(未提交,已删除)。变异用文件备份/恢复的方式就地改动,每次之后确认 git status 干净。诊断插桩和候选补丁在单独的输出目录里构建,已恢复;所有报告的结果都来自原始 PR 构建(index-CzdMsmE2.js)。脚本与原始结果:r2/harness、r2/results。

What this round adds

Round 1 (head 9f8c93e) and every browser test shipped in the PR run against a Playwright route-level fake daemon. This round removes every mock: Chromium → the daemon-served production Web Shell → a real qwen serve bundled from the merged tree → a real chats/<id>.jsonl. The session was seeded with 600 real turns (a deterministic OpenAI-compatible model behind the daemon): 1200 messages, 2.8 MB of JSONL. The daemon was restarted before browsing, so the browser exercises the cold-load path. Every expected value is computed from the JSONL on disk.

Results

Scenario group Passed
Long session: search + locate (600 real turns, cold-loaded) 25/25
Threshold on the replay path + narrow layout 4/4
Public navigateToMessage + conversationSearchThreshold (real record ids) 18/18
Older-daemon degradation (not covered in round 1) 6/6
Typed-live / tool-call / second-client probes 8/10
Total 61/63 — both failures are the same defect

Confirmed on a real daemon: a lower-case query finds upper-case text in turn #3, 570 turns outside the initially rendered window; Chinese; a token inside a code block; a+|b(c)*[d]?\.$^{2} matches literally and .* returns nothing; identical text in two turns yields two distinct hits, ArrowDown+Enter lands on turn #200 rather than #10, and the atRecordId the daemon receives is turn #200's uuid; 600 matches → exactly 200 rows plus the limit notice; the draft survives; Escape returns focus to the trigger; on the 1200-message session the threshold flips exactly between 1199 and 1200. The public API, driven with real uuids read from disk, returns located / not_found (unknown id, blank id, a system record id) / session_mismatch / cancelled (pre-aborted signal; superseded by a newer request) as documented, resolves identical-text records by identity, and leaves the host draft untouched after 12 calls.

Finding 1 (defect) — messages typed in this tab are counted and listed twice

Symptom. A brand-new session, every turn typed in the current tab:

turns typed 1 2 3 4 5 6
real messages 2 4 6 8 10 12
entry visible no no no yes yes yes
claimed (> 10) no no no no no yes

The entry appears at 8 messages, not at the 11th. In the same session with the same 8 messages, it disappears again after a reload (figure 1). The same cause duplicates search rows: a session with 9 user messages shows 9 rows when everything is replayed from disk; after two more turns typed in the tab it shows 13 rows (expected 11), the status reads 3 / 13, and each message typed in the tab appears twice (figure 2); a reload brings it back to 11. In a tool-call turn, the assistant text before the tool call is listed twice as well (figure 3) — for a coding agent that is the dominant turn shape.

Root cause (confirmed by instrumenting the scan on real data; instrumentation removed). The live user block typed in this tab is { kind: 'user', text } — it has neither sourceRecordIds nor promptId — while its persisted twin has both (promptId from daemonPromptId). As far as I can tell from DaemonSessionProvider.tsx, within a live turn only the final assistant record gets stamped (through assistant.done's branchPoint.assistantRecordUuid), which matches what I measured: the post-tool text is listed once, the pre-tool text twice. Consequently:

  • the unstampedByPrompt echo tier in scanConversation can never match — promptKey(block) is undefined on the live side, so the block is never registered — and the block stays in remainingLive, where messageCount + remainingLive.size adds it a second time. After k typed turns the count is 3k instead of 2k, and 3·4 = 12 > 10;
  • results in ConversationSearch.tsx de-duplicates by record id only (liveRecords), so the persisted hit and the live hit are both kept.

This is the premise the bot's round-1 reverse audit listed as unexplored (whether the daemon puts promptId on persisted user blocks) — it does; the missing side is the live one. The covering case in conversation-search.test.ts (counts the unpersisted live tail without double counting persisted blocks or local echoes) passes because it hand-sets promptId: 'prompt-1' on the local-user block. This is not a merge interaction: packages/sdk-typescript/src/daemon/ui/transcript.ts, which builds live blocks, is unchanged between the PR's merge-base and main.

Blast radius. No wrong navigation — both twin rows land on the same live message and flash it, with no extra requests — and no data is touched. But it contradicts the PR's headline claim ("hidden at 10 messages and appears at 11") on the most common path, and the duplicate rows are visible in every live session.

Fix direction. I tried the obvious 7-line patch (carry the store's echo match onto the hit and filter in the component) — it does not help, because the echo tier never matches in the first place. The live user block needs an identity: either backfill promptId onto the local echo when POST /session/:id/prompt answers 202 with it (the echo tier already in this PR then fixes the count, and a component-side filter removes the duplicate rows), or let the echo tier fall back to role + text + order for keyless live blocks. Either way, please add a case whose live block has no promptId — that is the real shape. Pre-tool assistant segments need the same treatment.

Observations (non-blocking)

  1. The 200-row cap keeps the oldest matches and drops the newest. [...older, ...live].slice(0, 200) plus a scan that starts at ordinal 0: with 600 matches the dialog shows turns pre-release: fix ci #1–希望可以加入类似claude code 的sub-agent系统 #200 and every live-window hit is cut. Searching a common word in a long conversation makes the most relevant (recent) results unreachable. Keeping the newest 200, or scanning newest-first, would fit the use case better.
  2. Every query re-reads the whole history. With 600 turns each debounced query costs 21 /transcript + 4 /turn-index requests, and 12 consecutive queries each paid it in full. On loopback the result is up in about 0.7 s; it will grow linearly with latency on a remote daemon. The author declared this tradeoff; caching the materialized text per snapshot while the dialog is open would remove it cheaply.
  3. One surviving mutant: the result.complete || result.messageCount > limit guard on the threshold probe (the probe side of R1-62) can be removed with all 219 tests green. Its search-side sibling is killed. The guard is close to unreachable with the real store (only legacy/no-client returns an incomplete, undercounted result, and the probe is already gated on mode === 'ready'), so this is untested defensive code, not a defect.
  4. Snippets show raw Markdown (the ```ts fence is visible). Cosmetic.

Round-1 leftovers

  • Keyboard / previous / next had no test → now covered by unit tests and a @smoke browser case, and re-confirmed here with ArrowDown+Enter on a real daemon. Closed.
  • The older-daemon path had never been driven in a browser → a transparent proxy that strips session_turn_navigation from /capabilities: the page loads, the entry sits under the legacy rail, the dialog says "Only loaded messages are available. Update Qwen Code to search all history.", only loaded messages are searched, and a deep-history needle honestly returns nothing (figure 6). Closed.
  • UI with more than 200 hits → observed; see observation 1.

Mutation matrix for the fixes pushed after round 1

Single-point reverts of the fixes in 427bcc6, c554b17, d6e494b, e162f3b and 09e0714, run against the PR's own 7 suites (219 tests per run). Control green; 9 of 10 mutants killed, each by the test written for it (R3-1, R3-15, R3-6, the 427bcc6 generation ordering, R1-60 IME, R1-62 search side, R1-63, R1-64, R2-5). The one survivor is observation 3. These fixes are pinned by their tests rather than merely passing. See figure 7.

Gates on the merged tree

  • Merge freshness: main moved 45 commits past merge-base fda4abad, 21 of them touching packages/web-shell; the 5 files both sides changed (App.tsx, App.module.css, App.test.tsx, README.md, visuals/screenshots.spec.ts) auto-merge cleanly.
  • pnpm install --frozen-lockfile 32 s, root npm run build 3 m 52 s, npm run bundle OK.
  • The PR's own Playwright specs: 18/18 (Linux x64, Chromium 149). Adjacent regression specs (history-viewport, turn-navigation-scroll-follow, smoke, session-overview): 50/50.
  • Full packages/web-shell vitest: 9490/9491. The single failure is a focus assertion in BranchPickerPopover.test.tsx, a file this PR does not touch; it passes 3/3 in isolation — load jitter from running Playwright in parallel.
  • tsc -p tsconfig.json --noEmit exit 0; eslint --max-warnings 0 on the 24 changed TS/TSX files exit 0; Prettier clean on every changed file.

Not covered

  • Real Windows/macOS desktops; WebKit/Firefox.
  • A real model (the fake one is deterministic; the defect is client-side and model-independent).
  • Search after compaction, subagent turns, image/attachment messages.
  • Whether a user message sent by a second client is also duplicated (I only asserted that its assistant reply appears exactly once).
  • A base-arm A/B on the real daemon — round 1 already proved the feature is absent at base; not repeated.
  • The bot's automated review of the current head was still running when I posted this.

Evidence

Figure 1 — same session, same 8 messages: entry visible when typed live, hidden after reload

threshold live vs reload

Figure 2 — 11 user messages, 13 rows: the two typed in this tab are listed twice

duplicate rows

Figure 3 — tool-call turn: the pre-tool assistant text is listed twice, the post-tool text once

tool call turn

Figure 4 — real daemon, 600 turns: lower-case query, upper-case hit 570 turns outside the rendered window, draft intact

deep history hit

Figure 5 — identity, not text: ArrowDown+Enter on the second of two identical-text hits lands on turn #200 with the flash highlight

located turn 200

Figure 6 — older daemon (capability stripped by a proxy): loaded-only notice, honest empty result

legacy loaded only

Figure 7 — scoreboard, live-threshold table and mutation matrix

scoreboard

More: 600 matches capped at 200 · public API landing on turn #3 with the timeline hidden

Methodology

Linux x86_64, Node 22.22.2. Merged-tree worktree with a real install, root build and bundle; the daemon runs with an isolated HOME/QWEN_HOME. Browser scenarios run the daemon-served production build (dist/web-shell); only the public-API and threshold groups use vite dev proxied to the same daemon, with a verification-only embedding host (uncommitted, since deleted). Mutants were applied in place with file backup/restore and git status confirmed clean after each. The diagnostic instrumentation and the candidate patch were built into separate output directories and reverted; every reported result comes from the pristine PR build (index-CzdMsmE2.js). Scripts and raw results: r2/harness, r2/results.

🤖 Generated with Claude Code — Claude Fable 5.1

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Addressed the real-daemon Finding 1 in 6fd5e61c1c6b530bc9923599866aeb9ff8320e86.

Messages typed in the current tab now use their existing prompt-admission/turn identity to pair the local echo with its persisted record. Pre-tool assistant messages use the exact containing turn's prompt ID when replay does not carry it. Search removes the paired history row only when the corresponding live search result is present, preserving distinct same-text turns and unpersisted messages. Assistant text is checked within the identified prompt so a retained later fragment cannot hide a different earlier fragment.

The real-daemon check also exposed a stale threshold probe: a still-streaming assistant could be counted before its record identity arrived, and the cached count was not refreshed when the block count stayed unchanged. Identity/completion changes now rerun the probe; text-only streaming deltas do not.

Validation on this change:

  • Root build, typecheck and bundle; focused ESLint/Prettier; normal commit hook.
  • 204 focused unit tests; 18 existing Chromium search/external-navigation scenarios.
  • Real qwen serve with a deterministic synthetic model and synthetic workspace: 2/4/6/8/10 messages keep the search entry hidden, the 11th user message shows it while its assistant reply is still pending; six live user prompts produce six hits; pre-tool and post-tool assistant text each produce one hit; navigation flashes the correct message; reload preserves the result counts. Repeated successfully against the final bundled daemon-served production UI, without a Vite proxy or mocked daemon routes.

The new real-daemon review's non-blocking observations 1–4 (latest-first cap, caching, defensive test coverage and Markdown snippets) remain deferred. No SDK or daemon protocol changes were needed. The new-head product CI and visual capture are running; no new-head CI pass is claimed yet. Local/tracking/remote SHA parity is confirmed with 0 0 divergence and a clean working tree.

10 real messages: search entry hidden

10 real messages: search entry hidden

11th user message, assistant pending: search entry visible

11th user message, assistant pending: search entry visible

Pre-tool assistant text: exactly one result

Pre-tool assistant text: exactly one result

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Refreshed CR and CI against 6fd5e61c1c6b530bc9923599866aeb9ff8320e86. The new review confirms R3-1 is fixed. R4-1/R4-2 are additional coverage/readability Suggestions; their deferrals are recorded in their threads under the repository rule for long-running review rounds. No new demonstrated product defect was found.

Fixed the current CI blocker in 4df4241bc8f40ddc2f8738b24d349334ed6905e5: Lint & Static required the newer main lint gate. Merged upstream main 2800e9bb4f4536f421116a8e1e12740113e1063f without conflicts. This is a pure baseline merge; the committed tree matches the merge preflight, with no handwritten product changes. Re-ran the repository freshness-check logic using gh-backed API reads against the pushed SHA: stale gate files: none.

Validation on the merged tree:

  • Root build, typecheck and bundle passed with the pinned pnpm 11.24.0 build workflow.
  • Seven focused search/navigation/App suites: 1,230 tests passed.
  • Existing Chromium search and external-navigation scenarios: 18/18 passed.
  • Final bundled, daemon-served production UI with synthetic data and a deterministic model: 16 scenario assertions passed, including exactly 10 messages hidden / the 11th user message visible before its assistant reply, live and pre/post-tool result deduplication, correct navigation flash, and consistent results after reload; no browser exceptions.
  • Two read-only merge reviews, diff checks and the normal commit hook passed. The hook did not change the validated tree.

Local HEAD, tracking branch, and remote branch match (0 0 divergence); working tree is clean and GitHub reports MERGEABLE. The new-head CI and visual capture are running, not yet claimed as passed. macOS/Windows unit and CLI integration lanes are skipped, not counted as passes.

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Rechecked exact head 4df4241bc8f40ddc2f8738b24d349334ed6905e5 after the baseline merge.

Product CI and visual capture both completed successfully on this exact SHA. Lint/static checks, Ubuntu unit tests, no-AK integration, Linux/Windows desktop builds, and Web Shell E2E smoke passed. macOS/Windows unit and CLI integration lanes were skipped, not counted as passes. The prior lint-gate freshness blocker is cleared.

Checked all 103 review threads: no unresolved Critical threads, and no new published review findings since the prior update. The separate automated review is still running; this is not a claim that it has completed or that maintainer approval has been granted. Previously documented Suggestion deferrals remain unchanged.

No additional code change was needed. Local/tracking/remote SHA parity is confirmed with 0 0 divergence, the working tree is clean, and GitHub reports MERGEABLE. Existing local verification remains tied to this unchanged tree: 1,230 tests, 18 browser scenarios, and 16 real-daemon production assertions; those were not redundantly rerun this pass.

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Rechecked the R5 review against the current branch and fixed the two demonstrated regressions:

  • R5-1 / R5-2: streaming transcript updates now use the existing animation-frame snapshot, and the timeline search trigger keeps its identity across text-only updates and search typing. The regression test failed before the fix because the memoized timeline rendered again in both cases; it now remains stable. Identity and completion updates remain enabled so the exact message threshold can still refresh.
  • R5-4: when a paginated search resolves to a live message, both successful live exits now publish that resolved selection. The regression reproduced a live return with a stale historical selection; the selected location now matches the returned live location, releasing the old selected page through the existing store logic.

R5-3 follows the updated entry-placement requirement in #12231: the search icon is hidden whenever the timeline is hidden. No fallback icon was added. R5-5 and previously deferred Suggestions remain follow-up work under the repository's repeated-review scope rule.

Validation: root build, typecheck and bundle passed; 1,252 focused and App tests passed (including the new red-to-green regressions), 18 Chromium search/public-navigation scenarios passed, and the newly built production bundle passed 16 real-daemon assertions with zero page errors. These cover 10 hidden / 11 visible before the assistant reply, live and pre-/post-tool deduplication, correct result navigation, and reload consistency. Model responses and workspace content were synthetic; daemon persistence and tool execution were real. Full diff received two clean self-audit passes and independent review. Skipped or pending remote checks are not counted as passes.

Fresh sanitized production screenshots:

11 messages: entry is visible while assistant is pending

Pre-tool assistant search returns exactly one result

Delivered as 9826676 with normal pre-commit hooks. Local HEAD, tracking branch and remote branch match, divergence is 0/0, and the worktree is clean. New-head CI is still pending; previous-head green checks are not reused as proof for this commit.

@samuelhsin

samuelhsin commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator Author

Rechecked current head 9826676, all 108 review threads, and the latest comments. No new actionable finding has arrived since the R12 fix. R5-1/2 and R5-4 are fixed in the current code; R5-3 remains the explicit entry-placement requirement documented in #12231. Deferred Suggestions remain unchanged. No additional code changes were needed.

Exact-head remote verification is now complete: product CI and visual capture both succeeded for this SHA. Executed product checks passed: Ubuntu unit tests, lint/static checks, no-AK integration, Linux/Windows Desktop Shell, and Web Shell E2E smoke. macOS/Windows unit lanes and CLI integration were skipped, not passed. Latest recheck: the automatic review has started and is currently executing its Run review step on this exact SHA. It has not published a new result yet, so this is not a claim that the new review has completed.

The PR is open and mergeable. Local HEAD, tracking branch and remote branch match, divergence is 0/0, and the worktree is clean. Freshness verification reports no stale gate files. Previous local test and synthetic production evidence remains in the R12 report; no tests were unnecessarily rerun on an unchanged tree.

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Fixed the disconnect race reported in R5-4. A search waiting on an independently started boundary load could restore a cleared selection after disconnect; reconnecting with the same client owner also reproduced it. Search now verifies both the client owner and the connection epoch before continuing. The same race was reproduced and fixed in the full-window fallback path, which captures its epoch before trying the normal walk so recovery cannot bind to a newer connection.

Three regression cases cover staying disconnected, reconnecting with the same client, and reconnecting during fallback. Each originally failed. Removing the walk guard makes both boundary-wait cases fail again; removing the fallback epoch guard makes the fallback case fail again. Restoring the guards passed all 147 search/navigation/page-table tests. Independent review and two self-audit passes found no further required changes. Existing Suggestions remain deferred under the repeated-review scope rule.

Additional validation: root build, typecheck and bundle passed; 18 Chromium search/public-navigation tests passed; the newly built production daemon passed 16 synthetic search scenarios with zero browser errors (10/11 threshold, live and tool-message deduplication, navigation and reload). The browser run is normal-flow regression evidence; the exact disconnect interleavings were verified by the real-store regression tests and guard mutations above. No UI changes or new screenshot claims are involved.

Delivered as 0c73062. Normal commit hooks passed without changing the validated tree; normal push succeeded. Local, tracking and remote SHAs match with 0/0 divergence and a clean worktree. New-head remote CI is pending and is not counted as passed.

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Rechecked head 0c73062 and all 109 review threads. No new review findings have arrived since the disconnect fix; automatic review is still running.

The current-head CI smoke run failed in mobile history search restores results and draft focus on mobile WebKit: the input-history search field was visible but not focused at the first focus assertion, including both retries. The other 139 smoke cases passed. Other executed product checks and visual capture passed; skipped jobs are not counted as passes.

Reproduction attempt on this exact head: the targeted mobile-WebKit case passed once, then passed five consecutive serial repetitions locally (6/6 total). The composer and this test are unchanged relative to the merged main baseline; the last commit only changed navigation lifecycle guards and their tests. This does not rule out a Linux-specific or timing-dependent defect, and the cause is not yet confirmed. No speculative source change was made.

Requested a rerun of the failed job on the same SHA: CI attempt 2 is in progress. The original failure is not considered fixed until remote verification completes. Local/tracking/remote SHAs match, divergence is 0/0, and the worktree remains clean.

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Round 7 confirmed the disconnect fixes and posted no new blocking findings. The previously failed mobile-history focus smoke also passed on the unchanged 0c73062 head in CI attempt 2. Existing review Suggestions remain deferred.

GitHub then reported a merge conflict after main advanced. Merged freshly fetched main 94104d5, retaining both sets of added App imports and both README sections (conversation search/message navigation and URL navigation). These were the only conflict hunks. The feature diff against main remains 29 files; this merge adds no further search features. Independent review checked the auto-merged App, message list, viewport and exported API interactions with upstream URL navigation.

Validation on the merged tree: frozen dependency installation and all 1,978 lockfile policy checks succeeded; root build, typecheck and bundle passed; 1,279 tests across nine search/navigation/App suites passed; 29 Chromium search/message-navigation/URL-navigation tests and the targeted mobile-WebKit history test passed. The newly built production daemon also passed 16 synthetic threshold/deduplication/navigation/reload assertions with zero browser errors. Existing search behavior and public message navigation remain intact alongside upstream URL navigation.

Delivered merge commit 391cba6. Normal commit hooks passed and preserved the validated tree exactly. Local HEAD, tracking and remote branch match, divergence is 0/0, and the worktree is clean. New-head CI is pending; prior-head successes are not counted as new-head verification.

@ytahdn ytahdn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

结论:静态审查(不跑测试)通过,未发现 Critical / Important。核心异步/取消状态机与搜索正确性站得住。

我独立核对过的关键点:

  • 触发搜索的 effect 有 250ms 防抖 + clearTimeout,且 onProgress/then/catch/finally 全部走 current 守卫;关闭对话框或换词时的旧请求不会覆盖新结果,取消也不会把视口卡在 loading。ci-bot 关注的 stale-response 竞态确实被处理。
  • 结果高亮为纯字面量、大小写不敏感:createConversationSearchSnippet 先转义正则元字符再用 /iu 匹配,渲染走 JSX 文本切片(无 dangerouslySetInnerHTML / innerHTML),不存在正则注入或 XSS。
  • 默认阈值 10(10 条隐藏、11 条出现)在常规 1:1 场景下与 README/描述一致。
  • 跳转复用既有历史视口 + 缓存预算;分页/淘汰与取消/回滚路径经两个模块分别追踪,未见破坏实时转录或泄漏状态。

非阻塞建议(Suggestion,不影响合并):

  1. 命中数超过 200 时的截断方向:scanConversation 从最旧往最新填充 hits,到 200 即停并置 truncated;UI 里 [...older, ...live].slice(0, 200) 在 older 已占满 200 时会把当前可见窗口里的 live 命中也切掉。虽然有 truncated 提示、且需要单条会话 >200 次命中才触发(属极端场景),但这与本功能“在长会话里找回刚看过的那条”的初衷相反。可考虑改为最新优先扫描,或在 merge 时始终保留 live 命中。
  2. 阈值计数在“块 vs 消息”上不一致:liveCount 按块计、持久化 count 对相邻同 recordId 去重、且 liveCount>limit 时跳过精确扫描。行为有界且与搜索实际展示自洽,但概念口径建议统一。
  3. 本轮 head 前进为一次 “Merge upstream main to resolve navigation integration conflicts”。核对合并后:本 PR 的核心搜索实现(ConversationSearch.tsx、turn-navigation-store.ts、transcript-page-table.ts、useMessageNavigation.ts)与合并前逐字节一致,上述结论在新 head 上仍成立;App.tsx / MessageList.tsx / useTranscriptViewport.ts 的差异来自上游(idle-settle streaming、shared-URL 导航等),本 PR 的搜索入口挂载与 scrollToSearchHit 管线在合并后保持完好。

Verdict: static review (no test run) passes — no Critical or Important issues. The async/cancel state machine and search correctness hold up.

What I independently verified:

  • The search-triggering effect is debounced (250ms + clearTimeout) and every onProgress/then/catch/finally is guarded by a current flag, so a stale response can't overwrite a newer query and cancellation can't leave the viewport stuck loading — the ci-bot's stale-response race concern is genuinely handled.
  • Highlighting is literal and case-insensitive: createConversationSearchSnippet escapes regex metacharacters before matching with /iu and renders via JSX text slicing (no dangerouslySetInnerHTML / innerHTML), so there is no regex injection or XSS.
  • The default-10 threshold boundary (hidden at 10, shown at 11) matches the README in the normal 1:1 case.
  • Jump-to-message reuses the existing historical viewport and cache budget; pagination/eviction and cancel/rollback paths were traced across two modules with no live-transcript corruption or leaked state.

Non-blocking suggestions:

  1. Truncation direction over the 200-cap: scanConversation fills hits oldest-first and sets truncated at 200, and the UI merge [...older, ...live].slice(0, 200) drops on-screen live matches once older fills the budget. There is a truncation banner and it needs >200 matches in one conversation to bite (an edge case), but this is opposite to the feature's goal of finding a recently-seen answer in a long conversation. Consider scanning newest-first, or always keeping live matches in the merge.
  2. Threshold counting is inconsistent between blocks and logical messages (liveCount counts blocks; persisted count dedups adjacent same-recordId blocks; the precise scan is skipped when liveCount>limit). Bounded and self-consistent with what the search shows, but worth unifying.
  3. This round's head advanced to a "Merge upstream main to resolve navigation integration conflicts". After the merge the core search implementation (ConversationSearch.tsx, turn-navigation-store.ts, transcript-page-table.ts, useMessageNavigation.ts) is byte-identical to what I reviewed, so the conclusions above still hold; the App.tsx / MessageList.tsx / useTranscriptViewport.ts deltas come from upstream (idle-settle streaming, shared-URL navigation, etc.) and this PR's search-entry mount and scrollToSearchHit plumbing survived intact.

Reviewed head 391cba6 (now MERGEABLE). CI on the current push is still running with no failing checks; I did not wait on the in-flight results.

@samuelhsin

Copy link
Copy Markdown
Collaborator Author

Rechecked 391cba6. The new human static review approved this exact head with no Critical/Important findings; the two non-blocking suggestions about the 200-result cap and block/message counting remain deferred under the repeated-review scope rule. Automated review is still running, so no completed automated-review verdict is claimed.

Ubuntu unit, lint/static, no-AK integration, Linux/Windows Desktop Shell and visual capture succeeded. macOS/Windows unit lanes and CLI integration were skipped, not passed. Web Shell smoke was cancelled by the job's 20-minute execution limit: the GitHub annotation explicitly reports the timeout, and the log reached the 169th/169 listed test before cancellation without a test assertion failure. This is not a successful smoke result.

Requested a rerun of that smoke job on the same SHA: https://github.com/QwenLM/qwen-code/actions/runs/35863645738/attempts/2 . No speculative code or test-timeout changes were made. PR remains open and mergeable; local/tracking/remote SHAs match at 0/0 divergence and the worktree is clean.

@wenshao

wenshao commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

Round 3 (delta, real daemon): round 2's Finding 1 is fixed, and I confirmed it with an A/B on the same stack. No new defects found. From the verification side this is merge-ready. Verified head 391cba626bf35bf428b62e2b6ff792bdba5bc651, tested as it would land: merged onto main f8ce07463b (merge tree 61be03ae09), macOS arm64, Node 24, Chromium 1228. Earlier rounds: round 1, round 2. This comment covers only what changed since then: 6fd5e61 (live echo de-dup), 9826676 (animation-frame snapshot and streaming identity), 0c73062 (reconnect cancellation), and the two main merges.

中文说明

第 3 轮(增量,真实 daemon):第 2 轮的发现 1 已修复,并在同一套真实环境上用 A/B 对照证实。没有发现新缺陷。从验证角度看可以合并。 验证 head 391cba626b,按实际落地的样子测试:合并到 main f8ce07463b 之上(merge tree 61be03ae09),环境为 macOS arm64、Node 24、Chromium 1228。本评论只写第 2 轮之后的改动:6fd5e61(实时回声去重)、9826676(按动画帧取快照、流式身份)、0c73062(重连时取消导航),以及两次合入 main。

环境

链路为 Chromium → daemon 自带的生产版 Web Shell → 从合并树打包的真实 qwen serve → 真实 chats/<id>.jsonl。模型是一个确定性的 OpenAI 兼容假模型,本轮给它加了慢速流式模式:约 8 秒内分 16 段发出。没有使用 mock daemon。唯一的网络干预是在 N5 和断网探针里对 /transcript 请求加延迟,内容不做任何改动。

A/B:发现 1 的修复确实起作用

两个臂只差一处:用 vite build 构建 web-shell 时是否回退 6fd5e61。daemon、会话播种、脚本全部相同。同一构建路径在未改动的源码上复现出 head 的 bundle 哈希 index-wEIMpmvY.js,逐字节一致,说明构建路径本身没有引入差异。

探针 回退 6fd5e61 head 391cba6
新会话、本页输入 4 轮(8 条消息)时入口是否可见 可见(图 1) 隐藏;10 条隐藏,12 条出现
回放 9 条 + 本页输入 2 条,搜 "Question #" 的行数 13 行,#10/#11 各重复一次(图 2) 11 行,无重复
本页输入的用户消息 / 工具调用前的助手文本 各 2 行(P1a、P2a 失败) 各 1 行(10/10)
本页输入 3 个工具调用轮次,搜 "PRE-TOOL-NEEDLE" 6 行(图 3) 3 行

刷新后两个臂都恢复为 11 行,这与第 2 轮的判断一致:问题只出在本页实时输入的块上。

针对第 2 轮之后新代码的探针(head 全部通过)

  • N1 流式(6/6):回复还在流式输出(约 8 秒)时打开弹框,每 250 毫秒采样一次,行数始终不超过 1。弹框一直开着,最后一段里的文字也能搜到。刷新后仍是 1 行,实时与回放一致。
  • N2 工具调用轮次计数(5/5):3 个工具调用轮次(9 条消息)全部在本页输入且不刷新时入口隐藏,第 11 条出现,与磁盘上的 11 条持久化消息一致。每一步实时与刷新后的结果都相同。
  • N3/N4 相同文本(3/3):同一条用户消息连发两次、两个工具调用轮次里工具调用前的文本完全相同,每种情况都是 2 行,实时与刷新后一致。回声去重是按身份匹配的,不会把不同轮次里文本相同的消息合并。
  • N5 导航途中重启 daemon(4/4):被延迟的导航不会在之后跳到旧目标。需要说明: 重启后 Web Shell 会退回 /,时间轴和入口一起消失。这是 Web Shell 对 daemon 重启的一贯表现:没有任何搜索操作时重启也一样,所以与本 PR 无关。重新打开会话后,搜索和定位都正常。
  • 导航途中断网 5 秒:会话保持,视口保持 live,没有跳到旧目标,也没有页面错误。弹框保持打开,并提示 "This result could not be located. Search again to refresh it."(图 4)。恢复联网后再次点击同一行,约 300 毫秒定位到第 3 轮。

关于 0c73062 的说明(不阻塞)

在我能构造的两个真实场景里(断网、重启 daemon),回退 0c73062 的构建与 head 行为完全相同(3/3 次)。已有守卫已经处理了这两种情况,所以这个修复在真实环境里我没能证明它是必需的。它针对的是更窄的竞态;在单元测试层面它是被钉住的,回退后有 3 个用例失败。

回归与门禁(合并树)

  • 600 轮、1200 条消息的冷加载长会话套件 26/26,阈值组 3/3,390px 窄屏 1/1。第一次运行时有两项失败,都是我脚本的问题,已查清:S1a 在首帧就检查时间轴,而时间轴实际存在(count=1);S11a 复用了一个已被上一次运行多加过一轮的会话。修正脚本后重跑全部通过。
  • 单元测试:PR 的 7 个测试文件 234/234。变异:回退 0c73062 后 3 个失败,回退 6fd5e61 后 8 个失败,其中包括"不带 ID 的 live 用户块和工具调用前片段只计一次"这个用例,正是第 2 轮要求补的真实形态。
  • PR 自带的 Playwright:21/21(chromium,搜索、消息导航、history-viewport)。
  • pnpm install --frozen-lockfile、根目录 build、bundle 均通过;web-shell tsc --noEmit 退出码 0;24 个改动的 TS/TSX 文件 eslint --max-warnings 0 退出码 0;Prettier 检查通过。

第 2 轮观察项的现状(作者已声明延后,不阻塞)

  • 200 行上限仍保留最旧的匹配:600 条命中时显示第 1–200 轮。
  • 每次查询仍会重读整段历史:600 轮时是 21 次 /transcript 加 4 次 /turn-index,本机约 0.55 秒。

未覆盖

  • Windows/Linux 桌面、WebKit/Firefox 下的真实 daemon。
  • 本轮没有重跑公开 navigateToMessage API 组和旧 daemon 降级组;第 2 轮已通过,且之后的改动没有涉及这两部分的对外行为。
  • 真实模型、会话压缩(compaction)之后的搜索。
  • 发布时 review-pr 仍在运行;合并状态为 BLOCKED,原因是 bot 之前给出的 CHANGES_REQUESTED 需要处理。

脚本与原始结果:r3/harness、r3/results。

Setup

Chromium → the daemon-served production Web Shell → a real qwen serve bundled from the merged tree → real chats/<id>.jsonl. The model is a deterministic OpenAI-compatible fake, extended this round with a slow-stream mode (16 chunks over ~8 s). There is no mock daemon. The only network intervention is a delay, never a rewrite, on /transcript in the reconnect probes.

A/B: the Finding 1 fix is what makes the difference

The two arms differ only in the web-shell vite build input: head versus head with 6fd5e61 reverted. Daemon, seeding and scripts are identical. The same build path run on untouched source reproduces head's bundle byte for byte (index-wEIMpmvY.js), so the build path itself adds no difference.

Probe 6fd5e61 reverted head 391cba6
New session, 4 turns typed in the tab (8 messages): entry visible? yes (fig. 1) no; hidden at 10, shown at 12
9 replayed + 2 typed user messages, query "Question #" 13 rows, #10/#11 twice (fig. 2) 11 rows, no duplicates
Typed user message / pre-tool assistant text 2 rows each (P1a, P2a fail) 1 row each (10/10)
3 tool-call turns typed live, query "PRE-TOOL-NEEDLE" 6 rows (fig. 3) 3 rows

After a reload both arms show 11 rows, which matches round 2's diagnosis: only blocks typed live in this tab were affected.

Fig. 1: 8 messages typed live. The reverted arm shows the icon; head hides it.
threshold A/B

Fig. 2: duplicate rows. Reverted: 3 / 13. Head: 1 / 11.
duplicate rows A/B

Fig. 3: pre-tool assistant text in 3 live tool turns. Reverted: 6 rows. Head: 3.
pre-tool rows A/B

New probes against the post-round-2 code (all pass on head)

  • N1, streaming (6/6): I opened the dialog while a reply was still streaming (~8 s) and sampled every 250 ms. The dialog never listed more than 1 row. Text from the last chunk is found while the dialog stays open. After a reload there is still 1 row, so live and replay agree.
  • N2, counting tool-call turns (5/5): with 3 tool turns (9 messages) typed live and no reload, the entry is hidden. It appears at the 11th message, which matches the 11 persisted messages on disk. Live and replay agree at every step.
  • N3/N4, identical text (3/3): the same user prompt sent twice, and identical pre-tool text in two tool turns, give 2 rows each, both live and after a reload. The echo matching goes by identity, not text: it doesn't merge equal text across turns.
  • N5, daemon restart during a deep-history navigation (4/4): the delayed navigation never lands on the stale target later. Caveat: a daemon restart sends Web Shell back to /, and the rail and entry go with it. Web Shell behaves the same way when no search is involved, so this is not this PR. After reopening the session, search and navigation work.
  • 5 s offline blip during a navigation: the session is kept, the viewport stays live, there is no stale jump and no page error. The dialog stays open and says "This result could not be located. Search again to refresh it." (fig. 4). Clicking the same row once back online lands on turn 如何自定义密钥文件 .env可能与其他文件冲突 #3 in about 300 ms.

Fig. 4: dialog after the offline blip.

offline blip dialog

Note on 0c73062 (non-blocking)

In both real scenarios I could build (offline blip, daemon restart), a build with 0c73062 reverted behaves exactly like head (3/3 runs), because older guards already handle those cases. So I could not show that this guard matters on a real stack. It targets a narrower race, and the unit tests pin it: reverting it turns 3 cases red.

Regression and gates on the merged tree

  • 600-turn / 1200-message cold-load suite 26/26, threshold group 3/3, 390 px layout 1/1. The first pass had two failures, both in my harness: S1a checked the timeline on the first frame even though it is present (count=1), and S11a reused a session that the previous run had already extended by a turn. With the harness fixed, everything passes.
  • Unit tests: the PR's 7 test files, 234/234. Mutation: reverting 0c73062 turns 3 cases red and reverting 6fd5e61 turns 8 red. One of the 8 is "counts live users without IDs and pre-tool assistant fragments with only prompt identity once", the real-world shape round 2 asked for.
  • The PR's own Playwright specs: 21/21 (chromium; conversation search, message navigation, history viewport).
  • pnpm install --frozen-lockfile, root build and bundle pass; web-shell tsc --noEmit exits 0; eslint --max-warnings 0 on the 24 changed TS/TSX files exits 0; Prettier is clean.

Round-2 observations (deferred by the author, non-blocking)

Not covered

  • Windows/Linux desktops; WebKit/Firefox against a real daemon.
  • I did not re-run the public navigateToMessage API group or the older-daemon group this round. Both passed in round 2, and nothing since has changed their external behaviour.
  • A real model; search after compaction.
  • review-pr was still running when I posted, and the merge state is BLOCKED by the bot's earlier CHANGES_REQUESTED.

Scripts and raw results: r3/harness, r3/results.

@wenshao

wenshao commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

Round 3 addendum: real model (qwen3.8-max) + fine-grained mutation. No new defects; this supports the merge-ready verdict in round 3. Same head 391cba626b and the same merge tree as round 3 (main f8ce07463b, tree 61be03ae09), macOS arm64, Node 24, Chromium 1228. Round 3 lists "a real model" under Not covered. This comment covers that gap and one more axis that round 3 didn't test. It does not repeat round 3's A/B.

中文说明

第 3 轮补充:真实模型(qwen3.8-max)+ 细粒度变异。没有发现新缺陷,支持第 3 轮"可合并"的结论。 与第 3 轮是同一个 head 391cba626b、同一棵合并树(main f8ce07463b,tree 61be03ae09),macOS arm64,Node 24,Chromium 1228。第 3 轮把"真实模型"列为未覆盖,本评论只补这一块,外加第 3 轮没做的另一个方向。第 3 轮的 A/B 不再重复。

真实模型会话

Chromium → daemon 自带的生产版 Web Shell → 由合并树打包的真实 qwen serve → qwen3.8-max(开启思考),HOME/QWEN_HOME 隔离。在当前标签页里输入了 7 轮:3 轮是"先说一句 PRE-FIG-n → 调工具 →(再说一句 MID-FIG-n → 再调工具)→ 最后说 POST-FIG-n",4 轮是纯文本。每轮结束后,另开一个浏览器上下文冷加载同一会话。期望值全部从磁盘上的 chats/<id>.jsonl 算出,与模型实际说了什么无关(图 1)。

  • 阈值: JSONL 里的消息数依次为 3/5/9/11/14/16/18。实时标签页和冷回放都在 9 条时隐藏、11 条时显示,每轮都一致。模型的思考块和纯工具调用不计入,与持久化计数一致。
  • 搜索行数: 6 个搜索词(FIG、PRE-FIG、MID-FIG、POST-FIG、PLAIN-FIG、R3Q),在第 7 轮时,实时、冷回放、JSONL 三方行数完全相等(18/6/2/6/8/7),刷新后也与 JSONL 相等。真实模型下的工具前文字、工具之间的文字都没有重复,也没有遗漏(图 2)。
  • 第二客户端发来的用户消息(第 2 轮报告里列为未覆盖):搜索对话框打开时,另一个客户端通过 API 发一条用户消息,实时显示 1 行,刷新后仍是 1 行。
  • 折叠工具组里的命中: 真实模型下,工具前、工具之间的文字会被折叠进 "Processed … · 2 tool calls · 3 thoughts" 分组。选中 MID-FIG-3 这条结果后,分组自动展开,滚动到目标并高亮,高亮的正好是两次工具调用之间的那句话(图 3);PRE-FIG-1 同样可以。
  • 全程没有页面错误。

细粒度变异矩阵(第 2 轮之后的 3 个修复提交)

第 3 轮是按整个提交回退的。这里对 6fd5e61、9826676、0c73062 做了 11 个单点变异,跑 PR 自带的 8 个测试文件(包括 App.test.tsx,每次 1264 个测试)。对照组全绿;每个变异体都确认文件确实被改动,11 个哈希互不相同;跑完后 git status 干净。杀死 8 个,存活 3 个:

  • N6(测试缺口):把 navigation.totalTurns 从阈值探针 effect 的依赖里删掉,所有测试仍然全绿。它的同伴 liveIdentity 依赖(N5)有测试钉住。
  • N7(测试缺口):9826676 在 live alias 分支(本页输入的用户记录)上加的 finishLiveLocation(alias) 改回直接 return alias,所有测试仍然全绿;同一提交在 live block 分支上的同类改动(N8)则被测试杀死。视口滚动用的是返回的位置,不受影响,所以这不是用户可见的缺陷(我没能构造出可见症状);但 store 里发布的 selected 状态在这条分支上没有测试守护。
  • N11(近似等价):viewport walk 的守卫只保留 sessionEpoch 检查、去掉客户端检查,测试不会变红。原因是每次 owner 变化和每条断开路径都会让 sessionEpoch += 1,只有 too_large 这条 legacy 分支换了客户端却不递增,而搜索不走那条分支。这属于纵深防御,不需要补测试。

建议(不阻塞):为 N6、N7 各加一个用例:一个是 totalTurns 变化而 live 块身份不变时重新探测阈值;另一个是通过 alias 定位本页输入的用户消息后,store 发布 status: 'ready' 的选中状态。

其他

  • R2 的 F1 用确定性模型在本树上复测:本页输入的 live/工具/第二客户端探针 10/10(R2 为 8/10);9 条回放 + 2 条本页输入 → 11 行,无重复(R2 为 13 行);新会话在 10 条消息时隐藏、12 条时显示(R2 在 8 条时就显示)。
  • PR 自带的 Playwright 用例(conversation-search + message-navigation):18/18。
  • 当前 head 上已执行的 CI 全部通过(包括重跑的 web-shell smoke);review-pr 仍在运行。
  • 200 条上限保留的是最旧匹配这一点仍然存在(R2 的观察 1,作者已声明延后处理),本轮没有重测。

未覆盖

  • Windows/Linux、WebKit/Firefox;真实模型下的会话压缩和子代理轮次;超过 200 条命中的真实模型会话。
  • 真实模型只测了 1 个会话、7+1 轮,不是统计意义上的覆盖。

Real-model session

Chromium → the daemon-served production Web Shell → a real qwen serve bundled from the merged tree → qwen3.8-max with thinking enabled, isolated HOME/QWEN_HOME. I typed 7 turns in the live tab. Three were tool turns: say PRE-FIG-n → call a tool → (say MID-FIG-n → call another tool) → say POST-FIG-n. Four were plain text. After every turn, a second browser context cold-loaded the same session. Every expected value comes from the on-disk chats/<id>.jsonl, so the oracle doesn't depend on what the model chose to say (fig. 1).

  • Threshold: JSONL message counts per turn were 3/5/9/11/14/16/18. Both the live tab and the cold replay hid the entry at 9 and showed it at 11, and they agreed at every step. Thinking blocks and tool-only parts don't count, which matches the persisted count.
  • Rows: for 6 needles (FIG, PRE-FIG, MID-FIG, POST-FIG, PLAIN-FIG, R3Q), live, cold replay and JSONL truth were identical at turn 7 (18/6/2/6/8/7), and after a reload they still matched the JSONL. With a real model, pre-tool and between-tool assistant text is neither duplicated nor lost (fig. 2).
  • User message from a second client (listed as not covered in round 2): with the dialog open, another client posts a user prompt through the API. It shows as 1 row live and 1 row after a reload.
  • Hits inside a collapsed tool group: with a real model, pre-tool and between-tool text is folded into a "Processed … · 2 tool calls · 3 thoughts" group. Choosing the MID-FIG-3 result expands the group, scrolls to the target and highlights exactly the sentence between the two tool calls (fig. 3). PRE-FIG-1 behaves the same way.
  • No page errors during the run.

Fig. 1: scoreboard: live vs. cold replay vs. JSONL truth, and the mutation matrix.
scoreboard

Fig. 2: real-model session, query pre-fig: 6 results (3 user prompts + 3 pre-tool assistant sentences), matching the JSONL.
dialog

Fig. 3: after choosing the MID-FIG-3 hit, the collapsed group is expanded and the sentence between the ReadFile and Glob calls is highlighted.
collapsed group hit

Fine-grained mutation (the 3 fix commits after round 2)

Round 3 reverted whole commits. Here I made 11 single-point mutants of 6fd5e61, 9826676 and 0c73062 and ran the PR's own 8 test files (including App.test.tsx; 1264 tests per run). The control run is fully green. Each mutant was verified to actually change the file, the 11 mutant hashes are all distinct, and git status was clean afterwards. 8 killed, 3 survived:

  • N6 (test gap): removing navigation.totalTurns from the threshold-probe effect's dependencies leaves every test green. Its sibling, the liveIdentity dependency (N5), is pinned by a test.
  • N7 (test gap): reverting 9826676's finishLiveLocation(alias) back to return alias on the live-alias branch (a user record typed in this tab) leaves every test green. The same change on the live-block branch (N8) is killed. The viewport scrolls using the returned location, so this is not a user-visible defect; I couldn't produce a visible symptom. But the selected state the store publishes on this branch has no test guarding it.
  • N11 (near-equivalent): keeping only the sessionEpoch check in the viewport-walk guard and dropping the client check doesn't turn any test red. Every owner change and every disconnect path runs sessionEpoch += 1; only the too_large legacy branch swaps the client without incrementing it, and search doesn't use that branch. This is defence in depth and doesn't need a test.

Suggestion (non-blocking): add one test for each of N6 and N7. For N6: the threshold is re-probed when totalTurns changes but live block identities don't. For N7: after an alias-located typed user message, the store publishes a status: 'ready' selection.

Other

  • Round 2's Finding 1, re-run on this tree with the deterministic model: live/tool/second-client probes 10/10 (round 2: 8/10); 9 replayed + 2 typed messages give 11 rows with no duplicates (round 2: 13); a new session hides the entry at 10 messages and shows it at 12 (round 2 showed it at 8).
  • The PR's own Playwright specs (conversation-search + message-navigation): 18/18.
  • Every executed CI check on the current head passes, including the web-shell smoke rerun; review-pr is still running.
  • The 200-row cap still keeps the oldest matches (round 2's observation 1; the author has deferred it). I did not re-measure it in this round.

Not covered

  • Windows/Linux; WebKit/Firefox; compaction and subagent turns under a real model; a real-model session with more than 200 hits.
  • The real-model run is one session of 7+1 turns. It is not statistical coverage.

Scripts and raw results: pr12234/r3-realmodel.

@wenshao
wenshao dismissed a stale review September 23, 2026 16:47

fixed

@samuelhsin
samuelhsin added this pull request to the merge queue Sep 23, 2026
Merged via the queue into QwenLM:main with commit abd2ece Sep 23, 2026
69 of 71 checks passed
qwen-code-dev-bot added a commit that referenced this pull request Sep 23, 2026
Reconcile main's in-conversation search timeline action (#12234) with this PR's MCP App page budgets and session-scoped App context.

MessageList keeps the McpAppSessionContext provider tree and re-adds the timeline action prop on SessionTimeline; TranscriptViewport passes both the App session id and the navigation-suppressed timeline action; transcript-page-table fits the anchored page within budget before resolving main's exact-record block lookup.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

review/self-reported The linked issue was opened by the PR author (self-reported)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(web-shell): search within the current conversation and jump to matching content

3 participants