Repository navigation
feat(web-shell): search within the current conversation - #12234
Conversation
E2E verification reportVerified on macOS with Chromium using synthetic daemon fixtures; no real conversation data, screenshots, credentials, or model calls were used.
Browser reproduction command (from PLAYWRIGHT_PORT=15231 npx playwright test --config playwright.config.ts --project=chromium client/e2e/web-shell.conversation-search.spec.tsBaseline: the upstream paginated-history suite passed 3/3 before implementation. The globally installed CLI (0.24.0-preview.0) could start an isolated daemon, but did not contain Web Shell assets, so browser baseline and feature verification used the repository's Web Shell test harness. Windows/Linux and real-provider end-to-end calls were not tested. During verification, Escape focus restoration, cancelled navigation cleanup, and mixed persisted/live message counting were corrected and re-tested. The final full build also passed after a message-kind type narrowing fix. 中文在 macOS Chromium 上使用合成 daemon 数据验证,不使用真实会话、原始截图、凭据或模型调用。
基线:实现前历史分页浏览器测试 3/3 通过。全局 CLI 0.24.0-preview.0 能启动隔离 daemon,但不含 Web Shell 静态资源,因此浏览器基线及功能验证使用仓库测试设施。未验证 Windows/Linux 或真实模型调用。 验证过程中修正并复测了 Escape 焦点恢复、取消定位后的状态清理,以及持久化/实时消息混合计数。消息类型收窄修正后,完整构建重新通过。 |
Review and CI follow-upAddressed the actionable findings on this PR:
Validation: 97 focused unit tests, 11 browser tests, and 2 visual scenarios passed. English/Chinese search behavior, timeline alignment/visibility, fixed dialog height, historical navigation, and draft preservation are covered. Root The initial E2E report above is marked historical because the final entry moved to the timeline and now hides with it. No review threads were resolved and no approval was submitted. 中文已处理:合入最新 main 修复 CI 规则过期;修正历史搜索首次请求缺少快照导致的真实 HTTP 400;补齐自动截图场景;将搜索专用文案移出只读导出共享翻译包,使导出体积回到原有上限内,未放宽检查。 真实 daemon 已验证完整扫描 24 条合成消息并命中 1 条。97 项相关单测、11 项端到端测试、2 项浅深主题截图场景通过。旧验证报告已标注为历史记录;未解决 review thread、未提交批准。 The issue and PR description now also track the requested public embedding API for exact persisted-message navigation. That additional API is explicitly marked as not implemented yet; the current patch implements internal in-conversation navigation. |
Public message navigation implemented and verifiedCommit The host opens the session/workspace first; the API finds the exact persisted user/assistant record, loads historical content outside the live window, and activates scrolling/highlighting. It also works with the timeline hidden and leaves the draft unchanged. Cancellation, stale owners, unsupported history, readiness, missing records, and failures return distinct statuses. Local validation on macOS Chromium, with synthetic data only:
Four screenshots are embedded in the PR description, showing the highlighted historical target and preserved synthetic draft. Screenshots are on a separate evidence branch and do not add image binaries to this code PR. The existing cross-session content-search endpoint still returns session+snippet without a stable message record ID. Host search integration, multiple-message results and archived-session search are not claimed as completed here. New-head CI is pending. 中文:已实现公开消息定位接口,并将 issue/PR 的“尚未实现”更新为实际契约。直接调用公开 API 的四种浏览器场景均验证旧消息加载、高亮、草稿保留和缺失记录行为;四张脱敏截图已放入 PR 正文。本地完整构建、类型检查、打包及回归通过,新提交 CI 待完成。 |
|
Verdict: merge-ready — 1109/1109 scripted assertions passed, 0 failed. Verified head: 中文摘要结论:可以合并(merge-ready)——1109/1109 项脚本断言全部通过,验证 head 为
Central claim and A/B proofCentral claim: above a configurable message threshold (default 10), a search icon appears below the session timeline and opens a dialog that performs literal, case-insensitive search over user/assistant messages — including persisted history outside the loaded live window — and navigates to the matching message with a temporary highlight. The PR's own browser e2e suites (real Chromium, real HTTP/WebSocket fetches intercepted only at the Playwright route layer — the full client and SDK wire code executes) were run unmodified at head, then copied verbatim into the base worktree and run against the base build. Same spec files, same fake-daemon scenario, both arms.
Every cell whose oracle requires the feature flips from red at base to green at head; every cell that asserts absence/hidden states is green on both arms — the diff is exactly what makes the difference. Witnesses: UI evidence from the head runs: Mutation matrix (vacuity check)Five single-point mutants of the new logic, run against the PR's own unit suites. All killed by the assertion that exists to catch them; M1 is the in-file positive control proving the harness can go red.
Witness: Targeted gates
Reviewer Test Plan walk-through
Findings
Not covered
Evidence01-ab-head-e2e — head 02-ab-base-e2e — base 03-mutation-cap100 — mutation M4 (hits cap 200→100) turns the bounds test red 04-search-dialog-history-hit — case-insensitive 05-history-hit-located-flash — historical viewport opened at the hit with flash highlight; unsent draft preserved 06-external-nav-located — host-driven 07-timeline-search-entry — 14px search entry aligned with the timeline rail MethodologyLocal maintainer round on Linux aarch64 (Orange Pi 6 Plus), Node 22, Chromium headless-shell 149 (Playwright 1.61.1). Head worktree at |
Follow-up review: navigation cancellation and keyboard coverageAddressed the non-blocking coverage suggestion in the maintainer verification report: a real browser test now exercises ArrowUp/ArrowDown selection, wrap-around at both ends, previous/next buttons, and Enter-to-select, then asserts that the selected historical message is visibly highlighted. Also reproduced and fixed an independent cancellation defect: Commit: Validation: 23/23 hook tests, 16/16 browser tests, root typecheck, root build/bundle, changed-file lint/format, and the real pre-commit hook passed. The two original regression cases failed before the fix and passed afterward. No visible layout changed; the existing synthetic screenshots in the PR remain representative. All executed CI checks passed on the preceding head 中文:补齐维护者建议的键盘与前后按钮浏览器验证,并修复“被拒绝的定位请求误取消有效定位”的竞态。先复现两项失败,再修复并验证 23 项接口单测、16 项浏览器测试及完整构建/类型检查/打包;无界面变化,现有脱敏截图仍适用。上一提交 CI 已全绿,新提交等待 CI。 |
|
Rechecked head
中文:本轮复查当前提交的已执行 CI 全部通过,无新增评审意见或待处理线程。未发现新的确定缺陷,因此不追加无必要代码改动;已更新正文中过时的 CI 等待状态。代码无合并冲突,GitHub 整体合并状态仍为 BLOCKED,未执行合并或批准。 |
|
Reviewed all 63 inline comments on
Validation at this commit:
The remaining 58 Suggestion comments are deferred from this correctness follow-up: API/product choices (including per-message vs per-occurrence search and visibility rules), other non-blocking behavioral, accessibility/performance/UI suggestions, fixture/test expansion, and documentation/cleanup. They are not being marked fixed. Existing public API shape and timeline-linked entry visibility remain as specified. Maintainer product/API decisions still require maintainer judgment; this report is not a merge approval. Local browser evidence is macOS Chromium with synthetic transport, not native IME hardware or Windows/Linux E2E. CI status for the new head is reported separately; previous green CI belongs to the previous commit. 中文:本轮仅处理 5 条 Critical,已补复现及回归验证;其余 58 条 Suggestion 明确后置,不声称已修复。保留原公开 API 与时间轴显隐约定,不扩展功能范围。R1-76 确认并修复的是取消搜索后丢失原历史阅读页,而非其证据未证明的成功定位误跳。 |
|
Rechecked head Continued through the deferred comments and fixed three concrete interaction issues in
The existing browser keyboard scenario now checks those ARIA relationships, matches message 1 exactly (cannot pass on message 10), and is tagged Verification:
The previous five Critical fixes remain in place. Other deferred suggestions and maintainer decisions about public API/product scope remain open. This is not a merge approval; new-head CI is distinct from the green checks on 中文:本轮继续修复 R1-64 选中项漂移、R1-63 不完整结果无法重试、R1-83 读屏语义缺失,并加强已有键盘 smoke 测试。40 项相关测试与 16 个浏览器场景通过;上一提交已执行 CI 全绿,新提交 CI 单独跟进。 |
|
Rechecked the exact previous head Fixed the current CI blockers in
Validation on the merged tree:
New-head CI is pending and is not covered by these local results. Previous-head failed runs: CI, visual capture. 中文:本轮同步主干解决 pnpm CI action、过期 lint gate 和旧导出预算阻塞,并修正截图测试遗漏的输入框角色。1,211 项单测、16 个浏览器场景、暗亮两项截图、构建/类型检查/bundle 均通过;新提交 CI 单独跟进。 |
|
Rechecked review round 2 against its exact head Addressed the confirmed correctness findings in
Verification:
Deferred R2-11/R2-12/R2-13/R2-14/R2-15 (trigger-render optimization and additional test-mutation coverage) with the existing non-blocking Suggestions, to keep repeated review rounds focused on demonstrated correctness defects. These are recorded, not silently treated as resolved. Product/public-API maintainer decisions remain open; this is not a merge approval. New-head CI is tracked separately. |
|
Reviewed round 3 against its exact head Fixed three demonstrated state-consistency defects in
R3-6 and R3-15 were labeled Suggestions, but were included because runtime reproduction confirmed incorrect state and a blocked history continuation, not merely missing test coverage. The rest of the new Suggestions remain deferred under the convergence policy: R3-2 through R3-5, R3-7 through R3-14, and R3-16 through R3-23, plus the previously deferred R2-11/R2-13/R2-15. No new public API or unrelated cleanup is included. Validation on the final tree:
New-head CI is tracked separately from the preceding head's green results. Existing maintainer/product/API review decisions remain open; this is not a merge approval. |
|
Rechecked exact head No new published actionable CR findings or additional confirmed correctness defects were found in this pass. Re-read the latest cancellation/recovery changes and their callers; R3-1 remains fixed despite its thread still being marked unresolved. R3-6 and R3-15 recovery fixes also remain intact. Previously recorded Suggestion deferrals remain unchanged. Fresh local verification: all 165 tests in the five focused navigation/search/viewport/page-table suites passed. Both the CI run and visual capture run completed successfully on this exact SHA: Ubuntu unit tests, lint/static checks, no-AK integration, Linux/Windows desktop builds, Web Shell E2E smoke, and visual capture. macOS/Windows unit tests and CLI integration were skipped, not counted as passes. On the subsequent recheck, the separate automated Local HEAD, tracking branch and remote branch match, divergence is |
|
Round 2 (delta, real daemon) — the feature works end to end against a real 中文说明第 2 轮(增量,真实 daemon)—— 功能在真实 本轮新增了什么第 1 轮(head 结果
真实 daemon 上确认可用的:小写查询命中第 3 轮里的大写文本(在初始渲染窗口之外 570 轮);中文;代码块内的 token; 发现 1(缺陷)—— 在本标签页输入的消息会被重复计数、重复列出现象。 在全新会话里,所有轮次都在当前标签页输入:
入口在 8 条消息时就出现,而不是第 11 条。同一会话、同样 8 条消息,刷新之后入口又消失(图 1)。同一原因还导致搜索结果重复:9 条用户消息的会话全部从磁盘回放时是 9 行;在本页再发 2 轮后变成 13 行(应为 11),状态栏显示 根因(已在真实数据上插桩确认,插桩已移除)。 本页输入的 live 用户块是
这正是 bot 第 1 轮 reverse-audit 标注为"未探索"的那个前提(daemon 是否给持久化用户块写 影响面。 没有错误导航 —— 两个孪生行点击后都落到同一条 live 消息并高亮,无多余请求,也不涉及数据。但它与 PR 的首要声明("10 条隐藏、11 条显示")在最常见的使用路径上不符,且重复行在每个实时会话里都可见。 修复方向。 我试了最直观的 7 行补丁(把 store 的回声匹配结果带到 hit 上,组件里过滤)—— 不起作用,因为回声匹配本身就没命中。需要让 live 用户块拿到身份:要么在 观察(不阻塞)
第 1 轮遗留项
针对第 1 轮之后修复提交的变异矩阵对 合并树上的门禁
未覆盖
方法Linux x86_64,Node 22.22.2。合并树工作区:真实安装 + 根构建 + 打包;daemon 使用隔离的 What this round addsRound 1 (head Results
Confirmed on a real daemon: a lower-case query finds upper-case text in turn #3, 570 turns outside the initially rendered window; Chinese; a token inside a code block; Finding 1 (defect) — messages typed in this tab are counted and listed twiceSymptom. A brand-new session, every turn typed in the current tab:
The entry appears at 8 messages, not at the 11th. In the same session with the same 8 messages, it disappears again after a reload (figure 1). The same cause duplicates search rows: a session with 9 user messages shows 9 rows when everything is replayed from disk; after two more turns typed in the tab it shows 13 rows (expected 11), the status reads Root cause (confirmed by instrumenting the scan on real data; instrumentation removed). The live user block typed in this tab is
This is the premise the bot's round-1 reverse audit listed as unexplored (whether the daemon puts Blast radius. No wrong navigation — both twin rows land on the same live message and flash it, with no extra requests — and no data is touched. But it contradicts the PR's headline claim ("hidden at 10 messages and appears at 11") on the most common path, and the duplicate rows are visible in every live session. Fix direction. I tried the obvious 7-line patch (carry the store's echo match onto the hit and filter in the component) — it does not help, because the echo tier never matches in the first place. The live user block needs an identity: either backfill Observations (non-blocking)
Round-1 leftovers
Mutation matrix for the fixes pushed after round 1Single-point reverts of the fixes in Gates on the merged tree
Not covered
EvidenceFigure 1 — same session, same 8 messages: entry visible when typed live, hidden after reload Figure 2 — 11 user messages, 13 rows: the two typed in this tab are listed twice Figure 3 — tool-call turn: the pre-tool assistant text is listed twice, the post-tool text once Figure 4 — real daemon, 600 turns: lower-case query, upper-case hit 570 turns outside the rendered window, draft intact Figure 5 — identity, not text: Figure 6 — older daemon (capability stripped by a proxy): loaded-only notice, honest empty result Figure 7 — scoreboard, live-threshold table and mutation matrix More: 600 matches capped at 200 · public API landing on turn #3 with the timeline hidden MethodologyLinux x86_64, Node 22.22.2. Merged-tree worktree with a real install, root build and bundle; the daemon runs with an isolated 🤖 Generated with Claude Code — Claude Fable 5.1 |
|
Addressed the real-daemon Finding 1 in Messages typed in the current tab now use their existing prompt-admission/turn identity to pair the local echo with its persisted record. Pre-tool assistant messages use the exact containing turn's prompt ID when replay does not carry it. Search removes the paired history row only when the corresponding live search result is present, preserving distinct same-text turns and unpersisted messages. Assistant text is checked within the identified prompt so a retained later fragment cannot hide a different earlier fragment. The real-daemon check also exposed a stale threshold probe: a still-streaming assistant could be counted before its record identity arrived, and the cached count was not refreshed when the block count stayed unchanged. Identity/completion changes now rerun the probe; text-only streaming deltas do not. Validation on this change:
The new real-daemon review's non-blocking observations 1–4 (latest-first cap, caching, defensive test coverage and Markdown snippets) remain deferred. No SDK or daemon protocol changes were needed. The new-head product CI and visual capture are running; no new-head CI pass is claimed yet. Local/tracking/remote SHA parity is confirmed with 10 real messages: search entry hidden 11th user message, assistant pending: search entry visible Pre-tool assistant text: exactly one result |
|
Refreshed CR and CI against Fixed the current CI blocker in Validation on the merged tree:
Local HEAD, tracking branch, and remote branch match ( |
|
Rechecked exact head Product CI and visual capture both completed successfully on this exact SHA. Lint/static checks, Ubuntu unit tests, no-AK integration, Linux/Windows desktop builds, and Web Shell E2E smoke passed. macOS/Windows unit and CLI integration lanes were skipped, not counted as passes. The prior lint-gate freshness blocker is cleared. Checked all 103 review threads: no unresolved Critical threads, and no new published review findings since the prior update. The separate automated review is still running; this is not a claim that it has completed or that maintainer approval has been granted. Previously documented Suggestion deferrals remain unchanged. No additional code change was needed. Local/tracking/remote SHA parity is confirmed with |
|
Rechecked the R5 review against the current branch and fixed the two demonstrated regressions:
R5-3 follows the updated entry-placement requirement in #12231: the search icon is hidden whenever the timeline is hidden. No fallback icon was added. R5-5 and previously deferred Suggestions remain follow-up work under the repository's repeated-review scope rule. Validation: root build, typecheck and bundle passed; 1,252 focused and App tests passed (including the new red-to-green regressions), 18 Chromium search/public-navigation scenarios passed, and the newly built production bundle passed 16 real-daemon assertions with zero page errors. These cover 10 hidden / 11 visible before the assistant reply, live and pre-/post-tool deduplication, correct result navigation, and reload consistency. Model responses and workspace content were synthetic; daemon persistence and tool execution were real. Full diff received two clean self-audit passes and independent review. Skipped or pending remote checks are not counted as passes. Fresh sanitized production screenshots: Delivered as 9826676 with normal pre-commit hooks. Local HEAD, tracking branch and remote branch match, divergence is 0/0, and the worktree is clean. New-head CI is still pending; previous-head green checks are not reused as proof for this commit. |
|
Rechecked current head 9826676, all 108 review threads, and the latest comments. No new actionable finding has arrived since the R12 fix. R5-1/2 and R5-4 are fixed in the current code; R5-3 remains the explicit entry-placement requirement documented in #12231. Deferred Suggestions remain unchanged. No additional code changes were needed. Exact-head remote verification is now complete: product CI and visual capture both succeeded for this SHA. Executed product checks passed: Ubuntu unit tests, lint/static checks, no-AK integration, Linux/Windows Desktop Shell, and Web Shell E2E smoke. macOS/Windows unit lanes and CLI integration were skipped, not passed. Latest recheck: the automatic review has started and is currently executing its Run review step on this exact SHA. It has not published a new result yet, so this is not a claim that the new review has completed. The PR is open and mergeable. Local HEAD, tracking branch and remote branch match, divergence is 0/0, and the worktree is clean. Freshness verification reports no stale gate files. Previous local test and synthetic production evidence remains in the R12 report; no tests were unnecessarily rerun on an unchanged tree. |
|
Fixed the disconnect race reported in R5-4. A search waiting on an independently started boundary load could restore a cleared selection after disconnect; reconnecting with the same client owner also reproduced it. Search now verifies both the client owner and the connection epoch before continuing. The same race was reproduced and fixed in the full-window fallback path, which captures its epoch before trying the normal walk so recovery cannot bind to a newer connection. Three regression cases cover staying disconnected, reconnecting with the same client, and reconnecting during fallback. Each originally failed. Removing the walk guard makes both boundary-wait cases fail again; removing the fallback epoch guard makes the fallback case fail again. Restoring the guards passed all 147 search/navigation/page-table tests. Independent review and two self-audit passes found no further required changes. Existing Suggestions remain deferred under the repeated-review scope rule. Additional validation: root build, typecheck and bundle passed; 18 Chromium search/public-navigation tests passed; the newly built production daemon passed 16 synthetic search scenarios with zero browser errors (10/11 threshold, live and tool-message deduplication, navigation and reload). The browser run is normal-flow regression evidence; the exact disconnect interleavings were verified by the real-store regression tests and guard mutations above. No UI changes or new screenshot claims are involved. Delivered as 0c73062. Normal commit hooks passed without changing the validated tree; normal push succeeded. Local, tracking and remote SHAs match with 0/0 divergence and a clean worktree. New-head remote CI is pending and is not counted as passed. |
|
Rechecked head 0c73062 and all 109 review threads. No new review findings have arrived since the disconnect fix; automatic review is still running. The current-head CI smoke run failed in Reproduction attempt on this exact head: the targeted mobile-WebKit case passed once, then passed five consecutive serial repetitions locally (6/6 total). The composer and this test are unchanged relative to the merged main baseline; the last commit only changed navigation lifecycle guards and their tests. This does not rule out a Linux-specific or timing-dependent defect, and the cause is not yet confirmed. No speculative source change was made. Requested a rerun of the failed job on the same SHA: CI attempt 2 is in progress. The original failure is not considered fixed until remote verification completes. Local/tracking/remote SHAs match, divergence is 0/0, and the worktree remains clean. |
|
Round 7 confirmed the disconnect fixes and posted no new blocking findings. The previously failed mobile-history focus smoke also passed on the unchanged 0c73062 head in CI attempt 2. Existing review Suggestions remain deferred. GitHub then reported a merge conflict after main advanced. Merged freshly fetched main 94104d5, retaining both sets of added App imports and both README sections (conversation search/message navigation and URL navigation). These were the only conflict hunks. The feature diff against main remains 29 files; this merge adds no further search features. Independent review checked the auto-merged App, message list, viewport and exported API interactions with upstream URL navigation. Validation on the merged tree: frozen dependency installation and all 1,978 lockfile policy checks succeeded; root build, typecheck and bundle passed; 1,279 tests across nine search/navigation/App suites passed; 29 Chromium search/message-navigation/URL-navigation tests and the targeted mobile-WebKit history test passed. The newly built production daemon also passed 16 synthetic threshold/deduplication/navigation/reload assertions with zero browser errors. Existing search behavior and public message navigation remain intact alongside upstream URL navigation. Delivered merge commit 391cba6. Normal commit hooks passed and preserved the validated tree exactly. Local HEAD, tracking and remote branch match, divergence is 0/0, and the worktree is clean. New-head CI is pending; prior-head successes are not counted as new-head verification. |
ytahdn
left a comment
There was a problem hiding this comment.
结论:静态审查(不跑测试)通过,未发现 Critical / Important。核心异步/取消状态机与搜索正确性站得住。
我独立核对过的关键点:
- 触发搜索的 effect 有 250ms 防抖 + clearTimeout,且 onProgress/then/catch/finally 全部走 current 守卫;关闭对话框或换词时的旧请求不会覆盖新结果,取消也不会把视口卡在 loading。ci-bot 关注的 stale-response 竞态确实被处理。
- 结果高亮为纯字面量、大小写不敏感:createConversationSearchSnippet 先转义正则元字符再用 /iu 匹配,渲染走 JSX 文本切片(无 dangerouslySetInnerHTML / innerHTML),不存在正则注入或 XSS。
- 默认阈值 10(10 条隐藏、11 条出现)在常规 1:1 场景下与 README/描述一致。
- 跳转复用既有历史视口 + 缓存预算;分页/淘汰与取消/回滚路径经两个模块分别追踪,未见破坏实时转录或泄漏状态。
非阻塞建议(Suggestion,不影响合并):
- 命中数超过 200 时的截断方向:scanConversation 从最旧往最新填充 hits,到 200 即停并置 truncated;UI 里 [...older, ...live].slice(0, 200) 在 older 已占满 200 时会把当前可见窗口里的 live 命中也切掉。虽然有 truncated 提示、且需要单条会话 >200 次命中才触发(属极端场景),但这与本功能“在长会话里找回刚看过的那条”的初衷相反。可考虑改为最新优先扫描,或在 merge 时始终保留 live 命中。
- 阈值计数在“块 vs 消息”上不一致:liveCount 按块计、持久化 count 对相邻同 recordId 去重、且 liveCount>limit 时跳过精确扫描。行为有界且与搜索实际展示自洽,但概念口径建议统一。
- 本轮 head 前进为一次 “Merge upstream main to resolve navigation integration conflicts”。核对合并后:本 PR 的核心搜索实现(ConversationSearch.tsx、turn-navigation-store.ts、transcript-page-table.ts、useMessageNavigation.ts)与合并前逐字节一致,上述结论在新 head 上仍成立;App.tsx / MessageList.tsx / useTranscriptViewport.ts 的差异来自上游(idle-settle streaming、shared-URL 导航等),本 PR 的搜索入口挂载与 scrollToSearchHit 管线在合并后保持完好。
Verdict: static review (no test run) passes — no Critical or Important issues. The async/cancel state machine and search correctness hold up.
What I independently verified:
- The search-triggering effect is debounced (250ms + clearTimeout) and every onProgress/then/catch/finally is guarded by a
currentflag, so a stale response can't overwrite a newer query and cancellation can't leave the viewport stuck loading — the ci-bot's stale-response race concern is genuinely handled. - Highlighting is literal and case-insensitive: createConversationSearchSnippet escapes regex metacharacters before matching with /iu and renders via JSX text slicing (no dangerouslySetInnerHTML / innerHTML), so there is no regex injection or XSS.
- The default-10 threshold boundary (hidden at 10, shown at 11) matches the README in the normal 1:1 case.
- Jump-to-message reuses the existing historical viewport and cache budget; pagination/eviction and cancel/rollback paths were traced across two modules with no live-transcript corruption or leaked state.
Non-blocking suggestions:
- Truncation direction over the 200-cap: scanConversation fills hits oldest-first and sets truncated at 200, and the UI merge [...older, ...live].slice(0, 200) drops on-screen live matches once older fills the budget. There is a truncation banner and it needs >200 matches in one conversation to bite (an edge case), but this is opposite to the feature's goal of finding a recently-seen answer in a long conversation. Consider scanning newest-first, or always keeping live matches in the merge.
- Threshold counting is inconsistent between blocks and logical messages (liveCount counts blocks; persisted count dedups adjacent same-recordId blocks; the precise scan is skipped when liveCount>limit). Bounded and self-consistent with what the search shows, but worth unifying.
- This round's head advanced to a "Merge upstream main to resolve navigation integration conflicts". After the merge the core search implementation (ConversationSearch.tsx, turn-navigation-store.ts, transcript-page-table.ts, useMessageNavigation.ts) is byte-identical to what I reviewed, so the conclusions above still hold; the App.tsx / MessageList.tsx / useTranscriptViewport.ts deltas come from upstream (idle-settle streaming, shared-URL navigation, etc.) and this PR's search-entry mount and scrollToSearchHit plumbing survived intact.
Reviewed head 391cba6 (now MERGEABLE). CI on the current push is still running with no failing checks; I did not wait on the in-flight results.
|
Rechecked 391cba6. The new human static review approved this exact head with no Critical/Important findings; the two non-blocking suggestions about the 200-result cap and block/message counting remain deferred under the repeated-review scope rule. Automated review is still running, so no completed automated-review verdict is claimed. Ubuntu unit, lint/static, no-AK integration, Linux/Windows Desktop Shell and visual capture succeeded. macOS/Windows unit lanes and CLI integration were skipped, not passed. Web Shell smoke was cancelled by the job's 20-minute execution limit: the GitHub annotation explicitly reports the timeout, and the log reached the 169th/169 listed test before cancellation without a test assertion failure. This is not a successful smoke result. Requested a rerun of that smoke job on the same SHA: https://github.com/QwenLM/qwen-code/actions/runs/35863645738/attempts/2 . No speculative code or test-timeout changes were made. PR remains open and mergeable; local/tracking/remote SHAs match at 0/0 divergence and the worktree is clean. |
|
Round 3 (delta, real daemon): round 2's Finding 1 is fixed, and I confirmed it with an A/B on the same stack. No new defects found. From the verification side this is merge-ready. Verified head 中文说明第 3 轮(增量,真实 daemon):第 2 轮的发现 1 已修复,并在同一套真实环境上用 A/B 对照证实。没有发现新缺陷。从验证角度看可以合并。 验证 head 环境链路为 Chromium → daemon 自带的生产版 Web Shell → 从合并树打包的真实 A/B:发现 1 的修复确实起作用两个臂只差一处:用
刷新后两个臂都恢复为 11 行,这与第 2 轮的判断一致:问题只出在本页实时输入的块上。 针对第 2 轮之后新代码的探针(head 全部通过)
关于
|
| Probe | 6fd5e61 reverted |
head 391cba6 |
|---|---|---|
| New session, 4 turns typed in the tab (8 messages): entry visible? | yes (fig. 1) | no; hidden at 10, shown at 12 |
| 9 replayed + 2 typed user messages, query "Question #" | 13 rows, #10/#11 twice (fig. 2) | 11 rows, no duplicates |
| Typed user message / pre-tool assistant text | 2 rows each (P1a, P2a fail) | 1 row each (10/10) |
| 3 tool-call turns typed live, query "PRE-TOOL-NEEDLE" | 6 rows (fig. 3) | 3 rows |
After a reload both arms show 11 rows, which matches round 2's diagnosis: only blocks typed live in this tab were affected.
Fig. 1: 8 messages typed live. The reverted arm shows the icon; head hides it.

Fig. 2: duplicate rows. Reverted: 3 / 13. Head: 1 / 11.

Fig. 3: pre-tool assistant text in 3 live tool turns. Reverted: 6 rows. Head: 3.

New probes against the post-round-2 code (all pass on head)
- N1, streaming (6/6): I opened the dialog while a reply was still streaming (~8 s) and sampled every 250 ms. The dialog never listed more than 1 row. Text from the last chunk is found while the dialog stays open. After a reload there is still 1 row, so live and replay agree.
- N2, counting tool-call turns (5/5): with 3 tool turns (9 messages) typed live and no reload, the entry is hidden. It appears at the 11th message, which matches the 11 persisted messages on disk. Live and replay agree at every step.
- N3/N4, identical text (3/3): the same user prompt sent twice, and identical pre-tool text in two tool turns, give 2 rows each, both live and after a reload. The echo matching goes by identity, not text: it doesn't merge equal text across turns.
- N5, daemon restart during a deep-history navigation (4/4): the delayed navigation never lands on the stale target later. Caveat: a daemon restart sends Web Shell back to
/, and the rail and entry go with it. Web Shell behaves the same way when no search is involved, so this is not this PR. After reopening the session, search and navigation work. - 5 s offline blip during a navigation: the session is kept, the viewport stays live, there is no stale jump and no page error. The dialog stays open and says "This result could not be located. Search again to refresh it." (fig. 4). Clicking the same row once back online lands on turn 如何自定义密钥文件 .env可能与其他文件冲突 #3 in about 300 ms.
Fig. 4: dialog after the offline blip.
Note on 0c73062 (non-blocking)
In both real scenarios I could build (offline blip, daemon restart), a build with 0c73062 reverted behaves exactly like head (3/3 runs), because older guards already handle those cases. So I could not show that this guard matters on a real stack. It targets a narrower race, and the unit tests pin it: reverting it turns 3 cases red.
Regression and gates on the merged tree
- 600-turn / 1200-message cold-load suite 26/26, threshold group 3/3, 390 px layout 1/1. The first pass had two failures, both in my harness: S1a checked the timeline on the first frame even though it is present (count=1), and S11a reused a session that the previous run had already extended by a turn. With the harness fixed, everything passes.
- Unit tests: the PR's 7 test files, 234/234. Mutation: reverting
0c73062turns 3 cases red and reverting6fd5e61turns 8 red. One of the 8 is "counts live users without IDs and pre-tool assistant fragments with only prompt identity once", the real-world shape round 2 asked for. - The PR's own Playwright specs: 21/21 (chromium; conversation search, message navigation, history viewport).
pnpm install --frozen-lockfile, root build and bundle pass; web-shelltsc --noEmitexits 0;eslint --max-warnings 0on the 24 changed TS/TSX files exits 0; Prettier is clean.
Round-2 observations (deferred by the author, non-blocking)
- The 200-row cap still keeps the oldest matches: with 600 hits the dialog shows turns pre-release: fix ci #1–希望可以加入类似claude code 的sub-agent系统 #200.
- Each query still re-reads the whole history: 21
/transcript+ 4/turn-indexrequests at 600 turns, about 0.55 s on loopback.
Not covered
- Windows/Linux desktops; WebKit/Firefox against a real daemon.
- I did not re-run the public
navigateToMessageAPI group or the older-daemon group this round. Both passed in round 2, and nothing since has changed their external behaviour. - A real model; search after compaction.
review-prwas still running when I posted, and the merge state is BLOCKED by the bot's earlier CHANGES_REQUESTED.
Scripts and raw results: r3/harness, r3/results.
|
Round 3 addendum: real model (qwen3.8-max) + fine-grained mutation. No new defects; this supports the merge-ready verdict in round 3. Same head 中文说明第 3 轮补充:真实模型(qwen3.8-max)+ 细粒度变异。没有发现新缺陷,支持第 3 轮"可合并"的结论。 与第 3 轮是同一个 head 真实模型会话Chromium → daemon 自带的生产版 Web Shell → 由合并树打包的真实
细粒度变异矩阵(第 2 轮之后的 3 个修复提交)第 3 轮是按整个提交回退的。这里对
建议(不阻塞):为 N6、N7 各加一个用例:一个是 其他
未覆盖
Real-model sessionChromium → the daemon-served production Web Shell → a real
Fig. 1: scoreboard: live vs. cold replay vs. JSONL truth, and the mutation matrix. Fig. 2: real-model session, query Fig. 3: after choosing the Fine-grained mutation (the 3 fix commits after round 2)Round 3 reverted whole commits. Here I made 11 single-point mutants of
Suggestion (non-blocking): add one test for each of N6 and N7. For N6: the threshold is re-probed when Other
Not covered
Scripts and raw results: |
Reconcile main's in-conversation search timeline action (#12234) with this PR's MCP App page budgets and session-scoped App context. MessageList keeps the McpAppSessionContext provider tree and re-adds the timeline action prop on SessionTimeline; TranscriptViewport passes both the App session id and the navigation-suppressed timeline action; transcript-page-table fits the anchored page within budget before resolving main's exact-record block lookup.























What this PR does
Adds a compact search icon below the left Web Shell session timeline. The 14px icon uses a narrow transparent button sized for the navigation gutter. It opens a dialog for literal, case-insensitive searches of user and assistant messages, including code. Results show highlighted snippets and jump to the matching message with a temporary highlight, including messages outside the loaded live window. The search icon shares the timeline visibility and is hidden whenever the timeline is hidden.
The host option
conversationSearchThresholddefaults to10: the icon is hidden at 10 messages and appears at 11. Historical messages and newly admitted live messages count toward the threshold. Searches read persisted history page by page and retain at most 200 snippets; navigation reuses the existing historical viewport and its cache budget. The dialog supports result navigation, Escape, focus restoration, error/retry states, and English/Chinese labels.Why it's needed
In a long conversation, users often remember a keyword but cannot quickly find the original answer or code. The sidebar content search delivered in #10612 finds sessions; this change locates content inside the session already open. It is scoped separately from the CLI/VS Code requests in #6824 and #11111.
Reviewer Test Plan
How to verify
conversationSearchThreshold={20}in an embedding: the boundary moves to 20/21.Evidence (Before & After)
Screenshots / 界面截图
Synthetic fixtures only; no private conversation data. / 仅使用合成数据,不含私人会话。
Timeline-aligned search entry / 与时间轴刻度对齐的搜索入口
Stable-height search dialog, dark / 固定高度搜索弹框(深色)
Chinese, light / 中文浅色弹框
The search entry hides with the timeline. Input, matching results, and empty results keep the dialog height and vertical position stable. Latest verification: 11 browser tests passed, including alignment, hidden narrow-layout entry, and stable dialog geometry.
搜索入口跟随时间轴显隐;输入、有结果和无结果时弹框高度与垂直位置保持稳定。最新 11 项浏览器测试通过,覆盖图标对齐、窄屏隐藏入口及弹框尺寸稳定性。
Before: the conversation footer only offered scroll-to-bottom; finding an old passage required manual scrolling. After: synthetic browser tests find an assistant message outside the live replay, navigate to its persisted identity, observe the highlight, and confirm the unsent draft is unchanged. The 10/11 threshold and all eight theme/language/width combinations are covered by browser tests. All fixtures are synthetic; no private conversations or original user screenshots are included.
Earlier validation at
713d23d095: 97 focused unit tests, 11 browser tests, and 2 visual scenarios (four light/dark captures) passed. Root build, typecheck, bundle, and changed-file lint/format checks passed locally. All executed GitHub CI checks subsequently passed at9f8c93e302; skipped jobs were not executed.Tested on
Environment (optional)
Local Web Shell with the repository's synthetic daemon Playwright harness; Chromium at 1440px and 390px. No real model calls are required by these tests.
Risk & Scope
External message navigation / 对外消息定位能力
Implemented in this PR:
shellRef.current.navigateToMessage({ sessionId, recordId, signal? }), withWebShellMessageNavigationRequestandWebShellMessageNavigationResultexported from the package root. Both WebShell embedding components expose it through the existingWebShellApi.The host first opens the target session/workspace through the existing provider lifecycle, then invokes navigation with the stable persisted user/assistant transcript record ID. WebShell resolves that exact record, loads older history outside the live/rendered window, and activates the existing scroll/highlight behavior. The operation also works when the timeline/search icon is hidden, without opening the search dialog or meeting its message-count threshold.
Results distinguish
located,not_found,not_ready,session_mismatch,unsupported,cancelled, anderror.locatedmeans the target has been activated; React applies the scroll/highlight on the following render, so this is not an animation-completion signal. An accepted new request, AbortSignal, changed session/workspace/transcript owner, or unmount invalidates stale work. Rejected calls from stale API owners, already-aborted signals, mismatched sessions, or empty record IDs do not cancel an active valid navigation. Identical text does not substitute for record identity. Drafts and ongoing responses are untouched.Resolution currently scans persisted history until the exact record is found, retaining one hit and reusing the bounded history viewport. Very old targets incur linear reads; no server-side record index or new daemon endpoint is introduced.
Integration boundary: the existing cross-session search endpoint still returns one session plus one snippet, without a message record ID. Host search results must supply that ID before they can use exact navigation. This PR does not add multiple-message search results, archive search, or downstream host UI changes. A host must handle
not_readyand wait for the selected session/history viewport before trying again; navigation itself never changes the active session.Earlier validation at
9f8c93e302: root build, typecheck, bundle, and the real pre-commit hook passed locally; 1,008 App tests and 87 history/store tests passed. All executed GitHub CI checks at this commit passed (including Ubuntu unit tests, lint/static, integration, desktop checks, screenshots and E2E smoke).Verification: 23 hook tests cover cancellation/ownership/readiness/errors; store tests cover exact IDs, duplicate text, stopping at the target, and missing records. Four real embedding browser scenarios call the public API directly with the timeline hidden, assert historical content and visible highlight, preserve the draft, and verify a missing record does not change the displayed content. The original 11 search-browser scenarios plus a new ArrowUp/ArrowDown wrap-around, previous/next button, and Enter-to-locate scenario pass: 16 browser scenarios in total including the four public-API cases.
中文:本 PR 已实现通过 shellRef 对外定位消息。宿主打开目标会话后传入稳定的持久化记录 ID,WebShell 加载未渲染历史并滚动高亮,不依赖搜索弹框、图标显隐或阈值。返回明确状态,取消及切换会话/工作目录后旧请求失效,保留草稿。现有跨会话搜索仍没有消息记录 ID,也不支持归档搜索;宿主搜索接入是另外的改动,不能把本 PR 的公开定位能力等同于已完成整个宿主功能。
Public API browser evidence / 公开接口本地截图
Synthetic embedding harness; public API invoked directly, timeline and search trigger hidden. The highlighted target is outside the initial live window; the composer still contains the unsent synthetic draft. / 合成嵌入场景,直接调用公开接口,旧消息已加载高亮且草稿保留。
Linked Issues
Closes #12231
中文说明
本 PR 做了什么
在 Web Shell 左侧会话时间轴下方新增 14px 小尺寸搜索图标,按钮采用适配窄栏的透明样式,点击弹框按字面量、不区分大小写搜索用户和助手消息,包括代码。结果显示高亮摘要,点击后定位并短暂高亮对应消息,包含当前 live window 之外的历史内容。搜索入口跟随时间轴显隐,时间轴隐藏时也隐藏。
宿主配置
conversationSearchThreshold默认为10:10 条消息时隐藏,11 条时显示。计数包含历史消息和新接纳的实时消息。搜索分页读取持久化历史,最多保留 200 条摘要;定位复用现有历史浏览窗口和缓存预算。弹框支持结果导航、Escape、焦点恢复、错误与重试状态,以及中英文文案。为什么需要
长会话中,用户经常记得关键词,却难以找到原来的回答或代码。#10612 已实现的侧边栏内容搜索用于查找会话,本次用于定位已经打开的会话内部内容,与 #6824、#11111 的 CLI / VS Code 需求分开处理。
评审验证计划
如何验证
conversationSearchThreshold={20},显示边界应变为 20/21。证据(前后对比)
改动前,会话底部只有置底功能,找回旧内容需要手动滚动。改动后,合成数据浏览器测试能找到 live replay 之外的助手消息,按持久化身份定位,观察高亮,并确认未发送的草稿保持不变。浏览器测试覆盖 10/11 条边界和全部八种主题/语言/宽度组合。所有测试数据均为合成数据,不包含私人会话或用户原始截图。
此前提交
713d23d095验证:97 项针对性单测、11 项浏览器测试、2 项截图场景(浅深主题共四张)通过;根目录 build、typecheck、bundle 和改动文件 lint/format 本地通过。随后9f8c93e302的已执行 CI 检查全部通过;跳过项未执行。测试平台
环境(可选)
本地 Web Shell 配合仓库的合成 daemon Playwright 测试设施;Chromium 宽度为 1440px 和 390px。这些测试不需要真实模型调用。
风险与范围
关联 Issue
Closes #12231