Skip to content

/review presubmit overlap matching is exact-line only — multi-line ranges and semantic duplicates pass as noConflict #9219

Description

@wenshao

What happened?

Presubmit's existing-comment overlap detection matches (path, line) by exact line equality only. On a manual review of PR #9204 (2026-08-15, commit 40ad8fd) this missed a duplicate in two ways:

  • Range blindness. The drafted finding was a multi-line inline comment (start_line: 2554, line: 2562). An existing comment sat at line 2558 — inside the range. Exact-line comparison of the two line fields (2558 vs 2562) found no conflict, so the report bucketed it noConflict.
  • Semantic blindness. Two other confirmed findings were genuine duplicates of existing comments anchored elsewhere (one at a different file's line, one at a nearby line of the same file). Presubmit reported noConflict; deciding they were duplicates required manually fetching each existing comment's full body with separate gh api calls (the report renders only ~60-character body prefixes) and comparing by hand.

The duplicates were eventually caught before posting, but only through manual work outside the tooling.

What did you expect to happen?

  • Overlap matching treats a drafted multi-line comment's range (start_line..line) as an interval and flags an existing comment whose line falls inside it.
  • The report renders full bodies of the comments it buckets (or offers a --full-bodies mode), so ruling on semantic duplication does not require one gh api call per comment.
  • Optionally: a possibleDuplicate bucket for same-file proximity (same path, within N lines) that the operator rules on — distinct from the deterministic overlap bucket, which stays exact.

Client information

Client Information
$ qwen /about
qwen-code v0.21.12 (npm install), glm-5.3
Platform: Linux 6.12.63+deb13-amd64 · Node v22.22.2 · npm 10.9.7
Session: interactive /review skill run, worktree mode, high effort

Login information

N/A — API-key/model-service session.

Anything else we need to know?

This is the mirror image of #9208 (overlap-drop false positives: carried-id re-posts swallowed, same-line distinct claims dropped). Together they say location-equality is the wrong equivalence in both directions: too weak for range/semantic duplicates (this issue), too strong for same-line distinct claims (#9208). Any fix should consider both directions — e.g. range-intersection plus content-aware exemption for carried-id prefixes.

中文

发生了什么?

presubmit 的既有评论重叠检测只按 (path, line) 精确相等匹配。在 PR #9204 的手动审查中(2026-08-15,commit 40ad8fd)以两种方式漏检了重复:

  • 范围盲区。 草稿发现是多行行内评论(start_line: 2554、line: 2562)。一条既有评论位于 2558 行——在范围内。两个 line 字段(2558 vs 2562)精确比较无冲突,报告归入 noConflict。
  • 语义盲区。 另外两条确认发现是既有评论的真重复,但锚在别处(一条在另一文件的某行,一条在同文件邻近行)。presubmit 报 noConflict;判定为重复需要逐条 gh api 手工抓取既有评论全文(报告只渲染约 60 字符的 body 前缀)并人工比对。

重复最终在发布前被拦下,但只靠工具之外的手工工作。

期望的行为?

  • 重叠匹配将草稿多行评论的范围(start_line..line)视为区间,既有评论的行落在其中即标记。
  • 报告渲染所归档评论的完整 body(或提供 --full-bodies 模式),语义查重不再需要每条评论一次 gh api。
  • 可选:为同文件邻近(同 path、±N 行)提供 possibleDuplicate 桶,由操作者裁决——与保持精确的确定性 overlap 桶区分。

其他

这是 #9208(overlap 误杀:carried-id 重发被吞、同线不同主张被丢)的镜像。两者合起来说明位置相等在两个方向上都是错误的等价关系:对范围/语义重复太弱(本 issue),对同线不同主张太强(#9208)。任何修复应同时考虑两个方向——例如区间相交 + 对 carried-id 前缀的内容感知豁免。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    category/developmentDevelopment experiencepriority/P2Medium - Moderately impactful, noticeable problemscope/commandsCommand implementationtype/bugSomething isn't working as expected

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions