Skip to content

feat(managed-agent): Implement event replay (Stage D3) - #12840

Merged
wenshao merged 8 commits into
mainfrom
feat/managed-agent-event-replay
Sep 27, 2026
Merged

wenshao merged 8 commits into
mainfrom
feat/managed-agent-event-replay

Conversation

@wenshao

@wenshao wenshao commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

This is slice D3 of #12793, the last one. It implements event replay as the contract defines it and moves the public event query and stream, the WebShell event stream and the WebShell transcript to implemented (contract v1.17).

  • Events keep what they were accepted with. Every event now stores its schema_version and projection_version and a top-level item_id and content_part_id, on both surfaces. Replay returns the stored values, so a later projector version will not rewrite history. The one exception is a retraction during Harness recovery, which already rewrote the retracted text and is announced by stream.reconciled; the contract now says so.

    • The identity follows the rule that the Items projection already uses, so a text delta names the Part that the Snapshot holds. The existing data.contentPartId (part_<turn>_<type>) does not match the Snapshot's Part id (part_<turn>_<type>_<first sequence>), so it is left as it was rather than moved up.
    • The store assigns the identity when it appends an event. Only a text delta needs the event before it, which it reads by primary key.
    • A Flyway Java migration gives events written before this change their identity with the same rule, reading each Session in pages and writing each run of equal identities with one ranged update.
    • A retraction during Harness recovery empties deltas that later deltas may have continued, so the store derives the identities from the first retracted delta on again, as the rebuilt Items name them.
  • JSON pages are real pages. The event query returns has_more and a next_cursor to pass as after, and it accepts limits up to 1000, as does the WebShell transcript.

  • A replay floor answers expired cursors. Each Session stores a replay floor, and replay_floor_sequence reports it.

    • A JSON cursor below the floor gets 409 cursor_expired, whose envelope carries replay_floor_sequence and snapshot_through_sequence.
    • Both streams check the floor whenever they read the store: at the start, after the in-memory hub overflowed, and after an idle wait. An expired cursor gets one agent.session.resync_required frame without an id, and the stream closes.
    • The floor never rises above the Snapshot's covered sequence, so a client that reloads the Snapshot can resume. Nothing raises it in production yet; pruning belongs to the retention work, and the tests raise it through the store.
  • The contract describes the resync frame. v1.17 adds SessionResyncRequired and WebShellResyncRequired, referenced from each stream through x-qwen-resync-frame; a public client resumes after the snapshot_through_sequence that the Items list returns. It drops the 409 that the WebShell stream declared but never returned, since an expired stream cursor is answered in-band. It widens content_part_id to the 128 characters that part_id already allows, and it states what the WebShell transcript returns without a cursor.

  • Sessions advertise snapshots and resync.

  • The web-shell client handles the frame. It turns the resync frame into the existing stream-gap event, which reloads the transcript and resumes after it.

  • Tests cover the exit check. A new replay suite covers the following; every streaming and paging case requires each sequence exactly once and in order:

    • JSON paging through next_cursor;
    • Last-Event-ID taking precedence over after and resuming across catch-up pages into live events;
    • a catch-up racing a writer;
    • a stuck stream that falls 600 events behind, more than the 512 the hub keeps;
    • a floor raised past a lagging stream;
    • expired cursors on the JSON query and both streams.

    The suite sets the poll and heartbeat intervals to one minute, so live events must come through the hub. The contract test validates resync frames. The parity test checks that every Item and Part that an event names exists in the Snapshot. The identity rule, a cross-batch Part, single appends, retractions and the upgrade backfill have their own tests.

  • A bilingual design note records the decisions; the D1 note and the server README point to it.

Why it's needed

#12793 defines D3's exit check: these routes become implemented, with tests for Last-Event-ID resumption, the switch from catch-up to live, subscriber overflow and a floor advanced by the test, and no duplicates, gaps or wrong resets. Before this change a client paging events never saw has_more, could not learn that it had fallen behind a replay floor, and could not tell which projector version produced an event or which Snapshot Part a delta extends.

Reviewer Test Plan

How to verify

  • Run the Managed Agent server's tests. The contract test passes with 3 remaining gap lines, all for the lifecycle work, and ManagedEventReplayTest passes.
  • Page a Session's events with limit=7, following next_cursor: every sequence appears once and the last page has has_more: false. limit=1000 is accepted and limit=1001 is 400 invalid_limit.
  • Raise a Session's floor through ManagedAgentStore.advanceReplayFloor. A JSON read below it returns 409 cursor_expired with both watermarks, a read at it returns 200, and both streams send one agent.session.resync_required frame without an id and close.
  • Read a finished Session's events and Items: every item_id and content_part_id that an event carries names an Item and a Part in the Items list.
  • In the web-shell package, the provider turns a resync frame into one stream_gap and stops.

Evidence (Before & After)

N/A (no visible UI change).

Local results on Linux:

  • Managed Agent server: the full mvn test suite (135 tests after merging main) plus Checkstyle passes.
  • MariaDB and MySQL: ManagedAgentMySqlIT passes against mariadb:10.11.18, the image CI uses, and against mysql:8.4, including the backfill of pre-V14 events and the replay floor.
  • Upgrade cost: the packaged jar migrated 200,000 legacy events in 100 Sessions on MariaDB in under three seconds (V14 and V15 together).
  • Packaged jar: the Spring Boot jar migrates a MariaDB database to V15, and the history shows V15 as a JDBC migration, so Flyway finds the Java migration inside the jar.
  • Web Shell: typecheck, eslint, the managed component tests (73) including the generated-types freshness test, and the managed-progress and managed-workspace-w0d e2e specs (Chromium, the latter against WorkspaceBrowserFixtureMain) pass.
  • Mutations: each of the following fails the matching tests:
    • ignoring Last-Event-ID;
    • skipping the floor check, and sending the resync frame twice;
    • a hub that hides its overflow, and a hub that never delivers;
    • returning no next_cursor, and reporting has_more on a full last page;
    • starting a new Part for every delta, and ignoring the previous event on a single append;
    • skipping the identity rederivation after a retraction;
    • skipping the V15 backfill, merging V15 ranges across a gap, and losing V15's state at a page boundary;
    • ignoring the resync frame in the web-shell provider.
  • Audit: four rounds of review, each with one undirected and one adversarial reviewer. Rounds 1 to 3 led to the fixes in the follow-up commits (identity after a retraction, V15 ranged and paged updates, tests that reach the hub, contract wording); the reviewers also ran randomized histories, concurrent writers on MariaDB and MySQL, and large V15 upgrades. The last round reported no Critical issue.

Tested on

OS Status
🍏 macOS ⚠️
🪟 Windows ⚠️
🐧 Linux ✅

Environment (optional)

Java 25 running the Java 21 build with Maven 3.9.9; MariaDB 10.11.18 and MySQL 8.4 in Docker for the integration test; Node 22 with pnpm from packageManager.

Risk & Scope

  • Main risk or tradeoff:
    • V15 lists the Sessions that have events and reads each one in pages of 5,000, all in the migration's transaction, and updates each run of equal identities with one statement. A table with many short runs takes longer than the benchmark above.
    • Upgrade all replicas together: a replica still on the previous version after V14 writes events without an identity.
    • A client that ignores the resync frame reconnects with the same cursor and gets the frame again. No production stream sends it until the retention work raises a floor.
  • Not validated / out of scope:
    • Pruning events and deciding when the floor rises.
    • Keeping the floor at or below the Snapshot while a stream reconciliation rebuilds the Items, and ending the WebShell transcript's older pages at the floor; both are follow-ups for the retention work.
    • The Items list reads the current Snapshot for every page, so a client that needs several pages can see two Snapshot versions; pinning one is the remaining listItems work.
    • The WebShell Session's floor and Snapshot fields, which stay planned.
    • Paged Snapshot versions for listItems.
    • The lifecycle gap lines.
    • macOS and Windows were not run locally.
  • Breaking changes / migration notes:
    • Flyway V14 adds columns with defaults; V15 backfills event identity.
    • Responses only gain fields.
    • The WebShell stream's 409 response leaves the contract; the server never returned it.
    • The generated web-shell types add WebShellResyncRequired; the optional event fields were already there.
  • Merge note: the Stage H task contract (feat(managed-agent): Add the planned Stage H task contract (H0a) #12830) took v1.16 on main, so this change is v1.17 and applies its edits on top of main's specification.

Design: English | 简体中文. Both versions carry the same decisions, compatibility notes and follow-up work.

Linked Issues

Closes #12793 (slice D3; D1 was #12808, D2 was #12822). Part of #12380.

中文说明

本 PR 做了什么

这是 #12793 的 D3 切片,也是最后一个切片。它按契约实现事件回放,并把公共事件查询与事件流、WebShell 事件流与 WebShell transcript 改为 implemented(契约 v1.17)。

  • 事件保留被接受时的内容。 每条事件现在都在两个入口上保存 schema_version、projection_version 以及顶层的 item_id 与 content_part_id。回放返回已存储的值,因此以后的投影器版本不会改写历史。唯一的例外是 Harness 恢复期间的撤回:它本来就会改写被撤回的文本,并以 stream.reconciled 告知;契约现在写明了这一点。

    • 身份遵循 Items 投影已采用的规则,因此文本增量指向 Snapshot 中的那个 Part。现有的 data.contentPartId(part_<turn>_<type>)与 Snapshot 的 Part id(part_<turn>_<type>_<首个 sequence>)不一致,因此保持原样,没有直接上移。
    • 存储层在追加事件时确定身份。只有文本增量需要它之前的那条事件,按主键读取。
    • 一个 Flyway Java 迁移按同一规则为本次变更之前写入的事件补上身份,逐个 Session 分页读取,身份相同的连续事件用一条范围更新写入。
    • Harness 恢复期间的撤回会清空之后的增量可能接续过的增量,因此存储层从第一条被撤回的增量起重新推导身份,与重建后的 Items 的命名一致。
  • JSON 分页是真实分页。 事件查询返回 has_more 与可作为 after 传入的 next_cursor,并接受最大 1000 的 limit;WebShell transcript 同样如此。

  • 回放下限回应过期游标。 每个 Session 保存一个回放下限,replay_floor_sequence 返回它。

    • 低于下限的 JSON 游标收到 409 cursor_expired,错误信封带有 replay_floor_sequence 与 snapshot_through_sequence。
    • 两个事件流每次读取存储时都检查下限:开始时、内存 hub 溢出后、以及空闲等待之后。过期游标收到一帧不带 id 的 agent.session.resync_required,随后事件流关闭。
    • 下限不会超过 Snapshot 已覆盖的 sequence,因此重新读取 Snapshot 的客户端可以继续。生产环境中目前没有路径提升它;清理属于保留策略的工作,测试通过存储层提升它。
  • 契约描述 resync 帧。 v1.17 新增 SessionResyncRequired 与 WebShellResyncRequired,各事件流通过 x-qwen-resync-frame 引用;公共客户端从 Items 列表返回的 snapshot_through_sequence 之后继续。它删除了 WebShell 事件流声明过、但从未返回的 409,因为过期的事件流游标在流内回应。它把 content_part_id 放宽到 part_id 已允许的 128 个字符,并写明 WebShell transcript 在没有游标时返回的内容。

  • Session 声明 snapshots 与 resync。

  • web-shell 客户端处理该帧。 它把 resync 帧转换为已有的 stream gap 事件,从而重新读取 transcript 并在其后继续。

  • 测试覆盖验收条件。 新的回放测试覆盖以下场景;每个事件流与翻页场景都要求每个 sequence 恰好出现一次且有序:

    • 按 next_cursor 进行 JSON 翻页;
    • Last-Event-ID 优先于 after,并跨越历史补齐分页进入实时事件;
    • 与写入方竞争的历史补齐;
    • 落后 600 条事件、超过 hub 保留的 512 条的卡住的事件流;
    • 越过落后事件流的下限;
    • JSON 查询与两个事件流上的过期游标。

    该测试集把轮询与心跳间隔都设为一分钟,因此实时事件必须经由 hub。契约测试校验 resync 帧。一致性测试检查事件所指向的每个 Item 与 Part 都存在于 Snapshot 中。身份规则、跨批次的 Part、逐条追加、撤回以及升级回填都有各自的测试。

  • 中英文设计说明记录了各项决策;D1 设计说明与服务端 README 都已指向它。

为什么需要

#12793 为 D3 规定的验收条件是:这些路由改为 implemented,测试覆盖 Last-Event-ID 续传、历史补齐切换到实时、订阅溢出以及由测试推进的下限,且不重复、不跳序、不误重置。在本次变更之前,翻页读取事件的客户端永远看不到 has_more,无法得知自己已落后于回放下限,也无法分辨事件由哪个投影版本产生、增量扩展的是 Snapshot 中的哪个 Part。

评审测试计划

如何验证

  • 运行 Managed Agent 服务的测试。契约测试在剩余 3 行差异下通过,这 3 行都属于生命周期工作;ManagedEventReplayTest 通过。
  • 以 limit=7 按 next_cursor 翻页读取一个 Session 的事件:每个 sequence 只出现一次,最后一页 has_more: false。limit=1000 被接受,limit=1001 返回 400 invalid_limit。
  • 通过 ManagedAgentStore.advanceReplayFloor 提升一个 Session 的下限。低于下限的 JSON 读取返回带两个水位的 409 cursor_expired,在下限处读取返回 200,两个事件流都发送一帧不带 id 的 agent.session.resync_required 后关闭。
  • 读取一个已完成 Session 的事件与 Items:事件携带的每个 item_id 与 content_part_id 都指向 Items 列表中的 Item 与 Part。
  • 在 web-shell 包中,provider 把 resync 帧转换为一个 stream_gap 后停止。

证据(前后对比)

N/A(没有可见的 UI 变化)。

Linux 本地结果:

  • Managed Agent 服务: 完整 mvn test(合入 main 后共 135 个测试)加 Checkstyle 通过。
  • MariaDB 与 MySQL: ManagedAgentMySqlIT 在 CI 所用的 mariadb:10.11.18 以及 mysql:8.4 上通过,包括 V14 之前事件的回填与回放下限。
  • 升级开销: 打包后的 jar 在 MariaDB 上用不到三秒迁移了 100 个 Session 中的 20 万条旧事件(V14 与 V15 合计)。
  • 打包后的 jar: Spring Boot jar 把 MariaDB 数据库迁移到 V15,迁移历史中 V15 的类型为 JDBC,说明 Flyway 能在 jar 中找到这个 Java 迁移。
  • Web Shell: typecheck、eslint、managed 组件测试(73 个,含生成类型的过期检查)以及 managed-progress 与 managed-workspace-w0d e2e 用例(Chromium,后者针对 WorkspaceBrowserFixtureMain)通过。
  • 变异验证: 以下变异都会让对应测试失败:
    • 忽略 Last-Event-ID;
    • 跳过下限检查,以及把 resync 帧发送两次;
    • hub 隐瞒溢出,以及 hub 从不投递;
    • 不返回 next_cursor,以及满页的最后一页报告 has_more;
    • 每个增量都新建 Part,以及逐条追加时忽略上一条事件;
    • 撤回后跳过身份的重新推导;
    • 跳过 V15 回填、V15 越过缺口合并范围、V15 在分页边界丢失状态;
    • web-shell provider 忽略 resync 帧。
  • 审计: 共四轮评审,每轮包含一次无方向评审与一次对抗式评审。前三轮的发现由后续提交修复(撤回后的身份、V15 的范围更新与分页、真正经过 hub 的测试、契约措辞);评审者还运行了随机历史、MariaDB 与 MySQL 上的并发写入以及大规模 V15 升级。最后一轮没有报告 Critical 问题。

测试环境

系统 状态
🍏 macOS ⚠️
🪟 Windows ⚠️
🐧 Linux ✅

环境(可选)

Java 25 运行 Java 21 构建,Maven 3.9.9;集成测试使用 Docker 中的 MariaDB 10.11.18 与 MySQL 8.4;Node 22,pnpm 使用 packageManager 固定的版本。

风险与范围

  • 主要风险或取舍:
    • V15 在迁移事务内先列出有事件的 Session,再按每页 5000 条读取每个 Session,并对每一段身份相同的事件执行一条更新。短片段很多的表会比上面的基准更慢。
    • 所有副本需要一起升级:V14 之后仍运行旧版本的副本写入的事件没有身份。
    • 忽略 resync 帧的客户端会用同一个游标重连并再次收到该帧。在保留策略的工作提升下限之前,生产环境中的事件流不会发送它。
  • 未验证 / 不在范围内:
    • 清理事件以及下限何时提升。
    • 在事件流重整重建 Items 期间把下限保持在 Snapshot 之下或与之相等,以及让 WebShell transcript 的更早分页止于下限;两者都是保留策略工作的后续事项。
    • Items 列表的每一页都读取当前 Snapshot,因此需要多页的客户端可能看到两个 Snapshot 版本;固定同一版本是 listItems 剩余的工作。
    • WebShell Session 的下限与 Snapshot 字段,仍为 planned。
    • listItems 的 Snapshot 版本分页。
    • 生命周期差异行。
    • macOS 与 Windows 未在本地运行。
  • 破坏性变更 / 迁移说明:
    • Flyway V14 新增带默认值的列;V15 回填事件身份。
    • 响应只新增字段。
    • WebShell 事件流的 409 响应从契约中移除;服务端从未返回过它。
    • 生成的 web-shell 类型新增 WebShellResyncRequired;可选的事件字段原本就存在。
  • 合入说明: main 上的 Stage H task 契约(feat(managed-agent): Add the planned Stage H task contract (H0a) #12830)已使用 v1.16,因此本次变更为 v1.17,并在 main 的 spec 之上应用其修改。

设计文档:English | 简体中文。两个版本的决策、兼容性说明与后续工作一致。

关联 Issue

关闭 #12793(D3 切片;D1 为 #12808,D2 为 #12822)。属于 #12380。

Close the D3 gaps of the Managed Agent contract and move the public event
query and stream, the WebShell event stream and the WebShell transcript to
implemented (contract v1.16).

- Events store their schema and projection versions and the Item and Part
  they change (Flyway V14). The identity follows the Items projection, so a
  text delta names the Snapshot's Part rather than data.contentPartId. A Java
  migration (V15) backfills events written before V14 with the same rule.
- JSON event pages return has_more and next_cursor and accept limits up to
  1000, as does the WebShell transcript.
- A persisted per-Session replay floor answers expired cursors with
  409 cursor_expired in JSON and one agent.session.resync_required frame on
  both streams. The floor never exceeds the Snapshot's covered sequence;
  nothing raises it in production until the retention work.
- Sessions report the stored floor and advertise snapshots and resync.
- The web-shell client turns the resync frame into a stream gap, which
  reloads the transcript.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions github-actions Bot added the review/self-reported The linked issue was opened by the PR author (self-reported) label Sep 27, 2026
- Derive event identities again after a retraction, from the first
  retracted delta on, so emptied deltas name nothing and a kept delta that
  continued one names the Part the rebuilt Items give it.
- Write each run of equal identities in V15 with one ranged update; 200,000
  legacy events in 100 Sessions migrate in under three seconds on MariaDB.
- Make the replay tests reach the hub: poll the store every 30 seconds, wait
  for catch-up before live appends, and cover a full last JSON page and the
  single-append path.
- State in the contract where a client resumes after a resync and what the
  WebShell transcript returns without a cursor.
- Correct the design note: the hub keeps 512 events, the optional event
  fields already existed, upgrade all replicas together, and the floor and
  transcript caveats that the retention work must close.
@wenshao
wenshao requested a lite review from Copilot September 27, 2026 11:49

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@wenshao
wenshao requested a lite review from Copilot September 27, 2026 11:51

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

- Read each Session's events in V15 in pages of 5,000, carrying the run
  and the previous event across pages, so memory does not grow with a
  Session's length; cover a Part longer than a page and two runs of one
  tool Item separated by an event without identity.
- Keep the event query's old order of checks: a missing Session answers
  404 before an invalid limit answers 400.
- State in the contract that events replay as accepted except after a
  stream.reconciled event, which the retraction now also uses to announce
  rewritten identities.
- Set the replay tests' heartbeat to one minute too, so an idle stream does
  not read the store during a test.
@wenshao
wenshao requested a lite review from Copilot September 27, 2026 12:22

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Resolve the OpenAPI conflict with the Stage H task contract (H0a, #12830),
which took v1.16. The spec keeps main's version of every node and applies
the event replay changes on top as v1.17: the three operations it marks
implemented, the event, event list and Session schemas it extends, and the
two resync frame schemas. The design note now names the contract v1.17.
@wenshao
wenshao requested a lite review from Copilot September 27, 2026 12:31

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

- Tell public clients to reload the Items after a stream.reconciled event
  until the Snapshot reaches that event, since it is rebuilt after it.
- Assert that the resync frame is the only frame of an expired stream, and
  give the service-level replay streams the same one-minute intervals.
- Say in the design note that a client in the rebuild window finds its
  cursor expired again through either 409 or another resync frame, and
  name the retraction exception in the server README.
@wenshao
wenshao requested a lite review from Copilot September 27, 2026 12:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

Real-stack verification: event replay (D3), final head cf72a892b4

Verdict: ready to merge. I found no blocker at cf72a892b4.

The head moved four times while I ran this: 42bdee4104 → d760dd572e → 349fb3ba28 → merge of main 6bba5132bc → cf72a892b4. I re-ran whatever each push touched. One serious problem turned up along the way: V15 ran out of memory on a long Session and then hung instead of failing. It was present at 42bdee4104 and d760dd572e, and 349fb3ba28 fixes it. Between 349fb3ba28 and cf72a892b4 the main Java code is identical, and the two jars differ only in the OpenAPI file (374/375 entry CRCs match).

On cf72a892b4, CI's Runtime Broker and Managed Agent MariaDB / Java 21 job, the only one that runs this module's tests, is green; other jobs were still running when I posted.

The triage bot left one claim unverified: that a client below the floor recovers instead of looping. I checked it in a real browser. The PR's Web Shell resumes about 25 ms after the resync frame; main's client loops (figure 1).

Rig

  • Server: the packaged Spring jar of each head, on MySQL 8.4.7 and on MariaDB 10.11.18, the image CI uses.
  • Harness and model: the packaged qwen serve --profile hosted-harness, driven by a scripted OpenAI-compatible model.
  • Clients: the real Web Shell (vite) in Chromium through Playwright, and raw HTTP/SSE clients.
  • Checking: Ajv validates every response body and SSE frame against the head's spec.
  • Floor: raised with SQL, because nothing in production calls advanceReplayFloor yet.
  • Baseline: main's jar (e68822c815) is the "before" arm, and the oracle for the upgrade check.

Results (final head unless noted)

Area Result
CI-equivalent mvn -Pmysql-integration clean verify checkstyle:check on MariaDB 10.11.18 135 unit + 10 IT pass, 0 Checkstyle violations. Every earlier head was green too.
Web Shell Regenerating the types gives no diff; 73/73 managed tests pass; typecheck is clean.
Paging and limits Paging limit=7 through next_cursor returns all 13 sequences once, in order. The last page has has_more:false and next_cursor:null. limit 1000 → 200; 1001 and 0 → 400 invalid_limit, for both the event query and the WebShell transcript. On main, the first page already says has_more:false, so only 7 of the 13 events are reachable, and limit=1000 → 400.
Identity on real Turns Every event carries schema_version 1 and projection_version 1. item_id and content_part_id name Items and Parts of the Snapshot, and each Part's deltas rebuild its text.
Replay floor Floor 8, Snapshot 13: after=7 → 409 cursor_expired with replay_floor_sequence, snapshot_through_sequence and request_id. after=8 → 200. Last-Event-ID wins over after in both directions. Both streams send exactly one agent.session.resync_required frame with no id, then close. All of it is Ajv-valid against 1.17.0.
Stuck client, hub overflow A client with a 4 KiB receive buffer stops reading while 4,000 events are written. The hub drops its range, and the stream re-reads the store from sequence 983 in pages of 100 (seen in the MySQL general log). All 4,000/4,000 events arrive once, in order.
Floor raised past a stuck stream 975 contiguous events arrive, then exactly one resync frame, then the server closes the stream.
Catch-up racing writers 8,601 events from Last-Event-ID: 0 (81 pages) while 300 renames write: every sequence once, in order.
Web Shell end to end (figure 1) PR client: resync frame → one transcript reload → resumes at 9; the new reply shows 27 ms after release. main's client: 5 stream requests in 15 s (the held one plus a reconnect every 3 s), all with cursor 5, no reload, and the reply never shows.
V13 → V15 upgrade Checked against main's Snapshot: MySQL 1,634 events / 435 text Parts, MariaDB 1,606 / 448, 0 disagreements. V15 runs as a JDBC migration found inside the jar.
V15 memory (F1) One 1.2M-event Session under a 256 MB heap: OOM, then a hang, at 42bdee4104/d760dd572e; migrates in 86 s at 349fb3ba28.
Retraction identities A randomized probe of 3,400 cases finds 0 failures at d760dd572e and 349fb3ba28; 356 of 400 cases fail at 42bdee4104.
Independent mutants 17/17 applicable mutants are killed at 349fb3ba28, whose code is identical at cf72a892b4.
Merge with main Your resolution to 1.17.0 matches my independent trial merge exactly (0 differing JSON nodes).

F1: fixed during review — V15 ran out of memory on long Sessions and then hung (42bdee4104, d760dd572e)

What happened. V15 read each Session with one unbounded SELECT, and Connector/J buffers the whole result. I migrated one Session of 1,200,000 legacy events under a 256 MB heap.

  • With -XX:+ExitOnOutOfMemoryError: V15 died with OutOfMemoryError: Java heap space right after Migrating ... to version "15".
  • Without that flag (the default): the JVM did not exit. It hung inside rollback() on the half-read connection:
    • the main thread sat in Flyway's rollback(), inside MySQLNamedLockTemplate, so it held Flyway's MySQL named lock;
    • old gen stayed at 99.8 %, and it stayed that way for 15 minutes at 42bdee4104;
    • it ignored SIGTERM, so SIGKILL was needed;
    • V15 was never recorded, so every restart repeats it.

The fix. 349fb3ba28 pages the reads (5,000 rows at a time, carrying the run and the previous event across pages).

  • Same test at the fix: the same Session migrates in 86 s under 256 MB, with the same identity counts.
  • Output bytes: on one shared 10,637-event dump (longest Session 9,509 events), the V15 of all three heads writes MD5-identical identity columns.
  • Tests: your new page-spanning test kills both paging mutants I tried: forgetting the previous event at a page edge, and reading only the first page.

No action needed.

Verified: identity after a retraction (d9dc6b730c) and the contract text (349fb3ba28, cf72a892b4)

The probe is a store-level randomized test:

  • each case builds 2–19 deltas from three Harness generations;
  • one generation is retracted;
  • the Snapshot is rebuilt;
  • the check: every stored identity must name the rebuilt Snapshot, and each Part's deltas must rebuild its text.

Results:

  • After the fix (d760dd572e and 349fb3ba28): 3,400 cases over 4 seeds, 0 failures.
  • Before the fix (42bdee4104): 356 of 400 cases fail, for example #5 empty delta names part_..._reasoning_5.

This answers the bot's content_part_id question: stored identities now resolve against the rebuilt Snapshot, and the event schemas state the stream.reconciled exception.

Non-blocking notes

  1. A rollback leaves rows without identity.
    • main's jar starts on a V15 schema. Flyway logs "has a version (15) that is newer than the latest available migration (13)" and continues.
    • main's jar then writes events with NULL item_id/content_part_id.
    • After you roll forward again, V15 does not re-run, so those rows stay NULL even though their Items exist. Same on MySQL and MariaDB.
    • "Upgrade all replicas together" covers mixed replicas. One clause saying that a rollback has the same effect would cover this case too.
  2. Old Web Shell tabs loop on the resync frame (figure 1, main's client). This is documented, and harmless while nothing raises the floor. The retention slice should raise floors only once clients that handle the frame are deployed, or accept that old tabs need a reload.
  3. The packaged hosted harness can't produce multi-delta Parts today. It sends one aggregated text delta per Turn and no reasoning deltas. So multi-delta, reasoning, tool-call and retraction identities were checked another way: the V15 differential (main's materializer as oracle), the retraction probe, and your unit tests.
  4. Observation, not this PR: in one run the stuck client was blocked for more than 60 s. The server closed the stream after 978 contiguous events (Tomcat's default write timeout). There was no gap and no duplicate, and a client resumes with Last-Event-ID.

Figures

1. Web Shell: resync → reload → resume, PR client vs main's client
Web Shell resync A/B

2. Event pages, replay floor and streams on the real stack
Streams on the real stack

3. V13 → V15 upgrade differential and V15 memory
Upgrade and V15 memory

4. Retraction probe, independent mutants, merge with main
Retraction, mutants, merge

The harness, the retraction probe, the raw results and a summary of the numbers are in pr12840/; see results/summary.txt.

中文版

真实链路验证:事件回放(D3),最终 head cf72a892b4

结论:可以合并。 在 cf72a892b4 上没有发现阻断问题。

验证期间 head 变了四次:42bdee4104 → d760dd572e → 349fb3ba28 → 合入 main 的 6bba5132bc → cf72a892b4。每次推送改到的部分我都重跑了。途中发现一个严重问题:长 Session 上 V15 内存耗尽后挂住,而不是直接失败。这个问题在 42bdee4104 和 d760dd572e 上存在,349fb3ba28 已修复。349fb3ba28 与 cf72a892b4 的主 Java 代码完全相同,两个 jar 只有 OpenAPI 文件不同(375 个条目中 374 个 CRC 一致)。

cf72a892b4 上 CI 的 Runtime Broker and Managed Agent MariaDB / Java 21(唯一运行本模块测试的任务)已通过,发帖时其余任务仍在运行。

triage bot 留了一个未验证的问题:客户端落到下限以下后能否恢复、而不是反复重连。我在真实浏览器里验证了:PR 的 Web Shell 收到 resync 帧后约 25 ms 恢复,main 的客户端则反复重连(图 1)。

装置

  • 服务端: 每个 head 打包出的 Spring jar,分别跑在 MySQL 8.4.7 和 MariaDB 10.11.18(CI 用的镜像)上。
  • Harness 与模型: 打包的 qwen serve --profile hosted-harness,由脚本化的 OpenAI 兼容模型驱动。
  • 客户端: 通过 Playwright 在 Chromium 中打开真实 Web Shell(vite),另有原始 HTTP/SSE 客户端。
  • 校验: 所有响应体和 SSE 帧都用 Ajv 按该 head 的规范校验。
  • 下限: 用 SQL 抬高,因为生产代码中还没有调用 advanceReplayFloor 的地方。
  • 基线: main 的 jar(e68822c815)作为"改动前"一臂,也作为升级校验的对照基准。

结果(未注明时均为最终 head)

方面 结果
与 CI 等价的 mvn -Pmysql-integration clean verify checkstyle:check(MariaDB 10.11.18) 135 个单测 + 10 个 IT 通过,0 个 Checkstyle 违规;此前每个 head 也都是绿的。
Web Shell 重新生成类型无 diff;73/73 个 managed 测试通过;typecheck 干净。
分页与 limit 以 limit=7 沿 next_cursor 翻页,13 个序号各出现一次且有序;最后一页 has_more:false、next_cursor:null。limit 1000 → 200;1001 和 0 → 400 invalid_limit(事件查询和 WebShell transcript 都如此)。main 上第一页就报 has_more:false,13 个事件只能拿到 7 个;limit=1000 → 400。
真实 Turn 的身份 每个事件都带 schema_version 1 和 projection_version 1;item_id/content_part_id 都指向 Snapshot 中的 Item/Part;按 Part 分组的 delta 能拼回该 Part 的文本。
回放下限 下限 8、Snapshot 13 时:after=7 → 409 cursor_expired,带 replay_floor_sequence、snapshot_through_sequence 和 request_id;after=8 → 200;Last-Event-ID 在两个方向上都优先于 after;两条流都只发一个不带 id 的 agent.session.resync_required 帧,然后关闭。全部按 1.17.0 规范 Ajv 校验通过。
卡住的客户端、hub 溢出 一个接收缓冲只有 4 KiB 的客户端停止读取,期间写入 4,000 个事件。hub 丢弃了这段范围,流从序号 983 起按每页 100 条重新读库(见 MySQL general log)。4,000/4,000 个事件各到达一次且有序。
下限越过卡住的流 先收到 975 个连续事件,然后恰好一个 resync 帧,随后服务端关闭流。
追赶与写入竞争 从 Last-Event-ID: 0 追赶 8,601 个事件(81 页),同时有 300 次重命名在写:每个序号恰好一次且有序。
Web Shell 端到端(图 1) PR 客户端: resync 帧 → 重载一次 transcript → 从 9 续传;放行后 27 ms 显示新回复。main 客户端: 15 秒内 5 次流请求(被挂起的那次,加上每 3 秒一次重连),游标都是 5,从不重载,新回复一直不显示。
V13 → V15 升级 以 main 的 Snapshot 为基准:MySQL 1,634 个事件 / 435 个文本 Part,MariaDB 1,606 / 448,0 处不一致;V15 以 jar 内的 JDBC 迁移执行。
V15 内存(F1) 一个 120 万事件的 Session、256 MB 堆:42bdee4104/d760dd572e 上 OOM 后挂住;349fb3ba28 上 86 秒完成迁移。
回撤后的身份 随机化探针共 3,400 个用例:d760dd572e、349fb3ba28 上 0 失败;42bdee4104 上 400 个用例有 356 个失败。
独立变异 349fb3ba28 上 17 个可用变异体全部被杀(代码与 cf72a892b4 相同)。
与 main 的合并 你把规范解成 1.17.0,与我独立做的试合并完全一致(0 个 JSON 节点不同)。

F1:审查期间已修复 —— 长 Session 上 V15 内存耗尽后挂住(42bdee4104、d760dd572e)

现象。 V15 对每个 Session 执行一条不分页的 SELECT,Connector/J 会把整个结果集缓存在内存里。我用 256 MB 堆迁移一个含 120 万条旧事件的 Session:

  • 加 -XX:+ExitOnOutOfMemoryError 时: 紧接 Migrating ... to version "15" 之后报 OutOfMemoryError: Java heap space。
  • 不加这个参数(默认)时: JVM 不退出,而是卡在半读连接的 rollback() 里:
    • 主线程停在 Flyway 的 rollback(),位于 MySQLNamedLockTemplate 之内,所以一直占着 Flyway 的 MySQL 命名锁;
    • 老年代保持 99.8 %,在 42bdee4104 上就这样持续了 15 分钟;
    • 不响应 SIGTERM,必须 SIGKILL;
    • V15 没有记入 history,所以每次重启都会重复这个过程。

修复。 349fb3ba28 改为分页读取(每页 5,000 行,跨页延续 run 和前一事件)。

  • 修复后同样测试: 同一个 Session 在 256 MB 下 86 秒完成迁移,身份计数一致。
  • 输出字节: 在同一份 10,637 事件的 dump(最长 Session 9,509 个事件)上,三个 head 的 V15 写出的身份列 MD5 相同。
  • 测试: 新增的跨页测试能杀死我试的两个分页变异体:在页边界丢掉前一事件、只读第一页。

无需处理。

已验证:回撤后的身份(d9dc6b730c)与契约文字(349fb3ba28、cf72a892b4)

探针是 store 层的随机化测试:

  • 每个用例由三代 Harness 产生 2–19 个 delta;
  • 回撤其中一代;
  • 重建 Snapshot;
  • 检查:每个已存储的身份都必须指向重建后的 Snapshot,每个 Part 的 delta 都要能拼回它的文本。

结果:

  • 修复后(d760dd572e、349fb3ba28):4 个种子共 3,400 个用例,0 失败。
  • 修复前(42bdee4104):400 个用例有 356 个失败,例如 #5 empty delta names part_..._reasoning_5。

这也回答了 bot 关于 content_part_id 的问题:回撤后已存储的身份能在重建后的 Snapshot 中解析,事件 schema 也写明了 stream.reconciled 这个例外。

非阻断说明

  1. 回滚会留下没有身份的行。
    • main 的 jar 能在 V15 schema 上启动。Flyway 打出 "has a version (15) that is newer than the latest available migration (13)",然后继续运行。
    • 之后 main 写入的事件,item_id/content_part_id 都是 NULL。
    • 再次升级时 V15 不会重跑,所以这些行一直是 NULL,尽管对应的 Item 存在。MySQL 与 MariaDB 表现相同。
    • 设计文档里的"所有副本一起升级"覆盖了混合副本的情况;再加一句"回滚也是同样效果",就能覆盖这种情况。
  2. 旧 Web Shell 标签页收到 resync 帧会反复重连(图 1,main 客户端)。文档已写明;在没有代码抬高下限之前没有影响。retention 那一步应在能处理该帧的客户端上线后再抬高下限,或者接受旧标签页需要刷新。
  3. 打包的 hosted harness 目前产生不了多 delta 的 Part:每个 Turn 只发一个聚合后的文本 delta,也不发推理 delta。因此多 delta、推理、工具调用和回撤场景的身份改由三条途径验证:V15 差分(以 main 的物化结果为基准)、回撤探针,以及作者自己的单测。
  4. 观察(与本 PR 无关): 有一次卡住的客户端被阻塞超过 60 秒,服务端在发出 978 个连续事件后关闭了流(Tomcat 默认写超时)。没有缺号也没有重复,客户端可用 Last-Event-ID 续传。

图见上文英文部分;装置、探针、原始结果和数字汇总见 pr12840/ 下的 results/summary.txt。

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

The PublicEvent description said to resume "after it", which reads as
after the stream.reconciled event. A client that did so after reloading a
Snapshot already past that event would apply its deltas twice. Name the
snapshot_through_sequence that the Items list returned as the resume point.
@wenshao
wenshao requested a lite review from Copilot September 27, 2026 13:31

@qqqys qqqys left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent Critical-only review — head cf72a892

Stage D3 event replay: a shared identity rule, two migrations (one SQL, one Java backfill), a classified cursor-expiry error with a stream resync frame, the replay window in the store and service, the contract surface, and the WebShell client's use of it. 35 files, +2343/-156.

A replay feature can fail in two catastrophic directions — a migration that damages stored events, and a cursor that silently skips events a client never received — so I spent the budget on those two and on the identity rule both migrations and the live path share.

No historical blocker

Copilot did not run at any of the six heads (reviewer quota), no other review has been submitted, and there are no inline comments, so there are no review threads and no Critical has ever been filed. Triage stage 2 read the whole diff, went back into the base tree to check the identity rule and the positional record change specifically, and reports "No critical blockers, and no AGENTS.md violations"; stage 3's verdict is approve. The maintainer's real-stack verification at this exact head reports "ready to merge… no blocker at cf72a892b4", including the disclosure that an intermediate head's V15 ran out of memory on a long Session and then hung instead of failing — which is what the paging below addresses.

Neither migration can damage existing data

V14 is purely additive: schema_version and projection_version as INT NOT NULL DEFAULT 1, item_id and content_part_id as nullable VARCHAR(128), and replay_floor_sequence BIGINT NOT NULL DEFAULT 0 on the Session. Every existing row is populated by its default, nothing is dropped, narrowed or rewritten, and the floor starts at 0 with the comment stating why — "Nothing prunes events yet, so every Session starts at 0" — which means no cursor can be below the floor today.

V15 is a Java migration that backfills only the two new nullable columns. It reads each Session's events in sequence order, paged at 5,000 rows per query so memory does not grow with a Session's length, and writes in batches of 500. The paging is correct where it matters: the resume parameter is the last sequence actually read (sequence_id > ? with previousSequence), the previous identity and previousSequence are declared outside the page loop so a run of deltas spanning a page boundary still sees its predecessor, the contiguity test previousSequence == sequence - 1 ? previous : null honours the rule's documented contract exactly, and the trailing run is flushed after the loop as well as at each batch boundary. It never touches data_json, never deletes, and re-running the same rule over the same rows yields the same values.

The cursor path fails closed, and the resync frame cannot lose events

ReplayCursorExpired extends ApiException with 409 cursor_expired and carries replay_floor_sequence and snapshot_through_sequence in its detail, so a JSON read is told precisely where the floor is rather than receiving an empty page. An empty page is the failure mode that matters here — it is how a replay implementation silently skips events — and this design cannot produce one for an expired cursor.

On the stream side both loops catch it, send one agent.session.resync_required frame with action: reload_snapshot, and complete the emitter. The detail that makes this safe is in the comment: the frame carries no SSE id. An id would move the client's Last-Event-ID past events it never received, so a reconnect would resume beyond the gap; without one, a reconnect presents the same cursor and gets the same resync frame until the client reloads the Snapshot. The page-size comparison also stopped being a magic number — events.size() == ManagedAgentService.STREAM_PAGE — so the loop's continuation test cannot drift from the query's limit, which is the other way this loop could have gone wrong.

One home for the identity rule, and it is deterministic

EventIdentity is the single implementation used by both the accept path and the V15 backfill, with its versioning contract stated up front: "A different rule needs a new projection version, not a change to this one." The parts that could silently diverge are handled: turn.accepted falls back to StoreModels.inputItemId(turnId), the same shared function the query surface advertises, so a turn's input Item is spelled one way everywhere; toolItemId prefers an explicit itemId, then toolCallId, then callId, and only then a deterministic turnId:sequence:N, hashed through UUID.nameUUIDFromBytes so the same event always yields the same Item; a text delta continues the previous Part only when the type, the Item and a non-null previous Part all match, and an empty or missing text produces no identity at all, matching a projection that skips empty text; and string() accepts only real String values, so a malformed field reads as absent rather than becoming the literal "null".

CI

20 checks pass at this head and none has failed, including the lanes that exercise this code against real databases — Runtime Broker and Managed Agent MariaDB / Java 21, Hosted no-tool processes / MySQL 8.4 / Java 21, Real daemon E2E / Java 11 and the full Java matrix (ubuntu 11/17/21, windows, macos) — plus Test (ubuntu-latest, Node 22.x), Lint & Static, Integration Tests (no-AK, No Sandbox), Capture web-shell visuals, both Desktop Shell lanes, triage and the routing checks. web-shell E2E Smoke and review-pr were still pending; neither is a gate.

Scope

I read both migrations in full, the whole cursor-expiry and resync path, the whole identity rule, and the stream loop's page handling. I did not read ManagedAgentStore's 138 added lines, ManagedAgentService's rewrite, the 94-line OpenAPI delta, the WebShell client changes, or the ~1,000 lines of new tests beyond their names — for the identity rule's fidelity to the Items projection I relied on triage's base-tree comparison and on EventIdentityTest, ManagedEventIdentityMigrationTest and the LegacyEvents helper running green in the Java lanes, and I am naming that reliance rather than implying coverage I did not have. The known-gaps file losing 18 lines is consistent with gaps closed by implementation rather than by exemption.

Verdict: APPROVE — No Critical found and none was ever filed. Both migrations are additive or backfill-only, with the Java one paged and batched so a long Session cannot exhaust memory and resuming strictly after the last sequence read; an expired cursor answers a classified 409 or a resync frame with no SSE id, so no path silently skips events; and the identity rule lives in one deterministic place shared by the live path and the backfill, with its versioning contract stated rather than implied.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@wenshao

wenshao commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

@qwen-code /triage

@wenshao
wenshao added this pull request to the merge queue Sep 27, 2026
Merged via the queue into main with commit 36710ff Sep 27, 2026
77 of 78 checks passed

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Partially reviewed — gaps disclosed. Suggestions are inline.

Not reviewed: reverse audit of chunk 2 — stopped at the 5-round reverse-audit cap while the chunk was still reporting new findings (F17, round 5, verified).

Not reviewed: reverse audit of chunk 6 — stopped at the 5-round reverse-audit cap while the chunk was still reporting new findings (F18, round 5, verified).

Not reviewed: reverse audit of chunk 9 — stopped at the 5-round reverse-audit cap while the chunk was still reporting new findings (F16 round 4, F20 round 5, both verified).

Not explored to full depth (tool budget reached): "agent reverse-audit (round 2)": 无 —— chunk 6 全部 hunk 均已走读; afterSequence: -1 的运行时形态未跑动态探针系证据精度取舍而非预算所限,不确定性已在发现内披露。; "agent reverse-audit (round 5)": 无 — brief、findings 清单(137 行全文)、chunk 6(diff L1730–2124)及上述全部生产/测试源文件均完整读取,无因预算中止的检查。; "agent reverse-audit (round 1)": 无——计划内检查全部完成。; chunk 1: 无** —— 所有想查的检查均已完成(8/33 次调用)。; "agent 1b": 无(所有计划的核查均已完成)。, and 5 more.

中文说明

仅完成部分审查,审查缺口已披露。 建议见行内评论。

未审查(原文为英文):reverse audit of chunk 2 — stopped at the 5-round reverse-audit cap while the chunk was still reporting new findings (F17, round 5, verified).

未审查(原文为英文):reverse audit of chunk 6 — stopped at the 5-round reverse-audit cap while the chunk was still reporting new findings (F18, round 5, verified).

未审查(原文为英文):reverse audit of chunk 9 — stopped at the 5-round reverse-audit cap while the chunk was still reporting new findings (F16 round 4, F20 round 5, both verified).

未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 2)":无 —— chunk 6 全部 hunk 均已走读; afterSequence: -1 的运行时形态未跑动态探针系证据精度取舍而非预算所限,不确定性已在发现内披露。;"agent reverse-audit (round 5)":无 — brief、findings 清单(137 行全文)、chunk 6(diff L1730–2124)及上述全部生产/测试源文件均完整读取,无因预算中止的检查。;"agent reverse-audit (round 1)":无——计划内检查全部完成。;chunk 1:无** —— 所有想查的检查均已完成(8/33 次调用)。;"agent 1b":无(所有计划的核查均已完成)。,另有 5 条。

— glm-5.3-flash via Qwen Code /review (v0.24.6)

Comment on lines +39 to +42
ResultSet rows = statement.executeQuery("SELECT DISTINCT"
+ " tenant_id, session_id FROM managed_agent_event")) {
while (rows.next()) {
sessions.add(new String[] {rows.getString(1),

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-1: The session enumeration is the one unbounded read in this otherwise fully-paged migration — SELECT DISTINCT tenant_id, session_id is materialized whole into a List<String[]>, so memory grows with the number of sessions rather than with any session's length, and a real probe ran the migration out of heap at 700k distinct sessions.

Everything else in V15 pages deliberately (5000-row event reads, 500-row write batches, with the comment "memory does not grow with a Session's length"), but the session list that drives those pages is unbounded. On a deployment with millions of distinct sessions the DISTINCT result set is fully materialized on-heap inside V15's single transaction (MySQL Connector/J additionally materializes a driver-side copy), the migration dies with OutOfMemoryError, and the service cannot start against that database until the heap is enlarged. Probe witness on a file-based H2 with 700,000 distinct sessions (1.4M rows):

Witness:

unmodified V15, -Xmx128m → PROBE migration COMPLETED in 197920 ms
unmodified V15, -Xmx96m  → PROBE migration FAILED after 247485 ms with
                           java.lang.OutOfMemoryError: Java heap space
                           (V14 had completed; the heap died in V15's phase)
paging comparison arm (enumeration via LIMIT + (tenant_id, session_id) > ?
cursor), same -Xmx96m    → PROBE migration COMPLETED in 303389 ms

Page the enumeration the same way the event reads are paged, carrying the last seen (tenant_id, session_id) as the cursor:

SELECT DISTINCT tenant_id, session_id
FROM managed_agent_event
WHERE (tenant_id, session_id) > (?, ?)
ORDER BY tenant_id, session_id
LIMIT ?

The fix's premises: the backfill cursor must keep sequence_id > ? (the last sequence actually read) and the cross-page continuation rule previousSequence == sequence - 1 ? previous : null, and ranged updates must not merge across sequence gaps — all pinned by the 6,000-event single-Part case in ManagedEventIdentityMigrationTest. The paging comparison arm above is the acceptance proof: it completes at the same heap where the unmodified form OOMs.

中文说明

[Suggestion] R1-1:会话枚举是这个全程分页的迁移里唯一无界的读取——SELECT DISTINCT tenant_id, session_id 被整体物化进 List<String[]>,内存随会话数量(而非单个会话长度)增长;真实探针在 70 万个 distinct 会话时把迁移跑出了堆溢出。V15 的其余部分都刻意分页(事件读取每页 5000 行、写入每批 500 行,注释写明"内存不随 Session 长度增长"),但驱动这些分页的会话清单本身无界。在数百万 distinct 会话的部署上升级时,DISTINCT 结果集在 V15 的单一迁移事务内全部进堆(MySQL Connector/J 默认还会在驱动侧再物化一份),迁移以 OutOfMemoryError 失败,服务在该库上无法启动,只能加大堆重试。建议用与事件读取同样的方式分页枚举会话(以最后读到的 (tenant_id, session_id) 为游标);探针对照臂(游标分页版)在同一堆上限下成功完成,即为验收证明。

— glm-5.3-flash via Qwen Code /review (v0.24.6)

Comment on lines +261 to +262
MockHttpServletResponse webShellStream = mvc.perform(
post("/api/agent/web-shell/v1/events/stream")

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-2: The web-shell stream's overflow fallback and non-zero-cursor resume are tested nowhere — the replay suite pins them only for the public stream, while streamWebShell is a ~70-line near-duplicate delivery loop that every server-side test drives only at afterSequence: 0.

The suite's javadoc promises "every stream must deliver each sequence exactly once and in order", and the production client really resumes with afterSequence: request.lastEventId (java-managed-agent-provider.ts) — but grep over the whole tree shows every web-shell stream drive site (the contract test's two, the replay test's one, and the four unit-test stubs) passes afterSequence: 0, and the overflow and Last-Event-ID resume tests open only publicStream. A regression confined to the web-shell loop — typically fixing streamPublic's overflow handling without mirroring it to its duplicate — passes the entire suite, and web-shell clients then observe gaps or duplicates after a hub overflow.

Witness:

not run — 缺测试本身即证据,由全树 grep 建立(web-shell 流所有服务端驱动点
afterSequence: 0;无 web-shell overflow 场景)

Parameterize overflowFallsBackToTheStoreWithoutGapsOrDuplicates (and the resume test) over both transports, opening the web-shell stream via POST with a non-zero afterSequence. Two constraints: the web-shell route takes afterSequence in the JSON body (WebShellStreamRequest, not a query parameter or header), and the mutant that proves the pin is dropping streamWebShell's delivery.overflowed() → reconcile = true branch — the web-shell variant must go red while the public stays green.

中文说明

[Suggestion] R1-2:web-shell 事件流的溢出回退与非零游标续传完全没有测试——回放套件只钉住了公共流,而 streamWebShell 是一个约 70 行的近似复制的投递循环,服务端所有测试都只以 afterSequence: 0 驱动它。套件 javadoc 承诺"每个事件流必须按序恰好投递每个 sequence 一次",生产客户端也确实以 afterSequence: request.lastEventId 续传,但全树 grep 显示每一处 web-shell 流驱动点(契约测试两处、回放测试一处、单元测试四处桩)都传 0,溢出与 Last-Event-ID 续传测试只打开 publicStream。仅限于 web-shell 循环的回归(典型:修了 streamPublic 的溢出处理却没同步到副本)会全绿通过,此后 web-shell 客户端在 hub 溢出后会看到缺口或重复。建议把溢出与续传测试在两种传输上参数化;变异证明:删掉 streamWebShell 的 delivery.overflowed() → reconcile = true 分支,web-shell 变体必须变红。

— glm-5.3-flash via Qwen Code /review (v0.24.6)

};
get?: never;
put?: never;
/** @description Without a cursor, a Session that has a Snapshot returns all of its Items, the events up to the Snapshot other than input, text-delta and tool-call updates, and every event after it. Otherwise, and for an olderCursor, limit bounds the page of events. */

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-3: The no-cursor transcript description lists three excluded event kinds, but the implementation excludes four — item.reasoning.delta is missing from the prose (the sentence originates at openapi.json:762 and is generated into this file).

ManagedAgentStore.findControlEvents excludes exactly turn.accepted, item.output_text.delta, item.reasoning.delta and item.tool_call.updated, but the description says "other than input, text-delta and tool-call updates". The same document's own enum (L2936 ["input_text", "output_text", "reasoning"]) treats reasoning as a distinct category, so "text-delta" cannot be read as covering reasoning deltas. And the asymmetry bites: tail events after the Snapshot are unfiltered, so the same client sees reasoning deltas in the live stream but not on the initial page, with no sentence anywhere explaining it. Worst case: an integrator wrongly infers the session has no reasoning content.

Witness:

not run — 纯 prose/代码失配,两侧原文并排引用即证据:openapi.json:762 与
generated TS L45 列三类 vs ManagedAgentStore.java:719-723 的 NOT IN 四类清单
Suggested change
/** @description Without a cursor, a Session that has a Snapshot returns all of its Items, the events up to the Snapshot other than input, text-delta and tool-call updates, and every event after it. Otherwise, and for an olderCursor, limit bounds the page of events. */
/** @description Without a cursor, a Session that has a Snapshot returns all of its Items, the events up to the Snapshot other than input, text-delta, reasoning-delta and tool-call updates, and every event after it. Otherwise, and for an olderCursor, limit bounds the page of events. */

Fix it at the source (openapi.json:762 — add the fourth kind) and regenerate this file, keeping the sentence consistent with findControlEvents' exact four-kind list; the "every event after it" half is correct as written (tail events are unfiltered) and should not change.

中文说明

[Suggestion] R1-3:无游标 transcript 的描述只列了三种被排除的事件类型,而实现排除了四种——item.reasoning.delta 在散文中缺席(句子源头在 openapi.json:762,生成进本文件)。ManagedAgentStore.findControlEvents 精确排除四种类型,且同一文档自身的枚举把 reasoning 立为独立类别,"text-delta" 读不进 reasoning 增量。不对称会咬人:Snapshot 之后的尾部事件不过滤,同一客户端在实时流里看得到 reasoning 增量、在首屏却拿不到,规格中无任何解释。最坏情况:集成方误判该会话没有推理内容。请在源头(openapi.json:762)补上第四类并重新生成;"every event after it" 半句是对的,勿动。

— glm-5.3-flash via Qwen Code /review (v0.24.6)

Comment on lines +84 to +86
- The WebShell transcript states what it already returned: without a cursor, a
Session with a Snapshot gets all of its Items, the events up to the Snapshot
other than input, text-delta and tool-call updates, and every later event;

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-4: The same three-vs-four exclusion defect is replicated verbatim in this design doc and its Chinese sister (§4.1) — fixing the contract sentence (the openapi description this doc mirrors) will not touch these two files, so after the fix the tracked bilingual docs remain wrong.

The implementation (ManagedAgentStore.findControlEvents, L719-724) excludes four kinds; this sentence (and the zh-CN doc's "除输入、文本增量与工具调用更新以外") lists three, omitting reasoning deltas. Each doc's own §4.2 table lists item.output_text.delta and item.reasoning.delta as distinct event types, so "text-delta" cannot be read as covering reasoning in-document either.

Witness:

not run — 纯描述性 prose 失配,两侧原文并排引用即观察(EN L84-86 / ZH L73)
Suggested change
- The WebShell transcript states what it already returned: without a cursor, a
Session with a Snapshot gets all of its Items, the events up to the Snapshot
other than input, text-delta and tool-call updates, and every later event;
- The WebShell transcript states what it already returned: without a cursor, a
Session with a Snapshot gets all of its Items, the events up to the Snapshot
other than input, text-delta, reasoning-delta and tool-call updates, and every later event;

Amend all texts together — this doc, the zh-CN sister's corresponding sentence ("推理增量"), and the openapi description — and keep the "and every later event" half unchanged (tail events really are unfiltered).

中文说明

[Suggestion] R1-4:同样的"三写四"排除缺陷在本设计文档与其中文姊妹版(§4.1)逐字复现——修复契约句子(本文档所概述的 openapi 描述)不会触及这两个文件,修复后被跟踪的双语文档仍然是错的。实现排除四种类型,本句(及中文版"除输入、文本增量与工具调用更新以外")只列三种,漏掉推理增量;两份文档各自的 §4.2 表都把 item.output_text.delta 与 item.reasoning.delta 列为不同事件类型。请四处一起改(本文档、中文版对应句加"推理增量"、openapi 描述),并保持 "and every later event" 半句不变。

— glm-5.3-flash via Qwen Code /review (v0.24.6)

Comment on lines +65 to +66
// More events than SessionEventHub buffers for one Session.
private static final int OVERFLOW = 600;

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-5: The overflow test's activation premise — that SessionEventHub.CAPACITY is smaller than the 600 events this test appends — lives only in this comment, pinned by no assertion; raising the capacity past 601 silently disarms the test while the suite stays green.

CAPACITY is private static final int CAPACITY = 512 (SessionEventHub.java:16) and appears nowhere in the test tree. If a future tuning change raises it to 1024, the 600 appended events all enqueue contiguously, await hands over live, reconcile never fires — and overflowFallsBackToTheStoreWithoutGapsOrDuplicates still passes with ids == range(1, last), while the store-fallback path it exists to guard has lost its only test. (The floor test would go red in the same scenario, but shaped like a replay/floor regression — misleading rather than diagnostic.)

Witness:

not run — 前提缺失由 grep 建立(测试树无 CAPACITY 引用;常量为 private)

Widen CAPACITY to package-visible (the test is in the same package) and pin the premise at the test's start:

assertThat(SessionEventHub.CAPACITY).isLessThan(OVERFLOW);

The mutation that proves the pin: temporarily set CAPACITY to 1024 — this new assertion must turn red while the rest of the suite stays green. Do not change the capacity's value or the drop semantics themselves; the fix only adds the dependency pin.

中文说明

[Suggestion] R1-5:溢出测试的激活前提——SessionEventHub.CAPACITY 小于本测试追加的 600 条事件——只存在于注释里,没有任何断言钉住;把容量提到 601 以上会让该测试静默失效而套件保持绿色。CAPACITY 是 private 常量且测试树中无任何引用。若未来调优把它提到 1024,600 条事件全部连续入队、await 直接活交付、reconcile 永不触发,溢出测试照样通过,而它守护的 store 回退路径失去唯一测试。建议把 CAPACITY 放宽为包可见并在测试开头加 assertThat(SessionEventHub.CAPACITY).isLessThan(OVERFLOW);;变异证明:临时把 CAPACITY 改成 1024,新断言必须变红而其余用例保持绿。

— glm-5.3-flash via Qwen Code /review (v0.24.6)

Comment on lines +182 to +184
assertThat(emitter.ids())
.containsExactlyElementsOf(range(1, last));
assertThat(emitter.failed).isEmpty();

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-6: This test never asserts that emitter.resync stayed empty, although RecordingEmitter records resync frames in every test — the class javadoc's core mechanism distinction (hub overflow → silently fall back to the store; floor advance → exactly one resync) is pinned only on the floor side.

"Overflow produces ZERO resync frames" is asserted nowhere: grep over the sdk-java test tree shows the only .resync content assertion is the floor test's containsExactly. If a regression makes the hub-overflow branch also emit agent.session.resync_required — say someone "hardens" the overflow path with a resync — the whole suite stays green: the resync frame has no id: line so ids() is unaffected, and the exactly-once delivery assertion above passes while delivery stays exact. Afterwards every real client stall (overflow's normal trigger) forces a full snapshot reload instead of the cheap store catch-up this route promises.

Witness:

not run — 覆盖缺失由全文件通读 + grep 建立(.resync 断言仅 L213 一处)
Suggested change
assertThat(emitter.ids())
.containsExactlyElementsOf(range(1, last));
assertThat(emitter.failed).isEmpty();
assertThat(emitter.ids())
.containsExactlyElementsOf(range(1, last));
assertThat(emitter.failed).isEmpty();
assertThat(emitter.resync).isEmpty();

The mutant that proves the pin: make streamPublic's overflow branch emit one resync frame without stopping delivery — this new assertion must go red while the other tests stay green. (One precision note from verification: a regression that reuses the existing resync() helper — which also completes the stream — WOULD be caught by the delivery assertions; the uncovered shape is precisely the resync-that-keeps-delivering one.)

中文说明

[Suggestion] R1-6:本测试从未断言 emitter.resync 为空,尽管 RecordingEmitter 在每个测试里都记录 resync 帧——类 javadoc 声明的核心机制区分(hub 溢出→静默回退 store;下限推进→恰好一条 resync)只在下限一侧被钉住。"溢出产生零条 resync"无处断言:若回归让 hub 溢出分支也发出 resync 帧(例如有人给溢出路径"加固"一条 resync),整个套件保持绿色——resync 帧无 id: 行、ids() 不受影响、投递断言照常通过;此后每个真实的客户端停滞都会强制整快照重载。建议在 failed 断言旁加 assertThat(emitter.resync).isEmpty();;变异证明:让溢出分支发一帧但不停止投递,新断言必须变红。(验证精度说明:若回归复用现成的 resync() 辅助方法——它会顺带关流——投递断言能抓住;未覆盖的正是"发帧但不停投递"的形态。)

— glm-5.3-flash via Qwen Code /review (v0.24.6)

Comment on lines +36 to +37
projection versions and the Item and Part identity they were accepted with,
except that a `stream.reconciled` event announces retracted deltas. A

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Suggestion] R1-7: This README sentence promises events keep the Item and Part identity they were accepted with and compresses the retraction into a mere "announcement" — but this PR's own retraction path rewrites identities and empties text, contradicting the machine-checked contract this doc summarizes.

retractContinuationOutput empties the retracted deltas' text and reassignIdentity (ManagedAgentStore.java:1833-1841) UPDATEs item_id/content_part_id for every event from the first retracted one on — retracted deltas lose their identity and later deltas are re-pointed at rebuilt Parts. The openapi this change ships spells this out ("which now have empty text and no item_id or content_part_id, and that later deltas may name other Parts"), as do both design docs; only the README collapses it to "announces". An integrator who treats the README as authoritative and caches each event's identity as immutable diverges silently after a Harness-recovery retraction.

Witness:

not run — 纯 prose 失配:README L36-37 vs openapi.json:2823/3807 与
reassignIdentity(L1833-1846 只 UPDATE 两列身份)逐行对照
Suggested change
projection versions and the Item and Part identity they were accepted with,
except that a `stream.reconciled` event announces retracted deltas. A
projection versions and the Item and Part identity they were accepted with,
except that Harness recovery's retraction empties the retracted deltas' text
and identity and may rename the Parts that later deltas extend, announced by
`stream.reconciled`; reload the Snapshot and resume after its
`snapshot_through_sequence` when you see it. A

Keep the "schema and projection versions are saved as accepted" half exactly as written — versions really are immutable (reassignIdentity UPDATEs only the two identity columns); the rewrite must not contradict openapi.json:2823 or the design docs.

中文说明

[Suggestion] R1-7:README 这句承诺事件保持被接受时的 Item/Part 身份,并把撤回压缩成单纯的"通告(announces)"——但本 PR 自己的撤回路径会改写身份并清空文本,与它所概述的机器校验契约直接矛盾。retractContinuationOutput 清空被撤回增量的文本,reassignIdentity 对第一条被撤回事件起的全部事件 UPDATE 身份列——被撤回增量失去身份、后续增量被改指到重建后的 Part。本变更随附的 openapi 写明了这一点,双语文档也写了,唯有 README 崩塌成"通告"。以 README 为准并缓存事件身份的集成方在 Harness 恢复撤回后会静默分叉。请按契约措辞改写本句;"schema/projection 版本被接受时保存"半句保持原样(版本确实不可变)。

— glm-5.3-flash via Qwen Code /review (v0.24.6)

water-in-stone pushed a commit to water-in-stone/qwen-code that referenced this pull request Sep 29, 2026
…LM#12968)

Follow-ups to QwenLM#12840 from the review posted after it merged.

- Page V15's Session enumeration by (tenant_id, session_id) over the
  Session table, so the backfill's memory no longer grows with the number
  of Sessions; an upgrade test crosses the 1,000-Session page inside a
  tenant.
- Run the overflow and lagging-stream replay tests on both streams from a
  non-zero cursor, since each stream has its own delivery loop; assert that
  an overflow sends no resync frame and that the hub keeps fewer events
  than the test appends.
- Test that the WebShell session hook reloads the transcript after a stream
  gap and resubscribes after its lastSequence.
- Name the four event types the WebShell transcript leaves out before the
  Snapshot, in the contract, the generated types and the design notes.
- Say in the server README that a retraction empties the retracted deltas'
  text and identity and may rename later Parts.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

review/self-reported The linked issue was opened by the PR author (self-reported) scope/sdk scope/web-shell

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(managed-agent): Stage D public API contract, generated DTOs, Session query and event replay

3 participants