Repository navigation
fix(managed-agent): stop database amplification on session hot paths - #13217
Conversation
E2E / verification reportMethod: the change lives in the standalone Spring control plane Measured before → after (same instrumented harness on both runs):
Regression: full Revocation semantics pinned both ways: a zero window restores per-event/per-chunk checks (existing tests), and new tests prove a revocation within the window takes effect at the window's end. MySQL 中文方法:本变更位于独立的 Spring 控制面 实测前后对比(同一插桩框架两次运行):
回归: 撤销语义双向钉住:零窗口恢复逐事件/逐分片复检(既有测试),窗口内撤销至多延迟一个窗口并在窗口末端生效(新增测试)。 MySQL |
deb6067 to
0ff2939
Compare
|
Please do not rebase or force-push to an active PR as it invalidates existing review comments. Note for future reference, the bots always squash all changes into a single commit automatically as part of the integration. 中文请勿对活跃的 PR 执行 rebase 或 force-push,因为这会使已有的评审评论失效。另外,供日后参考:作为集成流程的一部分,机器人始终会自动将所有改动压缩(squash)为单个提交。 |
Review round R1 addressed (2 Criticals + 23 Suggestions, all accepted)All 25 findings are fixed in the follow-up commit. Summary by area: Criticals
Store/service changes
Test pins (mutant-verified locally)
Docs
Full module suite: 471 tests, 0 failures, 2 errors — the two pre-existing Post-commit audit rounds (undirected + adversarial pairs over the committed diff, with mutation probes):
|
Four hot paths in the Managed Agent Runtime Broker amplified database work far beyond request volume, two of them while holding row locks: - materializeNextBatch rewrote the whole items_json snapshot after every <=200-event batch under the session row lock. The snapshot is now rewritten at creation, on a terminal event, every SNAPSHOT_REFRESH_EVENTS (1000) covered events, or on catch-up once the previous snapshot is SNAPSHOT_REFRESH_MILLIS (5s) old; a session whose snapshot lags its drained progress is re-selected by findMaterializationTargets once the snapshot ages out, so an idle session still converges instead of staying stale forever. - SSE fan-out ran a read-grant SQL check before nearly every event per subscriber. Streams now recheck the grant at most once per events.read-grant-recheck-interval (default 5s; PT0S restores per-event checks); session.deleted still terminates immediately. - listPublicSessions/listWebShellSessions ran 2-4 extra queries per row. Pages are now assembled from a constant number of grouped batch queries that project only the turn summary columns, with the approval mode riding the page's own SELECT *, plus one workspace-close batch read when the page holds bound sessions; single-session views share the same assembly, and the singular store reads delegate to the batch twins so each selection rule has one spelling. - Publication authorization rescanned the session journal back to the last activation.changed per revision under the publication lock on every publish/seal/prefix/finish/verifyDispatch. Migration V34 carries the current activation on the journal head, maintained by commit() inside the parse pass every commit already runs; authorization reads the head row it already locks, reserve/renew read the tool.intent at its own revision, and pre-migration heads get one legacy scan plus a backfill. Because a rolling fleet can briefly run a binary that does not maintain the columns, trusting them is gated behind tool-publication.journal-head-authorization (default off) until the fleet drains. Related: artifact downloads re-verified content access per 64 KiB chunk; the re-verification is now throttled to artifacts.read-revalidation-interval (default 5s), which also defers the stage-1 DELETING observation to the window's end (stage-2 retirement still aborts per chunk through the read lease). Migration V35 adds an event-type index so latest-turn lookups read index ranges instead of a session's event history. Query budgets per endpoint are pinned by Issue13181QueryBudgetTest (with the artifact revalidation window pinned in ManagedArtifactReadIntegrationTest), the windowed/frozen-clock semantics are pinned by ManagedEventStreamServiceTest, and the fencing behavior across the head columns, the legacy scan, and the rolling-deploy gate is pinned by ToolPublicationStoreTest. Fixes #13181
0ff2939 to
1b2496d
Compare
|
Rebase note: the branch was rebased onto current Two integration points came out of the rebase:
Suite on the squashed tree: 498 tests (including upstream's new |
Round-3 adversarial follow-ups (no Criticals): - expiredActivationFencesThroughTheHead now proves the fences came from the head columns: renew reads exactly the intent's revision (3 statements), and verifyDispatch performs zero locked journal reads. - renewReadsTheIntentAtItsOwnRevision additionally pins zero locked journal reads, so a lock-mode upgrade on the intent read cannot slip past the statement-count pin. - Drop the unreachable null-check after queryForObject in requireLegacyActivation (the query throws on a missing row, and record_bytes is NOT NULL).
Round-4 adversarial follow-up (no Criticals): the 20-row bound-page loops asserted titles and actions but never the sessionArchive/Unarchive/Delete flags, so a regression computing retention once per page would have passed. One bound session now carries a completed CLOSE row and both surfaces assert sessionDelete per row.
CI (Runtime Broker and Managed Agent MariaDB) caught preservesTextOrderAcrossToolsAndReasoningInSnapshots racing the 10ms materializer: a scheduler tick can create the snapshot mid-sequence, and a non-terminal catch-up drain inside the 5s floor then legally leaves it behind. The trailing event is now terminal so the explicit drain always rewrites; the content-ordering assertion is unchanged.
|
Final audit trail and CI note:
|
…eterministic Same scheduler race as the text-order test fixed in b0d357a, caught by the MariaDB CI job on the following run: a 10ms-tick snapshot created mid-sequence plus a non-terminal catch-up inside the 5s floor legally lags. The trailing reasoning delta is now terminal so the explicit drain always rewrites. The contentPartId assertions are unchanged.
Two correctness findings and the accepted suggestions from the second review round of the query-amplification fix: - ManagedExtensionRecords.millisLenient rejects a JSON number or numeric string wider than a long before materializing the BigInteger, so an exponent-form expiresAt no longer costs a giant allocation. - The journal head's activation columns carry activation_head_revision, the journal revision they reflect; every commit re-stamps it, both authorization gates require stamp equality, and the backfill re-stamps, so a rolling fleet still running pre-V34 writers detects the skew and rescans instead of certifying stale columns. - requireEvidence answers a fenced activation from the locked head before touching the journal: a fenced renew now pays zero journal statements, pinned in the budget test. - Deferred snapshot rewrites are recorded with a snapshot_stale_since marker (V36) that findMaterializationTargets re-selects once aged, so a drained trickle converges instead of idling stale forever. - transcript tails past the snapshot are bounded by the caller's limit, keeping the newest events and reporting the truncation for paging. - The latest-environment-event batch reads one row per session via a MAX(sequence_id) derived table over the (session, turn) pairs. - Workspace close state has a single spelling (the singular predicate delegates to the new batch query), and the approval mode rides the page's SELECT * onto SessionRecord instead of a per-row probe. - The publication journal fixture is now shared between ToolPublicationStoreTest and Issue13181QueryBudgetTest as PublicationJournalFixture, ending the drift between the two copies.
Undirected plus adversarial audits of the R2 commit found one Critical and three Minor issues: - A commit carrying no activation change re-stamped the carried-forward head columns with the new journal revision, laundering a skew left by a pre-V34 writer into a fresh-looking head. The stamp now advances only when the preserved columns were current at the previous revision, keeping the skew detectable until the rescan heals it; pinned by noChangeCommitDoesNotRestampASkewedHead, and the always-advance mutant goes red. - The millisLenient width pre-check overflowed int on extreme exponent-form strings; the comparison now widens to long, pinned by the 1e2147483647 arm. - sameEpochReleasePreventsReserveAndRenew is parameterized over the journal-head-authorization flag, pinning the release fence on the shipped legacy scan and the head fast path alike; the suite default stays false, matching production. - The application.yml mirror test also asserts the flattened keys exist, so a renamed or dropped key cannot false-pass on the Java defaults. - The design doc (both languages) records that a transcript backward walk can repeat the first page's inlined control events at the page seam and that clients deduplicate on the event sequence.
The round-2 undirected and adversarial audits both found the same Major: the millisLenient pre-check guarded only the negative-scale direction, so a 12-byte "1e-100000000" expiresAt still forced toBigIntegerExact to expand 10^100000000 before dividing (measured ~56s of CPU). The scale is now bounded too (at most 19 decimal places, which a representable long never needs), and a test arm pins the input class. Two test-coverage Minors from the same round: - transcriptTailIsBoundedByTheCallerLimit now pins the load-bearing "newest kept" property with exact sequences (the tail ends at the session's last event; olderCursor names the oldest returned one) — the oldest-kept mutant goes red. - deferredSnapshotConvergesOnTheAgedOutReselection now ages only the deferral marker for the selection arm and pins that the tick's scan references snapshot_stale_since without touching managed_agent_snapshot, so a regression to per-row snapshot probing fails twice over.
Review round R2 addressed (2 Criticals + 38 Suggestions — 39 fixed, 1 declined)All findings are in the follow-up commit. Summary by area: Criticals
Publication authorization
Snapshot gating
List pages
Transcript
Docs and wording
Test internals
Full module suite: 503 tests, 0 failures, 2 errors — the two pre-existing Post-commit audit rounds (undirected + adversarial over the committed diff):
|
…5-V38 Main landed V34 (tool output collection) plus five managed-agent merges since the review head. Resolution: - Migrations renumbered: activation head columns V35, event-type index V36, snapshot deferral marker V37, journal sequence index V38; every code comment, test name, and both design-doc languages updated. - ToolPublicationStore: the head read keeps #13192's SQL-side writer_live CASE while also selecting the activation columns and stamp; the head gate and the legacy fallback now run on the DB-clock epoch (#13192's nowEpoch) like the rest of the method. - ManagedAgentService: #13112's per-row creator-submit capability is folded into the batch page assembly instead of paid per row — the registry gains createdSessions and a findReadable batch twin (both preserving the binary-safe CAST comparisons, the sargable plain IN kept alongside), and maySubmitWorkspaceTurn splits the row-shape checks from the registry reads so the singular mutation path is unchanged. Unbound pages still cost 3 queries; bound pages 4, with the creator batch only when submit-shaped sessions exist and the grant batch only for creator-owned ones. - ToolPublicationStoreTest keeps both sides' new authorization tests; HarnessCoordinatorTest takes main's boundCancellingStore helper with the SessionRecord approval-mode argument.
Four Criticals and the accepted Suggestions from the third review: - Restore the uncapped cursor-less transcript tail (merge-base behaviour): capping it at the caller's limit falsified the published webShellTranscript contract and broke the prefix invariant the stream resume relies on (a lagging snapshot's middle band was served by neither page nor stream), and the new backward walk re-served already-projected delta rows past the seam. The contract sentence is now pinned in ManagedAgentApiContractTest, and the budget suite pins the full tail (every event past the snapshot, hasMore=false, no cursor) instead of the cap. - The commit-side activation capture applies the same sessionKey/v scope check both read paths enforce, so a foreign-scoped activation.changed is rejected at commit instead of being promoted into the durable head columns a later flag flip would trust. - The head path's contiguity count carries the legacy walk's byte-length tripwire, and the single-revision locate is pinned against holes and overlapping ranges; both intent reads reject duplicate lines at one sequence. - Pins added: the phase arm of the commit width guard, the deferral marker surviving a fresh empty tick, the per-event attributes read in the materialization budget (plus a whole-batch statement total), the idle-stream revocation closing at the next wake for both stream loops, the Range path's in-window exposure at the shipped 5s default, the configuration-to-store flag wiring, and the four-column last-activation-wins rotation. - Housekeeping: the dead args field dropped from the extracted fixture's consumer, the provably dead width guard dropped from backfillActivation (its callers always pass contract-validated values), the yml-mirror test now neutralizes the ambient QWEN_MANAGED_AGENT_* environment, the properties test keeps the flattened-key assertions, and the event fixture gained a scope/version overload. - Docs: the admission design's recheck sentence now states the windowed behaviour and the real closure bound (window plus one poll interval) in both languages; §2 qualifies path 4's relief as opt-in until the fleet runs the V35 code, with the operator flip tracked as #13295; §6 states the locate's row cost honestly (constant statement count, distance-proportional rows, strictly lighter than the legacy walk); §8 composes the bound-page budgets; §9 records the chain proof's narrowed parse scope; the README's revalidation row states the real in-window exposure and its env-var table names all three new knobs.
…aths Audit of the R3 wave found one Minor: the new commit-side check compared the sessionKey field by field, so a key carrying extra fields passed the commit but reads as foreign to both read paths' closed-key comparison. The check now builds the closed three-field key and deep-equals it, exactly the read-side semantics; the commit-rejection test gains the extra-fielded arm. Also narrows the duplicate-intent test's comment to the revisions a path actually reads (a stray out-of-range claim is invisible to the head path's locate and unreachable through the commit validation).
Review round R3 addressed (4 Criticals + 20 Suggestions — all accepted)Criticals
Suggestions — test pins (each verified against its mutant)
Suggestions — housekeeping and docs
The round's fixed-ruling confirmations (R1-1, R1-2, R2-1, R2-2, R2-3, R2-8, R2-20, R2-35, R2-36, R2-37) are acknowledged with thanks — no action needed. Full module suite on this commit: 611 tests, 0 failures — the only errors are the three pre-existing, environment-bound failures reproduced on unmodified origin/main under the same load (the two |
The Critical plus the accepted Suggestions from the fourth review: - readIntent now answers a damaged journal with the session store's corruption fault (500 managed_session_journal_corrupt) instead of a client request fault (400): the locate, the contiguity count, and the verified-page checks all throw journalCorrupt, widened to package-private for the shared spelling. The fence tests assert the code and the 5xx status; the legacy path answered 500 already. - requireLegacyActivation applies the closed-key scope check before promoting a scanned activation into the head columns, so a pre-existing journal carrying a foreign-scoped activation.changed is refused instead of laundered into trusted state. - The commit-side scope check now covers every enveloped managed_session_event_v1 line except domain.committed ones (whose requireEnvelope keeps its pinned messages) — a misscoped or unknown-version event line is refused at commit instead of reaching the journal; bare event-subtype lines with no envelope stay inert and tolerated as before. - tailEvents pages at SNAPSHOT_REFRESH_EVENTS (the gate's lag bound), so a maximally lagging snapshot's transcript tail is one event read instead of ten; the full-tail contract pins hold with the wider page. - The empty materialization tick is pinned whole (4 statements), and the Range revocation test uses the shared readingPolicy helper. The nine deferred observations from the review are recorded in the PR thread.
The adversarial audit of the R4 wave found two Suggestions: - readIntent answered a binding naming a never-committed intent sequence with the 500 corruption fault under the head path, where the legacy walk answers 400. The locate now splits the cases: an ambiguous range or a hole inside the committed span is corruption (500), while a sequence beyond the committed span is the requester's fault (400 "Committed publication evidence is missing"), matching the legacy walk. Pinned by aNeverCommittedIntentSequenceIsAClientFault under both flag settings (red without the split, verified). - The commit-side scope check's domain.committed carve-out also exempt unknown-domain lines that requireEnvelope never sees; those now get the closed-key check too (pinned by unknownDomainLinesAreRejectedAtCommit, red under the kind-based carve-out, verified). Plus the tail-read test now exercises the multi-page arm: a tail longer than one 1000-event page is served completely across two reads.
…plification' into fix/13181-managed-agent-query-amplification
Main's #13301 persists Workspace session tool profiles and claimed migration V35; this branch's four migrations move to V36-V39 (activation columns, event-type index, snapshot deferral marker, journal sequence index) and SessionRecord carries both new fields (approvalMode, toolProfile).
Review round R4 addressed (1 Critical + 5 Suggestions, all accepted; 9 deferred observations recorded)Critical
Suggestions
Deferred (recorded in the review body, not requested this round): the nine round-4 deferred observations (wall-clock sleeps in the revocation test, the Full module suite on the final head: 633 tests, 0 real failures — the only failures are the pre-existing, environment-bound set that reproduces on unmodified origin/main under load (the two Heads-up on the latest push: main's #13301 (persisted Workspace session tool profiles) landed its own V35 while this round was in flight, so this branch's four migrations are renumbered to V36–V39 and |
|
@qwen-code /triage |
Real-stack verification — #13217 at
|
| Arms | main = 6136786c0c (includes the Druid pool swap #13363 and #13365) · PR = trial merge 8136a6be2c = PR head 4d3b1351c7 + main; its diff against main is exactly the PR's 41 files. |
| Stack per arm | Spring Managed Agent Server fat jar (JDK 21.0.12, isolated Maven repo, embedded Runtime Broker) · native MySQL 8.4.7 (performance_schema on, ROW binlog, time_zone=+00:00, JVM TZ=UTC) · packaged Hosted Harness dist/cli.js serve --profile hosted-harness built from main (the PR changes no TypeScript) · deterministic fake OpenAI model |
| Traffic | Only POST/GET /v1/agents/sessions… and /api/agent/web-shell/v1/…. Workspace-bound Sessions run through the real Broker worker and mount. |
| Measurement | Statement counts and rows examined come from performance_schema.events_statements_summary_by_digest, filtered to the arm's schema. List counts subtract a 2 s idle window. Snapshot bytes are managed_agent_snapshot row-event bytes parsed from the binary log. Snapshot lag is sampled straight from MySQL. Nothing is instrumented inside the app. |
| Host | 10-core macOS; load average 30–40 from other sessions throughout, so wall times are noisy and I make no latency claims. |
1. The three API-reachable hot paths (figure 1)
| Operation | main | PR |
|---|---|---|
GET /v1/agents/sessions — 20 unbound |
41 stmts · 54 rows | 3 · 40 |
GET /v1/agents/sessions — 20 Workspace-bound |
81 · 94 | 4 · 60 |
WebShell sessions/query — 20 unbound |
61 · 254 | 3 · 90 |
WebShell sessions/query — 20 Workspace-bound |
141 · 1146 | 6 · 131 |
GET /v1/agents/sessions/{id} / WebShell sessions/get |
3 / 4 | 3 / 3 |
| Burst: 5 Turns × 3000 deltas (15,025 events) | 565 rewrites · 84.7 MiB binlog · 1817 ms SQL | 26 · 4.6 MiB · 121 ms |
| 60 s stream: 1500 deltas @ 40 ms (1,505 events, items_json ≈ 178 KB) | 258 rewrites · 86.0 MiB · 862 ms | 14 · 4.7 MiB · 58 ms |
| SSE: 4 subscribers on that stream (2 public + 2 WebShell) | 6,943 managed_workspace_access reads (1.15 per subscriber-event) |
50 (0.008) |
Notes on these numbers:
- On the real stack a Workspace-bound WebShell page costs 7 statements per row on main, more than the 4 per row behind the 41–81 range in the PR description: per-row
approval_mode, close-state, create-command and registry reads, plus the turn and environment-event lookups. PR's 6 equals the bound in design §8 for a bound page with creator-owned sessions. - All four SSE subscribers on PR received 1505/1505 events in order.
- History-Turn wall time was 23–33 s on main and 19–32 s on PR. Both arms spend that time mostly in the per-delta Session Store commit path (the 1,505-event stream made 1,608 vs 1,614 journal commits), so I see no end-to-end latency change on this host.
2. Bounded staleness, measured (figure 2)
-
Snapshot /
GET …/items. During the 60 s stream the snapshot was refreshed every 268 ms (median) on main and every 5058 ms on PR (max 5168 ms). On PR the oldest event missing from the snapshot was at most 4.9 s old, and the lag peaked at 120 events (main: 11). Both arms converge on the terminal batch. In the trickle-then-pause case PR lagged by at most 22 events and converged about 5 s after ingestion paused, through thesnapshot_stale_sincereselection; the marker was observed set and then cleared. The WebShell transcript showed every token throughout, because it tails the events past the snapshot. -
Read-grant revocation (
can_readflipped 4 s into a 15 s stream):Arm Stream closed after Events delivered after revoke main 123 ms 0 PR (default 5 s) 4666 ms 45 PR + read-grant-recheck-interval=PT0S144 ms 0 The public and WebShell streams behave identically. A new subscription after revocation returns 404 in every arm.
PT0Salso restores the per-event query volume (6,905 reads, against 6,943 on main).
3. Upgrade, rollback and the rolling-window stamp (figure 3)
One MySQL database through four phases. Each phase is a fresh Spring + Harness process pair on the same deployment directories, as on a real host.
- main (V35) creates 41 Sessions, 29 of them with Turns.
- PR boots on that database. Flyway applies V36–V39 in 0.049 s. On the same rows, all six list and detail responses are byte-identical to main's (15–18 KB bodies), while statements per call fall from 41 / 81 / 61 / 141 / 3 / 4 to 3 / 4 / 3 / 6 / 3 / 3. New Turns completed on 6/6 pre-existing Sessions and 1/1 new Session.
- main again on the V39 schema (rollback, or the old half of a rolling fleet). Flyway logs
Successfully validated 39 migrationsand the server runs normally. 6/6 old Sessions plus 1 new completed. A 205-event stream reached both SSE subscribers in full. No SpringERRORlines in any phase. - PR again. 6/6 plus 1/1 completed.
Journal-head activation stamp on the 28 list Sessions that have journals:
- Before the upgrade: 28
null. - After PR commits: 6
current. - After old-binary commits: those 6 are
lagging. The old binary advancesjournal_revisionwithout touching the columns, which is exactly the rolling-window skew the stamp is designed to expose, here produced by real traffic. - After the next PR commit carrying an activation change: 6
currentagain. - In the ab-pr run, real Hosted Turns filled the V36 columns on all 31 heads, each with a current stamp.
The reader side (refusing a lagging or null head, then rescanning and backfilling) needs tool publication, so it is covered in §4 instead.
Scaled database. I copied the main-arm data ×100 into a V35 schema: 4,300 Sessions, 1,736,300 events, 2.1 GiB.
- V36–V39 applied in 4.87 s, of which the V37 event index took 4.74 s.
- I then rebuilt V37 on 2.14M rows while a writer inserted an event every ~5 ms. The 474 inserts during the 3.3 s build had a max latency of 6.4 ms and a mean of 0.54 ms, so InnoDB's online DDL did not block ingestion.
After each restart, the first Turn on a pre-existing Session waited 32–35 s for the previous Harness's writer lease. This happened in every phase with both binaries, so it is unrelated to the PR.
4. Paths the public API can't reach
Publication authorization (path 4) and the artifact re-validation window both need qwen.managed-agent.tool-publication.enabled (default false) plus an object store. They also need hosted-workspace-shell/1 Sessions, and the Managed Agent API only creates hosted-workspace-files/1, so neither path can be reached through the public API with shipped defaults. What I ran instead:
- The PR's own
Issue13181QueryBudgetTeston real MySQL. I patched only its fixture so that each test gets a fresh MySQL 8.4.7 schema migrated V1–V39 by Flyway (36-line patch, on the assets branch). Result: 30/30 pass on InnoDB, the same as on H2. That includes the publication budgets (verifyDispatchreads the locked journal 0 times,renewcosts 4 journal statements, a fenced renew costs 0) and the stamp casesstaleHeadStampRescansInsteadOfTrustingTheColumns,staleHeadStampRebackfillsAndRejoinsTheHeadPath,noChangeCommitDoesNotRestampASkewedHeadandpublicationAuthorizationRescansAndBackfillsPreMigrationHeads. - Full
managed-agent-serverunit suite on the trial merge (main with Druid, plus the PR): 633 run, 0 failures, 0 errors, 1 skipped. The two timing-sensitiveToolPublicationStoreTestcases named in the PR description passed here. - Plans on the scaled database (figure 4): a 20-row page drops from 41 / 81 / 61 / 121 statements to 3 / 4 / 3 / 5. Both new batch reads use the V37 index:
findLatestTurnsreads 39 index entries andfindLatestEnvironmentEventsreads 78 rows. When Sessions have 500 Turns each, the optimizer switches the environment read tomanaged_agent_event_turn_idxby itself (40 rows).
Observations (non-blocking)
findLatestTurnscost grows with Turns per listed Session (figure 4). It reads everyturn.acceptedindex entry of each listed Session and then takes MAX; MySQL uses no loose index scan here. Server time for a 20-row page: 0.61 ms at 1 Turn per Session, 1.16 ms at 50, 4.49 ms at 500, 35.95 ms at 5,000. main's per-row read costs 0.10–0.30 ms × 20 plus 20 round trips, so the crossover lies somewhere between 500 and 5,000 Turns per Session; at 500 PR is still ahead. I don't think this blocks the PR, but it is worth remembering for very long-lived Sessions. I tried a per-SessionORDER BY … DESC LIMIT 1rewrite and it was slower (117 ms, or 158 ms withFORCE INDEX), so I'm not proposing a fix.- R5-3 is real but small. On PR,
QwenHostedHarnessConnectorstill issues about 1.7 single-rowSELECT approval_modeper Workspace Turn (41 across 24 Turns). main's 904 such reads came mostly from per-row list lookups, which this PR removes. Next to roughly 1,600 Session Store commits per 1,500-event Turn, the remainder is noise, so a follow-up is fine. - API clients can see both documented trade-offs. In the trickle run, 9.8 s in, the WebShell transcript had 38 tokens while
GET …/itemsshowed 19. A revoked reader received 45 more events over 4.7 s. Both match the design.PT0Srestores the strict behaviour at main's query cost. Operators who rely on prompt revocation may want a line about this in the release notes.
Not covered
- No end-to-end tool publication or artifact download (see §4). MariaDB and Linux/Windows were not run locally; CI's MariaDB, MySQL 8.4 and Windows Java legs are green on the head.
- Concurrency stayed below the core count to avoid the JDK 21 virtual-thread pinning wedge described in fix(managed-agent): ≥8 concurrent Turns stall after the model answers on modest hardware (lock convoy in the store path) #13333, which is not related to this PR.
Evidence: figures, raw results.json and samples for every run, the rig (rig.mjs, scale.mjs, configs) and the MySQL fixture patch are on assets-pr13217.
中文版
真实栈验证 — #13217 @ 4d3b1351c7(试合并到 main 6136786c0c)
合并参考结论:支持合并。 在真实 MySQL 8.4.7 + Druid 部署上,只通过公开 API 与 WebShell API 驱动,公开 API 能触达的三条热路径都按 PR 所述大幅下降:
- 列表页:41–141 条语句 → 3–6 条。
- 快照:重写次数降为 1/18–1/22,binlog 字节约降为 1/18.5。
- SSE 授权查询:降为 1/139。
两项有界陈旧放宽都落在文档承诺的范围内;read-grant-recheck-interval=PT0S 能恢复 main 的撤销行为和查询量。升级路径方面:
- V36–V39 在有数据的库上顺利应用。
- 升级前后 6 个列表/详情接口的响应逐字节一致。
- 旧二进制能继续在迁移后的 schema 上服务。
- 旧二进制提交后
activation_head_revisionstamp 会落后,这正是设计所依赖的滚动窗口前提,并在真实流量下得到确认。
无阻断项。R5-3 的冗余查询确认存在,但影响很小。发布授权(第 4 条路径)和 artifact 下载节流在默认配置下公开 API 无法触达,改为在真实 MySQL 上运行 PR 自带的预算测试来覆盖(第 4 节)。
运行方式
| 两臂 | main = 6136786c0c(含 Druid 连接池替换 #13363、#13365)· PR = 试合并 8136a6be2c = PR head 4d3b1351c7 + main,相对 main 的 diff 恰为 PR 的 41 个文件 |
| 每臂的栈 | Spring Managed Agent Server fat jar(JDK 21.0.12,隔离 Maven 仓库,内嵌 Runtime Broker)· 本机 MySQL 8.4.7(开启 performance_schema,ROW binlog,time_zone=+00:00,JVM TZ=UTC)· 从 main 构建的打包 Hosted Harness dist/cli.js serve --profile hosted-harness(PR 不改 TypeScript)· 确定性假 OpenAI 模型 |
| 流量 | 只调用 /v1/agents/sessions… 和 /api/agent/web-shell/v1/…;Workspace 绑定会话真实经过 Broker worker 与挂载目录 |
| 计量 | 语句数和扫描行数来自 performance_schema.events_statements_summary_by_digest,按该臂的 schema 过滤;列表计数扣除 2 秒空闲窗口内的后台语句;快照字节取 binlog 中 managed_agent_snapshot 的行事件字节;快照滞后直接从 MySQL 采样。应用内部不做任何插桩 |
| 宿主 | 10 核 macOS,全程有其他会话占用,负载 30–40,耗时数据噪声大,因此不对延迟下结论 |
1. 公开 API 可触达的三条热路径(图 1)
| 操作 | main | PR |
|---|---|---|
GET /v1/agents/sessions,20 个未绑定会话 |
41 条语句 · 扫描 54 行 | 3 · 40 |
GET /v1/agents/sessions,20 个 Workspace 绑定会话 |
81 · 94 | 4 · 60 |
WebShell sessions/query,20 个未绑定会话 |
61 · 254 | 3 · 90 |
WebShell sessions/query,20 个 Workspace 绑定会话 |
141 · 1146 | 6 · 131 |
| 公开详情 / WebShell 详情 | 3 / 4 | 3 / 3 |
| 突发:5 个 Turn × 3000 delta(15,025 个事件) | 565 次重写 · binlog 84.7 MiB · SQL 1817 ms | 26 · 4.6 MiB · 121 ms |
| 60 秒流:1500 delta、间隔 40 ms(1,505 个事件,items_json ≈ 178 KB) | 258 次 · 86.0 MiB · 862 ms | 14 · 4.7 MiB · 58 ms |
| 该流上 4 个 SSE 订阅者(2 个公开流 + 2 个 WebShell 流) | managed_workspace_access 查询 6,943 次(每订阅者每事件 1.15 次) |
50 次(0.008) |
关于这些数字:
- 在真实栈上,main 的 Workspace 绑定 WebShell 列表页每行要 7 条语句(approval_mode、关闭状态、create-command、registry 各一次逐行读取,外加 turn 与环境事件查询),比 PR 描述中 41–81 区间所对应的每行 4 条还多。PR 的 6 条等于设计 §8 对「含创建者自有会话的绑定页」给出的上界。
- PR 下 4 个 SSE 订阅者都按序收到了全部 1505 个事件。
- 历史 Turn 的墙钟时间:main 23–33 秒,PR 19–32 秒。两臂的时间主要都花在每个 delta 一次的 Session Store 提交上(1,505 个事件的流分别提交了 1,608 / 1,614 次 journal),在这台机器上看不出端到端延迟的变化。
2. 实测有界陈旧(图 2)
-
快照 /
GET …/items:60 秒流式输出期间,main 的快照刷新间隔中位数为 268 ms,PR 为 5058 ms(最大 5168 ms)。PR 下未进入快照的最老事件最多 4.9 秒,最多落后 120 个事件(main 为 11 个)。两臂都在 terminal 批次时追平。「先慢流再暂停」场景里,PR 最多落后 22 个事件,在写入暂停约 5 秒后经snapshot_stale_since重新选中而收敛;实测 marker 先置位、后清除。WebShell transcript 因为会读取快照之后的尾部事件,全程显示所有 token。例如 9.8 秒时,transcript 已有 38 个 token,而GET …/items只有 19 个。 -
读权限撤销(15 秒的流进行到第 4 秒时把
can_read置为 FALSE):臂 流在撤销后关闭的时间 撤销后多送出的事件 main 123 ms 0 PR(默认 5 秒窗口) 4666 ms 45 PR + read-grant-recheck-interval=PT0S144 ms 0 公开流与 WebShell 流表现一致;撤销后发起的新订阅在所有臂都返回 404。
PT0S同时也恢复了逐事件的查询量(6,905 次,main 为 6,943 次)。
3. 升级、回滚与滚动窗口 stamp(图 3)
同一个 MySQL 库依次经历四个阶段。每个阶段都是一对新的 Spring + Harness 进程,部署目录保持不变,与真实主机重启一致。
- main(V35) 创建 41 个会话,其中 29 个有 Turn。
- PR 在该库上启动。 Flyway 用 0.049 秒应用 V36–V39。同一批数据下,6 个列表/详情接口的响应与 main 逐字节一致(15–18 KB),每次调用的语句数从 41 / 81 / 61 / 141 / 3 / 4 降到 3 / 4 / 3 / 6 / 3 / 3。6/6 个旧会话和 1/1 个新会话的新 Turn 全部完成。
- main 回到 V39 库上(即回滚,或滚动发布中仍是旧版本的那一半实例)。Flyway 输出
Successfully validated 39 migrations,服务正常运行。6/6 个旧会话加 1 个新会话全部完成;一个 205 个事件的流完整送达两个 SSE 订阅者。所有阶段的 Spring 日志都没有ERROR。 - PR 再次启动。 6/6 加 1/1 全部完成。
28 个带 journal 的列表会话上,journal head 激活 stamp 的变化:
- 升级前:28 个
null。 - PR 提交后:6 个
current。 - 旧二进制提交后:这 6 个变为
lagging。旧二进制推进了journal_revision,却不维护这些列;这正是 stamp 设计要识别的滚动窗口偏差,这里由真实流量产生。 - PR 下一次携带激活变更的提交后:这 6 个重新变为
current。 - 在 ab-pr 那一轮中,真实 Hosted Turn 把全部 31 个 head 的 V36 列都填上了,stamp 均为最新。
读侧行为(拒绝 lagging 或 null 的 head,然后回扫并回填)需要开启 tool publication,因此改在第 4 节覆盖。
放大库。 把 main 臂的数据复制 100 倍到一个 V35 schema:4,300 个会话、1,736,300 个事件、2.1 GiB。
- V36–V39 用 4.87 秒完成,其中 V37 事件索引占 4.74 秒。
- 随后在 214 万行上重建 V37,同时每约 5 ms 写入一个事件。建索引的 3.3 秒内完成了 474 次写入,最大延迟 6.4 ms、平均 0.54 ms,说明 InnoDB 的在线 DDL 没有阻塞事件写入。
每次重启后,旧会话上的第一个 Turn 都要等上一个 Harness 的 writer lease 过期,约 32–35 秒。这在每个阶段、两种二进制上都一样,与本 PR 无关。
4. 公开 API 触达不到的路径
发布授权(第 4 条路径)和 artifact 重新校验窗口都需要 qwen.managed-agent.tool-publication.enabled(默认 false)和对象存储,还需要 hosted-workspace-shell/1 类型的会话;而 Managed Agent API 只会创建 hosted-workspace-files/1 会话。因此在默认配置下,这两条路径都无法通过公开 API 触达。改为做了以下验证:
- 在真实 MySQL 上运行 PR 自带的
Issue13181QueryBudgetTest。 只改了它的 fixture,让每个用例都使用一个新建的 MySQL 8.4.7 schema,并由 Flyway 完整执行 V1–V39(36 行补丁,已放在 assets 分支)。结果是 InnoDB 上 30/30 全部通过,与 H2 相同。其中包括发布路径的预算(verifyDispatch持锁读 journal 0 次、renew4 条 journal 语句、fenced renew 0 条),以及 stamp 相关用例:staleHeadStampRescansInsteadOfTrustingTheColumns、staleHeadStampRebackfillsAndRejoinsTheHeadPath、noChangeCommitDoesNotRestampASkewedHead、publicationAuthorizationRescansAndBackfillsPreMigrationHeads。 - 在试合并上运行
managed-agent-server全量单测(含 Druid 的 main 加上本 PR):共 633 个,0 失败、0 错误、1 跳过。PR 描述中提到的两个时序敏感的ToolPublicationStoreTest用例在这里也通过了。 - 放大库上的执行计划(图 4):20 行列表页的语句数从 41 / 81 / 61 / 121 降到 3 / 4 / 3 / 5。两条新的批量查询都走 V37 索引:
findLatestTurns读 39 个索引项,findLatestEnvironmentEvents读 78 行。当每个会话有 500 个 Turn 时,优化器会自动把环境事件查询切换到managed_agent_event_turn_idx(读 40 行)。
观察(不阻断)
findLatestTurns的成本随列表中每个会话的 Turn 数增长(图 4)。它会读出每个会话的全部turn.accepted索引项再取 MAX,MySQL 在这里没有用 loose index scan。20 行列表页的服务端耗时:每会话 1 个 Turn 时 0.61 ms,50 个时 1.16 ms,500 个时 4.49 ms,5,000 个时 35.95 ms。main 的逐行查询每条 0.10–0.30 ms,共 20 条,另加 20 次往返,所以交叉点位于每会话 500 到 5,000 个 Turn 之间;500 个时 PR 仍更快。我认为这不阻断合并,但对超长生命周期的会话值得留意。我试过按会话ORDER BY … DESC LIMIT 1的改写,反而更慢(117 ms,加FORCE INDEX后 158 ms),所以没有提出修改方案。- R5-3 确实存在,但影响很小。 PR 下
QwenHostedHarnessConnector每个 Workspace Turn 仍会发出约 1.7 次SELECT approval_mode单行查询(24 个 Turn 共 41 次)。main 的 904 次主要来自列表页的逐行查询,本 PR 已经去掉。与每 1,500 个事件的 Turn 约 1,600 次 Session Store 提交相比,剩下的这些可以忽略,留作后续即可。 - 两项有文档说明的权衡,API 调用方都能直接看到。 慢流场景进行到 9.8 秒时,WebShell transcript 已有 38 个 token,而
GET …/items只有 19 个。被撤销读权限的订阅者在 4.7 秒内又收到了 45 个事件。两者都与设计一致;PT0S能以 main 的查询量为代价恢复严格行为。依赖及时撤销的运维方,可能需要在发布说明里看到这一点。
未覆盖
- 没有端到端驱动 tool publication 或 artifact 下载(见第 4 节)。本地没有跑 MariaDB 和 Linux/Windows;head 上 CI 的 MariaDB、MySQL 8.4 和 Windows Java 任务均为绿色。
- 并发量始终控制在核数以下,以避开 fix(managed-agent): ≥8 concurrent Turns stall after the model answers on modest hardware (lock convoy in the store path) #13333 中 JDK 21 虚拟线程钉住导致的卡死;该问题与本 PR 无关。
证据(图、每轮原始 results.json 与采样、装置 rig.mjs / scale.mjs / 配置、MySQL fixture 补丁)位于 assets-pr13217。
qqqys
left a comment
There was a problem hiding this comment.
Critical-only review at head 4d3b1351. No blocking issue found — approving.
Four rounds of this PR filed ten Criticals, so I did not treat the newest approval or the empty round-5 ledger as proof on its own. What I checked:
No Critical thread stands open at this head. I enumerated every review thread with isResolved == false through GraphQL — there are more than fifty, across the design docs, both store classes, the event-stream and artifact services and the new budget suite — and every one of them is graded [Suggestion]. Not one unresolved thread carries [Critical]. Round 5's ledger, whose sha is this exact commit, likewise files three findings and all three are sev:"S". So the R1-R4 Criticals are not merely unanswered; nothing blocking remains anchored here.
One of the re-pointed R3 Criticals I confirmed in the source myself. The write-time journal scope fence is present at ManagedExtensionRecordStore.java:190-196, building a sessionScope node from tenantId/workspaceId/sessionId with the comment that the closed-key check both read paths enforce now applies at write time too, so a misscoped line never enters the journal. I could not locate the second one's site (the cursor-less transcript-tail bound) inside this review's time box; I am relying on the round-5 pass at this head plus the absence of any open Critical thread for that one, and saying so rather than implying I traced it.
CI at this head really does cover the four new migrations. My first read of the checks was truncated at 100 of 493 and looked like almost nothing had run; paginating gives 26 successful, 455 skipped, 12 cancelled and zero failures, with Runtime Broker and Managed Agent MariaDB / Java 21 and Hosted process fault gates / MySQL 8.4 / Java 21 both green — those are the lanes that apply V36-V39 against a real database — plus the full Java matrix, Test (ubuntu-latest, Node 22.x) and Lint & Static. The Windows and macOS TypeScript lanes are skipped, which is this repo's normal routing and not a signal either way.
No migration version collision, verified against current main rather than assumed. Main's migration directory today tops out at V35__managed_session_tool_profile.sql, this head adds V36 through V39, and git merge-tree main head reports a clean merge with no conflicted paths.
Two non-blocking notes, neither a condition of this approval:
- V36 is claimed by another open PR. #13355 adds
V36__managed_task_event.sqland was, when I reviewed it earlier today, also clean against a main that stopped at V35. Only one of the two can keep the number: the Flyway uniqueness gate sees one branch at a time, so it cannot catch a cross-PR collision, and the loser fails at startup with "Found more than one migration with version 36". Whichever merges second needs renumbering to V40+. Worth settling before either lands. - The bounded-staleness cluster deserves a maintainer's eye even though it is graded Suggestion. The 5s SSE read-grant recheck window and the artifact revalidation window trade prompt revocation for the query savings this PR exists to deliver; the approving review records them as issue-sanctioned and reversible with
PT0S, and several open Suggestions (R1-13, R2-12, R1-24, R1-14) argue the committed contract and design docs still state a stricter bound than the code now guarantees. That is a documentation-versus-behaviour gap on a security-relevant latency, so it is worth closing even though it does not block.
Coverage, so this approval is not read as broader than it is. I verified the items above and read the migration set, the scope fence and the thread state. I did not independently audit the remaining production surface — the snapshot snapshot_stale_since deferral marker and re-selection disjunct, the batched session-page assembly, the journal-head activation stamp and its journal-head-authorization flag (default off, so the shipped path is the pre-existing journal scan), or the millisLenient digit-width pre-check. Round 5 examined this head and filed nothing blocking there, which is why I am approving rather than deferring; if a maintainer wants one part attested independently, the head-authorization stamp and its self-healing rescan is where I would look first.
yiliang114
left a comment
There was a problem hiding this comment.
Deep-review pass at 4d3b135 (commenting; qqqys's approval already stands).
Verified the four hot paths against the code, not the design text:
- Snapshot gating: the four rewrite conditions match §3 exactly, and the deferral marker writes the snapshot's
updated_at(so a fresh snapshot waits out the window; an aged one converges on the next tick).rewriteStaleSnapshotclears the marker on convergence and on recreation-by-retraction. - SSE recheck window: per-stream
ReadGrantwith aSystem.nanoTimewindow; initial grant check unchanged;session.deletedstill terminates immediately. - V36 activation columns: commit extracts
activation.changedduring the existing parse pass, blanks oversized values instead of rejecting (the readers treat NULL as absent — checked against the freshness rule), and the stamp-equality rule makes a pre-V36 writer's commit self-detecting (journal_revision bumps without the stamp → rescan + backfill). Rolling window is genuinely self-healing as designed. - producerBindingLocked: O(1) only under the flag; legacy scan + backfill otherwise.
requireEvidence's intent locate uses the V39 index — four statements regardless of journal depth, as §6 claims. - Migrations: V36–V39 are clean — main's head is V35, no collision.
Not line-verified: the QueryLedger budget test's exact per-endpoint counts (read the budgets, did not re-derive each query), and the E2E arms. Nothing blocking.
#13217 landed migrations V36 through V39 on main, so the task-event outbox moves to V40. ManagedExtensionRecordStore.apply() keeps this branch's per-line checks and takes #13217's additions: the last activation.changed payload rides back in ApplyResult for the journal head columns, and a non-Stage-H event line scoped to another Session or version is still refused as "Journal event scope conflicts", the message both publication read paths give it. An event past the declared count now names its invalid journal position as well. #13217's tests built journal lines the commit now refuses: an untyped {} in the commit marker's place, the unknown kinds tool.progress and checkpoint.saved, and envelopes without eventId and occurredAt. They now write real commit markers, known kinds and full envelopes through PublicationJournalFixture. duplicateIntentLinesAtOneSequenceAreFenced plants its duplicate in the stored revision, because the commit refuses a line out of its sequence.
Resolve ManagedWorkspaceRegistry.java: main's #13217 anchored its new createdSessions batch twin on isSessionCreator, which this PR deletes. Keep the batch twin (ManagedAgentService still calls it) and keep the deletion plus this PR's bindingCurrent probe; the merged tree has no remaining isSessionCreator reference. Co-authored-by: Qwen-Coder <[email protected]> Patrol-Run: qwen-pr-conflict/jmutwkfe753
Reconciles the extension-store verification this branch adds with main's #13217 (stop database amplification on session hot paths): * The store keeps the transaction-request form of apply() so the declared digests and ranges stay verifiable, and returns the new ApplyResult (receipts + the last activation's payload) that #13217's commit uses instead of re-parsing the journal. The session-scope message unifies on "Journal event scope conflicts", pinned already by five sites on main; the closed-shape and position checks the shared event rules already covered are not duplicated. * The migration V36 became V40 and V37 became V41: main landed its journal-activation V36 through snapshot-deferral V38 and journal- sequence V39 after this branch's numbering. * The shared PublicationJournalFixture now commits in the journal shapes the store verifies — a real genesis, a commit marker ending every transaction and canonical content digests — instead of the loose "{}" fillers and raw text digests both sides previously used. * #13217's own replay snapshots are reconciled to the store's closed event rules: phases, workers, lease fields and subjects completed where the commit still lands, and the oversized/mistyped-payload arms (id over maxIdBytes, a phase outside the vocabulary, a stringly typed or absent expiry, a misplaced second event) now pin the refusal where the stricter layer fires instead of assuming the commit stores the poison. The duplicate-intents case asserts the closed-domain/scope rule the line actually trips. Verified on this tree: full Java unit suite 641 green (the three most affected classes 133/133 first), the managed-runtime TypeScript suites 1933 green, and neutralizing both canonical digest guards turns the declared-digest battery red before the guards were restored.
…#13565) Move the merged-PR history of the Managed Agent dual-path proposal (QwenLM#12380) out of the issue body into a bilingual ledger under docs/design/. The issue body had come within 10 KB of GitHub's 256 KiB limit, so the body will keep only the delivery snapshot and the open PRs, and merged rows move here rewritten to their merged final state. The ledger carries every merged row the issue tracked up to its 2026-10-03 reconcile (adding the missing merge commit to the five earliest rows and normalising the Chinese state cells), rewrites the eight rows the issue still listed as open although their PRs had merged (QwenLM#13141, QwenLM#13166, QwenLM#13174, QwenLM#13210, QwenLM#13214, QwenLM#13217, QwenLM#13218, QwenLM#13247), and adds rows for the 48 managed-agent PRs merged between that reconcile and main 0c13502 that had no row yet. PRs closed without merging (QwenLM#13087, QwenLM#13336) sit in their own table. Later merges land at the next reconcile. Co-authored-by: wenshao <[email protected]>




What this PR does
Fixes the four database-amplification hot paths reported in the issue, plus the two related ones. Snapshot materialization no longer rewrites the whole
items_jsondocument under the session row lock after every ≤200-event batch — it rewrites at creation, on a terminal event, every 1000 covered events, or on catch-up once the previous snapshot is 5s old, and a deferred rewrite is recorded as asnapshot_stale_sincemarker on the consumer-progress row (migration V38) so the session is re-selected until the snapshot converges. SSE streams no longer run a read-grant SQL check before nearly every delivered event per subscriber — each stream re-verifies at most once per configurable recheck window (default 5s). The session list endpoints no longer run 2–4 extra queries per row — a page is assembled from a constant number of grouped batch queries that project only the summary columns, and the single-session views share that assembly. Tool-publication authorization no longer rescans the session journal backwards (one locked read per revision) on every publish/seal/prefix/finish — a new migration carries the current activation state on the journal head, maintained transactionally by commit inside the parse pass every commit already runs, with a scan-and-backfill fallback for pre-migration journals; reserve/renew read thetool.intentat its own revision instead of walking the journal, and a fenced activation is answered from the locked head before any journal read. Related: artifact downloads now re-validate content access at most once per revalidation window instead of per 64 KiB chunk; and two further migrations add an event-type index for latest-turn lookups and a(tenant_id, session_id, last_sequence)index for the intent range read.Because a rolling fleet can briefly run a binary that commits without maintaining the new head columns, trusting them is gated behind
qwen.managed-agent.tool-publication.journal-head-authorization(defaultfalse= the pre-existing journal scan); an operator flips it once every writer runs the V36 schema's code. The head columns carry anactivation_head_revisionstamp naming the journal revision they reflect, so a premature flip is still safe: a head last written by an old binary fails the stamp check, gets rescanned and re-stamped — the rolling window self-heals per session instead of certifying stale state. Activation-expiry parsing pre-checks the digit width before materializing any big integer, so a hostile exponent-formexpiresAtcosts nothing on any of the three read paths.Why it's needed
A code audit showed the broker's database work scaling with accumulated history rather than with change: snapshot writes grew quadratically with session length while blocking event ingestion on the same row lock, N subscribers × M events produced N×M permission queries on an unbounded executor, a default 20-row session page cost 41–81 queries, and a long turn delayed each 16 MB publication segment by hundreds to thousands of locked journal reads. The issue asks for heat to scale with change; this PR does that with pinned per-endpoint query budgets so none of it can silently regress.
Reviewer Test Plan
How to verify
Run
mvn -o testinpackages/sdk-java/managed-agent-server(install the sibling modules first:mvn -o install -DskipTests -Dgpg.skip=trueinpackages/sdk-java/qwencodeandpackages/sdk-java/runtime-broker). The newIssue13181QueryBudgetTestdrives the production stores/services over a statement-recording H2 (MySQL mode) DataSource and pins the measured before→after counts: snapshot rewrites per 9-batch burst materialization go from[1,1,1,1,1,1,1,1]to[1,0,0,0,0,0,0,0,1], and a 12-tick trickle records[1,0,…,0]; a 21-event workspace stream'smanaged_workspace_accessqueries go from 43 to 2; a 20-row page costs 3/4/3/3 queries where it cost 41/61/61/81;verifyDispatchwith the head 30 revisions past the last activation goes from 31 locked journal reads to 0;renewreads a constant 4 journal statements regardless of filler depth (the revision range read, the chain-contiguity count, and the verified page), and a fenced renew pays zero journal statements; a pre-migration head gets exactly one rescan that backfills the head. The bounded-staleness semantics are pinned both ways: zero-window restores per-event/per-chunk checks (existing revocation tests), and new tests prove revocation lands at the window's end, that a successful recheck re-anchors the window, and that the shipped 5s default is what the suite exercises. Note two pre-existing timing-sensitiveToolPublicationStoreTestcases (renewsTheOriginalClaimWhileScanningSlowObjectBytes,streamsLargeOutputAndReadsItsTailAfterStoreReplacement) fail on loaded machines on the base too — verified identical on an unmodified worktree.Evidence (Before & After)
N/A (non-UI; measured query counts are in the issue comment I'll post and in
Issue13181QueryBudgetTest's recorded output)Tested on
Environment (optional)
OpenJDK 21 + Maven (offline), H2 in MySQL mode running the real Flyway migrations (V34–V39 apply cleanly;
RuntimeBrokerFlywaySchemaTestpasses). MySQL*ITintegration suites need Docker and were not run locally — left to CI.Risk & Scope
PT0Srestores the old per-event/per-chunk behavior); the snapshot may lag the projection by up to 1000 covered events during bursts, rewrites unconditionally on a terminal event, and a deferred trickle snapshot converges within 5s via aged-out reselection. The artifact window also defers the stage-1DELETINGlifecycle observation to the window's end (stage-2 retirement still aborts per chunk through the read lease).journal-head-authorizationis enabled, which an operator should do after the fleet fully runs the V36 code — and the stamp makes even a premature flip self-healing (rolling-deploy rationale in the design doc).Linked Issues
Fixes #13181
中文说明
这个 PR 做了什么
修复 issue 报告的四条数据库放大热路径及两个相关项。快照物化不再在每批 ≤200 事件后、于会话行锁内全量重写
items_json—— 改为首次创建时、terminal 事件批次、每覆盖 1000 事件时,或追平且距上次快照写入满 5 秒时重写;被推迟的重写以snapshot_stale_since标记记录在 consumer-progress 行上(迁移 V38),会话因此被重新选中直至快照收敛。SSE 流不再于几乎每个投递事件前、按订阅者执行读授权 SQL —— 每条流在可配置的复检窗口(默认 5 秒)内最多复检一次。会话列表接口不再每行多跑 2–4 条查询 —— 页面由固定数量的分组批量查询装配且只投影摘要列,单会话视图复用同一装配。工具发布授权不再于每次 publish/seal/prefix/finish 倒扫会话 journal(每个 revision 一次持锁读)—— 新迁移把当前 activation 状态承载到 journal head,由 commit 在每次提交本就要做的解析遍历中顺带维护,迁移前的 journal 由"扫描一次 + 回填"回退覆盖;reserve/renew 按 intent 所在 revision 直接读取而非逐 revision 回扫,被围栏的 activation 在任何 journal 读取之前就由已加锁的 head 作答。相关项:artifact 下载的内容访问复检从每 64 KiB 分片一次改为每个复检窗口最多一次;另两个迁移分别新增事件类型索引(最新 turn 查询读索引范围)与(tenant_id, session_id, last_sequence)索引(intent 范围读)。由于滚动部署期间集群可能短暂运行不维护新 head 列的旧二进制,信任这些列由
qwen.managed-agent.tool-publication.journal-head-authorization(默认false,即沿用旧的 journal 扫描)门控;运维在全部写入方运行 V36 代码后打开。head 列带有activation_head_revision时间戳,标明它们反映到哪个 journal revision,因此提前打开开关同样安全:旧二进制最后写入的 head 无法通过时间戳检查,会被重扫并重新打戳 —— 滚动窗口按会话自愈,而不是把陈旧状态认证为有效。activation 过期时间的解析在任何大整数物化之前先做宽度预检,恶意的指数形态expiresAt在三条读取路径上都零开销。为什么需要
代码审计显示 broker 的数据库工作量随累积历史而非变化量增长:快照写入随会话长度平方增长并在同一行锁上阻塞事件摄入;N 订阅者 × M 事件在无界执行器上产生 N×M 条权限查询;默认 20 行会话页成本 41–81 条查询;长 turn 让每 16MB 发布分片拖后数百到数千次持锁 journal 读。issue 要求热度随变化量伸缩;本 PR 做到了,并为每个端点钉住查询预算,防止静默回归。
评审者测试计划
如何验证
在
packages/sdk-java/managed-agent-server运行mvn -o test(先安装兄弟模块:在packages/sdk-java/qwencode与packages/sdk-java/runtime-broker执行mvn -o install -DskipTests -Dgpg.skip=true)。新增的Issue13181QueryBudgetTest在记录语句的 H2(MySQL 模式)DataSource 上驱动生产 store/service,钉住实测的前后对比:9 批突发物化的快照重写从[1,1,1,1,1,1,1,1]变为[1,0,0,0,0,0,0,0,1],12 周期涓流记录为[1,0,…,0];21 事件工作区流的managed_workspace_access查询从 43 降为 2;20 行页的查询数为 3/4/3/3(原 41/61/61/81);head 领先 activation 30 个 revision 时verifyDispatch的持锁 journal 读从 31 降为 0;renew无论堆积多少 revision 恒定 4 条 journal 语句(revision 范围读 + 链式计数 + 校验页),被围栏的 renew 零 journal 语句;迁移前的 head 首次回扫一次并回填。有界陈旧语义双向钉住:零窗口恢复逐事件/逐分片检查(既有撤销测试),新测试证明撤销在窗口末端生效、成功复检会重新装配窗口、以及发布的 5 秒默认值正是套件所覆盖的。注意ToolPublicationStoreTest有两个既有的时序敏感用例(renewsTheOriginalClaimWhileScanningSlowObjectBytes、streamsLargeOutputAndReadsItsTailAfterStoreReplacement)在负载高的机器上于基线同样失败 —— 已在未修改的工作树上验证一致。证据(前后对比)
N/A(非 UI;实测查询计数见我将发布到 issue 的评论与
Issue13181QueryBudgetTest的记录输出)已验证平台
环境(可选)
OpenJDK 21 + Maven(离线),H2 MySQL 模式跑真实 Flyway 迁移(V34–V39 干净应用;
RuntimeBrokerFlywaySchemaTest通过)。MySQL*IT集成套件需要 Docker,本地未跑 —— 留给 CI。风险与范围
PT0S恢复旧的逐事件/逐分片行为);突发期间快照最多落后投影 1000 条已覆盖事件,terminal 事件批次必重写,被推迟的涓流快照经老化重选在 5 秒内收敛。artifact 窗口同样把 stage-1DELETING生命周期观察推迟到窗口末端(stage-2 retirement 仍经读取租约逐分片中止)。journal-head-authorization打开后才被信任,运维应在集群全部运行 V36 代码后打开 —— 且时间戳使提前打开也能自愈(滚动部署依据见设计文档)。关联 Issue
Fixes #13181