Repository navigation
feat(managed-agent): Restore safe retired tool output collection - #13225
Conversation
Co-authored-by: Qwen-Coder <[email protected]>
Co-authored-by: Qwen-Coder <[email protected]>
|
[codex] Restored O4-2 at
Final committed-head gate: build, typecheck, bundle, migrations, source-matching Java dependency installation, complete ordinary MySQL profiles and source-derived report checks passed. Broker unit reports: 593 total, 592 passed and one optional real-worker bundle test skipped; Broker MySQL: 7/7 passed. Server: 487/487 unit and 35/35 MySQL integration passed. Checkstyle 0. Two clean independent code-review passes and self-audit completed; fixture, GET and controlled performance execution snapshots are bound to the final Git objects. These local macOS/MySQL and controlled HTTPS results do not establish Linux identity guards, full Hosted producer/Broker/Runtime acceptance or current-revision credentialed cloud OSS. The historical maintainer OSS 36/37 run still lacks the delete-denied identity. Public Workspace DELETE remains unavailable, and #13192 remains an independent open PR. Those are deployment/readiness boundaries, not green checks inferred from admin-only CI. Current #13225 product CI is separately pending; physical GC is disabled and grace remains 24 hours. 中文:O4-2 已以 上表中的容量、诊断和 mutation 建议逐项明确延期到 #12380,部分采纳不算全部修复。永久阻塞轮询饱和、已排期宽限期缩短不会即时生效、C3 和 handoff 剩余 mutation 缺口继续保留。Linux/完整 Hosted、真实 OSS 删除拒绝身份、公开 Workspace DELETE 及独立开放的 #13192 不被本地绿灯替代。没有自动合并;物理 GC 关闭,部署宽限保持 24 小时。 |
Conflicts were limited to close-slice files that main has since
reworked: docs/design/workspace-session-reliable-close.{md,zh-CN.md}
and WorkspaceRecoveryStoreTest.java (operation schema fixture moved to
RUNNING/LEASED with lease_until). Resolution takes main's versions; this
PR now carries only the collection slice.
Migration numbering on the merged tree: V31 recovery, V32 close (SQL
bytes identical on both sides), V33 collection; Java migrations remain
V15/V29 only, so no Flyway version collision.
Verified on the merged tree: sdk-java managed-agent-server mvn test,
487 tests 0 failures, including the Flyway-backed
WorkspaceRecoveryStoreTest and the new collection tests.
|
Merged current origin/main (c15dd26) into this branch to clear the conflict, as merge commit 0c64229. The conflict was limited to close-slice files that main has since reworked: the two workspace-session-reliable-close design docs and WorkspaceRecoveryStoreTest.java (operation fixture moved to the RUNNING/LEASED + lease_until schema). Resolution takes main's versions, so this PR now carries only the collection slice. Numbering re-verified on the merged tree: V31 recovery, V32 session close (SQL bytes identical to main's), V33 tool output collection; Java migrations remain V15/V29 only — no Flyway version collision. Verified on the merged tree before pushing: sdk-java managed-agent-server full mvn test — 487 tests, 0 failures, including the Flyway-backed WorkspaceRecoveryStoreTest and the new collection tests (matches your recorded 487/487 gate). |
Brings AgentDefinition D8a (#13142) and session archive/delete (#13194). The textual merge was clean; the one semantic collision is the migration number: main took V33 for managed_agent_definitions, so this PR's tool output collection migration is renumbered V33 -> V34 with SQL bytes unchanged. The retention design docs (en + zh-CN) census now reads V31 recovery / V32 close / V33 agent definitions / V34 collection, and provisional V28/V33 collection histories require explicit reconciliation, matching the existing no-automatic-repair policy. Verified on the merged tree with a clean build: sdk-java managed-agent-server mvn test, 514 tests 0 failures, with the Flyway-backed suites exercising V31 through V34.
|
Second merge of origin/main (1abccdb) as 88d4ca9. Main took V33 for managed_agent_definitions (#13142), so the collection migration is renumbered V33 -> V34 with SQL bytes unchanged; the retention docs census now reads V31 recovery / V32 close / V33 definitions / V34 collection, with provisional V28/V33 collection histories requiring explicit reconciliation per the existing policy. Verified on the merged tree with a clean build: managed-agent-server mvn test, 514 tests 0 failures (Flyway-backed suites exercise V31 through V34). The earlier note stands: the Hosted MySQL gate failure signature predates this merge and appears on unrelated branches (e.g. codex/managed-writer-epoch run 37029292498, same test, same signature). |
Local real-environment verification of #13225 at
|
| Arm | Tree | Purpose |
|---|---|---|
| PR | b95e64cdf0 = 0c64229ae6 merged into main 623cfc67 + V33→V34 rename. git diff b95e64cdf0 88d4ca9466 -- packages/sdk-java is empty |
everything below |
| base | main 623cfc67 |
A/B for collection, F2 and GET retry |
| candidate | PR + 2-line patch (candidate-1mib-read-granularity.patch) |
F2 follow-up only |
- Gates and boot matrix: macOS arm64, JDK 21, MySQL 8.4.7.
- Every retirement scenario ran on Linux. On macOS the Workspace close path is unavailable: the durable worker identity reads
/etc/machine-idand/proc, sosession_close=false. The Linux stack was:- a colima VM (aarch64, 4 CPU / 8 GB) running a JDK 21 container and a MySQL 8.4.11 container;
- the Spring fat jar, the packaged Hosted Harness and a durable local-process Broker worker;
- a fake model, and an OSS double over TLS driven by the real
aliyun-sdk-oss3.18.4.
- Real Aliyun OSS: a temporary private, unversioned bucket in cn-hangzhou, deleted afterwards.
- Seams, stated plainly:
- A tap rewrites the Harness tool profile
hosted-workspace-files/1→hosted-workspace-shell/1, because public Shell admission is not on main yet. - The legacy-write-evidence, recovery-status and stale-read-lease cases are injected by SQL.
- Everything else goes through the public API: create, Turn, download,
POST /close,DELETE.
- A tap rewrites the Harness tool profile
- The Harness/worker bundle was built from main
623cfc67. The1abccdb26amerge only adds apackages/corechange (fix(core): stop misdiagnosing malformed tool-call args as max_tokens truncation #12982) that is unrelated to this path.
1. Landing and gates
- At
0c64229ae6, the trial merge into current main fails the repo's owncheck-flyway-migrations.js: two migrations claim version 33. Spring also refuses to boot on it. - The renumbered tree boots on an empty database and on a database that main has already migrated (V33 definitions → V34 collection).
- A database migrated by the PR head alone cannot be opened by main (checksum mismatch at V33). That is the provisional-history case the PR body already calls out.
- Unit: 514/514, Checkstyle 0.
-Pmysql-integration: 51/51, includingWorkspaceSessionRetentionMySqlIT16/16; the failsafe report checker exits 0.
2. Public lifecycle on Linux, and against real OSS
- Public calls:
POST /close202 → completed (554 ms);DELETE202 → completed (292 ms). - Collection: the row waited for
grace_perioduntil exactlyretired_at + grace(10 s). It was claimed at +10.5 s andCOLLECTEDat +10.9 s. - OSS deletes: 15 DELETEs, one per catalog key, only the deleted Session's keys; no prefix scans.
- Cleanup and quota: inline payloads cleared;
held13,641,895 → 0 andreleased_held_bytes= 13,641,895. - Sibling Session: untouched and still readable with the right sha256.
- On main: the same lifecycle leaves the output
RETIRINGwith all objects and quota held. - Real OSS: 8/8 objects of the deleted Session were removed (bucket 16 → 8 objects); the sibling's 8 were kept.
- Real bucket with versioning enabled: the collector failed closed for two attempts with
Tool publication bucket cannot enforce immutable objectsand deleted nothing (4/4 objects kept).
3. Blockers and siblings
Quarantine: I flipped one bit of a stored object. A public download detected it (503 artifact_unavailable) and production code quarantined that publication.
After DELETE:
- The quarantined publication kept 7/7 objects and its quota.
- The other publication in the same Session was collected.
The legacy_write_evidence_missing and recovery_protected rows were kept. A stale read lease held the row as reader_active until the next 60 s poll after it expired (collected at +71.2 s). Tenant held bytes dropped by exactly the two collected publications.
Unknown PUT (finding 3). I injected one PUT 500 on the wire. The next attempt returned, and all 11 objects are VERIFIED. The ledger still has UNKNOWN=1, so after DELETE the publication stays RETIRING / put_unresolved with 6,296,450 bytes held. Nothing in the code moves an attempt out of UNKNOWN.
4. Storage faults
One request per failed DELETE, whole-page retry: a DELETE 500 or connection reset on the 5th key produced exactly one request for that key; the SDK did not replay it. The whole page was deferred by 60 s, and generation 2 re-issued the page; the keys already gone answered 204.
Versioning, crash and multi-page:
- With versioning
Suspended, the collector issued 0 DELETEs. kill -9of Spring while a DELETE was held: the restarted process waited out the old claim (59.6 s), then finished as generation 2.- A 113-key publication was collected in two pages, one tick apart.
- Two live Spring instances with GC on, sharing one database: both collected (two owners seen), but every publication had one owner and generation 1, and each of the 93 keys was deleted once.
In every case the quota was released exactly once.
5. F2: download cost and the 120 s lease
A TCP relay in front of MySQL adds 5 ms to every client→server packet. On main, a 40 MiB download is cut at 24.2 MiB when the lease ends; on the PR it completes in 55.5 s. A 100 MiB download still gets cut at 85.8 MiB on the PR, about 0.73 MiB/s.
What remains is about 11 statements per 64 KiB response chunk:
ManagedArtifactService's 64 KiB write loop checks the lease and access each time;VerifiedStream.readcaps every read at 64 KiB and runs the full guard before and after.
The candidate raises those two constants to 1 MiB, the granularity the PR already uses in the store path. Results:
- Cost and speed: 38.8 statements/MiB; 40 MiB in 10.3 s and 128 MiB in 31.3 s at +5 ms.
- Revocation unchanged: revoke or DELETE mid-download still aborts the response, after at most 0.88 MiB.
It touches main's code rather than the PR's diff. It can land here or as a follow-up; otherwise the body's 100 MiB claim should be softened.
6. OSS GET retry and mutation testing
GET retry. I injected faults on the third object of a public download through the real SDK.
- Retried only on the PR: a 500 with an unrecognised error code, and a 502 with a non-XML body. On main these abort the response after one request; the PR retries once and the download completes with the right digest.
- Same on both arms: 500
InternalError, 503, a connection reset, six 500s in a row (4 requests, then abort) and 403 (no retry). - PUT: a PUT 500 results in two separate attempts, each with one wire request; the SDK replays nothing.
Mutation testing: 28 mutants of the new code, each run against the full unit suite (514 tests). 19 were killed. Of the 9 survivors:
- Not real gaps:
- M04, M06 and M21 are fences that cannot differ while one process owns a claim (
runOnceis synchronized; the owner is per process). The two-collector run above exercised the cross-process case. - M07 (the per-page probe) is redundant: the per-key probe makes the same check and is pinned by a test (M08); removing both (M29) is killed.
- M04, M06 and M21 are fences that cannot differ while one process owns a claim (
- Real gaps:
- M24–M27: the collector tests only exercise the
grace_periodand legacy blockers. Collecting a quarantined, recovery-protected,put_unresolvedorreader_activepublication passes the suite; only the real-stack runs in section 3 show those are retained. - M17: reverting the guarded wrapper to raw
read()also passes. It changes nothing on the public download path: the real stack measured 205.6 statements/MiB with the mutant, the same as the PR. But it guards Session-store and verification reads, which no bounded-query test measures.
- M24–M27: the collector tests only exercise the
Suggestions (none blocking)
- Coordinate V34 with feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210.
- Take the 2-line read-granularity patch, or reword the 100 MiB claim.
- Before turning on
gc-enabled, add a way to resolveUNKNOWNPUT attempts. For example, a guarded HEAD of the exact key with a digest comparison, or an operator command. Without it, every transient PUT error permanently leaks that publication's storage and tenant quota. It also adds to the already-deferred 60 s polling of permanent blockers. - Duplicate versioning probes: each deleted key costs two OSS requests (115 versioning probes for 113 DELETEs), because
deleteIfPresentprobes again after the per-page probe. Only the per-key probe is pinned by a test (M08); the page probe is redundant today (M07). - Add a parameterised collector test over every blocker kind (quarantined, recovery-protected,
reader_active,put_unresolved, …), so that M24–M27 fail.
Evidence (figures, harness, logs, candidate patch): assets-pr13225.
中文版
#13225 本地真实环境验证(88d4ca9466)
结论:在 GC 保持关闭(默认)的前提下,没有阻塞合入的问题。 0c64229ae6 上存在的 Flyway V33 撞号,已由 88d4ca9466 修复。新 head 的 packages/sdk-java 与我验证的树逐字节相同。在 Linux 上走真实的公开 close → DELETE 链路,回收器的行为与 PR 描述一致;假 OSS 和真实阿里云 OSS bucket 都验证过。合入或启用 GC 前需要关注三点:
- 与 feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210 的迁移顺序。 在飞的 feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210 同样新增
V34__managed_session_creator.sql,后合入的一方必须改号。 - F2 的表述。 存储侧 SQL 放大确已修复(每 MiB 711 → 205 条)。但正文"5 ms SQL 延迟下 100 MiB 约 108 s 完成"在本装置上未能复现:120 s 读取租约到期时,100 MiB 只送出了 85.8 MiB。下文的两行后续改动可把成本降到每 MiB 39 条,128 MiB 31 s 完成。
- 启用 GC 前:一次瞬时 PUT 失败会让该 publication 永远无法回收。 这符合设计文档("unknown … 永久阻止自动回收,直到外部解决"),但目前没有任何解决路径,该 publication 的存储和租户配额永不释放。
环境
| 臂 | 树 | 用途 |
|---|---|---|
| PR | b95e64cdf0 = 0c64229ae6 合入 main 623cfc67 + V33→V34 改号;git diff b95e64cdf0 88d4ca9466 -- packages/sdk-java 为空 |
以下全部 |
| base | main 623cfc67 |
回收、F2、GET 重试对照 |
| 候选 | PR + 两行补丁(candidate-1mib-read-granularity.patch) |
仅 F2 后续 |
- 门禁与启动矩阵: macOS arm64,JDK 21,MySQL 8.4.7。
- 所有退役相关场景都在 Linux 上运行。 macOS 上 Workspace close 不可用:durable worker 身份读取
/etc/machine-id和/proc,session_close=false。Linux 栈的组成:- colima VM(aarch64,4 CPU / 8 GB)中的 JDK 21 容器和 MySQL 8.4.11 容器;
- Spring fat jar、打包后的 Hosted Harness,以及 durable local-process Broker worker;
- 假模型,以及由真实
aliyun-sdk-oss3.18.4 经 TLS 访问的 OSS 替身。
- 真实阿里云 OSS: cn-hangzhou 临时私有、未开版本控制的 bucket,用完已删除。
- 缝合点(如实说明):
- tap 把 Harness 的 tool profile 由
hosted-workspace-files/1改写为hosted-workspace-shell/1,因为 main 尚未开放公开 Shell 准入。 - 历史写入证据缺失、恢复状态、残留读租约三种情况通过 SQL 注入。
- 其余全部走公开 API:创建、Turn、下载、
POST /close、DELETE。
- tap 把 Harness 的 tool profile 由
- Harness/worker bundle 由 main
623cfc67构建。1abccdb26a合并只多了一个packages/core改动(fix(core): stop misdiagnosing malformed tool-call args as max_tokens truncation #12982),与本链路无关。
1. 合入与门禁
0c64229ae6试合并到当前 main 后,仓库自带的check-flyway-migrations.js报两个迁移都占 version 33,Spring 也无法启动。- 改号后的树在空库上能启动,在已由 main 迁移的库上也能启动(V33 definitions → V34 collection)。
- 只跑过 PR head 的库,main 无法打开(V33 checksum mismatch)。这是 PR 正文已说明的临时迁移历史情形。
- 单测: 514/514,Checkstyle 0。
-Pmysql-integration: 51/51(含WorkspaceSessionRetentionMySqlIT16/16),failsafe 报告检查器退出码 0。
2. Linux 公开生命周期与真实 OSS
- 公开调用:
POST /close202 → completed(554 ms);DELETE202 → completed(292 ms)。 - 回收时机: 宽限期严格等到
retired_at + grace(10 s);+10.5 s 被领取,+10.9 sCOLLECTED。 - OSS 删除: 15 次 DELETE,每个 catalog key 一次,只删被删会话的 key,没有按前缀扫描。
- 清理与配额: inline 已清空;
held13,641,895 → 0,released_held_bytes= 13,641,895。 - 兄弟会话: 未受影响,下载摘要正确。
- main 上: 同样的生命周期结束后,输出一直停在
RETIRING,对象和配额全部保留。 - 真实 OSS: 被删会话的 8/8 对象被删除(bucket 16 → 8),兄弟会话的 8 个保留。
- 真实 bucket 开启版本控制: 回收器失败关闭,两轮报
Tool publication bucket cannot enforce immutable objects,没有删除任何对象(4/4 保留)。
3. 阻塞与兄弟
隔离: 我翻转了存储对象的一个比特。公开下载检测到损坏(503 artifact_unavailable),生产代码随即隔离了该 publication。
删除后:
- 被隔离的 publication 保留 7/7 对象和配额;
- 同一会话中的另一个 publication 被回收。
legacy_write_evidence_missing 和 recovery_protected 均保留。残留读租约使该行停在 reader_active;租约过期后,在下一次 60 s 轮询时回收(+71.2 s)。租户 held 恰好减少两个已回收 publication 的量。
未决 PUT(第 3 点)。 我在链路上注入了一次 PUT 500。下一次尝试成功返回,11 个对象全部 VERIFIED。但账本里仍留下 UNKNOWN=1,删除后该 publication 永远停在 RETIRING / put_unresolved,6,296,450 字节一直被占用。代码里没有任何路径会把尝试移出 UNKNOWN。
4. 存储故障
失败的 DELETE 只发一次,整页重试: 第 5 个 key 的 DELETE 遇到 500 或连接重置时,该 key 只发出了一次请求,SDK 没有重放。整页延后 60 s,由第 2 代重新执行;已删掉的 key 返回 204。
版本控制、崩溃与多页:
- 版本控制
Suspended时,回收器没有发出任何 DELETE。 - DELETE 被挂起时
kill -9Spring:重启后的进程等旧 claim 到期(59.6 s),再以第 2 代完成。 - 113 个 key 的 publication 分两页回收,两页之间相隔一个 tick。
- 两个开启 GC 的 Spring 实例共享一个库:两个 owner 都参与了回收,但每个 publication 只有一个 owner、都是第 1 代,93 个 key 各删一次。
每种情况下,配额都恰好只释放一次。
5. F2:下载成本与 120 s 租约
在 MySQL 前放了一个 TCP 中继,给每个客户端→服务器包加 5 ms 延迟。main 上 40 MiB 下载在租约到期时被截断在 24.2 MiB;PR 上 55.5 s 完成。但 PR 上 100 MiB 仍在 85.8 MiB 处被截断,吞吐约 0.73 MiB/s。
剩余开销约为每 64 KiB 响应块 11 条 SQL:
ManagedArtifactService的 64 KiB 写循环每次都检查租约和访问权限;VerifiedStream.read每次最多读 64 KiB,并在读前、读后各跑一遍完整守卫。
候选补丁把这两个常量提高到 1 MiB,与 PR 在存储路径已采用的粒度一致。结果:
- 成本与速度: 每 MiB 38.8 条;+5 ms 下 40 MiB 用时 10.3 s,128 MiB 用时 31.3 s。
- 撤权语义不变: 下载途中撤权或 DELETE 仍会中止响应,之后最多再送出 0.88 MiB。
这两处是 main 的代码,不在 PR diff 内,可在本 PR 一并修改,也可作为后续 PR;否则建议改写正文中关于 100 MiB 的表述。
6. OSS GET 重试与变异测试
GET 重试。 我通过真实 SDK,在公开下载的第 3 个对象上注入故障。
- 只有 PR 会重试的: 错误码不在白名单里的 500,以及非 XML 响应体的 502。main 上这两种情况发一次请求就中止响应;PR 重试一次后完整下完,摘要正确。
- 两臂一致的: 500
InternalError、503、连接重置、连续六次 500(4 次请求后中止)、403(不重试)。 - PUT: 一次 PUT 500 产生两次独立尝试,每次只发出一个请求,SDK 没有重放。
变异测试: 对新代码构造了 28 个变异体,每个都跑完整单测(514 个)。杀死 19 个。存活的 9 个:
- 不是真缺口:
- M04、M06、M21 是在"一个进程持有一个 claim"时不可能触发的围栏(
runOnce是 synchronized 的,owner 按进程区分)。上文的双实例运行覆盖了跨进程的情形。 - M07(页级探针)是冗余的:每 key 探针做同样的检查且被测试钉住(M08);两者都去掉(M29)会被杀死。
- M04、M06、M21 是在"一个进程持有一个 claim"时不可能触发的围栏(
- 真实缺口:
- M24–M27: 回收器测试只覆盖了
grace_period和历史写入证据两种阻塞。把被隔离、恢复保护、put_unresolved或reader_active的 publication 回收掉,单测依然全绿;只有第 3 节的真实栈运行证明了它们会被保留。 - M17: 把受守卫的包装层改回原始
read(),单测同样全绿。它对公开下载没有影响:真实栈上带该变异体测得每 MiB 205.6 条,与 PR 相同。但这个包装层保护着 Session Store 资源读取和校验读取,而这两条路径没有任何有界查询测试。
- M24–M27: 回收器测试只覆盖了
建议(均不阻塞)
- 与 feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210 协调 V34。
- 采纳两行读取粒度补丁,或改写 100 MiB 的表述。
- 打开
gc-enabled之前,增加处理UNKNOWNPUT 尝试的手段,例如对确切 key 做受保护的 HEAD 并比对摘要,或提供运维命令。否则每一次瞬时 PUT 错误都会永久泄漏该 publication 的存储和租户配额,也会加重已延期的"永久阻塞每 60 s 轮询"问题。 - 重复的版本控制探针: 每删除一个 key 要发两个 OSS 请求(113 次 DELETE 对应 115 次版本探针),因为
deleteIfPresent在页级探针之后又探了一次。目前只有每 key 的探针被测试钉住(M08),页级探针是冗余的(M07)。 - 为回收器补一个覆盖全部阻塞类型(quarantined、recovery_protected、
reader_active、put_unresolved等)的参数化测试,使 M24–M27 能被杀死。
证据(图、装置脚本、日志、候选补丁):assets-pr13225。
|
@qwen-code /triage |
Base PR #13225 (O4-2 collector) squash-merged into main as fa795e0. Resolve the stack onto current main: - Drop branch-side V33__managed_tool_output_collection.sql (byte-identical to main's V34 after the pre-merge renumber; main's V33 is managed_agent_definitions from #13142). Sequence stays V30-V34, unique. - workspace-session-reliable-close docs (both languages) and WorkspaceRecoveryStoreTest: take main's version; O4-3's own commits never touched them, branch side was old-stack residue. - managed-tool-output-retention docs (both languages): main's refreshed version plus O4-3 commit 9d9d016's first-paragraph delta (32-candidate tick bound, grace semantics, publication-scoped cleanup). Co-authored-by: Qwen-Coder <[email protected]>
This PR's base moved to main after #13225 merged; no PR will target the merged codex/managed-tool-result-o4-2 branch again. Co-authored-by: Qwen-Coder <[email protected]>
…nLM#13090) * feat(managed-agent): Collect retired tool output safely Co-authored-by: Qwen-Coder <[email protected]> * chore(managed-agent): Sync retention formatting and verify collection Co-authored-by: Qwen-Coder <[email protected]> * codex: address PR review feedback (QwenLM#13087) * codex: address PR review feedback (QwenLM#13087) * test(managed-agent): Add tool output collection deployment gates Co-authored-by: Qwen-Coder <[email protected]> * codex: address PR review feedback (QwenLM#13090) * docs(managed-agent): Preserve deployment gate identifiers * fix(managed-agent): Advance eligible tool output collection * codex: address PR review feedback (QwenLM#13090) * chore(ci): Drop merged O4-2 base branch from SDK Java trigger This PR's base moved to main after QwenLM#13225 merged; no PR will target the merged codex/managed-tool-result-o4-2 branch again. Co-authored-by: Qwen-Coder <[email protected]> * docs(managed-agent): Refresh operations runbook for merged O4-2 Audit found the runbook still pointed at the pre-merge baseline SHA and the old collection V33, contradicting the lifecycle design doc in the same diff. State the merged stack without a volatile SHA, use the current V30/V31/V32/V33(definitions)/V34(collection) sequence in both languages, and say the SDK Java workflow targets main now that the stacked base branch is merged. Co-authored-by: Qwen-Coder <[email protected]> --------- Co-authored-by: Qwen-Coder <[email protected]> Co-authored-by: yiliang114 <[email protected]> Co-authored-by: Shaojin Wen <[email protected]>






What this PR does
O4-2 collects eligible, permanently retired tool-output publications using durable claims, exact catalog keys and fenced SQL confirmation. Each tick checks at most 32 due candidates and deletes at most one 100-key page. Inline cleanup and quota release occur only after all objects are confirmed. Physical GC remains disabled.
This revision restores the slice on current main, carries forward the bounded progress and expired-claim repairs, reduces guarded download SQL work, and restores explicit transient GET recovery without implicit PUT or DELETE replay. Unsupported storage adapters still fail closed.
Why it's needed
Retirement retains storage and quota until collection can prove complete write closure and the absence of live readers and recovery protection. The previous branch also collided with current migrations and could not safely land. Small HTTPS reads amplified repeated SQL guards enough to truncate a 40 MiB response at the fixed two-minute budget; disabling the SDK retry hook also prevented the existing explicit GET retry policy from seeing transient responses.
Reviewer Test Plan
How to verify
Confirm that close, archive and runtime reclamation retain output, while completed Session deletion permanently retires it. Collection must retain incomplete, legacy, quarantined and recovery-protected publications, preserve sibling resources and quota on every failed or stale claim, and release quota once after the last confirmed page. Inline-only pages must work without OSS; physical keys must retain their unversioned-bucket checks.
Verify full and range downloads return exact lengths and digests. Under fragmented HTTPS reads and 5 ms SQL latency, the tested 40 MiB and 100 MiB responses complete within the original nonrenewable 120-second lease. Revoking access, retiring the Session or resuming beyond that lease must return no further bytes. Transient GET retries check their original guards before each wire request; permanent responses, interrupted waits and returned-body failures retain their original failure behavior, and a failed PUT sends once.
Final clean commit
bea89fcdefc840314738965634a76f827adbcbe5passed build, typecheck, bundle, 33 unique migrations, complete Broker/Server ordinary MySQL profiles and the source-derived report check, with zero Checkstyle violations. Broker unit reports contain 593 cases (592 passed, one optional real-worker bundle case skipped); Broker MySQL integration passed 7/7. Server unit and integration cases passed 487/487 and 35/35; the Recovery fixture now uses a genuinely claimed lease and executes the eight previously failing branches. Independent verification includes 58 SDK fault fixtures, all 20 Recovery cases and NULL/expired lease rejection controls, bound byte-for-byte to this commit. Two clean independent review passes found no Critical. Exact counts and limitations are in the current top-level report.Evidence (Before & After)
A controlled, source-matching MySQL/actual OSS SDK HTTPS fixture measured 8 MiB download SQL cost at 1122.5 statements/MiB before and 197.75 after. At 5 ms SQL latency, the original 40 MiB response ended at approximately 120 seconds after returning only 16 MiB; the repaired 40 MiB and 100 MiB responses completed in approximately 43 and 108 seconds with exact bytes and digests. These are already-admitted artifact fixtures, not full Hosted producer execution or credentialed cloud OSS measurements. There is no TUI change.
Tested on
Environment
macOS arm64, Node 22.14, Java 21.0.12.1, Maven 3.9.16 and an owned MySQL 8.4.11 instance. The HTTPS fixture uses OSS SDK 3.18.4 with real TLS fragmentation and normal certificate verification. Only isolated test resources were created and cleaned.
Risk & Scope
Design: English · 简体中文.
Linked Issues
Supersedes closed #13087 because GitHub rejected reopening it. Related to #12380. O4-1 #13084 is merged; O4-3 #13090 remains stacked above this PR. This does not close the umbrella issue or change #13019 recovery policy.
中文说明
本 PR 的改动
O4-2 通过持久 claim、catalog 精确 key 和带屏障的 SQL 确认,回收符合条件且已永久退役的工具输出 publication。每个 tick 最多检查 32 个到期候选,并最多删除一页 100 个 key。全部对象确认后才清理 inline payload 并释放配额。物理 GC 保持关闭。
本次将分片恢复到当前 main,带入有界候选推进和过期 claim 修复,降低受保护下载的 SQL 成本,并恢复显式瞬态 GET 重试,同时不引入隐式 PUT 或 DELETE 重放。不支持的存储适配器仍拒绝操作。
为什么需要
退役后仍需保留存储与配额,直到回收能证明写入已闭合,且没有活跃读者或恢复保护。原分支还与当前迁移冲突,无法安全落地。HTTPS 小块读取放大了重复 SQL 守卫,导致 40 MiB 响应在固定两分钟预算下被截断;关闭 SDK retry hook 也让已有显式 GET 重试策略无法看到瞬态响应。
评审测试计划
如何验证
确认 close、archive 和 runtime 回收继续保留输出,而完成 Session 删除会永久退役。回收必须保留未完成、历史证据缺失、隔离及恢复保护的 publication,在失败或过期 claim 下保持同会话其他资源和配额,并仅在最后一页确认后释放一次配额。纯 inline 页不依赖 OSS,物理 key 继续执行未启用版本控制的桶检查。
确认完整与 range 下载的长度和 digest 精确。在碎片化 HTTPS 读取、每条 SQL 5 ms 的条件下,被测 40 MiB 和 100 MiB 响应在原不可续期 120 秒租约内完成。撤销权限、退役 Session,或暂停至租约过期后恢复,都不得继续返回字节。瞬态 GET 重试每次发送前检查原守卫;永久失败、等待中断和已返回 body 的读取失败保持原行为,失败 PUT 只发送一次。
最终干净提交
bea89fcdefc840314738965634a76f827adbcbe5通过 build、typecheck、bundle、33 个迁移无重号、完整 Broker/Server 普通 MySQL profile 和源码导出的报告检查,Checkstyle 为零。Broker 单元报告含 593 项(592 项通过,1 个可选真实 worker bundle 用例跳过),Broker MySQL 集成 7/7 通过。Server 单元 487/487、集成 35/35 通过,修正 Recovery fixture 使用真实领取的租约,原 8 个失败分支现已执行。独立验证包含 58 个 SDK 故障 fixture、完整 Recovery 20 项及 NULL/过期租约拒绝控制,源码逐字节绑定最终提交。独立审查两轮无 Critical。准确计数及验证限制见本次顶层报告。前后证据
受控且匹配源码的 MySQL/实际 OSS SDK HTTPS fixture 测得 8 MiB 下载 SQL 成本由每 MiB 1122.5 条降至 197.75 条。每条 SQL 5 ms 时,原 40 MiB 响应约 120 秒后结束,仅返回 16 MiB;修后 40 MiB、100 MiB 分别约 43、108 秒完成,字节及 digest 正确。这里使用已接纳的 Artifact fixture,不代表完整 Hosted 生产链路或真实云 OSS 测量。没有 TUI 变化。
测试系统
环境
macOS arm64、Node 22.14、Java 21.0.12.1、Maven 3.9.16,使用任务拥有的 MySQL 8.4.11。HTTPS fixture 使用 OSS SDK 3.18.4,实际 TLS 分块,并正常校验证书。仅创建和清理隔离测试资源。
风险与范围
设计:English · 简体中文。
关联 Issue
替代 GitHub 拒绝重开的已关闭 #13087。关联 #12380。O4-1 #13084 已合并;O4-3 #13090 继续堆叠于本 PR。不关闭总 issue,不改变 #13019 恢复策略。