Skip to content

feat(managed-agent): Restore safe retired tool output collection - #13225

Merged
wenshao merged 6 commits into
mainfrom
codex/managed-tool-result-o4-2
Oct 3, 2026
Merged

wenshao merged 6 commits into
mainfrom
codex/managed-tool-result-o4-2

Conversation

@doudouOUC

@doudouOUC doudouOUC commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does

O4-2 collects eligible, permanently retired tool-output publications using durable claims, exact catalog keys and fenced SQL confirmation. Each tick checks at most 32 due candidates and deletes at most one 100-key page. Inline cleanup and quota release occur only after all objects are confirmed. Physical GC remains disabled.

This revision restores the slice on current main, carries forward the bounded progress and expired-claim repairs, reduces guarded download SQL work, and restores explicit transient GET recovery without implicit PUT or DELETE replay. Unsupported storage adapters still fail closed.

Why it's needed

Retirement retains storage and quota until collection can prove complete write closure and the absence of live readers and recovery protection. The previous branch also collided with current migrations and could not safely land. Small HTTPS reads amplified repeated SQL guards enough to truncate a 40 MiB response at the fixed two-minute budget; disabling the SDK retry hook also prevented the existing explicit GET retry policy from seeing transient responses.

Reviewer Test Plan

How to verify

Confirm that close, archive and runtime reclamation retain output, while completed Session deletion permanently retires it. Collection must retain incomplete, legacy, quarantined and recovery-protected publications, preserve sibling resources and quota on every failed or stale claim, and release quota once after the last confirmed page. Inline-only pages must work without OSS; physical keys must retain their unversioned-bucket checks.

Verify full and range downloads return exact lengths and digests. Under fragmented HTTPS reads and 5 ms SQL latency, the tested 40 MiB and 100 MiB responses complete within the original nonrenewable 120-second lease. Revoking access, retiring the Session or resuming beyond that lease must return no further bytes. Transient GET retries check their original guards before each wire request; permanent responses, interrupted waits and returned-body failures retain their original failure behavior, and a failed PUT sends once.

Final clean commit bea89fcdefc840314738965634a76f827adbcbe5 passed build, typecheck, bundle, 33 unique migrations, complete Broker/Server ordinary MySQL profiles and the source-derived report check, with zero Checkstyle violations. Broker unit reports contain 593 cases (592 passed, one optional real-worker bundle case skipped); Broker MySQL integration passed 7/7. Server unit and integration cases passed 487/487 and 35/35; the Recovery fixture now uses a genuinely claimed lease and executes the eight previously failing branches. Independent verification includes 58 SDK fault fixtures, all 20 Recovery cases and NULL/expired lease rejection controls, bound byte-for-byte to this commit. Two clean independent review passes found no Critical. Exact counts and limitations are in the current top-level report.

Evidence (Before & After)

A controlled, source-matching MySQL/actual OSS SDK HTTPS fixture measured 8 MiB download SQL cost at 1122.5 statements/MiB before and 197.75 after. At 5 ms SQL latency, the original 40 MiB response ended at approximately 120 seconds after returning only 16 MiB; the repaired 40 MiB and 100 MiB responses completed in approximately 43 and 108 seconds with exact bytes and digests. These are already-admitted artifact fixtures, not full Hosted producer execution or credentialed cloud OSS measurements. There is no TUI change.

Tested on

OS Status
🍏 macOS ✅ Local build, types, Java/MySQL and controlled HTTPS fixtures
🪟 Windows ⚠️ Not tested
🐧 Linux ⚠️ Current PR CI pending; local results do not establish Linux Hosted acceptance

Environment

macOS arm64, Node 22.14, Java 21.0.12.1, Maven 3.9.16 and an owned MySQL 8.4.11 instance. The HTTPS fixture uses OSS SDK 3.18.4 with real TLS fragmentation and normal certificate verification. Only isolated test resources were created and cleaned.

Risk & Scope

  • The fixed 120-second read lease remains; these measurements do not guarantee arbitrary size or latency. Public output still checks guards before and after bounded reads. Remaining collector capacity and diagnostic Suggestions are deferred under proposal(serve): Define Managed Agent dual-path architecture and staged delivery #12380 after multiple review rounds, including permanent-blocker polling saturation; GC enablement requires further capacity evidence.
  • Real cloud OSS deletion-denied identity, complete Linux Hosted foreground Shell acceptance, public Workspace DELETE admission, and production enablement are outside this repair. Closed Session lifecycle does not itself retire output. Remaining mutation/coverage suggestions are recorded explicitly rather than counted as fixed. fix(managed-agent): Preserve writer and publication epoch deadlines #13192 is an independent, still-open writer-epoch change, not an inherited merged fix.
  • Current main retains recovery V31; close is renumbered to V32 and collection to V33 without changing SQL bytes. Provisional close V28/V30/V31 or collection V28 database histories require explicit reconciliation or recreation. No automatic Flyway repair is performed. Upgrade every Java publication writer before deployment collection; retain the 24-hour grace policy.

Design: English · 简体中文.

Linked Issues

Supersedes closed #13087 because GitHub rejected reopening it. Related to #12380. O4-1 #13084 is merged; O4-3 #13090 remains stacked above this PR. This does not close the umbrella issue or change #13019 recovery policy.

中文说明

本 PR 的改动

O4-2 通过持久 claim、catalog 精确 key 和带屏障的 SQL 确认,回收符合条件且已永久退役的工具输出 publication。每个 tick 最多检查 32 个到期候选,并最多删除一页 100 个 key。全部对象确认后才清理 inline payload 并释放配额。物理 GC 保持关闭。

本次将分片恢复到当前 main,带入有界候选推进和过期 claim 修复,降低受保护下载的 SQL 成本,并恢复显式瞬态 GET 重试,同时不引入隐式 PUT 或 DELETE 重放。不支持的存储适配器仍拒绝操作。

为什么需要

退役后仍需保留存储与配额,直到回收能证明写入已闭合,且没有活跃读者或恢复保护。原分支还与当前迁移冲突,无法安全落地。HTTPS 小块读取放大了重复 SQL 守卫,导致 40 MiB 响应在固定两分钟预算下被截断;关闭 SDK retry hook 也让已有显式 GET 重试策略无法看到瞬态响应。

评审测试计划

如何验证

确认 close、archive 和 runtime 回收继续保留输出,而完成 Session 删除会永久退役。回收必须保留未完成、历史证据缺失、隔离及恢复保护的 publication,在失败或过期 claim 下保持同会话其他资源和配额,并仅在最后一页确认后释放一次配额。纯 inline 页不依赖 OSS,物理 key 继续执行未启用版本控制的桶检查。

确认完整与 range 下载的长度和 digest 精确。在碎片化 HTTPS 读取、每条 SQL 5 ms 的条件下,被测 40 MiB 和 100 MiB 响应在原不可续期 120 秒租约内完成。撤销权限、退役 Session,或暂停至租约过期后恢复,都不得继续返回字节。瞬态 GET 重试每次发送前检查原守卫;永久失败、等待中断和已返回 body 的读取失败保持原行为,失败 PUT 只发送一次。

最终干净提交 bea89fcdefc840314738965634a76f827adbcbe5 通过 build、typecheck、bundle、33 个迁移无重号、完整 Broker/Server 普通 MySQL profile 和源码导出的报告检查,Checkstyle 为零。Broker 单元报告含 593 项(592 项通过,1 个可选真实 worker bundle 用例跳过),Broker MySQL 集成 7/7 通过。Server 单元 487/487、集成 35/35 通过,修正 Recovery fixture 使用真实领取的租约,原 8 个失败分支现已执行。独立验证包含 58 个 SDK 故障 fixture、完整 Recovery 20 项及 NULL/过期租约拒绝控制,源码逐字节绑定最终提交。独立审查两轮无 Critical。准确计数及验证限制见本次顶层报告。

前后证据

受控且匹配源码的 MySQL/实际 OSS SDK HTTPS fixture 测得 8 MiB 下载 SQL 成本由每 MiB 1122.5 条降至 197.75 条。每条 SQL 5 ms 时,原 40 MiB 响应约 120 秒后结束,仅返回 16 MiB;修后 40 MiB、100 MiB 分别约 43、108 秒完成,字节及 digest 正确。这里使用已接纳的 Artifact fixture,不代表完整 Hosted 生产链路或真实云 OSS 测量。没有 TUI 变化。

测试系统

OS 状态
🍏 macOS ✅ 本地构建、类型检查、Java/MySQL 和受控 HTTPS fixture
🪟 Windows ⚠️ 未测试
🐧 Linux ⚠️ 当前 PR CI 待执行;本地结果不代表 Linux Hosted 验收

环境

macOS arm64、Node 22.14、Java 21.0.12.1、Maven 3.9.16,使用任务拥有的 MySQL 8.4.11。HTTPS fixture 使用 OSS SDK 3.18.4,实际 TLS 分块,并正常校验证书。仅创建和清理隔离测试资源。

风险与范围

  • 保留固定 120 秒读取租约,这些测量不保证任意大小或延迟。公开输出仍在有界读取前后检查守卫。多轮评审后,剩余 collector 容量和诊断建议延期到 proposal(serve): Define Managed Agent dual-path architecture and staged delivery #12380,包括永久阻塞轮询饱和;启用 GC 前仍需容量证据。
  • 真实云 OSS 删除拒绝身份、完整 Linux Hosted 前台 Shell、公开 Workspace DELETE 接纳及生产启用不属于本次修复。关闭会话本身不会退役输出。剩余 mutation/覆盖建议明确记录,不算已修复。fix(managed-agent): Preserve writer and publication epoch deadlines #13192 是独立且仍开放的 writer epoch 改动,不是已合入并继承的修复。
  • 当前 main 保留 recovery V31;close 改为 V32、collection 改为 V33,SQL 字节不变。临时 close V28/V30/V31 或 collection V28 数据库历史需显式核对或重建,不自动 Flyway repair。部署回收前必须升级全部 Java publication writer,保持 24 小时宽限策略。

设计:English · 简体中文。

关联 Issue

替代 GitHub 拒绝重开的已关闭 #13087。关联 #12380。O4-1 #13084 已合并;O4-3 #13090 继续堆叠于本 PR。不关闭总 issue,不改变 #13019 恢复策略。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

[codex] Restored O4-2 at bea89fcdefc840314738965634a76f827adbcbe5 in replacement #13225, based on main 9478f28731f45ff6df66062e7e05d29e92906f19. GitHub rejected reopening #13087 and changing its closed base, so #13225 supersedes it. O4-3 #13090 remains stacked on the same repaired branch. No merge or GC enablement was performed.

Disposition Action
Fixed SQL+Java migration census is 33 unique versions: recovery V31 retained, close V32, collection V33; rename only, SQL bytes unchanged. Provisional database histories require explicit reconciliation/recreation, never automatic repair.
Fixed / carried forward Bounded due-candidate progress and original grace deadline scheduling, inline-only pages without OSS, original owner/generation defer after expired renewal, and target-publication cleanup regression coverage. Current main reader/writer/deletion protections are preserved.
Fixed F2 download SQL amplification: fill bounded verification reads instead of guarding every raw TLS fragment; remove a duplicate SQL liveness probe while retaining fresh retirement predicates and original before/after guards. Actual SDK/TLS 8 MiB cost reduced from 1122.5 to 197.75 SQL/MiB. At minimum 5 ms/SQL, baseline 40 MiB truncated at 16 MiB after 120 seconds; fixed 40 MiB/100 MiB completed with exact digests in 42.6/108.1 seconds. Range, real 126-second process pause, permission revocation and retirement controls retain no-more-bytes behavior. One original nonrenewable read lease, no auth cache.
Fixed GET SDK retry-hook regression: the strategy never allows SDK replay; it classifies transient GET failures for the adapter's existing bounded guarded loop. Permanent diagnostics retain original OSSException, every retried wire request is guarded, PUT/DELETE remain single SDK attempts, returned body failures do not reopen. 58 real-SDK fault fixtures, including 9 recovery assertions and protected failure controls, passed.
Fixed Recovery fixture now claims a real DB-clock lease before completing retirement; all original 20 cases pass, and matching-owner/generation NULL or expired leases remain rejected. The initial final gate's eight fixture failures were reproduced independently before this test-only fix; no production fence was relaxed.
Deferred after multiple review rounds R2-1/2/4/5/6 and R1-11: shared observer scheduling, additional throughput/index/lock capacity and physical probe amortization remain deployment follow-ups under #12380. Candidate limits do not prove bounded fleet capacity.
Deferred R1-3/R2-9/R3-2: publication warning identity, success logs and DELETING observer visibility. The existing #12380 disposition and separate state/blocker SQL queries remain the ledger.
Partial / deferred R1-8/R1-9/R2-10/11/12/13: sibling cleanup coverage strengthened; remaining quota-field/adapter/live-negative/observer mutation coverage, lease-reap and test-helper refactors remain explicit follow-ups, not claimed fixed.
Deferred / deployment restriction Historical F1 permanent-blocker polling saturation and C3 failure/error mutation pinning. Shortened grace does not recalculate already scheduled deadlines. These concerns are recorded under #12380; GC remains off. M21/M25/R4/R5 remaining mutation pins from the handoff are not silently counted as completed.

Final committed-head gate: build, typecheck, bundle, migrations, source-matching Java dependency installation, complete ordinary MySQL profiles and source-derived report checks passed. Broker unit reports: 593 total, 592 passed and one optional real-worker bundle test skipped; Broker MySQL: 7/7 passed. Server: 487/487 unit and 35/35 MySQL integration passed. Checkstyle 0. Two clean independent code-review passes and self-audit completed; fixture, GET and controlled performance execution snapshots are bound to the final Git objects.

These local macOS/MySQL and controlled HTTPS results do not establish Linux identity guards, full Hosted producer/Broker/Runtime acceptance or current-revision credentialed cloud OSS. The historical maintainer OSS 36/37 run still lacks the delete-denied identity. Public Workspace DELETE remains unavailable, and #13192 remains an independent open PR. Those are deployment/readiness boundaries, not green checks inferred from admin-only CI. Current #13225 product CI is separately pending; physical GC is disabled and grace remains 24 hours.

中文:O4-2 已以 bea89fcd 在当前 main 上恢复为替代草稿 #13225,接替 GitHub 拒绝重开的 #13087;O4-3 #13090 继续使用同一已修复分支。已修复迁移撞号、带入候选推进/inline/过期 claim 修复、下载 SQL 放大和 GET 重试 hook 回退,并修正真实领取租约的 Recovery fixture。完整本地最终门禁通过:Broker 592 个单元通过、1 个可选 bundle 用例跳过,7 个 MySQL 集成通过;Server 单元 487/487、MySQL 集成 35/35,Checkstyle 为零。58 个 SDK 故障 fixture、20 个 Recovery 用例及受控 HTTPS 成本/权限/退役/真实租约暂停证据均绑定最终源码。

上表中的容量、诊断和 mutation 建议逐项明确延期到 #12380,部分采纳不算全部修复。永久阻塞轮询饱和、已排期宽限期缩短不会即时生效、C3 和 handoff 剩余 mutation 缺口继续保留。Linux/完整 Hosted、真实 OSS 删除拒绝身份、公开 Workspace DELETE 及独立开放的 #13192 不被本地绿灯替代。没有自动合并;物理 GC 关闭,部署宽限保持 24 小时。

Conflicts were limited to close-slice files that main has since
reworked: docs/design/workspace-session-reliable-close.{md,zh-CN.md}
and WorkspaceRecoveryStoreTest.java (operation schema fixture moved to
RUNNING/LEASED with lease_until). Resolution takes main's versions; this
PR now carries only the collection slice.

Migration numbering on the merged tree: V31 recovery, V32 close (SQL
bytes identical on both sides), V33 collection; Java migrations remain
V15/V29 only, so no Flyway version collision.

Verified on the merged tree: sdk-java managed-agent-server mvn test,
487 tests 0 failures, including the Flyway-backed
WorkspaceRecoveryStoreTest and the new collection tests.
@wenshao

wenshao commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Merged current origin/main (c15dd26) into this branch to clear the conflict, as merge commit 0c64229.

The conflict was limited to close-slice files that main has since reworked: the two workspace-session-reliable-close design docs and WorkspaceRecoveryStoreTest.java (operation fixture moved to the RUNNING/LEASED + lease_until schema). Resolution takes main's versions, so this PR now carries only the collection slice.

Numbering re-verified on the merged tree: V31 recovery, V32 session close (SQL bytes identical to main's), V33 tool output collection; Java migrations remain V15/V29 only — no Flyway version collision.

Verified on the merged tree before pushing: sdk-java managed-agent-server full mvn test — 487 tests, 0 failures, including the Flyway-backed WorkspaceRecoveryStoreTest and the new collection tests (matches your recorded 487/487 gate).

Brings AgentDefinition D8a (#13142) and session archive/delete (#13194).
The textual merge was clean; the one semantic collision is the migration
number: main took V33 for managed_agent_definitions, so this PR's tool
output collection migration is renumbered V33 -> V34 with SQL bytes
unchanged. The retention design docs (en + zh-CN) census now reads V31
recovery / V32 close / V33 agent definitions / V34 collection, and
provisional V28/V33 collection histories require explicit
reconciliation, matching the existing no-automatic-repair policy.

Verified on the merged tree with a clean build: sdk-java
managed-agent-server mvn test, 514 tests 0 failures, with the
Flyway-backed suites exercising V31 through V34.
@wenshao

wenshao commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Second merge of origin/main (1abccdb) as 88d4ca9. Main took V33 for managed_agent_definitions (#13142), so the collection migration is renumbered V33 -> V34 with SQL bytes unchanged; the retention docs census now reads V31 recovery / V32 close / V33 definitions / V34 collection, with provisional V28/V33 collection histories requiring explicit reconciliation per the existing policy. Verified on the merged tree with a clean build: managed-agent-server mvn test, 514 tests 0 failures (Flyway-backed suites exercise V31 through V34). The earlier note stands: the Hosted MySQL gate failure signature predates this merge and appears on unrelated branches (e.g. codex/managed-writer-epoch run 37029292498, same test, same signature).

@wenshao

wenshao commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Local real-environment verification of #13225 at 88d4ca9466

Verdict: no blocker for landing with GC disabled (the default). The Flyway V33 collision that existed at 0c64229ae6 is fixed by 88d4ca9466. Its packages/sdk-java tree is byte-identical to the tree I verified. The collector did what the PR describes on the real public close → DELETE path on Linux, against both the OSS double and a real Aliyun OSS bucket. Check three things before merging or enabling GC:

  1. Migration order with feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210. The open feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210 also adds V34__managed_session_creator.sql. Whichever PR lands second has to be renumbered.
  2. F2 claim. The store-side SQL amplification is fixed (711 → 205 statements per MiB). The body's "100 MiB completes in ~108 s at 5 ms SQL latency" did not reproduce here. 100 MiB was cut at 85.8 MiB when the 120 s read lease ran out. A two-line follow-up (below) brings the cost down to 39 statements per MiB, and 128 MiB then completes in 31 s.
  3. Before enabling GC: a single transient PUT failure makes a publication permanently uncollectable. This follows the design doc ("unknown … permanently blocks automatic collection until externally resolved"). No resolution path exists yet, so the publication's storage and tenant quota are never released.

Setup

Arm Tree Purpose
PR b95e64cdf0 = 0c64229ae6 merged into main 623cfc67 + V33→V34 rename. git diff b95e64cdf0 88d4ca9466 -- packages/sdk-java is empty everything below
base main 623cfc67 A/B for collection, F2 and GET retry
candidate PR + 2-line patch (candidate-1mib-read-granularity.patch) F2 follow-up only
  • Gates and boot matrix: macOS arm64, JDK 21, MySQL 8.4.7.
  • Every retirement scenario ran on Linux. On macOS the Workspace close path is unavailable: the durable worker identity reads /etc/machine-id and /proc, so session_close=false. The Linux stack was:
    • a colima VM (aarch64, 4 CPU / 8 GB) running a JDK 21 container and a MySQL 8.4.11 container;
    • the Spring fat jar, the packaged Hosted Harness and a durable local-process Broker worker;
    • a fake model, and an OSS double over TLS driven by the real aliyun-sdk-oss 3.18.4.
  • Real Aliyun OSS: a temporary private, unversioned bucket in cn-hangzhou, deleted afterwards.
  • Seams, stated plainly:
    • A tap rewrites the Harness tool profile hosted-workspace-files/1 → hosted-workspace-shell/1, because public Shell admission is not on main yet.
    • The legacy-write-evidence, recovery-status and stale-read-lease cases are injected by SQL.
    • Everything else goes through the public API: create, Turn, download, POST /close, DELETE.
  • The Harness/worker bundle was built from main 623cfc67. The 1abccdb26a merge only adds a packages/core change (fix(core): stop misdiagnosing malformed tool-call args as max_tokens truncation #12982) that is unrelated to this path.

1. Landing and gates

landing

  • At 0c64229ae6, the trial merge into current main fails the repo's own check-flyway-migrations.js: two migrations claim version 33. Spring also refuses to boot on it.
  • The renumbered tree boots on an empty database and on a database that main has already migrated (V33 definitions → V34 collection).
  • A database migrated by the PR head alone cannot be opened by main (checksum mismatch at V33). That is the provisional-history case the PR body already calls out.
  • Unit: 514/514, Checkstyle 0.
  • -Pmysql-integration: 51/51, including WorkspaceSessionRetentionMySqlIT 16/16; the failsafe report checker exits 0.

2. Public lifecycle on Linux, and against real OSS

lifecycle

  • Public calls: POST /close 202 → completed (554 ms); DELETE 202 → completed (292 ms).
  • Collection: the row waited for grace_period until exactly retired_at + grace (10 s). It was claimed at +10.5 s and COLLECTED at +10.9 s.
  • OSS deletes: 15 DELETEs, one per catalog key, only the deleted Session's keys; no prefix scans.
  • Cleanup and quota: inline payloads cleared; held 13,641,895 → 0 and released_held_bytes = 13,641,895.
  • Sibling Session: untouched and still readable with the right sha256.
  • On main: the same lifecycle leaves the output RETIRING with all objects and quota held.
  • Real OSS: 8/8 objects of the deleted Session were removed (bucket 16 → 8 objects); the sibling's 8 were kept.
  • Real bucket with versioning enabled: the collector failed closed for two attempts with Tool publication bucket cannot enforce immutable objects and deleted nothing (4/4 objects kept).

3. Blockers and siblings

blockers

Quarantine: I flipped one bit of a stored object. A public download detected it (503 artifact_unavailable) and production code quarantined that publication.

After DELETE:

  • The quarantined publication kept 7/7 objects and its quota.
  • The other publication in the same Session was collected.

The legacy_write_evidence_missing and recovery_protected rows were kept. A stale read lease held the row as reader_active until the next 60 s poll after it expired (collected at +71.2 s). Tenant held bytes dropped by exactly the two collected publications.

Unknown PUT (finding 3). I injected one PUT 500 on the wire. The next attempt returned, and all 11 objects are VERIFIED. The ledger still has UNKNOWN=1, so after DELETE the publication stays RETIRING / put_unresolved with 6,296,450 bytes held. Nothing in the code moves an attempt out of UNKNOWN.

4. Storage faults

faults

One request per failed DELETE, whole-page retry: a DELETE 500 or connection reset on the 5th key produced exactly one request for that key; the SDK did not replay it. The whole page was deferred by 60 s, and generation 2 re-issued the page; the keys already gone answered 204.

Versioning, crash and multi-page:

  • With versioning Suspended, the collector issued 0 DELETEs.
  • kill -9 of Spring while a DELETE was held: the restarted process waited out the old claim (59.6 s), then finished as generation 2.
  • A 113-key publication was collected in two pages, one tick apart.
  • Two live Spring instances with GC on, sharing one database: both collected (two owners seen), but every publication had one owner and generation 1, and each of the 93 keys was deleted once.

In every case the quota was released exactly once.

5. F2: download cost and the 120 s lease

f2

A TCP relay in front of MySQL adds 5 ms to every client→server packet. On main, a 40 MiB download is cut at 24.2 MiB when the lease ends; on the PR it completes in 55.5 s. A 100 MiB download still gets cut at 85.8 MiB on the PR, about 0.73 MiB/s.

What remains is about 11 statements per 64 KiB response chunk:

  • ManagedArtifactService's 64 KiB write loop checks the lease and access each time;
  • VerifiedStream.read caps every read at 64 KiB and runs the full guard before and after.

The candidate raises those two constants to 1 MiB, the granularity the PR already uses in the store path. Results:

  • Cost and speed: 38.8 statements/MiB; 40 MiB in 10.3 s and 128 MiB in 31.3 s at +5 ms.
  • Revocation unchanged: revoke or DELETE mid-download still aborts the response, after at most 0.88 MiB.

It touches main's code rather than the PR's diff. It can land here or as a follow-up; otherwise the body's 100 MiB claim should be softened.

6. OSS GET retry and mutation testing

retry-mutation

GET retry. I injected faults on the third object of a public download through the real SDK.

  • Retried only on the PR: a 500 with an unrecognised error code, and a 502 with a non-XML body. On main these abort the response after one request; the PR retries once and the download completes with the right digest.
  • Same on both arms: 500 InternalError, 503, a connection reset, six 500s in a row (4 requests, then abort) and 403 (no retry).
  • PUT: a PUT 500 results in two separate attempts, each with one wire request; the SDK replays nothing.

Mutation testing: 28 mutants of the new code, each run against the full unit suite (514 tests). 19 were killed. Of the 9 survivors:

  • Not real gaps:
    • M04, M06 and M21 are fences that cannot differ while one process owns a claim (runOnce is synchronized; the owner is per process). The two-collector run above exercised the cross-process case.
    • M07 (the per-page probe) is redundant: the per-key probe makes the same check and is pinned by a test (M08); removing both (M29) is killed.
  • Real gaps:
    • M24–M27: the collector tests only exercise the grace_period and legacy blockers. Collecting a quarantined, recovery-protected, put_unresolved or reader_active publication passes the suite; only the real-stack runs in section 3 show those are retained.
    • M17: reverting the guarded wrapper to raw read() also passes. It changes nothing on the public download path: the real stack measured 205.6 statements/MiB with the mutant, the same as the PR. But it guards Session-store and verification reads, which no bounded-query test measures.

Suggestions (none blocking)

  1. Coordinate V34 with feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210.
  2. Take the 2-line read-granularity patch, or reword the 100 MiB claim.
  3. Before turning on gc-enabled, add a way to resolve UNKNOWN PUT attempts. For example, a guarded HEAD of the exact key with a digest comparison, or an operator command. Without it, every transient PUT error permanently leaks that publication's storage and tenant quota. It also adds to the already-deferred 60 s polling of permanent blockers.
  4. Duplicate versioning probes: each deleted key costs two OSS requests (115 versioning probes for 113 DELETEs), because deleteIfPresent probes again after the per-page probe. Only the per-key probe is pinned by a test (M08); the page probe is redundant today (M07).
  5. Add a parameterised collector test over every blocker kind (quarantined, recovery-protected, reader_active, put_unresolved, …), so that M24–M27 fail.

Evidence (figures, harness, logs, candidate patch): assets-pr13225.

中文版

#13225 本地真实环境验证(88d4ca9466)

结论:在 GC 保持关闭(默认)的前提下,没有阻塞合入的问题。 0c64229ae6 上存在的 Flyway V33 撞号,已由 88d4ca9466 修复。新 head 的 packages/sdk-java 与我验证的树逐字节相同。在 Linux 上走真实的公开 close → DELETE 链路,回收器的行为与 PR 描述一致;假 OSS 和真实阿里云 OSS bucket 都验证过。合入或启用 GC 前需要关注三点:

  1. 与 feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210 的迁移顺序。 在飞的 feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210 同样新增 V34__managed_session_creator.sql,后合入的一方必须改号。
  2. F2 的表述。 存储侧 SQL 放大确已修复(每 MiB 711 → 205 条)。但正文"5 ms SQL 延迟下 100 MiB 约 108 s 完成"在本装置上未能复现:120 s 读取租约到期时,100 MiB 只送出了 85.8 MiB。下文的两行后续改动可把成本降到每 MiB 39 条,128 MiB 31 s 完成。
  3. 启用 GC 前:一次瞬时 PUT 失败会让该 publication 永远无法回收。 这符合设计文档("unknown … 永久阻止自动回收,直到外部解决"),但目前没有任何解决路径,该 publication 的存储和租户配额永不释放。

环境

臂 树 用途
PR b95e64cdf0 = 0c64229ae6 合入 main 623cfc67 + V33→V34 改号;git diff b95e64cdf0 88d4ca9466 -- packages/sdk-java 为空 以下全部
base main 623cfc67 回收、F2、GET 重试对照
候选 PR + 两行补丁(candidate-1mib-read-granularity.patch) 仅 F2 后续
  • 门禁与启动矩阵: macOS arm64,JDK 21,MySQL 8.4.7。
  • 所有退役相关场景都在 Linux 上运行。 macOS 上 Workspace close 不可用:durable worker 身份读取 /etc/machine-id 和 /proc,session_close=false。Linux 栈的组成:
    • colima VM(aarch64,4 CPU / 8 GB)中的 JDK 21 容器和 MySQL 8.4.11 容器;
    • Spring fat jar、打包后的 Hosted Harness,以及 durable local-process Broker worker;
    • 假模型,以及由真实 aliyun-sdk-oss 3.18.4 经 TLS 访问的 OSS 替身。
  • 真实阿里云 OSS: cn-hangzhou 临时私有、未开版本控制的 bucket,用完已删除。
  • 缝合点(如实说明):
    • tap 把 Harness 的 tool profile 由 hosted-workspace-files/1 改写为 hosted-workspace-shell/1,因为 main 尚未开放公开 Shell 准入。
    • 历史写入证据缺失、恢复状态、残留读租约三种情况通过 SQL 注入。
    • 其余全部走公开 API:创建、Turn、下载、POST /close、DELETE。
  • Harness/worker bundle 由 main 623cfc67 构建。1abccdb26a 合并只多了一个 packages/core 改动(fix(core): stop misdiagnosing malformed tool-call args as max_tokens truncation #12982),与本链路无关。

1. 合入与门禁

landing

  • 0c64229ae6 试合并到当前 main 后,仓库自带的 check-flyway-migrations.js 报两个迁移都占 version 33,Spring 也无法启动。
  • 改号后的树在空库上能启动,在已由 main 迁移的库上也能启动(V33 definitions → V34 collection)。
  • 只跑过 PR head 的库,main 无法打开(V33 checksum mismatch)。这是 PR 正文已说明的临时迁移历史情形。
  • 单测: 514/514,Checkstyle 0。
  • -Pmysql-integration: 51/51(含 WorkspaceSessionRetentionMySqlIT 16/16),failsafe 报告检查器退出码 0。

2. Linux 公开生命周期与真实 OSS

lifecycle

  • 公开调用: POST /close 202 → completed(554 ms);DELETE 202 → completed(292 ms)。
  • 回收时机: 宽限期严格等到 retired_at + grace(10 s);+10.5 s 被领取,+10.9 s COLLECTED。
  • OSS 删除: 15 次 DELETE,每个 catalog key 一次,只删被删会话的 key,没有按前缀扫描。
  • 清理与配额: inline 已清空;held 13,641,895 → 0,released_held_bytes = 13,641,895。
  • 兄弟会话: 未受影响,下载摘要正确。
  • main 上: 同样的生命周期结束后,输出一直停在 RETIRING,对象和配额全部保留。
  • 真实 OSS: 被删会话的 8/8 对象被删除(bucket 16 → 8),兄弟会话的 8 个保留。
  • 真实 bucket 开启版本控制: 回收器失败关闭,两轮报 Tool publication bucket cannot enforce immutable objects,没有删除任何对象(4/4 保留)。

3. 阻塞与兄弟

blockers

隔离: 我翻转了存储对象的一个比特。公开下载检测到损坏(503 artifact_unavailable),生产代码随即隔离了该 publication。

删除后:

  • 被隔离的 publication 保留 7/7 对象和配额;
  • 同一会话中的另一个 publication 被回收。

legacy_write_evidence_missing 和 recovery_protected 均保留。残留读租约使该行停在 reader_active;租约过期后,在下一次 60 s 轮询时回收(+71.2 s)。租户 held 恰好减少两个已回收 publication 的量。

未决 PUT(第 3 点)。 我在链路上注入了一次 PUT 500。下一次尝试成功返回,11 个对象全部 VERIFIED。但账本里仍留下 UNKNOWN=1,删除后该 publication 永远停在 RETIRING / put_unresolved,6,296,450 字节一直被占用。代码里没有任何路径会把尝试移出 UNKNOWN。

4. 存储故障

faults

失败的 DELETE 只发一次,整页重试: 第 5 个 key 的 DELETE 遇到 500 或连接重置时,该 key 只发出了一次请求,SDK 没有重放。整页延后 60 s,由第 2 代重新执行;已删掉的 key 返回 204。

版本控制、崩溃与多页:

  • 版本控制 Suspended 时,回收器没有发出任何 DELETE。
  • DELETE 被挂起时 kill -9 Spring:重启后的进程等旧 claim 到期(59.6 s),再以第 2 代完成。
  • 113 个 key 的 publication 分两页回收,两页之间相隔一个 tick。
  • 两个开启 GC 的 Spring 实例共享一个库:两个 owner 都参与了回收,但每个 publication 只有一个 owner、都是第 1 代,93 个 key 各删一次。

每种情况下,配额都恰好只释放一次。

5. F2:下载成本与 120 s 租约

f2

在 MySQL 前放了一个 TCP 中继,给每个客户端→服务器包加 5 ms 延迟。main 上 40 MiB 下载在租约到期时被截断在 24.2 MiB;PR 上 55.5 s 完成。但 PR 上 100 MiB 仍在 85.8 MiB 处被截断,吞吐约 0.73 MiB/s。

剩余开销约为每 64 KiB 响应块 11 条 SQL:

  • ManagedArtifactService 的 64 KiB 写循环每次都检查租约和访问权限;
  • VerifiedStream.read 每次最多读 64 KiB,并在读前、读后各跑一遍完整守卫。

候选补丁把这两个常量提高到 1 MiB,与 PR 在存储路径已采用的粒度一致。结果:

  • 成本与速度: 每 MiB 38.8 条;+5 ms 下 40 MiB 用时 10.3 s,128 MiB 用时 31.3 s。
  • 撤权语义不变: 下载途中撤权或 DELETE 仍会中止响应,之后最多再送出 0.88 MiB。

这两处是 main 的代码,不在 PR diff 内,可在本 PR 一并修改,也可作为后续 PR;否则建议改写正文中关于 100 MiB 的表述。

6. OSS GET 重试与变异测试

retry-mutation

GET 重试。 我通过真实 SDK,在公开下载的第 3 个对象上注入故障。

  • 只有 PR 会重试的: 错误码不在白名单里的 500,以及非 XML 响应体的 502。main 上这两种情况发一次请求就中止响应;PR 重试一次后完整下完,摘要正确。
  • 两臂一致的: 500 InternalError、503、连接重置、连续六次 500(4 次请求后中止)、403(不重试)。
  • PUT: 一次 PUT 500 产生两次独立尝试,每次只发出一个请求,SDK 没有重放。

变异测试: 对新代码构造了 28 个变异体,每个都跑完整单测(514 个)。杀死 19 个。存活的 9 个:

  • 不是真缺口:
    • M04、M06、M21 是在"一个进程持有一个 claim"时不可能触发的围栏(runOnce 是 synchronized 的,owner 按进程区分)。上文的双实例运行覆盖了跨进程的情形。
    • M07(页级探针)是冗余的:每 key 探针做同样的检查且被测试钉住(M08);两者都去掉(M29)会被杀死。
  • 真实缺口:
    • M24–M27: 回收器测试只覆盖了 grace_period 和历史写入证据两种阻塞。把被隔离、恢复保护、put_unresolved 或 reader_active 的 publication 回收掉,单测依然全绿;只有第 3 节的真实栈运行证明了它们会被保留。
    • M17: 把受守卫的包装层改回原始 read(),单测同样全绿。它对公开下载没有影响:真实栈上带该变异体测得每 MiB 205.6 条,与 PR 相同。但这个包装层保护着 Session Store 资源读取和校验读取,而这两条路径没有任何有界查询测试。

建议(均不阻塞)

  1. 与 feat(managed-agent): broker authentication and broker-provisioned writer credentials #13210 协调 V34。
  2. 采纳两行读取粒度补丁,或改写 100 MiB 的表述。
  3. 打开 gc-enabled 之前,增加处理 UNKNOWN PUT 尝试的手段,例如对确切 key 做受保护的 HEAD 并比对摘要,或提供运维命令。否则每一次瞬时 PUT 错误都会永久泄漏该 publication 的存储和租户配额,也会加重已延期的"永久阻塞每 60 s 轮询"问题。
  4. 重复的版本控制探针: 每删除一个 key 要发两个 OSS 请求(113 次 DELETE 对应 115 次版本探针),因为 deleteIfPresent 在页级探针之后又探了一次。目前只有每 key 的探针被测试钉住(M08),页级探针是冗余的(M07)。
  5. 为回收器补一个覆盖全部阻塞类型(quarantined、recovery_protected、reader_active、put_unresolved 等)的参数化测试,使 M24–M27 能被杀死。

证据(图、装置脚本、日志、候选补丁):assets-pr13225。

@wenshao
wenshao marked this pull request as ready for review October 2, 2026 23:32
@wenshao

wenshao commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

@qwen-code /triage

@wenshao
wenshao added this pull request to the merge queue Oct 3, 2026
Merged via the queue into main with commit fa795e0 Oct 3, 2026
83 of 85 checks passed
wenshao added a commit that referenced this pull request Oct 3, 2026
Base PR #13225 (O4-2 collector) squash-merged into main as fa795e0.
Resolve the stack onto current main:
- Drop branch-side V33__managed_tool_output_collection.sql (byte-identical
  to main's V34 after the pre-merge renumber; main's V33 is
  managed_agent_definitions from #13142). Sequence stays V30-V34, unique.
- workspace-session-reliable-close docs (both languages) and
  WorkspaceRecoveryStoreTest: take main's version; O4-3's own commits never
  touched them, branch side was old-stack residue.
- managed-tool-output-retention docs (both languages): main's refreshed
  version plus O4-3 commit 9d9d016's first-paragraph delta (32-candidate
  tick bound, grace semantics, publication-scoped cleanup).

Co-authored-by: Qwen-Coder <[email protected]>
wenshao added a commit that referenced this pull request Oct 3, 2026
This PR's base moved to main after #13225 merged; no PR will target the
merged codex/managed-tool-result-o4-2 branch again.

Co-authored-by: Qwen-Coder <[email protected]>
JadeCong pushed a commit to CloudEngineHub/qwen-code that referenced this pull request Oct 3, 2026
…nLM#13090)

* feat(managed-agent): Collect retired tool output safely

Co-authored-by: Qwen-Coder <[email protected]>

* chore(managed-agent): Sync retention formatting and verify collection

Co-authored-by: Qwen-Coder <[email protected]>

* codex: address PR review feedback (QwenLM#13087)

* codex: address PR review feedback (QwenLM#13087)

* test(managed-agent): Add tool output collection deployment gates

Co-authored-by: Qwen-Coder <[email protected]>

* codex: address PR review feedback (QwenLM#13090)

* docs(managed-agent): Preserve deployment gate identifiers

* fix(managed-agent): Advance eligible tool output collection

* codex: address PR review feedback (QwenLM#13090)

* chore(ci): Drop merged O4-2 base branch from SDK Java trigger

This PR's base moved to main after QwenLM#13225 merged; no PR will target the
merged codex/managed-tool-result-o4-2 branch again.

Co-authored-by: Qwen-Coder <[email protected]>

* docs(managed-agent): Refresh operations runbook for merged O4-2

Audit found the runbook still pointed at the pre-merge baseline SHA and
the old collection V33, contradicting the lifecycle design doc in the
same diff. State the merged stack without a volatile SHA, use the
current V30/V31/V32/V33(definitions)/V34(collection) sequence in both
languages, and say the SDK Java workflow targets main now that the
stacked base branch is merged.

Co-authored-by: Qwen-Coder <[email protected]>

---------

Co-authored-by: Qwen-Coder <[email protected]>
Co-authored-by: yiliang114 <[email protected]>
Co-authored-by: Shaojin Wen <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants