Skip to content

feat(managed-agent): Collect retired tool output safely - #13087

Closed
doudouOUC wants to merge 3 commits into
codex/managed-tool-result-o4-1from
codex/managed-tool-result-o4-2
Closed

doudouOUC wants to merge 3 commits into
codex/managed-tool-result-o4-1from
codex/managed-tool-result-o4-2

Conversation

@doudouOUC

@doudouOUC doudouOUC commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Superseded by #13225 on current main after GitHub rejected reopening this closed PR. Original historical description follows.

GitHub 拒绝重开,本 PR 由基于当前 main 的 #13225 替代。以下保留原历史说明。

What this PR does

O4-2 collects eligible retired foreground Shell publications through durable claims and at most 100 exact catalog objects per tick. Storage deletion runs outside SQL transactions; cursor confirmation and final inline cleanup release held/used quota once, retaining identity and audit tombstones. Physical GC stays disabled by default. Failed pages receive a persisted retry delay so one inaccessible object cannot block healthy publications.

Why it's needed

Retirement alone retains storage and quota indefinitely. Safe reclamation must survive lost delete responses, worker replacement and SQL confirmation failure without deleting uncatalogued keys or releasing quota early.

Reviewer Test Plan

How to verify

Verify 100/100/remainder pagination, preservation of objects outside the catalog, and retained quota/inline data until the final confirmation. Lose a delete response, interrupt a page, fail SQL confirmation, and replace the claim owner: retries must repeat the same exact keys, an old generation must not confirm, and quota must be released once. A continuously failing publication must yield to a healthy one. Disabled GC and legacy missing evidence must cause no deletion.

The actual scheduled blocking-delete regression now proves that other live Sessions continue materializing. Claim-expiry regressions retain quota until a replacement confirms every key; the production OSS adapter rejects versioning enabled after construction and preserves deletion failures. The full default Java gate passed 388 tests / 52 classes and Checkstyle; the full ordinary MySQL profile passed those 388 unit tests plus all 25 integration tests / 2 classes, without selectors or skips. Independent review found no remaining Critical. Repository build, typecheck and bundle passed on the clean committed head. Exact commands, totals and limitations are in the separate validation report.

Evidence (Before & After)

N/A — internal storage behavior without a TUI or Web Shell flow change. Previously retired publications retained all objects; enabled fixture collection now deletes only catalog keys and clears inline copies after complete confirmation.

Tested on

OS Status
🍏 macOS ✅ tested
🪟 Windows ⚠️ not tested
🐧 Linux ⚠️ not tested

Environment (optional)

macOS 26.6.2 arm64, Java 21.0.12.1, Maven 3.9.16, MySQL 8.4.11 and H2; Node 22.14.0 for repository checks.

Risk & Scope

  • Main risk or tradeoff: an unresolved PUT or unsafe publication continues consuming quota. One dedicated single-thread output scheduler per instance processes bounded pages and uses generation fencing; retries conservatively repeat deletes.
  • Not validated / out of scope: real MySQL/OSS faults, process kills and large closures are O4-3 deployment gates. Public history erasure, active Session TTL and shared/background/MCP/media output are outside scope.
  • Breaking changes / migration notes: depends on O4-1 feat(managed-agent): Protect Session-owned tool output retirement #13084; collection migration is V28 after O4-1 V27; recheck latest main migration numbers before landing. Upgrade all Java publication writers and satisfy O4-3 gates before enabling GC. No public API change.

Linked Issues

Related to #12380; depends on #13084; #13037 and #12894 are merged prerequisites. Does not close the umbrella issue or change #13019 recovery policy.

中文说明

这个 PR 做了什么

O4-2 通过持久 claim 回收符合条件的已退役前台 Shell publication,每个 tick 最多处理 catalog 中的 100 个精确对象。存储删除在 SQL 事务外执行;游标确认及最终 inline 清理一次性释放 held/used 配额,并保留身份及审计墓碑。物理 GC 默认关闭。失败页面持久记录重试延迟,避免单个不可访问对象阻塞健康 publication。

为什么需要

仅退役会无限保留存储和配额。安全回收必须承受删除应答丢失、worker 替换及 SQL 确认失败,不能删除目录外的 key 或提前释放配额。

评审测试计划

如何验证

验证 100/100/余量分页、目录外对象保留,以及最终确认前配额和 inline 数据保持不变。注入删除应答丢失、页面中断、SQL 确认失败及 claim owner 替换:重试必须重复同一组精确 key,旧 generation 不能确认,配额只能释放一次。持续失败的 publication 必须让健康 publication 得到处理。GC 关闭和历史证据缺失时不得删除。

实际调度的阻塞删除回归现已证明其他活跃 Session 继续物化。claim 到期回归保留配额,直到接管者确认全部 key;生产 OSS adapter 拒绝构造后开启版本控制,并保留删除异常。完整 Java 默认门禁通过 388 项测试 / 52 个类及 Checkstyle;完整普通 MySQL profile 通过这 388 项单元测试以及全部 25 项集成测试 / 2 个类,无选择器或跳过。独立审查未发现剩余 Critical。 仓库 build、typecheck、bundle 已针对干净的已提交 HEAD 通过。准确命令、数量和限制见单独验证报告。

前后证据

N/A — 内部存储行为,没有 TUI 或 Web Shell 流程变更。此前已退役 publication 保留全部对象;启用的 fixture 回收现在只删除 catalog key,完整确认后才清空 inline 副本。

测试系统

OS 状态
🍏 macOS ✅ 已测试
🪟 Windows ⚠️ 未测试
🐧 Linux ⚠️ 未测试

环境

macOS 26.6.2 arm64、Java 21.0.12.1、Maven 3.9.16、MySQL 8.4.11 和 H2;仓库检查使用 Node 22.14.0。

风险与范围

  • 主要风险或取舍:未解决 PUT 或不安全 publication 继续占额。每实例独立单线程输出调度器处理有界页面并以 generation 隔离;重试保守重复删除。
  • 未验证或范围外:真实 MySQL/OSS 故障、进程死亡及大闭包是 O4-3 部署门禁。公开历史擦除、活跃 Session TTL 和共享/后台/MCP/媒体输出不在范围内。
  • 兼容性及迁移:依赖 O4-1 feat(managed-agent): Protect Session-owned tool output retirement #13084;清理迁移为 O4-1 V27 之后的 V28;落地前再次核对 main 的最新迁移号。启用 GC 前升级全部 Java publication writer 并通过 O4-3 门禁。不改变公开 API。

关联 Issue

关联 #12380,依赖 #13084;#13037 和 #12894 已合并。不关闭总 issue,也不改变 #13019 恢复政策。

Design: English · 简体中文.

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

O4-2 E2E/test report: focused Java tests passed 61/61 (Collector 16, Retention 9, Publication 31, Artifact API 5), zero failures/errors/skips. Independent test-engineer passed 20/20 before the final fairness regression. Verified exact-key pagination, lost delete response, partial page retry, worker generation fencing, SQL confirmation rollback, one-time quota release, default-off/legacy protection and failure fairness. Independent review is clean.

Environment: macOS, Java 21.0.8, Maven 3.9.14, H2 MySQL-mode, controlled memory storage. No real OSS or process-kill/large-output deployment evidence is claimed; O4-3 gates remain required before deployment GC.

中文:针对性测试 61/61 通过,失败/错误/跳过均为零;独立 test-engineer 在最终公平性回归前通过 20/20。精确 key 分页、应答丢失、部分页面重试、worker generation、SQL 回滚、一次性配额释放、默认关闭/历史保护及失败公平性已验证。范围为 macOS Java 21 + H2 + 受控内存存储;真实 OSS、进程死亡及大输出证据仍由 O4-3 补齐。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

Final verification update: synced O4-1 formatting and corrected collector block formatting without changing lexical tokens. O4-2 full Maven verify passes at c19df01: 61 focused tests, zero failures/errors/skips, compilation/packaging/Checkstyle green. GC remains disabled by default; deployment gates are recorded in O4-3.

中文:已同步 O4-1 格式修正,并修正 collector 代码块格式,词法内容不变。O4-2 在 c19df01 的完整 Maven verify 通过:61 项针对性测试,失败/错误/跳过均为零,编译/打包/Checkstyle 通过。GC 继续默认关闭,部署门禁在 O4-3 记录。

@wenshao wenshao left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed. Suggestions are inline.

Not explored to full depth (tool budget reached): "## The diff": 无(约 24 次工具调用,全部检查完成;未运行需要真实 MySQL/OSS 的 failsafe IT,与本机基础设施限制一致,已在上面用 H2 MySQL-mode 全量单测兜底披露)。; "agent 5": none — 所有检查在预算内完成,无被中止的验证项。; "agent 3c": 无——所有计划内的检查均已完成,未因工具预算中止任何核对。; "agent reverse-audit (round 1)": 无——设定的各层走查均已完成;唯一未闭合点(broker 侧退役附着闸门)属 diff 之外的前置边界,已在 fence-interplay 回执中具名,而非预算截断。; "agent 1e": 无 — 所有计划检查均已完成,未触及工具预算上限(约 46 次中使用了 8 次)。, and 2 more.

中文说明

已审查。 建议见行内评论。

未探索到全部深度(达到工具调用预算):"## The diff":无(约 24 次工具调用,全部检查完成;未运行需要真实 MySQL/OSS 的 failsafe IT,与本机基础设施限制一致,已在上面用 H2 MySQL-mode 全量单测兜底披露)。;"agent 5":none — 所有检查在预算内完成,无被中止的验证项。;"agent 3c":无——所有计划内的检查均已完成,未因工具预算中止任何核对。;"agent reverse-audit (round 1)":无——设定的各层走查均已完成;唯一未闭合点(broker 侧退役附着闸门)属 diff 之外的前置边界,已在 fence-interplay 回执中具名,而非预算截断。;"agent 1e":无 — 所有计划检查均已完成,未触及工具预算上限(约 46 次中使用了 8 次)。,另有 2 条。

— glm-5.3-flash via Qwen Code /review (v0.24.6)

@doudouOUC
doudouOUC changed the base branch from codex/managed-tool-result-o4-1 to main October 1, 2026 03:47
@doudouOUC
doudouOUC marked this pull request as ready for review October 1, 2026 03:47
@doudouOUC
doudouOUC changed the base branch from main to codex/managed-tool-result-o4-1 October 1, 2026 03:48
@doudouOUC
doudouOUC marked this pull request as draft October 1, 2026 03:48
@doudouOUC
doudouOUC force-pushed the codex/managed-tool-result-o4-1 branch 3 times, most recently from 21ad30b to 3ee4304 Compare October 1, 2026 10:04
doudouOUC added a commit that referenced this pull request Oct 1, 2026
@doudouOUC
doudouOUC force-pushed the codex/managed-tool-result-o4-2 branch from c19df01 to abad13c Compare October 1, 2026 10:33
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Please do not rebase or force-push to an active PR as it invalidates existing review comments. Note for future reference, the bots always squash all changes into a single commit automatically as part of the integration.

中文

请勿对活跃的 PR 执行 rebase 或 force-push,因为这会使已有的评审评论失效。另外,供日后参考:作为集成流程的一部分,机器人始终会自动将所有改动压缩(squash)为单个提交。

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

[codex] Review handling for abad13c. Rebased O4-2 on the repaired O4-1, preserving database epoch/writer/read protection and V27; collection uses V28. This slice adds bounded exact-key deletion and conservative quota release. Deployment GC remains off until O4-3 gates pass.

Thread Action Reason
R1-1 not taking Keeping the adapter-level unversioned check: deleteIfPresent is callable independently of the collector, and its guard matches the existing PUT/read contract. Page-level checking remains too; optimizing round trips cannot remove that fail-closed boundary.
R1-2 fixed The collector uses a dedicated named single-thread scheduler, preserving synchronized non-overlap. A real scheduled blocking-delete regression verifies that MessageMaterializer advances while the page is paused.
R1-3 partially addressed The scheduled collector now logs its worker identity and the throwable, preserving the failure cause and stack. Existing durable gc_blocker state still identifies pending publications; this change does not add new retry/error state.
R1-4 fixed A two-object regression expires the claim during the first delete, proves that the second delete does not happen and all quota remains held, then verifies replacement replay. An independent source-copy mutation removing the renewal expiry fence fails the permanent regression; the unmodified control passes.
R1-5 fixed A one-object regression expires the claim after physical deletion but before SQL confirmation, verifies DELETING and held quota, and checks that retry replays the key. An independent source-copy removal of the expiry conjunct fails the permanent regression; the unmodified control passes.
R1-6 fixed The remaining-object check now uses a PK-prefix LIMIT 1 existence query. It retains the exact quota-release boundary without scanning every remaining row on each page.
R1-7 fixed The production OSS adapter has focused mocked-client tests: flip versioning after construction to pin delete-time rejection, check bucket/exact-key argument order, and preserve permission failures.
R1-8 deferred The current schema has exactly three explicit quota buckets, with end-to-end accounting assertions. No fourth bucket is being added; extracting hypothetical future helpers would widen this slice without fixing current behavior.
R1-9 deferred The only production adapter implements deletion; unsupported implementations fail closed and retain quota. Making the internal interface abstract would require unrelated fake-adapter churn. Compile-time tightening can be handled as follow-up.
R1-10 not taking Permanent Session retirement blocks every restore/read/writer path before intentionally-cleared resource payloads reach verification. Retaining resource identity, receipt and digest tombstones follows the design; a new resource lifecycle/export API is outside this collector slice.
R1-11 not taking Keeping the database owner/generation/expiry fence before each physical delete. A wall-clock throttle can permit a resumed stale worker to issue further deletion calls; this slice favors the checked invariant over reducing database round trips.
R1-12 partially addressed The prior fairness fix defers a failing publication for 60 seconds, allowing other publications to progress; the regression verifies that behavior. Failure diagnostics are improved. The failed publication intentionally retains its full quota until every deletion is confirmed; dead-letter or leaked-byte accounting is a separate policy decision.
R1-13 fixed The actual scheduled collector path is covered, and a direct tick failure regression verifies that the exception is contained and a WARN retains the throwable while quota stays held.
R1-14 fixed Resource cleanup now uses the Session resource primary-key prefix derived by ManagedSessionStore.sessionScopeKey, retaining tenant/session identity checks. The fixture uses the same actual Session hash, and rollback/cleanup accounting regressions remain intact.

Validation and environment evidence are reported separately below the PR description. Deferred items remain recorded here for follow-up; they do not claim those concerns are fixed.

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

[codex] O4-2 validation on clean committed head abad13c77c09f9e28fda7cbe7ac4a76b4f287604, based on repaired O4-1 3ee43046ab6777c6a6ac028ede942a78d5e2aee3.

  • Repository npm run build, npm run typecheck, npm run bundle: passed on this exact committed head.
  • Full module mvn verify checkstyle:check: 52 source classes / 52 fresh reported classes, 388 tests, 0 failures/errors/skips, Checkstyle 0; missing/extra/stale reports 0.
  • Full ordinary mvn -Pmysql-integration verify checkstyle:check with only ordinary MySQL configuration and a fresh isolated schema: first attempt passed, no test selector or skip. Surefire 388 tests; Failsafe 2 classes / 25 tests (17 ManagedAgentMySqlIT + 8 ToolPublicationRecoveryMySqlIT), 0 failures/errors/skips. Unmodified source-derived node scripts/check-failsafe-reports.js non-hosted packages/sdk-java/managed-agent-server passed; missing/extra/stale reports 0.
  • Focused collector / production OSS adapter / existing scheduling suite: 29 cases passed. Actual Spring scheduler regression reproduced the old materializer stall (1→1 while delete blocks), then verified 1→6 with the dedicated production scheduler and final COLLECTED / held=0.
  • Independent copies of source: 8 controls pass and 8 permanent regression executions kill renewal expiry, confirmation expiry, shared scheduler, lost throwable, version Enabled/Suspended, wrong exact-key arguments and swallowed storage failure mutations.
  • Two clean full self-audit passes and independent exact-head review: no remaining Critical or new high-value Suggestion.

Environment: macOS 26.6.2 arm64, Node 22.14.0, Java 21.0.12.1, Maven 3.9.16, MySQL 8.4.11 (owned isolated local instance), H2 and mocked OSS transport. Real OSS deletion/permissions, O4 process/capacity gates and full Hosted foreground Shell acceptance are separate O4-3 deployment evidence; this report does not claim they ran. GC remains disabled by default and deployment policy remains 24 hours. Collection migration is V28 after O4-1 V27; recheck current main versions before landing.

14 original threads are individually handled with reasons, including partial acceptance and deferrals. The original O4-1 combined MySQL attempt had an unchanged-main unit-test flake reported there; the current O4-2 full combined profile passed on its first run.

@doudouOUC
doudouOUC marked this pull request as ready for review October 1, 2026 10:35
@doudouOUC
doudouOUC force-pushed the codex/managed-tool-result-o4-2 branch from abad13c to 6ff843b Compare October 1, 2026 11:44
@doudouOUC

Copy link
Copy Markdown
Collaborator Author

[codex] O4-2 validation on clean committed head 6ff843bc140e36ee282f108fb1f2ca6bb92d113b, based on repaired O4-1 9f5e66f2a37b1755164233d2d1c9d60e5c46dfbe.

  • Repository npm run build, npm run typecheck, npm run bundle: passed on this exact committed head.
  • Full module mvn clean verify checkstyle:check: 52 source classes / 52 fresh reported classes, 388 tests, 0 failures/errors/skips, Checkstyle 0; missing/extra/stale reports 0.
  • Full ordinary mvn -Pmysql-integration verify checkstyle:check with only ordinary MySQL configuration and a fresh isolated schema: first attempt passed, no test selector or skip. Surefire 388 tests; Failsafe 2 classes / 25 tests (17 ManagedAgentMySqlIT + 8 ToolPublicationRecoveryMySqlIT), 0 failures/errors/skips. Unmodified source-derived node scripts/check-failsafe-reports.js non-hosted packages/sdk-java/managed-agent-server passed; missing/extra/stale reports 0.
  • Focused collector / production OSS adapter / existing scheduling suite: 29 cases passed. Actual Spring scheduler regression reproduced the old materializer stall (1→1 while delete blocks), then verified 1→6 with the dedicated production scheduler and final COLLECTED / held=0.
  • Independent copies of source: 8 controls pass and 8 permanent regression executions kill renewal expiry, confirmation expiry, shared scheduler, lost throwable, version Enabled/Suspended, wrong exact-key arguments and swallowed storage failure mutations.
  • Two clean full self-audit passes and independent exact-head review: no remaining Critical or new high-value Suggestion.

Environment: macOS 26.6.2 arm64, Node 22.14.0, Java 21.0.12.1, Maven 3.9.16, MySQL 8.4.11 (owned isolated local instance), H2 and mocked OSS transport. Real OSS deletion/permissions, O4 process/capacity gates and full Hosted foreground Shell acceptance are separate O4-3 deployment evidence; this report does not claim they ran. GC remains disabled by default and deployment policy remains 24 hours. Collection migration is V28 after O4-1 V27; recheck current main versions before landing.

14 original threads are individually handled with reasons, including partial acceptance and deferrals. The original O4-1 combined MySQL attempt had an unchanged-main unit-test flake reported there; the current O4-2 full combined profile passed on its first run.

Rebased onto the independently reproduced Hosted lifecycle correction: CLOSE remains SEALED and completed DELETE requires DELETED, as the unchanged permanent-retirement transaction establishes. All three O4-2 patch bodies and ten shipping file blobs are byte-identical to the previously reviewed revision; only this inherited test/document correction changed. Fresh checks above apply to 6ff843b, while source-copy scheduler/claim/OSS mutation evidence was executed before rebase against the identical source. Existing 14 replies are deduplicated and remain resolved; no duplicate inline reply is posted.

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

[codex] Thanks — agreed on the CI evidence gap; recording the non-blocking suggestions without widening this shipping batch. Rechecked stage 1, stage 2 and stage 3 against current head 6ff843bc140e36ee282f108fb1f2ca6bb92d113b.

Item Decision and reason
Remote CI Agreed. This stacked base does not trigger ordinary Java/build/test PR CI. The latest report is local execution evidence, not remote CI approval. After #13084 lands, retarget/rebase this PR onto main, recheck migration numbering and require the actual Java CI before landing; repeat for #13090 after this layer lands.
Tick/claim configuration Deferred. This slice deliberately uses one dedicated worker, at most 100 exact keys per page, a 1-second delay and a 60-second fenced claim/retry delay. Operational throttling settings need measured deployment requirements and their own validation; GC remains off. The existing runbook describes stopping collection on every instance.
Expired Session leases Not moving this cleanup. Leases protect the Session closure, and the final confirmation deletes only rows for that same Session already expired by database time. It cannot revoke a live lease; moving it to retirement would not clean leases that expire later.
Permanently denied deletion / dead-letter accounting Agreed that fairness is not eventual reclamation. A permanently denied key retains its publication and the full held quota; there is no dead-letter or automatic leaked-byte release implementation here. This is a follow-up slice, explicitly carried into #13090's deployment decision. Before any GC enablement, the deployment owner must approve the sustained-delete-denial and held-quota handling policy; successful capacity tests alone do not settle it. The existing O4-3 failure-handling section requires escalation of sustained collection_retry and forbids manually releasing quota to hide it. #13090 remains draft and GC remains off while real OSS and full Hosted acceptance are unavailable.
Resource verification ordering Confirmed and recorded here: every current verifyStoredResource call is reached through the permanent-retirement guard (requireReadGrant / findHeadForUpdate), and collection requires that retirement tombstone. Therefore the retained MYSQL_INLINE metadata with cleared bytes is unreachable through a live-resource read. Any future export/restore/replay consumer must preserve this ordering; a dedicated documentation refinement is deferred rather than reopening the already dispositioned R1-10 thread.

No new fix commit or CI rerun is claimed. All 14 original inline threads remain resolved; the maintainer's approval and real CI are still pending.


中文:已对照当前 6ff843bc 核实三条新 triage 评论。确认堆叠基线不会触发普通 Java CI,本地测试报告不能代替远端门禁;#13084 合入后再将本层改到 main,核对迁移编号并等待实际 CI,随后同样处理 #13090。暂不扩大本轮配置面;Session 租约清理只涉及该 Session 在数据库时间下已过期的行。永久删除失败会保留 publication 和全部 held 配额,公平重试并不解决这笔滞留:dead-letter/泄漏记账作为后续切片,在任何 GC 启用前须由部署负责人明确持续删除失败及配额处理策略。#13090 仍为草稿、GC 关闭,真实 OSS 与完整 Hosted 验收未完成。已记录资源校验必须先经过永久退役栅栏的调用顺序;后续新增导出、恢复或重放路径必须保持该约束。14 个原行内线程均已解决,本评论没有声称新增提交、重跑或远端 CI 通过。

@wenshao

wenshao commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

Real-stack verification of O4-2 at 6ff843bc

Verdict: no blocker in O4-2's own code. Every step of the Reviewer Test Plan held on a real stack: MySQL 8.4.7, the PR's Spring jar, real Shell calls through the packaged Hosted Harness and Broker worker, a fake OSS with fault injection, and one leg on a real Aliyun OSS bucket. Pagination is bounded, only catalog keys are deleted, quota is released once and only after the last page, retries repeat the same keys, and a stale JVM can neither renew nor confirm. Before GC is enabled, three things need attention:

  1. Nothing can retire a Shell publication on this stack today. The public and WebShell DELETE refuse every Workspace Session with 409 workspace_unavailable, and O2 Shell publications only exist on Workspace Sessions. I had to drive retirement through a seam (below). This is not an O4-2 defect, but merging adds no runtime behaviour, and the end-to-end path from API delete to collection still has no owner.
  2. F1: blocked candidates use up the collector's one-per-second claim slot. With 151 real publications, an eligible publication was collected 95.2 s after it became eligible. With a small candidate patch the same run took 0.96 s. With the default 24 h grace, every retired publication is re-evaluated about 1440 times.
  3. Test gap (M12): widening the Session-resource cleanup to the whole Session (which would clear messages and checkpoints) passes all 41 unit tests. A +29-line test pins it.

What ran

Item Result
Heads Started on abad13c7; the PR was force-pushed to 6ff843bc mid-run. git diff abad13c7 6ff843bc touches only README.md and one assertion in HostedHarnessMySqlIT (SEALED → DELETED). src/main, runtime-broker, qwencode, packages/cli and packages/core are byte-identical, so every real-stack result below applies to 6ff843bc.
Java gate, non-Hosted (abad13c7) mvn -Pmysql-integration clean verify checkstyle:check on MySQL 8.4.7: 388/388 surefire, 25/25 failsafe, Checkstyle 0. check-failsafe-reports.js non-hosted: ok.
Trial merge with main 0a5f518b Clean merge. 389/389 + 25/25, Checkstyle 0. V27/V28 do not collide with main (main is at V26).
Hosted lane (6ff843bc, -Phosted-harness-mysql) The changed HostedHarnessMySqlIT passed 2/2. 14/16 Hosted ITs passed; the other two are Linux-only (HostedWorkspaceConcurrencyIT fails fast with "MySQL acceptance requires physical Linux identity"; HostedWorkspaceStorageGuardMySqlIT is skipped). Neither is touched by the stack. In the first attempt the surefire stage hit ManagedAgentServerIntegrationTest.ignoresLateEnvironmentResultFromAnOlderTurn once at load 39. That test is untouched by the stack and passed 3/3 on its own.
Real stack Jar abad13c7 (≡ 6ff843bc) plus O4-1 base jar 3ee43046 plus main jar 93efe355. Harness and worker dist built at e47fa595, whose TS differs from the head only in one core test file. 300+ real Shell publications, from 70 KiB to 1 GiB.
Retirement seam retire.sh writes the DELETE operation row plus the DELETING status that beginOperation would write. The production SessionLifecycleCoordinator then runs completeOperation → lockDeletion → retire.

Reviewer Test Plan, item by item

Claim Real-stack result
100/100/remainder pages, exact keys, objects outside the catalog kept 252-row publication: pages 100/100/52, one claim generation per page. 255 DELETEs = 255 catalog keys, 0 outside the catalog. A same-prefix decoy, a sibling-prefix decoy and a live Session's objects all survived. 1 GiB publication: 1027 objects in 11 pages, at most 100 per tick, done in 14 s.
Quota and inline data kept until the final confirmation, then released once Held stayed at 259,312,090 until page 3, then 0. released = collected = 262,761,401; the tenant total dropped by exactly that. Only the 2 tool-result manifest copies were cleared; 33 other inline resources (messages, checkpoints) were kept.
Lost delete response +1.9 s collection_retry with quota held. +62.2 s: the same keys were replayed (3 × 204-missing, 6 × 204-removed), then COLLECTED.
Interrupted page / process death SIGKILL with the 50th DELETE in flight. The replacement JVM waited out the 60 s claim, replayed the page as gen 2 and finished at gen 3; quota was released once. The dead JVM's delete landed at +435 s as 204-missing.
SQL confirmation failure Journal-head lock held by another session, lock wait 3 s: CannotAcquireLockException, rollback. Inline copies and quota were kept, and the retry after 60 s re-deleted the same 9 keys.
Replacement owner; an old generation must not confirm Two JVMs on one DB. ① A hung DELETE ends at the SDK's 50 s socket timeout, before the 60 s claim expires, so A defers itself. ② A was SIGSTOPped for ~63 s; its first SQL after SIGCONT was the renew, which returned affected=0, and A stopped. ③ A's confirm was held 75 s in a SQL relay while B collected at gen 2. A's FOR UPDATE then saw COLLECTED and committed no writes; its defer returned affected=0, and the row was unchanged.
A failing publication yields P1 got 403 on every key and retried at +1 / +62 / +122 s, keeping 8,406,850 held bytes. Healthy P2 was collected at +2.5 s.
Versioning, disabled GC, legacy evidence Versioning switched on after startup (fake bucket Enabled and Suspended, real bucket Enabled): 0 deletes. gc-enabled=false: 0 deletes; the observer still reports. A row written by the main jar and then migrated: legacy_write_evidence_missing, kept. O4-1 base jar: the retired output stays RETIRING, because that jar has no collector; after upgrading the same DB, V28 applied in 0.37 s and the output was collected on the first tick.
Real OSS (temporary private bucket, deleted afterwards) 102 real objects deleted in 2 pages; page 1 took 5.5 s (≈55 ms per object). Decoys and the live Session were kept. DeleteObject on a key that never existed returned HTTP 204, so replay is safe on real OSS. Versioning Enabled: fails closed.

collection pages
faults and takeover
fail closed and gating
real OSS

Finding 1: retirement is unreachable for Workspace Sessions

DELETE /v1/agents/sessions/{id} on a Workspace Session returned 409 workspace_unavailable on all 30 attempts (ManagedAgentService.java:684-697 via SessionLifecycleService.java:83, and again at ManagedAgentStore.java:578-579). A legacy Session can be deleted, but a Hosted Workspace Shell turn on it fails with "Hosted Workspace Broker scope does not match the saved Session", so it never has a publication. Open #13135 enables close, not delete, for files-profile Sessions. Suggestion: name the Workspace delete path as an explicit prerequisite in #13090's deployment gate, next to the dead-letter policy.

Finding 2: blocked candidates starve eligible ones

claim() takes LIMIT 1. A blocked row, including one that is still in grace, is pushed now + 60 s and the tick ends. Measured with 151 real publications, grace 120 s, E retired at t0 and 150 more at t0 + 60 s:

head 6ff843bc candidate
E collected after becoming eligible 95.2 s 0.96 s
blocked re-evaluations 57 / 59 per minute (one per tick, each taking the tenant row FOR UPDATE) 151 in total, exactly one per publication

Grace-only smoke run (grace 5 s): collected at +62 s, because the first evaluation inside grace added 60 s. With the default 24 h grace, each retired publication costs about 1440 ticks. Once more than about 60 publications per instance are in grace, the collector is saturated, and eligible publications wait about N seconds each. Per the code, permanently blocked rows (legacy_write_evidence_missing, truncated captures under not_accepted_complete, quarantine) keep the same 60 s cadence forever.

Candidate patch (+49/−33, mostly re-indentation). A grace blocker gets gc_next_at = retired_at + grace, and up to 32 candidates are tried per tick, so a blocked row no longer ends the tick. The 42 collector, retention, scheduling and adapter tests pass, and Checkstyle reports 0. This is not a safety issue and GC is off, so it can be done here or in #13090's capacity gate, but before enabling GC.

blocked candidates

Finding 3: mutation matrix

There were 23 single mutants and 3 combined mutants against 41 unit tests (control green). 12 of the 23 were killed. The survivors:

  • M12: the cleanup at :167 widened to every resource of the Session. On the real stack the code is correct (2 cleared, 33 kept), but no test would notice the WHERE clause widening.
  • M13: used bytes not zeroed. Minor, because held bytes are still zeroed.
  • Candidate test (+29): passes on the head and kills M12 and M13.

Other survivors, none of which is a safety gap:

  • Removing the confirm identity fences (owner, generation, cursor), individually or all together, still passes every test. The same holds for the renew owner and generation fences. The state and expiry conjuncts carry the existing tests. A stale confirm can only follow a fully deleted page, and the catalog no longer changes after retirement, so neither an early release nor a double release follows. I reproduced the rejection on MySQL (②/③ above).
  • The rest are race-only rechecks (M16, M21), the page-level versioning check (M19; the adapter re-checks), and lease cleanup (M22).

mutation

Landing notes

Not covered

  • Linux-only Hosted ITs.
  • Real-OSS permission denial (the only AK available is the account owner's).
  • Multi-instance throughput beyond two JVMs.
  • Windows.

Rig, harness, candidates and raw logs: assets-pr13087@b8414fb8/pr13087

中文版

O4-2 真实环境验证(6ff843bc)

结论:O4-2 自身代码没有阻塞项。 评审测试计划的每一条都在真实栈上成立。真实栈包括:MySQL 8.4.7、本 PR 的 Spring jar、经打包 Hosted Harness 与 Broker worker 执行的真实 Shell、带故障注入的假 OSS,外加一轮真实阿里云 OSS 临时 bucket。结果是:分页有界,只删除 catalog 中的 key,配额只在最后一页确认后释放且只释放一次,重试重复同一组 key,过期 JVM 既不能续约也不能确认。启用 GC 之前需要处理三件事:

  1. 当前栈上没有任何真实路径能让 Shell publication 退役。 公开 API 和 WebShell 的 DELETE 对所有 Workspace Session 都返回 409 workspace_unavailable,而 O2 Shell publication 只存在于 Workspace Session 上。因此我只能通过一个缝来驱动退役(见下文)。这不是 O4-2 的缺陷,但它意味着合入后运行时行为不变,而“从 API 删除到回收”这条端到端路径目前没有归属。
  2. F1:被阻塞的候选占用了每秒一次的 claim 名额。 用 151 个真实 publication 测试,一个已可回收的 publication 在变为可回收后 95.2 s 才被回收;用候选补丁后同一实验只需 0.96 s。在默认 24 h 宽限期下,每个退役 publication 约会被重新评估 1440 次。
  3. 测试缺口(M12): 把 Session 资源清理范围扩大到整个 Session(这会清掉消息和 checkpoint),41 个单测仍然全绿。一个 +29 行的测试可以钉住它。

跑了什么

  • head 变化:从 abad13c7 开始验证,途中 PR 被 force-push 到 6ff843bc。两者差异只有 README 和 HostedHarnessMySqlIT 中的一条断言(SEALED → DELETED)。src/main、runtime-broker、qwencode、packages/cli 和 packages/core 逐字节相同,所以下面的真实栈结论适用于 6ff843bc。
  • 非 Hosted 门禁(abad13c7):388/388 单测 + 25/25 集成测试,Checkstyle 0。
  • 与 main 0a5f518b 试合并:389/389 + 25/25,Checkstyle 0。
  • Hosted 通道(6ff843bc):被改动的 HostedHarnessMySqlIT 2/2 通过,Hosted IT 共 14/16 通过。另外 2 个是仅限 Linux 的测试,本栈没有改动它们。第一次运行时,单测阶段的 ManagedAgentServerIntegrationTest.ignoresLateEnvironmentResultFromAnOlderTurn 在负载 39 下失败过一次;该测试未被本栈改动,单独重跑 3/3 通过。
  • 退役缝:retire.sh 写入与 beginOperation 相同的 DELETE 操作行和 DELETING 状态,之后由生产代码 SessionLifecycleCoordinator 执行 completeOperation → lockDeletion → retire。

评审测试计划逐项结果

  • 分页与精确 key:252 行的 publication 分 100/100/52 三页,每页一个新的 claim generation。255 次 DELETE 正好等于 255 个 catalog key,目录外删除为 0。同前缀诱饵、兄弟前缀诱饵和未退役 Session 的对象都保留。1 GiB publication 有 1027 个对象,分 11 页,每 tick 不超过 100 次删除,14 s 完成。
  • 配额与 inline 数据:held 在第 3 页前一直保持 259,312,090,之后归零,released = collected = 262,761,401,tenant 总量正好减少这么多。只清除了 2 个 manifest 副本,其余 33 个 inline 资源(消息、checkpoint)都保留。
  • 删除应答丢失:+1.9 s 进入 collection_retry,配额保留;+62.2 s 重放同一组 key,之后进入 COLLECTED。
  • 进程中途被杀:第 50 个 DELETE 在途时 SIGKILL。替换 JVM 等旧 claim 过期后以 gen 2 重放整页,最终在 gen 3 完成,配额只释放一次;死进程的那次删除在 +435 s 才落地,结果为 204-missing。
  • SQL 确认失败:另一个会话锁住 journal head,确认因锁等待超时而回滚,inline 副本和配额都保留;60 s 后重试,重新删除同样的 9 个 key。
  • 替换 owner / 旧 generation 不能确认:两个 JVM 共用一个数据库。
    • ① 挂起的 DELETE 在 SDK 的 50 s socket 超时时结束,早于 60 s claim 过期,所以 A 会自己 defer;
    • ② SIGSTOP A 约 63 s,恢复后它执行的第一条 SQL 是 renew,返回 affected=0,A 随即退出;
    • ③ 把 A 的确认语句在 SQL 中继里扣住 75 s,期间 B 以 gen 2 完成回收。A 的 FOR UPDATE 读到 COLLECTED,提交时没有任何写入;它的 defer 返回 affected=0,行状态不变。
  • 失败的 publication 让路:P1 的每个 key 都返回 403,在 +1/+62/+122 s 重试,并保留 8,406,850 字节的 held 配额;健康的 P2 在 +2.5 s 被回收。
  • 版本控制、GC 关闭与旧数据:
    • 启动后开启版本控制(假 bucket 的 Enabled 和 Suspended、真实 bucket 的 Enabled)时,都是 0 次删除;
    • gc-enabled=false 时 0 次删除,观察器仍正常上报;
    • main jar 写入、再经迁移的旧行被 legacy_write_evidence_missing 挡住,未删除;
    • O4-1 基线 jar 没有回收器,退役输出一直停在 RETIRING。同一个库升级后,V28 迁移用时 0.37 s,退役输出在第一个 tick 就被回收。
  • 真实 OSS:
    • 102 个真实对象分两页删除,第一页 5.5 s(约 55 ms/对象),诱饵和未退役 Session 的对象都保留;
    • 对从未存在的 key 执行 DeleteObject 返回 HTTP 204,因此在真实 OSS 上重放是安全的;
    • 开启版本控制后失败关闭;
    • 临时 bucket 已删除。

发现 1:Workspace Session 无法退役

DELETE 30 次尝试全部返回 409,拦截点在 ManagedAgentService.requireLegacyWorkspace 和 ManagedAgentStore.beginOperation:578。legacy Session 可以删除,但它跑不了 Hosted Workspace Shell,因此不会有 publication。#13135 只为 files profile 的 Session 开放了 close,没有开放 delete。建议在 #13090 的部署门禁里把 “Workspace Session 删除路径” 和 dead-letter 策略一起列为前置条件。

发现 2:被阻塞的候选拖住可回收的候选

claim() 只取 LIMIT 1。被阻塞的行(包括仍在宽限期内的行)会被推后 now + 60 s,然后这个 tick 就结束了。

  • head:回收延迟 95.2 s;每分钟 57/59 次无效评估,每次都要对 tenant 行加 FOR UPDATE。
  • 候选补丁:回收延迟 0.96 s;每个被阻塞的 publication 只评估一次。

在 24 h 宽限期下,每个 publication 约要消耗 1440 个 tick。每个实例同时处于宽限期的 publication 超过约 60 个时,回收器就会饱和。按代码,永久被阻塞的行(旧数据、截断的 capture、隔离)会一直按 60 s 节奏被重新评估。候选补丁:宽限期阻塞时把 gc_next_at 设为 retired_at + grace,每个 tick 最多尝试 32 个候选;42 个相关测试通过,Checkstyle 0。这不是安全问题,可以在本 PR 或 #13090 的容量门禁中处理,但应在启用 GC 之前完成。

发现 3:变异测试

共 23 个单变异体和 3 个组合变异体,杀死 12 个单变异体。

  • M12(资源清理范围扩大到整个 Session)存活:真实栈上行为正确,但没有测试钉住。
  • M13(used 未归零)存活,影响较小。
  • 候选测试 +29 行,在 head 上通过,可以杀死 M12 和 M13。
  • 确认和续约的身份栅栏单独去掉或全部去掉都不会让测试失败,但分析后不构成安全缺口,并且我已在 MySQL 上复现了它们的拒绝行为。

合入提示

未覆盖

  • 仅限 Linux 的 Hosted IT;
  • 真实 OSS 的权限拒绝;
  • 两个以上实例的吞吐;
  • Windows。

@yiliang114

Copy link
Copy Markdown
Collaborator

Readiness review of unchanged O4-2 head 6ff843bc140e36ee282f108fb1f2ca6bb92d113b: no remaining actionable correctness blocker in the collector's own change. The independent real-stack report supports exact-key paging, retained decoy keys, claim fences, once-only quota release, real Shell output through 1 GiB and the real OSS two-page case. This readiness check reads the exact current code; it does not claim a new test run. Hosted acceptance in that report was 14/16, with Linux-only boundaries explicitly retained.

The three new findings are dispositioned as follows:

All 14 original threads have recorded dispositions. Bounded exact-head review found no new actionable defect. The PR may now be Ready for review; O4-1 must still land before the stack targets main, and its currently advancing head must be integrated and verified then. V27/V28 must be reconciled against the final landing order. No merge or GC enablement is authorized or performed by this readiness transition.

@yiliang114
yiliang114 marked this pull request as ready for review October 1, 2026 18:45
@doudouOUC
doudouOUC marked this pull request as draft October 2, 2026 03:03
@doudouOUC
doudouOUC deleted the branch codex/managed-tool-result-o4-1 October 2, 2026 05:49
@doudouOUC doudouOUC closed this Oct 2, 2026
@wenshao

wenshao commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Handoff record for whoever reopens/supersedes this O4-2 slice, now that O4-1 (#13084) has merged at e1167c1f9c:

Migration renumbering is mandatory before any new head. main now carries V28__managed_hook_admission_index.sql (#13136), V29__managed_hook_admission_backfill.java (Java migration, #13136), and V30__managed_tool_output_retention.sql (#13084). This branch's V28__managed_tool_output_collection.sql collides with SQL V28 directly, and a naive shift to V29/V30 also fails — V31 is the next free version. Check both migration directories (src/main/resources/db/migration/ and src/main/java/db/migration/); renumbering only against the SQL directory twice produced broken heads during O4-1 (the duplicate-V28 and duplicate-V29 merge trees were each fatal at Flyway startup while git-merging cleanly).

Deferral ledger approved into O4-1, owned by this slice (from the approval on #13084, review 5388632631):

  1. F2 / R2-3 remedy (b) — per-chunk lease.check() + guard.run() cost is deliberately unamortized (measured 327 → ~690 read statements/MiB, 73 → 538 write statements/MiB; the fixed 120 s read budget caps a download at ~28 MiB at 5 ms DB RTT and oversized artifacts fail as committed-200 + truncated body). Accepted only for the observe-only slice; the cost redesign must land here before any deployment enablement, not merely before merge.
  2. Statement-counting bound over the artifact content path (the acceptance criterion the R2-3 finding named) — to be added with the cost work; nothing in CI currently fails if per-chunk re-checks grow.
  3. M21 / M25 negative pins (accepted_complete derivation, admission-object PUT ledger) — the verifier's +2-line M21 candidate from the feat(managed-agent): Protect Session-owned tool output retirement #13084 report thread still applies.
  4. R4 / R5 guard-wiring mutants — adapter-internal guard is pinned, the two production wiring sites are not; the verifier's +21-line candidate (pr13084/r3/harness/candidate-test-retry-guard.patch in the assets branch) kills R4, and R5 needs the same test through the download path. Also consider a per-test @Timeout on the retry tests (mutant R3 hangs the suite via backoff overflow instead of failing).

@doudouOUC

Copy link
Copy Markdown
Collaborator Author

[codex] Restored O4-2 at bea89fcdefc840314738965634a76f827adbcbe5 in replacement #13225, based on main 9478f28731f45ff6df66062e7e05d29e92906f19. GitHub rejected reopening #13087 and changing its closed base, so #13225 supersedes it. O4-3 #13090 remains stacked on the same repaired branch. No merge or GC enablement was performed.

Disposition Action
Fixed SQL+Java migration census is 33 unique versions: recovery V31 retained, close V32, collection V33; rename only, SQL bytes unchanged. Provisional database histories require explicit reconciliation/recreation, never automatic repair.
Fixed / carried forward Bounded due-candidate progress and original grace deadline scheduling, inline-only pages without OSS, original owner/generation defer after expired renewal, and target-publication cleanup regression coverage. Current main reader/writer/deletion protections are preserved.
Fixed F2 download SQL amplification: fill bounded verification reads instead of guarding every raw TLS fragment; remove a duplicate SQL liveness probe while retaining fresh retirement predicates and original before/after guards. Actual SDK/TLS 8 MiB cost reduced from 1122.5 to 197.75 SQL/MiB. At minimum 5 ms/SQL, baseline 40 MiB truncated at 16 MiB after 120 seconds; fixed 40 MiB/100 MiB completed with exact digests in 42.6/108.1 seconds. Range, real 126-second process pause, permission revocation and retirement controls retain no-more-bytes behavior. One original nonrenewable read lease, no auth cache.
Fixed GET SDK retry-hook regression: the strategy never allows SDK replay; it classifies transient GET failures for the adapter's existing bounded guarded loop. Permanent diagnostics retain original OSSException, every retried wire request is guarded, PUT/DELETE remain single SDK attempts, returned body failures do not reopen. 58 real-SDK fault fixtures, including 9 recovery assertions and protected failure controls, passed.
Fixed Recovery fixture now claims a real DB-clock lease before completing retirement; all original 20 cases pass, and matching-owner/generation NULL or expired leases remain rejected. The initial final gate's eight fixture failures were reproduced independently before this test-only fix; no production fence was relaxed.
Deferred after multiple review rounds R2-1/2/4/5/6 and R1-11: shared observer scheduling, additional throughput/index/lock capacity and physical probe amortization remain deployment follow-ups under #12380. Candidate limits do not prove bounded fleet capacity.
Deferred R1-3/R2-9/R3-2: publication warning identity, success logs and DELETING observer visibility. The existing #12380 disposition and separate state/blocker SQL queries remain the ledger.
Partial / deferred R1-8/R1-9/R2-10/11/12/13: sibling cleanup coverage strengthened; remaining quota-field/adapter/live-negative/observer mutation coverage, lease-reap and test-helper refactors remain explicit follow-ups, not claimed fixed.
Deferred / deployment restriction Historical F1 permanent-blocker polling saturation and C3 failure/error mutation pinning. Shortened grace does not recalculate already scheduled deadlines. These concerns are recorded under #12380; GC remains off. M21/M25/R4/R5 remaining mutation pins from the handoff are not silently counted as completed.

Final committed-head gate: build, typecheck, bundle, migrations, source-matching Java dependency installation, complete ordinary MySQL profiles and source-derived report checks passed. Broker unit reports: 593 total, 592 passed and one optional real-worker bundle test skipped; Broker MySQL: 7/7 passed. Server: 487/487 unit and 35/35 MySQL integration passed. Checkstyle 0. Two clean independent code-review passes and self-audit completed; fixture, GET and controlled performance execution snapshots are bound to the final Git objects.

These local macOS/MySQL and controlled HTTPS results do not establish Linux identity guards, full Hosted producer/Broker/Runtime acceptance or current-revision credentialed cloud OSS. The historical maintainer OSS 36/37 run still lacks the delete-denied identity. Public Workspace DELETE remains unavailable, and #13192 remains an independent open PR. Those are deployment/readiness boundaries, not green checks inferred from admin-only CI. Current #13225 product CI is separately pending; physical GC is disabled and grace remains 24 hours.

中文:O4-2 已以 bea89fcd 在当前 main 上恢复为替代草稿 #13225,接替 GitHub 拒绝重开的 #13087;O4-3 #13090 继续使用同一已修复分支。已修复迁移撞号、带入候选推进/inline/过期 claim 修复、下载 SQL 放大和 GET 重试 hook 回退,并修正真实领取租约的 Recovery fixture。完整本地最终门禁通过:Broker 592 个单元通过、1 个可选 bundle 用例跳过,7 个 MySQL 集成通过;Server 单元 487/487、MySQL 集成 35/35,Checkstyle 为零。58 个 SDK 故障 fixture、20 个 Recovery 用例及受控 HTTPS 成本/权限/退役/真实租约暂停证据均绑定最终源码。

上表中的容量、诊断和 mutation 建议逐项明确延期到 #12380,部分采纳不算全部修复。永久阻塞轮询饱和、已排期宽限期缩短不会即时生效、C3 和 handoff 剩余 mutation 缺口继续保留。Linux/完整 Hosted、真实 OSS 删除拒绝身份、公开 Workspace DELETE 及独立开放的 #13192 不被本地绿灯替代。没有自动合并;物理 GC 关闭,部署宽限保持 24 小时。

qwen-code-dev-bot pushed a commit to jjj138138/qwen-code that referenced this pull request Oct 3, 2026
…nLM#13225)

* feat(managed-agent): Collect retired tool output safely

Co-authored-by: Qwen-Coder <[email protected]>

* chore(managed-agent): Sync retention formatting and verify collection

Co-authored-by: Qwen-Coder <[email protected]>

* codex: address PR review feedback (QwenLM#13087)

* codex: address PR review feedback (QwenLM#13087)

---------

Co-authored-by: Qwen-Coder <[email protected]>
Co-authored-by: Shaojin Wen <[email protected]>
JadeCong pushed a commit to CloudEngineHub/qwen-code that referenced this pull request Oct 3, 2026
…nLM#13090)

* feat(managed-agent): Collect retired tool output safely

Co-authored-by: Qwen-Coder <[email protected]>

* chore(managed-agent): Sync retention formatting and verify collection

Co-authored-by: Qwen-Coder <[email protected]>

* codex: address PR review feedback (QwenLM#13087)

* codex: address PR review feedback (QwenLM#13087)

* test(managed-agent): Add tool output collection deployment gates

Co-authored-by: Qwen-Coder <[email protected]>

* codex: address PR review feedback (QwenLM#13090)

* docs(managed-agent): Preserve deployment gate identifiers

* fix(managed-agent): Advance eligible tool output collection

* codex: address PR review feedback (QwenLM#13090)

* chore(ci): Drop merged O4-2 base branch from SDK Java trigger

This PR's base moved to main after QwenLM#13225 merged; no PR will target the
merged codex/managed-tool-result-o4-2 branch again.

Co-authored-by: Qwen-Coder <[email protected]>

* docs(managed-agent): Refresh operations runbook for merged O4-2

Audit found the runbook still pointed at the pre-merge baseline SHA and
the old collection V33, contradicting the lifecycle design doc in the
same diff. State the merged stack without a volatile SHA, use the
current V30/V31/V32/V33(definitions)/V34(collection) sequence in both
languages, and say the SDK Java workflow targets main now that the
stacked base branch is merged.

Co-authored-by: Qwen-Coder <[email protected]>

---------

Co-authored-by: Qwen-Coder <[email protected]>
Co-authored-by: yiliang114 <[email protected]>
Co-authored-by: Shaojin Wen <[email protected]>
wenshao added a commit to wenshao/qwen-code that referenced this pull request Oct 7, 2026
…#13565)

Move the merged-PR history of the Managed Agent dual-path proposal
(QwenLM#12380) out of the issue body into a bilingual ledger under
docs/design/. The issue body had come within 10 KB of GitHub's 256 KiB
limit, so the body will keep only the delivery snapshot and the open
PRs, and merged rows move here rewritten to their merged final state.

The ledger carries every merged row the issue tracked up to its
2026-10-03 reconcile (adding the missing merge commit to the five
earliest rows and normalising the Chinese state cells), rewrites the
eight rows the issue still listed as open although their PRs had merged
(QwenLM#13141, QwenLM#13166, QwenLM#13174, QwenLM#13210, QwenLM#13214, QwenLM#13217, QwenLM#13218, QwenLM#13247), and
adds rows for the 48 managed-agent PRs merged between that reconcile and
main 0c13502 that had no row yet. PRs closed without merging (QwenLM#13087,
QwenLM#13336) sit in their own table. Later merges land at the next reconcile.

Co-authored-by: wenshao <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants