Repository navigation
feat(managed-agent): Add W0e terminal recovery fences - #12839
Conversation
Baseline reproduction and design verification
中文说明基线复现与设计验证
|
Local real-environment verification: head
|
| Design claim | How I checked | Result |
|---|---|---|
§1: a host crash pins the LOST generation behind the unsettled call (#12670) |
Existing ProcessCrashFaultGateTest (FG3): 6/6 pass, including aHostCrashPinsTheLostGenerationBehindTheUnsettledCall (7.3 s), which reproduces the author's baseline |
✅ |
§1: local-process ownership lives in memory, a restarted Broker sees UNKNOWN even after the worker exited, the deadline is 4 × 30 s leases, and a timeout leaves the binding READY |
P3: after an orderly close the worker exits in ~0.4 s and the binding stays READY. After restart, every warm returns runtime_broker_reconcile_timeout at 4.07–4.37× the lease (12/12, with 2 s and 5 s leases), and no worker starts. EmbeddedRuntimeBroker.LEASE = 30 s |
✅ |
| §1: orderly shutdown terminates workers without retiring bindings, and a Broker crash can leave a worker alive | P3, plus FG3 theProductionProvisionerCannotAdoptAfterARestart. The worker has no parent-death watch |
✅ |
§1: reconcileExecution answers IN_FLIGHT for an orphaned EXECUTING record |
FG3 and P2b | ✅ |
| §4: Session validation and execution insertion are separate, so a late insertion can land after loss | P2b: a stale Broker inserted a new row into the generation after another Broker committed LOST, with no parking needed |
✅ stronger than stated |
| §5: the stdin boot barrier means an incomplete boot never admits tools | P5 on the bundled worker. An empty or partial boot exits 1 in ~0.36 s. When a launcher holding the only write end is SIGKILLed mid-boot, the worker is gone 90 ms later with no ready line. A launcher that stays alive but silent keeps a worker PID with no identity until the 30 s boot timeout, a window the §5 per-seed lock must cover | ✅ |
§6: the busy check is keyed only on runtimeSessionId, and the placement key includes cwd, capability, isolation and provisioner |
Code: JdbcToolExecutionRepository.hasActiveByRuntimeSession, JdbcRepositorySupport.requestKey |
✅ |
| §6: old binaries cannot read the new enum value | P4 (details under F3) | ✅ |
| Docs | Prettier is clean and all relative links resolve. EN and zh-CN match: 10/10 headings, 36/36 table rows, identical code identifiers and issue references | ✅ |
F1: a worker dying under a live Broker already reuses the placement next to an escaped writer
When a worker dies while its Broker keeps running, the next operation on that lease goes through invalidateBinding → failBinding. That retires the generation as FAILED without checking for unsettled executions, and the next warm provisions generation 2 on the same placement.
P1 (legacy placement, production provisioner, one Broker JVM): a Shell call started a writer that double-forked into its own session (parent PID 1). Then the whole worker tree was SIGKILLed. The results:
- The call went
UNKNOWNand generation 1 wentFAILED. - Generation 2 ran a new call in the same directory.
- The escaped writer landed 16 more writes after the replacement's first write, identically in 3/3 runs.
The design already wants this blocked. §3 says "general worker-crash reclamation needs a killable isolation domain", the §3 table says "Worker PID gone … → Block", and §9 says the correct result for worker-only death with Shell descendants is explicit blocking. But the current code already recovers from that case automatically, and the design never says so:
- §1 describes only the restart paths.
- §4 scopes its
FAILED/releaseUnusableSessionrule to "this recovery flow". - No slice owns changing this path.
- FG3
aWorkerKilledMidExecutionLeavesItUnknownWithoutEvidencepins the replacement ("a new one serves new work").
Suggestion:
- Add this path to §1.
- Assign it to a slice in §7.
- State the user-visible consequence. If it blocks until a stop proof exists, every legacy worker crash (OOM, SIGKILL) leaves that placement unusable until W0e-3.
- List that FG3 gate among those that must flip, and add an "Old writer" case without a restart to §8.
F2: a stale Broker admits into, then retires, a pinned LOST generation
This needs two live Brokers sharing the store, the concurrent-Broker case that §4 and §5 already consider. Broker B committed LOST and answered runtime_broker_runtime_lost (#12670's pin). Broker A still held a Session context for that generation:
- P2: A's
releasereturned409 runtime_session_busy. By thenreleaseUnusableSessionhad already runinvalidateBinding, so the binding wasFAILED. B's nextwarmprovisioned generation 2 and ran a call while generation 1's call was stillUNKNOWN. - P2b: A's
createwith a new key was accepted, adding a row on generation 1 after the loss commit. Its dispatch then retired the binding the same way. - Reconcile polls do not do this.
requireAnswerableBindingrejectsLOSTbefore the liveness check, so the poll returnsIN_FLIGHT, then409 runtime_execution_evidence_unavailable, and the binding staysLOST.
Root cause: claimOperation accepts any isActive() binding, including LOST, and compareAndSet refuses only reactivation. So LOST → FAILED is legal from any stale context. The late insertion confirms §4's premise; the retirement is new.
Suggestion:
- In W0e-1, let
LOSTleave only through the evidence path: nofailBindingfromLOST, and no retirement while unsettled rows remain. - Add a §8 row: "a Broker with a stale Session context cannot insert into or retire a
LOSTgeneration". - F1 and F2 share this fix point, so one W0e-1 change can close both.
F3 (minor): the settled/active split is also a set of SQL literals, and old Brokers misreport abandoned calls
§2 frames the split around isSettled(). But the JDBC pin and busy check are SQL literals, execution_state <> 'SETTLED' in hasActiveByBinding and hasActiveByRuntimeSession, while the in-memory repositories use !isSettled().
P4: I set a settled row to ABANDONED and had the merged Broker read it:
release→runtime_session_busy. The SQL pin fails closed, which is good.getandcancelthrowNo enum constant …ABANDONED. Over HTTP,sendErrormaps that to400 runtime_broker_invalid_request(from code; not run over HTTP).reconcile→503 runtime_execution_reconcile_failed.- A retry with the same key →
409 runtime_execution_conflict("execution identity is already in use").
This supports the design's statement that mixed old and new Brokers are unsupported. Suggestion: name the SQL predicates in the §6 table. Also note that during a rollout an old Broker reports an abandoned call as an invalid request or identity conflict, not as a terminal result.
Not verified
- Workspace-bound (W0c-3) placements. From code,
WorkspaceRuntimeTransport.releaseclears the storage holder only after the worker acknowledgesactivateWorkspace(false), andreleaseUnusableSessionnever calls the transport. So a replacement's acquire should getworkspace_busy(retryable); that holder is the second boundary §2 says legacy placements lack. I did not run the Spring + MySQL stack, and the rig's managed transport skips holder checks. - MySQL/MariaDB repositories, Linux, Windows, and a real host reboot.
Evidence: wenshao/qwen-code@2ed8809 → pr12839/ holds the four cards, the probe harness (harness/W0eProbeTest.java, boot-launcher.pl) and the raw logs of every run.
中文版
本地真实环境验证:head 353b6526e8
结论: 作为设计文档可以合入。本 PR 不改运行时行为,文档检查干净,我能实测的基线论断在真实进程上全部成立。建议合入前补进 F1 和 F2:两者都是当前代码中已经绕过本设计新引入的物理复用门禁的路径,会改变 W0e-1 的范围和需要翻转的 FG3 门禁。F3 是措辞精度问题。
环境: macOS、JDK 21、Node 24。构建 PR head(pnpm 安装,然后 npm run build && npm run bundle,得到 dist/cli.js 0.24.6)。沿用现有 Stage F 装置:真实 Broker JVM、打包后的 managed-runtime-worker、TCP 上的文件 H2。从基线 e0b8bea 到当前 main b4ca6ba9,runtime-broker、worker 入口、EmbeddedRuntimeBroker 和 WorkspaceRuntimeTransport 均未改动,所以结论同样适用于 main。新探针是可直接放入的 @Tag("fault-gate") 测试类,每个都通过断言当前行为而通过。共跑了三轮完整探针,另加三轮 P2b,各轮结果一致。
已核实的论断
| 设计论断 | 核实方式 | 结果 |
|---|---|---|
§1:宿主崩溃后,未终结调用把 LOST 代数钉住(#12670) |
现有 ProcessCrashFaultGateTest(FG3)6/6 通过,含 aHostCrashPins…(7.3 s),复现了作者的基线 |
✅ |
§1:local-process 归属只在内存中,重启后的 Broker 即使 worker 已退出也看到 UNKNOWN;截止期为 4 × 30 s 租约;超时后绑定保持 READY |
P3: 有序关闭后 worker 约 0.4 s 退出,绑定仍是 READY。重启后每次 warm 都在租约的 4.07–4.37 倍时返回 runtime_broker_reconcile_timeout(12/12,租约 2 s 和 5 s),也不启动新 worker。EmbeddedRuntimeBroker.LEASE = 30 s |
✅ |
| §1:有序关闭会终止 worker 但不退役绑定;Broker 崩溃可能留下存活 worker | P3,以及 FG3 theProductionProvisionerCannotAdoptAfterARestart;worker 没有父进程死亡检测 |
✅ |
§1:对孤立的 EXECUTING 记录,reconcileExecution 返回 IN_FLIGHT |
FG3 与 P2b | ✅ |
| §4:Session 校验与执行插入分离,迟到插入可落在丢失之后 | P2b: 另一个 Broker 提交 LOST 之后,旧 Broker 仍向该代数插入了新行,而且无需暂停 |
✅ 比文中描述更强 |
| §5:stdin boot 屏障保证不完整 boot 永不准入工具 | 在打包 worker 上跑 P5:空或半截 boot 约 0.36 s 以 exit 1 退出;持有唯一写端的启动器在 boot 中途被 SIGKILL,worker 90 ms 后退出且没有 ready 行;启动器存活但不写入时,worker PID 在没有身份记录的情况下存活到 30 s boot 超时,这个窗口需要 §5 的按 seed 锁覆盖 | ✅ |
§6:busy 检查只用 runtimeSessionId;placement key 包含 cwd、capability、isolation 和 provisioner |
代码:hasActiveByRuntimeSession、requestKey |
✅ |
| §6:旧二进制读不了新枚举值 | P4(详见 F3) | ✅ |
| 文档 | Prettier 干净,相对链接全部可达;中英文一致:标题 10/10、表格行 36/36,代码标识与 issue 引用相同 | ✅ |
F1:Broker 存活时 worker 死亡,已会在逃逸写入者旁复用 placement
worker 死亡而 Broker 仍在运行时,该租约上的下一次操作会走 invalidateBinding → failBinding,在不检查未终结执行的情况下把该代数退役为 FAILED,下一次 warm 随即在同一 placement 上创建第 2 代。
P1(legacy placement、生产 provisioner、单个 Broker JVM):一个 Shell 调用启动了一个写入者,它双重 fork 进入独立会话(父进程为 PID 1),然后整棵 worker 进程树被 SIGKILL。结果:
- 调用变为
UNKNOWN,第 1 代变为FAILED。 - 第 2 代在同一目录执行了新调用。
- 逃逸写入者在替代代数首次写入之后又写入 16 次,3/3 轮一致。
设计本身希望阻断这种情况:§3 写明"通用 worker 崩溃回收需要可整体终止的隔离域",§3 表格写"worker PID 消失 … → 阻断",§9 写明带 Shell 后代的纯 worker 死亡的正确结果是明确阻断。但当前代码已经在自动做这种恢复,而设计没有提到:
- §1 只描述了重启路径。
- §4 把
FAILED/releaseUnusableSession规则限定在"本恢复流程"。 - 没有切片负责修改这条路径。
- FG3
aWorkerKilledMidExecutionLeavesItUnknownWithoutEvidence固定的正是这种替代("a new one serves new work")。
建议:
- 在 §1 补上这条路径。
- 在 §7 指定由哪个切片修改它。
- 写明用户可见的后果:如果阻断到有停止证明为止,每次 legacy worker 崩溃(OOM、SIGKILL)都会让该 placement 一直不可用,直到 W0e-3。
- 把该 FG3 门禁列入必须翻转的清单,并在 §8 增加不涉及重启的"旧写入者"用例。
F2:旧 Broker 会向已钉住的 LOST 代数准入新调用,并将其退役
需要两个共享存储的存活 Broker,这正是 §4 和 §5 已经考虑的并发 Broker 场景。Broker B 提交 LOST 并返回 runtime_broker_runtime_lost(#12670 的钉住)。Broker A 仍持有该代数的 Session 上下文:
- P2: A 的
release返回409 runtime_session_busy。此时releaseUnusableSession已经执行了invalidateBinding,绑定已变为FAILED。B 的下一次warm创建了第 2 代并执行了新调用,而第 1 代的调用仍是UNKNOWN。 - P2b: A 用新 key 的
create被接受,在丢失提交之后向第 1 代新增了一行;它的 dispatch 随后以同样方式退役了该绑定。 - 对账轮询不会这样。
requireAnswerableBinding在存活检查之前就拒绝LOST,轮询先返回IN_FLIGHT,再返回409 runtime_execution_evidence_unavailable,绑定保持LOST。
根因: claimOperation 接受任何 isActive() 绑定(包括 LOST),而 compareAndSet 只拒绝重新激活,所以任何旧上下文都能合法地执行 LOST → FAILED。迟到插入印证了 §4 的前提;退役是新发现。
建议:
- W0e-1 中让
LOST只能经证据路径离开:不允许从LOST执行failBinding,存在未终结行时也不允许退役。 - 在 §8 增加一行:"持有旧 Session 上下文的 Broker 不能向
LOST代数插入,也不能将其退役"。 - F1 和 F2 共享这个修复点,一次 W0e-1 改动即可同时关闭两者。
F3(次要):结算/活跃的区分也写在 SQL 字面量里,旧 Broker 会误报已放弃的调用
§2 围绕 isSettled() 阐述这一区分,但 JDBC 的钉住与 busy 检查是 SQL 字面量:hasActiveByBinding 和 hasActiveByRuntimeSession 里的 execution_state <> 'SETTLED';内存仓储则用 !isSettled()。
P4: 把一条已结算的行改为 ABANDONED,由已合入的 Broker 读取:
release→runtime_session_busy。SQL 钉住失败时关闭,这是好的。get和cancel抛出No enum constant …ABANDONED。经 HTTP 时,sendError会把它映射为400 runtime_broker_invalid_request(依据代码,未经 HTTP 实跑)。reconcile→503 runtime_execution_reconcile_failed。- 同 key 重试 →
409 runtime_execution_conflict("execution identity is already in use")。
这支持了设计中"不支持新旧 Broker 混用"的说法。建议: 在 §6 表格中写明这些 SQL 谓词;并注明升级期间旧 Broker 会把已放弃的调用报告为无效请求或身份冲突,而不是终态结果。
未验证
- Workspace 绑定(W0c-3)的 placement。按代码,
WorkspaceRuntimeTransport.release只在 worker 确认activateWorkspace(false)后才清除存储 holder,而releaseUnusableSession从不调用 transport,所以替代代数的 acquire 应得到workspace_busy(可重试);这个 holder 正是 §2 所说 legacy placement 缺少的第二层边界。我没有跑 Spring + MySQL 全栈,装置的 managed transport 也跳过了 holder 检查。 - MySQL/MariaDB 仓储、Linux、Windows,以及真实宿主机重启。
证据: wenshao/qwen-code@2ed8809 → pr12839/ 下有四张证据图、探针装置(harness/W0eProbeTest.java、boot-launcher.pl)以及每一轮的原始日志。
W0e-1 implementation verificationImplementation commit: Results
Environment: macOS, Java 21, Node.js, real Broker JVMs and bundled worker processes, file-backed H2 and isolated MySQL 8.4.11. MySQL checks used the existing integration profiles. No model invocation was needed. Windows, Linux and MariaDB were not run locally. Before / after and regression evidence
Review feedback and auditsThe F1/F2/F3 real-environment feedback is addressed. F1's ordinary worker-loss path now fences LOST and blocks replacement; F2's stale Broker cannot admit new rows or retire LOST through a generic failure/release path. The prior replacement fault-gate expectations have been changed accordingly. F3's SQL terminal predicates and unsupported old-binary behavior are recorded in both design languages. Review findings were reproduced, fixed and reverified. The last two independent open-ended full reviews found no new confirmed defect; each covered the complete implementation diff, all new files, the final bilingual supplement and downstream callers. Coordinator self-review also found no new issue. These were native Codex reviews: the repository Workflow runner was unavailable in this session, so this is not an automated Qwen gate receipt or maintainer approval. Limits and earlier failuresOne earlier Stage F startup attempt did not reach its parked attestation request within 45 seconds. The unchanged case passed an isolated rerun and the final full run; the original timeout remains unexplained and is retained in the local report. A separate run was intentionally interrupted because it overlapped a build. The first MySQL invocation lacked test JVM properties and never reached the database; the corrected invocation ran both real MySQL profiles successfully. This verifies W0e-1 evidence consumption and safe blocking. It does not certify a production writer-stop source, actual host reboot, escaped-writer containment, production restart adoption or Workspace crash-holder reclamation. Those remain W0e-2/3; complete Workspace execution capability remains disabled. After worker-only loss, a production placement deliberately remains blocked until independent stop proof exists. Upgrade all Broker readers before writing ABANDONED; mixed old/new binaries are unsupported. 中文说明W0e-1 实现验证实现提交: 结果
环境:macOS、Java 21、Node.js、真实 Broker JVM 与打包 worker 进程、文件 H2 和隔离 MySQL 8.4.11。MySQL 使用现有集成 profile,无需模型调用。本地未运行 Windows、Linux 或 MariaDB。 前后对照与回归证据
评审反馈与审计已处理 F1/F2/F3 真实环境反馈:F1 的普通 worker 失联路径现在标记 LOST 并阻止替换;F2 的旧 Broker 无法准入新行,也不能通过通用失败/释放路径退役 LOST。原有自动替换故障门禁已改为相应阻断断言。F3 的 SQL 终态谓词与旧二进制不兼容行为已写入双语设计。 审计发现均已复现、修复并重新验证。最后两轮独立无方向完整审查未发现新的可确认缺陷,均覆盖全部实现差异、新文件、最终双语补充及下游调用方;协调者自审也未发现新问题。本次使用原生 Codex 审查,仓库 Workflow runner 在当前会话不可用,因此不代表自动化 Qwen gate receipt 或 maintainer 批准。 限制与早期失败较早一次 Stage F 启动测试在 45 秒内未到达暂停的 attestation 请求;未修改该案例,单独重跑和最后完整运行均通过。原超时原因仍未确定,已保留在本地报告。另一次运行因与 build 重叠而主动中断。首次 MySQL 调用缺少测试 JVM 参数,未进入数据库;修正调用后已实际跑通两个 MySQL profile。 本次证明 W0e-1 的证据消费与安全阻断,不证明生产停写来源、真实宿主重启、逃逸写入者隔离、生产重启接管或 Workspace 崩溃 holder 回收。这些仍属于 W0e-2/3,完整 Workspace 执行能力继续关闭。仅 worker 失联后,生产 placement 会保持阻断,直到独立停写证明成立。写入 ABANDONED 前需要升级所有 Broker 读取方,不支持旧新二进制混用。 |
Round 2 real-environment verification: head
|
| Suite | Result |
|---|---|
runtime-broker |
374 run, 0 failures, 1 skip. The skip is the real-worker case; with -Dqwen.runtime.worker.bundle it passes 15/15 |
Stage F fault gates (-Pfault-gates) |
28/28 |
managed-agent-server |
112/112 |
| TS provider suites | 44/44 |
These match the author's numbers.
Round-1 findings
| Round 1 | Status on 727aafb4 |
|---|---|
| F1: worker death under a live Broker reused the placement beside an escaped writer | ✅ fixed (B1). The binding goes LOST and the call ABANDONED. reconcile, get and a same-key retry all return the ABANDONED receipt; release is refused and no replacement worker starts. The escaped writer still runs, but nothing new is placed beside it. |
F2: a stale Broker admitted into, then retired, a LOST generation |
✅ fixed (C1). The stale release returns 503 runtime_reconciliation_required and the stale create returns 409 runtime_admission_closed. The binding stays LOST. |
| F3: SQL pin literals and old-Broker behaviour | ✅ addressed. The JDBC predicates are now NOT IN ('SETTLED','ABANDONED'), and both design languages document the old-Broker behaviour. |
F6 (blocker): Flyway V14 collides with main
#12840 landed V14__managed_event_replay.sql and V15__managed_event_identity.java on main at 14:18Z. This PR adds V14__runtime_loss_evidence.sql. git merge-tree reports a clean merge because the filenames differ, but Flyway refuses duplicate versions.
On the test merge 727aafb4 + 36710ff6, the managed-agent-server suite has 137 run and 72 errors across 9 classes, all from FlywayException: Found more than one migration with version 14: every Spring context fails.
Renaming the PR's file to V16__runtime_loss_evidence.sql on the same merge makes mvn clean test pass 137/137; the upgrade tests pin pre-change targets 11/13 only. PR CI's "SDK Java" job ran at 13:20Z, before #12840 merged, so its green check cannot see this. §6a's "Flyway V14" wording needs the same change.
F4 (blocker): one failed attestation of a healthy worker becomes a permanent, tenant-wide LOST
Every warm/acquire on a live binding calls provisioner.confirm, which is an HTTP attestation. Any failure there now runs invalidateBinding, which sets READY → LOST with no loss evidence and no longer releases the worker. A LOST binding can leave only through recoverLost, which requires stop proof, and no production stop-proof producer exists. I injected one fault on one /attest exchange with the rig's FaultProxy; the worker process was never touched.
- A1 (one connection reset). The base head recovers:
FAILED, worker released, generation 2 on the nextwarm. On727aafb4:- The binding goes
LOSTand every laterwarmreturns503 runtime_broker_runtime_lost(retryable: true), even though 4 further attest exchanges reach the same live worker without a fault. - The worker stays alive and is never released.
createon the existing Session returns409 runtime_admission_closed.- A new placement for another Workspace of the same tenant returns
409 runtime_placement_recovery_required; tenant-b is unaffected.
- The binding goes
- A3 (the worker is only slow: one attest response arrives after the 10 s rig timeout; production uses 30 s): same result.
- A2 (managed-agent-server's default
isolationClass: session, so each Hosted Session is its own placement). After one fault on harness-1, the existing harness-2 placement keeps working. But every new Session (harness-3, harness-4) gets409 runtime_placement_recovery_required. On the base head all wereREADY.
§3 row 1 says a network failure or timeout keeps the result unresolved and blocks reuse. Here a timeout against a process the Broker still owns, and can re-attest, becomes an irreversible loss. The tenant-wide legacy guard then turns it into a tenant outage.
Suggestion:
- Record loss only when the provisioner observes the owned process has exited;
observe()already producesJOURNAL_LOSTevidence for exactly that case. - On an attestation error from an alive owned process, fail the call and drop only the local route, so the next
warmre-attests. - Let a
LOSTbinding without loss evidence return toREADYwhen the exact identity re-attests. This is what §8's "Worker alive after Broker restart" row asks for.
F5 (decide before enabling): upgrade locks out tenants with any historical crash
Before this PR, every worker crash or attestation failure wrote a FAILED binding that keeps its seed. blocksPlacement now treats seeded FAILED rows like LOST, which §6a states deliberately.
U1: the baseline Broker handled one worker crash the old way, leaving gen1 FAILED, gen2 READY. After the new initializer ran on that data (rows preserved), a new placement for tenant-a returned 409 runtime_placement_recovery_required; tenant-b was READY. Nothing in production can clear this.
Before any deployment that has enabled the Runtime Broker upgrades, it needs either an operator procedure (verify the host, then clear or retire those rows) or a narrower historical check. Also add the consequence to the rollout notes: every tenant with a pre-upgrade crash cannot create new placements after upgrade.
N1 (minor)
runtime_broker_runtime_lost and runtime_reconciliation_required are 503 with retryable: true, yet without stop proof they never clear. A caller that honours retryable will loop. runtime_placement_recovery_required is already 409 with retryable: false.
Not verified
- MySQL/MariaDB locally. CI's MariaDB and MySQL jobs passed on the pre-feat(managed-agent): Implement event replay (Stage D3) #12840 merge ref.
- The Spring holder path under F4.
- Linux and Windows.
- The TS
terminal/reasonmetadata has no production caller yet (only the provider and its tests), which is fine for W0e-1.
Evidence: wenshao/qwen-code@9c3f17f → pr12839/r2/ holds the three cards, harness/W0eR2ProbeTest.java (runs unchanged on both heads), the raw logs of all runs, and results/suites-summary.txt, which covers the test-merge and V16 runs.
中文版
第二轮真实环境验证:head 727aafb482(W0e-1 实现)
结论: 暂不宜合入。W0e-1 的屏障本身有效:第一轮的 F1、F2 在真实进程中已修复,PR head 上各测试套件全部通过。有两项阻塞合入:
- F6(机械性): 本 PR 的 Flyway
V14与 main(feat(managed-agent): Implement event replay (Stage D3) #12840)的V14冲突。与当前 main 试合并后,managed-agent-server 测试137 个中 72 个报错;把 PR 文件改名为V16__…即可修复(137/137)。 - F4(行为): 对 Broker 仍持有的存活 worker 做一次失败或过慢的 attestation,就会让绑定永久变为
LOST。legacy placement 下这会阻止整个租户新建 placement;在默认isolationClass: session下,该租户的每个新 Hosted Session 都会被拒绝。
F5(升级悬崖)需要在任何启用 Runtime Broker 的部署升级之前做出决策。N1 为次要问题。
环境: 与第一轮相同:macOS、JDK 21、打包的 dist/cli.js、真实 Broker JVM、TCP 上的 H2。新写了一个只记日志的探针文件,原样在仅设计的 head 353b6526 和 727aafb4 上各跑一遍做 A/B。head 上重复跑了 3 轮,结果一致。U1 探针先用基线 Broker 自己的 classpath 在基线 schema 的数据库上运行,再在同一份数据上运行新的 schema 初始化器和新 Broker。
测试套件(PR head)
| 套件 | 结果 |
|---|---|
runtime-broker |
374 个,0 失败,1 个跳过。跳过的是真实 worker 用例,加上 -Dqwen.runtime.worker.bundle 后 15/15 通过 |
Stage F 故障门禁(-Pfault-gates) |
28/28 |
managed-agent-server |
112/112 |
| TS provider 测试 | 44/44 |
与作者的数据一致。
第一轮发现的状态
| 第一轮 | 在 727aafb4 上 |
|---|---|
| F1:Broker 存活时 worker 死亡,会在逃逸写入者旁复用 placement | ✅ 已修复(B1)。绑定变为 LOST,调用变为 ABANDONED;reconcile、get 和同 key 重试都返回 ABANDONED 回执;释放被拒绝,也不会启动替代 worker。逃逸写入者仍在运行,但它旁边不会再放置新工作。 |
F2:旧 Broker 会向 LOST 代数准入新调用并将其退役 |
✅ 已修复(C1)。旧 Broker 的 release 返回 503 runtime_reconciliation_required,create 返回 409 runtime_admission_closed,绑定保持 LOST。 |
| F3:SQL 钉住字面量与旧 Broker 行为 | ✅ 已处理。JDBC 谓词改为 NOT IN ('SETTLED','ABANDONED'),双语设计文档都已写明旧 Broker 的行为。 |
F6(阻塞):Flyway V14 与 main 冲突
#12840 于 14:18Z 向 main 合入了 V14__managed_event_replay.sql 和 V15__managed_event_identity.java;本 PR 新增的是 V14__runtime_loss_evidence.sql。由于文件名不同,git merge-tree 显示合并无冲突,但 Flyway 不接受重复版本号。
在试合并 727aafb4 + 36710ff6 上,managed-agent-server 测试共 137 个、9 个类中 72 个报错,全部来自 FlywayException: Found more than one migration with version 14:所有 Spring 上下文都起不来。
在同一试合并上把 PR 文件改名为 V16__runtime_loss_evidence.sql 后,mvn clean test 137/137 通过;升级测试只固定了迁移前目标 11/13。PR CI 的 "SDK Java" 任务在 13:20Z 运行,早于 #12840 合入,所以它的绿灯看不到这个问题。§6a 中的 "Flyway V14" 字样也需要同步修改。
F4(阻塞):对健康 worker 的一次 attestation 失败会变成永久的、租户级的 LOST
存活绑定上的每次 warm/acquire 都会调用 provisioner.confirm,这是一次 HTTP attestation。现在这里的任何失败都会执行 invalidateBinding:在没有丢失证据的情况下把 READY → LOST,而且不再释放 worker。LOST 绑定只能经 recoverLost 离开,这需要停写证明,而生产环境中没有任何停写证明的来源。我用装置的 FaultProxy 只对一次 /attest 交互注入故障,完全没有碰 worker 进程。
- A1(一次连接重置)。基线 head 能恢复:
FAILED、释放 worker,下一次warm得到第 2 代。在727aafb4上:- 绑定变为
LOST,之后每次warm都返回503 runtime_broker_runtime_lost(retryable: true),尽管之后又有 4 次 attest 交互无故障地到达了同一个存活 worker。 - 原 worker 一直存活且从未被释放。
- 在已有 Session 上
create返回409 runtime_admission_closed。 - 同租户另一个 Workspace 的新 placement 返回
409 runtime_placement_recovery_required;tenant-b 不受影响。
- 绑定变为
- A3(worker 只是慢:一次 attest 响应超过装置的 10 s 超时才到达;生产环境为 30 s):结果相同。
- A2(managed-agent-server 默认
isolationClass: session,每个 Hosted Session 各占一个 placement)。harness-1 发生一次故障后,已有的 harness-2 placement 仍可用,但每个新 Session(harness-3、harness-4)都得到409 runtime_placement_recovery_required。基线 head 上它们全部是READY。
§3 第 1 行写的是:网络失败或超时时结果保持未决、阻断复用。而这里,针对 Broker 仍持有、并且可以重新 attest 的进程的一次超时,变成了不可逆的丢失;租户级 legacy 守卫再把它放大成整个租户的故障。
建议:
- 只有在 provisioner 观察到自己持有的进程已经退出时才记录丢失;
observe()在这种情况下本来就会产出JOURNAL_LOST证据。 - 存活进程的 attestation 出错时,只让本次调用失败并丢弃本地路由,让下一次
warm重新 attest。 - 允许没有丢失证据的
LOST绑定在同一身份重新 attest 成功后回到READY,这正是 §8 "Broker 重启后 worker 存活"一行的要求。
F5(启用前需决策):升级会把有过崩溃记录的租户锁住
本 PR 之前,每次 worker 崩溃或 attestation 失败都会写入一条保留 seed 的 FAILED 绑定。blocksPlacement 现在把带 seed 的 FAILED 行当作 LOST 处理,这是 §6a 有意为之。
U1: 基线 Broker 按旧方式处理了一次 worker 崩溃,留下 gen1 FAILED、gen2 READY。新初始化器在这份数据上运行后(行保留完好),tenant-a 的新 placement 返回 409 runtime_placement_recovery_required,tenant-b 为 READY。生产环境中没有任何机制能清除这种状态。
已经启用 Runtime Broker 的部署在升级前,需要一套运维流程(先核实宿主,再清除或退役这些行),或者一个更窄的历史检查。另外应在升级说明中写明后果:升级前有过崩溃的租户,升级后都无法新建 placement。
N1(次要)
runtime_broker_runtime_lost 和 runtime_reconciliation_required 是 503 且 retryable: true,但没有停写证明时永远不会解除。遵循 retryable 的调用方会一直重试。runtime_placement_recovery_required 已经是 409 且 retryable: false。
未验证
- 本地未跑 MySQL/MariaDB。CI 的 MariaDB 和 MySQL 任务在 feat(managed-agent): Implement event replay (Stage D3) #12840 合入前的 merge ref 上通过。
- F4 下的 Spring holder 路径。
- Linux 与 Windows。
- TS 的
terminal/reason元数据目前没有生产调用方(只有 provider 本身及其测试),对 W0e-1 来说没问题。
证据: wenshao/qwen-code@9c3f17f → pr12839/r2/ 下有三张证据图、harness/W0eR2ProbeTest.java(两个 head 上原样可跑)、每一轮的原始日志,以及包含试合并和 V16 运行结果的 results/suites-summary.txt。
|
Follow-up to the round-2 verification, on head
W0e-1 remains draft. W0e-3 still needs dedicated Linux physical reboot acceptance. 中文摘要:F4 已修复并用真实子进程复验:只有显式确认原进程仍存活的本地 provisioner 才允许在一次认证传输失败后重试;进程死亡、身份冲突和其他 provisioner 仍保持封闭。F5 的历史 |
Round 3 real-environment verification: head
|
| Finding | Status on 9e42875b |
|---|---|
| F6: Flyway V14 collision | ✅ Fixed (renamed to V16). The test merge with current main passes 139/139. |
F4: one failed/slow attestation → permanent LOST + tenant block |
✅ Fixed. A1 (one reset /attest): warm #1 returns 503, the binding stays READY, warms #2–#4 succeed, and a new call runs. Release works, and tenant-a/workspace-b and tenant-b are both READY. A3 (12 s slow response): same. A2 (session isolation): harness-1 recovers, and new harness-3/harness-4 are READY. |
| F1 / F2 | ✅ Still hold. B1: LOST, call ABANDONED, no replacement. C1: stale release 503, stale create 409 runtime_admission_closed. |
F5: historical seeded FAILED rows |
Documented. The README inventory query run on my U1 upgrade database returns [tenant-a=1], which matches the guard's 409 for tenant-a (tenant-b READY). |
N1: permanent 503 with retryable: true |
Author's rationale accepted; no change. |
New checks, all passing:
- A5 (hung but alive worker). SIGSTOP the worker: each warm returns 503 after the 10 s request timeout and the binding stays
READY. After SIGCONT, warm succeeds and a call settlessuccess. On353b6526the same stop producedFAILED, generation 2 was started beside the stopped worker, and the call on the old Session endedUNKNOWN. - C2 (deferred dispatch from feat(serve): add gated Hosted Workspace file tool turns #12831). A call reserved with
prepareExecutionbecomesABANDONEDat loss. The stale Broker'sstartExecutiongets409 runtime_admission_closed, a same-keypreparereturns theABANDONEDreceipt, and nothing ran. This needed two local rig ops, included asrig-prepare-start-ops.patch.
F7 (new): a transient fault while a managed-context placement starts blocks the whole Workspace
A4: managed-context rig (W0c-3 boot v2, one placement per Hosted Session). One /attest exchange is reset during the first start of harness-2's placement:
- On both heads, harness-2 returns
503 runtime_provision_failedand then409 runtime_broker_recovery_blocked. Managed startup failures blocked their own placement before this PR. - New in this PR: harness-3, a new Session in the same Workspace, gets
409 runtime_placement_recovery_required. On353b6526it wasREADY. - The existing harness-1 placement keeps working.
- 3/3 runs identical.
blocksPlacement now counts RECOVERY_BLOCKED for every managed placement in the Workspace. But this particular blocked row cannot hide a writer. The provisioner destroyForcibly()s its own unadopted process when attestation fails during start, before the binding records a lease and before any Session context is installed, so no tool could have run. Nothing in production clears the row, so one network blip permanently refuses new Sessions in that Workspace. That is the F4 pattern again: the round-3 fix covers the confirm path only.
Suggestion: do not count, as an unreclaimed domain, a blocked binding that never recorded a lease and whose process the provisioner destroyed itself. Alternatively, let start failures record that proof on the row. Either way, add an A4-style gate.
Not verified
- MySQL/MariaDB locally (CI jobs are green on this head).
- Linux, Windows, and a real host reboot (W0e-3).
- The Spring holder path with the A4 fault.
Evidence: wenshao/qwen-code@605b68f → pr12839/r3/ holds the cards, the probe sources, the local rig patch, the raw logs of all runs and results/suites-summary.txt.
中文版
第三轮真实环境验证:head 9e42875b12
结论: 第二轮的两个阻塞项都已修复,并在真实进程中验证:F6 用与当前 main 的试合并验证,F4 用与第二轮相同的故障注入验证。第一轮的 F1、F2 依然成立,从 main 合入的延迟派发路径(#12831)也同样被屏障挡住。新发现一个与 F4 同类的问题(F7),出在 managed-context 启动路径上。建议在本 PR 中修复,因为正是本 PR 新增的 Workspace 级守卫让它变成永久阻断;如果选择推迟,请把它记为启用 Workspace 绑定执行之前的前提条件。
环境: 与第二轮相同:macOS、JDK 21、该 head 自己打包的 dist/cli.js、真实 Broker JVM、TCP 上的 H2。第二轮的探针原样运行,另加新的 A4、A5、C2 探针。head 上跑了 3 轮,所有关键结果一致(归一化后的结果行在各轮之间哈希相同)。A4、A5 也在仅设计的 head 353b6526 上跑了 A/B。
测试套件
runtime-broker:381 个,0 失败,2 个条件跳过;真实 worker 用例加上 bundle 后 15/15 通过。- 故障门禁:28/28。TS provider 测试:44/44。
managed-agent-server在与当前 main0a136f89的试合并上:139/139(mvn clean test)。迁移现在是 main 的 V14、V15 加上本 PR 的 V16。- PR CI:该 head 上所有 Java、MariaDB、MySQL 任务都为绿。
此前发现的状态
| 发现 | 在 9e42875b 上 |
|---|---|
| F6:Flyway V14 冲突 | ✅ 已修复(改名为 V16)。与当前 main 的试合并 139/139 通过。 |
F4:一次失败/过慢的 attestation → 永久 LOST + 租户级阻断 |
✅ 已修复。A1(重置一次 /attest):warm #1 返回 503,绑定保持 READY,warm #2–#4 成功,新调用正常执行;release 成功,tenant-a/workspace-b 和 tenant-b 都是 READY。A3(响应慢 12 s):结果相同。A2(session 隔离):harness-1 恢复,新的 harness-3/harness-4 为 READY。 |
| F1 / F2 | ✅ 依然成立。B1: LOST,调用 ABANDONED,没有替代 worker。C1: 旧 Broker 的 release 返回 503,create 返回 409 runtime_admission_closed。 |
F5:带 seed 的历史 FAILED 行 |
已写入文档。在我的 U1 升级数据库上执行 README 的清点查询,返回 [tenant-a=1],与守卫对 tenant-a 返回的 409 一致(tenant-b 为 READY)。 |
N1:永久性 503 却标 retryable: true |
接受作者的理由,不做修改。 |
新增检查,全部通过:
- A5(worker 卡住但仍存活)。 对 worker 发 SIGSTOP 后,每次 warm 在 10 s 请求超时后返回 503,绑定保持
READY;SIGCONT 之后 warm 成功,调用以success完成。在353b6526上,同样的暂停导致FAILED,并在被暂停的 worker 旁边启动了第 2 代,旧 Session 上的调用最终为UNKNOWN。 - C2(feat(serve): add gated Hosted Workspace file tool turns #12831 的延迟派发)。 用
prepareExecution预留的调用在丢失时变为ABANDONED;旧 Broker 的startExecution返回409 runtime_admission_closed,同 key 的prepare返回ABANDONED回执,没有任何东西被执行。为此在本地装置中加了两个操作,见rig-prepare-start-ops.patch。
F7(新):managed-context placement 启动时的一次瞬时故障会阻断整个 Workspace
A4: managed-context 装置(W0c-3 boot v2,每个 Hosted Session 一个 placement)。在 harness-2 的 placement 首次启动时重置一次 /attest 交互:
- 两个 head 上 harness-2 都先返回
503 runtime_provision_failed,再返回409 runtime_broker_recovery_blocked。managed 启动失败阻断自身 placement,在本 PR 之前就已如此。 - 本 PR 新增的变化:harness-3(同一 Workspace 的新 Session)得到
409 runtime_placement_recovery_required。在353b6526上它是READY。 - 已有的 harness-1 placement 仍可正常使用。
- 3/3 轮一致。
blocksPlacement 现在会对该 Workspace 内的所有 managed placement 计入 RECOVERY_BLOCKED。但这一条被阻断的行不可能掩盖任何写入者:启动阶段 attestation 失败时,provisioner 会对自己尚未接管的进程执行 destroyForcibly(),这发生在绑定记录 lease 之前,也在安装任何 Session 上下文之前,所以不可能有工具运行过。生产环境中没有机制能清除这一行,因此一次网络抖动就会让该 Workspace 永久拒绝新 Session。这又是 F4 的模式:第三轮的修复只覆盖了 confirm 路径。
建议: 对于从未记录过 lease、且进程已被 provisioner 自己销毁的被阻断绑定,不要把它计为未回收的执行域;或者让 start 失败时把这一证明写到该行上。无论哪种做法,都补一个 A4 式的门禁测试。
未验证
- 本地未跑 MySQL/MariaDB(该 head 上 CI 的相应任务为绿)。
- Linux、Windows,以及真实宿主机重启(W0e-3)。
- 带 A4 故障的 Spring holder 路径。
证据: wenshao/qwen-code@605b68f → pr12839/r3/ 下有证据图、探针源码、本地装置补丁、每一轮的原始日志和 results/suites-summary.txt。
|
[codex] Round 3 follow-up for review at
The new H2 and in-memory regression covers pre-admission startup, original-placement refusal, post-attestation fencing, and unknown provisioner fencing. Independent process verification reproduced F7 before the fix and confirmed the repaired path. Final local gates passed: runtime-broker 382 tests (0 failures, 2 conditional skips), managed-agent-server 139 tests, Checkstyle, repository build, typecheck, and bundle. No inline review threads needed replies or resolution (0/0). I did not rerun MySQL/MariaDB locally; CI will validate the new SHA. |
|
[codex] CI follow-up: the Lint & Static job stopped at its freshness gate because On the new commit, local build, typecheck, bundle, Broker/Server tests and Checkstyle pass. Source ESLint passed on the preceding merge commit when local Maven-generated |
Round 4 real-environment verification: head
|
| Guard | Mutant | Result |
|---|---|---|
| Release/retirement refused with loss evidence alone | hasStoppedWriters() = loss only; JDBC stop gate removed; in-memory stop gate removed; requireRelease accepts loss |
all 4 killed (unit) |
| Service cached-release precheck accepts loss | survives, but masked: completeSessionRelease re-checks requireRelease under the binding lock. The B1/C1 probes on this mutant still get 503 runtime_reconciliation_required. |
|
SETTLED result survives recovery |
Query filter and abandon() both admit SETTLED (in-memory; JDBC) |
both killed |
| F7 exemption reverted | killed (new ManagedContextRecoveryTest) |
|
| Admission fence ignores binding state; identity failure no longer fences | both killed | |
Local canRetryFailedConfirm() = true |
killed by the gates (ContextInstallationFaultGateTest) |
|
Local canRetryFailedConfirm() = false (F4 fix reverted) |
survives unit + all 28 gates. The A1/A3 probes on this mutant show the round-2 bug again: the binding is LOST forever, and tenant-a/workspace-b gets 409 runtime_placement_recovery_required. |
|
| F7 exemption widened to legacy rows | survives. Low reachability: A6 shows legacy start faults stay PROVISIONING. |
The F4 unit tests exercise the service's decision through a fake provisioner, but nothing pins LocalProcessRuntimeProvisioner.canRetryFailedConfirm() = isUsable(lease). Suggestion: add an A1-style fault gate. Warm a legacy binding, proxy.schedule("attest", RESET), warm again, then assert that the binding is still READY, that the retry succeeds, and that a new placement in the same tenant is READY. The rig already has everything this needs; my probe (W0eR2ProbeTest#a1…) can serve as a starting point.
Merge-order note: three open PRs claim Flyway V16
This PR (V16__runtime_loss_evidence.sql), #12881 (V16__managed_session_operation.sql) and #12855 (V16__managed_extension_record.sql) each add a V16. Git merges them cleanly, but whichever lands second will fail every Spring context with Found more than one migration with version 16, the same failure as round-2 F6. Renumber at merge time, and re-run the server suite on the post-merge tree.
Evidence: wenshao/qwen-code@3253276 → pr12839/r4/ holds the cards, mutate.sh / mutate-probe.sh, every mutant's exact diff and verdict, and the probe logs (head and M11/M5).
中文版
第四轮真实环境验证:head 752bff4906
结论: 从我这边看已没有阻塞项。
- 第三轮的 F7 已修复,此前所有修复在真实进程中依然成立。该 head 已包含当前 main,所有测试套件全部通过。
- 我还用变异测试回答了 triage 机器人的遗留问题:测试套件是否钉住了两条承重声明?答案是两条都已钉住。
- 有一个测试缺口建议在合入前补上: 第二轮 F4 在唯一生产 provisioner 中的修复没有被任何测试钉住。把它回退后,382 个单测和 28 个故障门禁全部照样通过。
环境: 与前几轮相同的装置(macOS、JDK 21、该 head 自己打包的 worker、真实 Broker JVM、TCP 上的 H2),探针文件相同,另加 A6。变异体在一个独立的干净 worktree 中逐个应用,每次运行后都恢复源码。
测试套件
runtime-broker:382 个,0 失败。真实 worker 用例 15/15,故障门禁 28/28。managed-agent-server:139/139(mvn clean test)。TS provider 测试:44/44。- 两次合并 main 都没有冲突块(
git show --remerge-diff)。迁移是 main 的 V14、V15 加上本 PR 的 V16。
状态
- F7 已修复(A4)。 harness-2 首次 managed-context 启动时重置一次
/attest后,harness-2 仍然阻断它自己的 placement:409 runtime_broker_recovery_blocked,这在本 PR 之前就是如此。harness-3(同一 Workspace 的新 Session)现在是READY;在9e42875b上它得到409。 - A6(新增)。 同样的故障发生在 legacy placement 首次启动时是可重试的:该行保持
PROVISIONING,下一次 warm 为READY,不产生任何阻断。 - 与第三轮相比未变:
- A1/A2/A3:瞬时或过慢的 attestation 保持
READY,重试成功。 - A5:被 SIGSTOP 的 worker 保持
READY,SIGCONT 后恢复。 - B1:worker 死亡 →
LOST+ABANDONED,不启动替代 worker。 - C1:旧 Broker 的
release返回 503,create返回 409。 - C2:丢失后的延迟
start→ 409。 - U1:README 清点查询返回
[tenant-a=1]。
- A1/A2/A3:瞬时或过慢的 attestation 保持
变异检查
每个变异体都是对一处守卫的单行修改。所有变异体都跑单测(382 个),存活者再跑故障门禁(28 个)。未变异基线干净。
| 守卫 | 变异体 | 结果 |
|---|---|---|
| 只有丢失证据时拒绝释放与退役 | hasStoppedWriters() 只看丢失证据;去掉 JDBC 停写门禁;去掉内存停写门禁;requireRelease 接受丢失证据 |
4 个全部被杀(单测) |
| 服务层缓存释放预检接受丢失证据 | 存活,但被下层掩盖:completeSessionRelease 会在绑定锁下重新执行 requireRelease。在该变异体上跑 B1/C1 探针,仍得到 503 runtime_reconciliation_required。 |
|
SETTLED 结果在恢复后保持不变 |
查询过滤与 abandon() 同时放行 SETTLED(内存;JDBC) |
两个都被杀 |
| F7 豁免被回退 | 被杀(新增的 ManagedContextRecoveryTest) |
|
| 准入屏障忽略绑定状态;身份失败不再触发屏障 | 两个都被杀 | |
本地 canRetryFailedConfirm() = true |
被故障门禁杀死(ContextInstallationFaultGateTest) |
|
本地 canRetryFailedConfirm() = false(回退 F4 修复) |
单测和全部 28 个门禁都存活。在该变异体上跑 A1/A3 探针,第二轮的问题重新出现:绑定永久 LOST,tenant-a/workspace-b 得到 409 runtime_placement_recovery_required。 |
|
| F7 豁免扩大到 legacy 行 | 存活。可达性低:A6 表明 legacy 启动故障会保持 PROVISIONING。 |
F4 的单测通过一个假 provisioner 测试服务层的判断逻辑,但没有任何测试钉住 LocalProcessRuntimeProvisioner.canRetryFailedConfirm() = isUsable(lease)。建议: 补一个 A1 式的故障门禁:先 warm 一个 legacy 绑定,执行 proxy.schedule("attest", RESET),再 warm 一次,然后断言绑定仍为 READY、重试成功、同租户的新 placement 为 READY。装置已经具备所需的一切,可以参考我的探针(W0eR2ProbeTest#a1…)作为起点。
合入顺序提示:三个未合入的 PR 都使用 Flyway V16
本 PR(V16__runtime_loss_evidence.sql)、#12881(V16__managed_session_operation.sql)和 #12855(V16__managed_extension_record.sql)都新增了 V16。git 合并不会报冲突,但第二个合入的 PR 会让所有 Spring 上下文因 Found more than one migration with version 16 失败,与第二轮 F6 相同。合入时请重新编号,并在合入后的代码上重跑服务端测试。
证据: wenshao/qwen-code@3253276 → pr12839/r4/ 下有证据图、mutate.sh / mutate-probe.sh、每个变异体的完整 diff 和结论,以及探针日志(head 以及 M11/M5)。
|
[codex] Round 4 follow-up at
Final local verification after the commit: Broker unit suite 382 tests, full real-process fault gates 29 tests, zero failures; Checkstyle, repository build, typecheck, and bundle passed. No inline review threads required replies or resolution (0/0). CI is rerunning on |
Round 5 real-environment verification: head
|
|
@qwen-code /triage |
…ences main's W0e terminal recovery fences (#12839) wrapped cancelExecution in a safeStage/requireOwnedExecution pre-check and widened the in-lock guard from isSettled() to isTerminal(), which collides with this branch's Tool v3 carve-out on the same guard. Resolution keeps main's structure and re-applies the carve-out as a third conjunct of the UNKNOWN group, so a runtimeProtocol 3 record with no local invocation still falls through to the physical transport cancel. The carve-out stays reachable because UNKNOWN is not terminal, and main's terminal fence keeps precedence: a SETTLED or ABANDONED record returns before it.
Resolve the conflicts with the durable Session lifecycle (D4, #12881) and the W0e recovery fences (#12839): - SessionCapabilities keeps D4's per-Session session_lifecycle and adds tasks after it, in the order of the schema; the controllers take both the lifecycle and the task services. - The OpenAPI version becomes 1.19.0 after D4's 1.18.0, and the description keeps both entries. - The Stage H migration moves to V18. main now holds two V16 migrations (D4 and W0e), one of which has to move to V17. - W0e's new Broker state ABANDONED maps to outcome_unknown, and the shared fixtures gain its Broker case; the Broker answers it as an unknown execution. - The deletion tests use D4's delete operation: the H2 test checks a Session being deleted, and the MySQL race case deletes through begin, claim and complete on another connection.
The durable Session lifecycle (#12881) and the W0e recovery fences (#12839) both added a Flyway V16, so main now holds two migrations with version 16, and Flyway refuses to start: every Spring context and every migration test fails with "Found more than one migration with version 16". W0e merged first, so the lifecycle migration moves to V17, the next free version, and the D4 design, the server README and the comments and tests that name the migration follow it. No database can hold the lifecycle migration as V16: Flyway has refused every migration on main since both landed.
wenshao
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed.
1 Suggestion-level finding(s) this review confirmed are already reported on this PR and are not repeated:
- releaseUnusableSession and the cached releaseSession gate answer a retryable 503 runtime_reconciliation_required that cannot clear without stop proof in this slice — already reported as N1 (comment 5856965272) and settled there…
Not reviewed: build-and-test — "Integration Tests (CLI, No Sandbox)" was skipped in CI and its suite did not run locally.
Not explored to full depth (tool budget reached): "agent reverse-audit (round 1)": did not verify which RuntimeBrokerHttpServer routes back each of the three BrokerManagedRuntimeProvider call sites that now return { outcome: 'unknown', te…; "agent reverse-audit (round 2)": 未走查 EN §6 表格之后的散文段(文档第 262–293 行,落在 diff 行 346 之后,属 chunk 3 范围)——我只把 §6a/§7 作为上下文读取以判定 §4/§6 的断言,未对其独立定论。; "agent reverse-audit (round 2)": §5「Production local-worker identity for #12766」的具体断言未逐条对代码核验——文档 Status 行与 §6a 均把生产 worker 接管标为后续工作(W0e-2/W0e-3),按简报中「文档自陈的路线图/未实现表面」排除项处理,未投入预算。; "agent reverse-audit (round 5)": 未对 mapBinding / updateBinding / setBinding 做 javac + javap 静态字节码尺寸测量(base 与 head 双侧),因共享 worktree 内禁止 mvn / gradle 构建且无现成 target/classes 可作 classpath;故本…; "agent reverse-audit (round 4)": 未实际执行 runtime-broker 测试套件验证发现 1/发现 2 的 mutation 变绿结论——brief 明确禁止在本共享 worktree 内 git checkout / git stash /就地构建或跑 mvn / gradle ,两条 mutation 均由源码路径推导( execution…, and 23 more.
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.
中文说明
仅完成部分审查,审查缺口已披露。
本轮确认的 1 条建议级发现已在 PR 上报告过,不再重复发布(列表见上方英文部分)。
未审查(原文为英文):build-and-test — "Integration Tests (CLI, No Sandbox)" was skipped in CI and its suite did not run locally.
未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 1)":did not verify which RuntimeBrokerHttpServer routes back each of the three BrokerManagedRuntimeProvider call sites that now return { outcome: 'unknown', te…;"agent reverse-audit (round 2)":未走查 EN §6 表格之后的散文段(文档第 262–293 行,落在 diff 行 346 之后,属 chunk 3 范围)——我只把 §6a/§7 作为上下文读取以判定 §4/§6 的断言,未对其独立定论。;"agent reverse-audit (round 2)":§5「Production local-worker identity for #12766」的具体断言未逐条对代码核验——文档 Status 行与 §6a 均把生产 worker 接管标为后续工作(W0e-2/W0e-3),按简报中「文档自陈的路线图/未实现表面」排除项处理,未投入预算。;"agent reverse-audit (round 5)":未对 mapBinding / updateBinding / setBinding 做 javac + javap 静态字节码尺寸测量(base 与 head 双侧),因共享 worktree 内禁止 mvn / gradle 构建且无现成 target/classes 可作 classpath;故本…;"agent reverse-audit (round 4)":未实际执行 runtime-broker 测试套件验证发现 1/发现 2 的 mutation 变绿结论——brief 明确禁止在本共享 worktree 内 git checkout / git stash /就地构建或跑 mvn / gradle ,两条 mutation 均由源码路径推导( execution…,另有 23 条。
未审查:反向审计——在 5 轮的反审轮数上限内未收敛。
— kimi-k3 via Qwen Code /review (v0.24.7)
| ```sql | ||
| SELECT tenant_id, COUNT(*) AS failed_bindings | ||
| FROM qwen_runtime_binding | ||
| WHERE binding_state = 'FAILED' AND provision_seed_ciphertext IS NOT NULL |
There was a problem hiding this comment.
[Critical] R1-1: [certifies-falsely] [new-surface] The new pre-upgrade inventory query counts only seeded FAILED rows, so it reports an all-clear for databases the new placement guard will hard-block
The guard this change adds refuses a new placement for LOST rows unconditionally and for RECOVERY_BLOCKED rows conditionally (RuntimeBindingRecord.blocksPlacement:323-339; requireRecoverablePlacement selects all three states), but the runbook query the same commit adds counts only binding_state = 'FAILED' AND provision_seed_ciphertext IS NOT NULL and then says "A nonempty result is a rollout blocker for this database". Pre-upgrade code already wrote LOST (base RuntimeBrokerService:1291) and RECOVERY_BLOCKED (base :1447), and a LOST row written before this change has loss_evidence_json NULL, so recoverLost returns it unchanged forever and requireSafeReplacement blocks every CAS escape. An operator stops admission, runs the documented query, gets zero rows, resumes traffic, and every new placement for that tenant/workspace then fails 409 runtime_placement_recovery_required with retryable=false and no migration path. Two further mismatches in the same paragraph: findOrCreate returns the slot's active binding before the guard runs, so the promised "including when a later generation is READY" case is never blocked; and for managed-context rows the fence is per-workspace, not per-tenant, so "blocks affected tenants" is wider than the code for exactly the deployment this runbook addresses.
Witness:
set difference with predicates parsed from source and document, plus a real H2 database holding one pre-upgrade-style LOST row: `GUARD SQL predicate states: ['FAILED','LOST','RECOVERY_BLOCKED']` / `README operator query states: ['FAILED']` / `MISSING from README query: ['LOST','RECOVERY_BLOCKED']`; `PROBE-R6 README query rows = 0` / `guard query rows = 1` / `recoverLost attempt 1|2|3 = LOST` / `sibling findOrCreate = HTTP 409 runtime_placement_recovery_required retryable=false` / `LOST->RELEASED via CAS = IllegalArgumentException: Lost binding requires atomic recovery`. Applying the fix moves the operator's answer from 0 to 1.
Suggested fix: Widen the query to the states the guard actually reads — WHERE binding_state IN ('LOST','RECOVERY_BLOCKED','FAILED') grouped by tenant and state — and state per branch what blocks: LOST blocks unconditionally and a pre-upgrade LOST row can never be reclaimed; FAILED blocks only with a seed; RECOVERY_BLOCKED blocks except for an unadmitted local managed startup. Also scope the "blocks affected tenants" sentence to the Workspace for managed rows and to the tenant for legacy rows, and drop or qualify "including when a later generation is READY", since the guard is only consulted when a new binding row must be created.
One premise this fix must not break: The RECOVERY_BLOCKED arm is conditioned on the request, not the row (RuntimeBindingRecord.java:326-333: !request.isManagedContext() || !LocalProcessRuntimeProvisioner.KIND.equals(request.getProvisionerKind()) || lease != null || attestationGeneration > 0), so a widened WHERE can only over-approximate it — the prose must not claim the query reproduces the guard exactly. provision_seed_ciphertext IS NOT NULL is a faithful proxy for provisionSeed != null because mapSeed returns non-null only when all three seed columns are present (JdbcRuntimeBindingRepository.java:776-789).
中文说明
[Critical] R1-1:本次新增的升级前自检 SQL 只统计带 seed 的 FAILED 行,而它要预测的那个放置守卫实际读取三种状态:LOST 无条件阻断,RECOVERY_BLOCKED 有条件阻断(RuntimeBindingRecord.blocksPlacement:323-339,requireRecoverablePlacement 的 IN 列表也是这三种)。升级前的代码本来就会写 LOST(基线 RuntimeBrokerService:1291)和 RECOVERY_BLOCKED(基线 :1447),而升级前写入的 LOST 行 loss_evidence_json 必为 NULL,于是 recoverLost 永远在 getLossEvidence() == null 处原样返回,requireSafeReplacement 又禁止任何 CAS 逃逸。运维按文档停准入、跑这条查询、得到 0 行、按“非空才是上线阻塞”恢复流量之后,该租户/工作区每次新放置都会拿到 409 runtime_placement_recovery_required、retryable=false,且本模块明确不提供迁移。同一段还有两处与代码不符:findOrCreate 会先返回 slot 上仍 active 的绑定再执行守卫,所以“即使后一代是 READY 也阻断”那一句描述的情形根本不会被阻断;而对 managed-context 行,围栏范围是 workspace 而非 tenant,因此“blocks affected tenants”对这份 runbook 面向的部署反而说得过宽。
修复建议:把查询扩到守卫真正读取的状态——WHERE binding_state IN ('LOST','RECOVERY_BLOCKED','FAILED') 并按 tenant 与 state 分组——同时逐支说明阻断条件:LOST 无条件阻断且升级前写入的 LOST 行永远无法回收;FAILED 仅在有 seed 时阻断;RECOVERY_BLOCKED 除“未准入的本地 managed 启动”外都阻断。另外把“blocks affected tenants”按代码收窄(managed 行为该 Workspace,legacy 行为整租户),并删掉或限定“including when a later generation is READY”——守卫只在必须新建绑定行时才被咨询。
修复不得违反的既有事实:修复不得声称这条查询能精确复现守卫:RECOVERY_BLOCKED 那一支的条件挂在请求上而不是行上(RuntimeBindingRecord.java:326-333),所以放宽后的 WHERE 只能过近似。而 provision_seed_ciphertext IS NOT NULL 确实是 provisionSeed != null 的忠实代理,因为 mapSeed 只在三列全非空时才返回非空(JdbcRuntimeBindingRepository.java:776-789)。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)
| // A local managed worker without a lease or attested generation never | ||
| // admitted a Session; failed startup still blocks its own slot. | ||
| boolean unreclaimed = state == State.LOST | ||
| || state == State.RECOVERY_BLOCKED |
There was a problem hiding this comment.
[Critical] R1-2: [fails-closed] [regression] A non-managed durable placement that lands in RECOVERY_BLOCKED before it ever records a lease blocks its whole tenant permanently, with no exit transition
blocksPlacement exempts only managed-context local workers from the RECOVERY_BLOCKED arm (!request.isManagedContext() || short-circuits the exemption), and it matches on tenantId alone for non-managed requests, skipping the workspaceId comparison. Reachable with a pure client-input error: a durable, non-managed placement (storageId == null, provisionerKind == "local-process", so requiresDurableIdentity() is true) persists its resource handle, then LocalProcessRuntimeProvisioner.start throws the non-retryable 400 "capabilityDigest is not a sha256 digest."; RuntimeBrokerService:1050-1057 takes !retryable and calls blockRecoveryQuietly, writing RECOVERY_BLOCKED while lease == null and attestationGeneration == 0 — a binding that never admitted a Session and cannot hide a writer. Every later request in that tenant, with any other isolation key or scope and a perfectly valid capability digest, then gets 409 runtime_placement_recovery_required with retryable=false. There is no exit: RELEASED is written only by recoverLost, which requires state == LOST; requireSafeReplacement forbids RECOVERY_BLOCKED -> FAILED; invalidateBinding writes LOST only from READY. Base had no runtime_placement_recovery_required at all and would simply have minted a new generation, so this converts a retryable input error into a permanent tenant-wide outage that only a database edit clears.
Witness:
three probe runs on unmodified PR source: a non-managed durable binding driven to RECOVERY_BLOCKED with lease == null and attestationGeneration == 0, then findOrCreate for a different isolation key in the same tenant -> `409 runtime_placement_recovery_required retryable=false`; a different tenant -> PROVISIONING; and no transition out of RECOVERY_BLOCKED reachable from any production writer (enumerated: blockRecovery/blockRecoveryQuietly, failBinding refused by requireSafeReplacement, invalidateBinding READY-only, recoverLost LOST-only).
Suggested fix: Align the exemption with the code's own comment ("A local managed worker without a lease or attested generation never admitted a Session") — the "never admitted a Session" test is lease == null && attestationGeneration == 0, which has nothing to do with isManagedContext(). Drop the !request.isManagedContext() || disjunct so any local-process binding that never recorded a lease and never attested stops blocking its slot. If non-managed placements are meant to stay conservative, give RECOVERY_BLOCKED an exit (e.g. allow RECOVERY_BLOCKED -> FAILED for a row with lease == null && attestationGeneration == 0, and make that FAILED shape not block placement), because as shipped the state is terminal.
One premise this fix must not break: The unconditional state == State.LOST arm must stay (a LOST binding may have admitted Sessions, and only loss/stop evidence may clear it), and the state == State.FAILED && provisionSeed != null arm must stay (it is what catches historical seeded rows — see the separate coverage finding). requireSafeReplacement's "Unproved runtime loss cannot free a placement" clause (RuntimeBindingRecord.java:350-353) forbids routing the fix through failBinding, so an exit needs its own rule.
Fix acceptance — A repository-contract case on both implementations: drive a non-managed local-process binding to RECOVERY_BLOCKED with no lease and no attestation, then assert findOrCreate for a different isolation key in the same tenant returns a new generation instead of throwing; reverting the fix reddens it. ManagedContextRecoveryTest's existing managed-context arms must stay green. Please prove it by mutation: apply the fix, then revert it and confirm that test goes red.
中文说明
[Critical] R1-2:blocksPlacement 的 RECOVERY_BLOCKED 豁免只给 managed-context 的本地 worker(!request.isManagedContext() || 直接短路掉豁免),而对非 managed 请求它只按 tenantId 匹配、跳过 workspaceId 比较。一个纯客户端输入错误就能触发:一个 durable 但非 managed 的放置(storageId == null、provisionerKind == "local-process",因此 requiresDurableIdentity() 为真)在持久化 resource handle 之后,LocalProcessRuntimeProvisioner.start 抛出不可重试的 400 capabilityDigest is not a sha256 digest.;RuntimeBrokerService:1050-1057 命中 !retryable 并调用 blockRecoveryQuietly,写下 RECOVERY_BLOCKED,而此时 lease == null、attestationGeneration == 0——这条绑定从未准入过任何 Session,不可能掩盖任何写入者。此后同租户内任何其它请求(不同 isolation key、不同 scope、capability digest 完全合法)都会得到 409 runtime_placement_recovery_required、retryable=false。而且没有出口:RELEASED 只由 recoverLost 写入且要求状态为 LOST;requireSafeReplacement 禁止 RECOVERY_BLOCKED -> FAILED;invalidateBinding 只从 READY 写 LOST。基线完全没有 runtime_placement_recovery_required,会直接开新一代,所以这把一次可重试的输入错误变成了只能改库才能解除的整租户中断。
修复建议:让豁免条件与代码自己的注释一致(“A local managed worker without a lease or attested generation never admitted a Session”)——“从未准入 Session”的判据是 lease == null && attestationGeneration == 0,与 isManagedContext() 无关。删掉 !request.isManagedContext() || 这一析取项即可。若确实想让非 managed 放置保守阻断,则必须同时给 RECOVERY_BLOCKED 一个出口(例如允许 lease == null && attestationGeneration == 0 的 RECOVERY_BLOCKED -> FAILED,且该形态不 blocksPlacement),因为按现在的实现它是终态。
修复不得违反的既有事实:无条件阻断的 state == State.LOST 那一支必须保留(LOST 意味着可能已准入过 Session,只有丢失/停写证据能放行),state == State.FAILED && provisionSeed != null 那一支也必须保留(它负责拦住历史 seeded 行,另有独立的覆盖缺口发现)。requireSafeReplacement 的 “Unproved runtime loss cannot free a placement”(RuntimeBindingRecord.java:350-353)禁止把出口做成 failBinding,所以需要单独的规则。
修复验收:上方 “Fix acceptance” 一句点名的测试即验收标准(测试名与标识符在两种语言中逐字保留);请用变异验证——先应用修复,再回退它,确认该测试变红。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)
| : safeStage(() -> provisioner.reconcile(recovered.getRequest(), | ||
| recovered.getProvisionSeed(), recovered.getResourceHandle(), | ||
| recovered.getLease())).toCompletableFuture() | ||
| .orTimeout(operationLeaseDuration.toMillis(), TimeUnit.MILLISECONDS) |
There was a problem hiding this comment.
[Critical] R1-3: [fails-closed] [new-surface] reclaimLostBindingNow runs an external reconcile under an operation lease it never renews, so a slow observation guarantees the evidence is discarded
The new evidence-gathering step calls provisioner.reconcile(...) and bounds it with orTimeout(operationLeaseDuration), but the timer is armed after claimOperation and after the first recoverLost transaction, and nothing in reclaimLostBindingNow renews the lease — while both sibling paths do (provisionBinding:920-921, startReconciliation:1191-1193). So whenever the observation takes longer than the remaining lease, the timeout necessarily fires after the lease has expired: compareAndSet returns null on a dead claim, the second recoverLost returns null at its hasLiveOperationAt gate, the observation and its evidence are thrown away, and the caller gets 503 runtime_broker_runtime_lost. .exceptionally(error -> null) also collapses a TimeoutException and a genuine "no stop proof" into the same null, so nothing distinguishes them. This is reachable in the shipped product, not only under a short test lease: EmbeddedRuntimeBroker.LEASE is 30s for both the operation and the dispatch lease, and the shipped observe path attests with a 30s timeout (LocalProcessRuntimeProvisioner.READY_TIMEOUT, HttpRuntimeTransport:46, ManagedAgentProperties requestTimeout), so one hung-worker attestation can consume the entire unrenewed lease. Recovery then never advances for that generation, and README:10-12's standing promise that the service "renews them while external work is in flight" is falsified by this new call site.
Witness:
probe with in-memory repositories, operationLeaseDuration=1s and a provisioner returning complete loss+stop evidence after 1450ms. Unmodified PR: `[R11] elapsed=1142ms reconciles=1` / `warm -> FAILED RuntimeBrokerException status=503 code=runtime_broker_runtime_lost retryable=true` / `binding state=LOST lossEvidence=null stopEvidence=null` / `execution state=PREPARED` / `activeSessions=1`. With the fix (BindingRenewal started + orTimeout on operationDeadlineNanos()): `elapsed=1616ms` / `binding state=RELEASED lossEvidence=persisted stopEvidence=persisted` / `execution state=ABANDONED` / `activeSessions=0`. One variable: whether the observation lands inside a live lease.
Suggested fix: Renew the operation claim across the external observation the way provisionBinding and startReconciliation do (BindingRenewal renewal = new BindingRenewal(claimed); renewal.start(); around the reconcile leg, stopped in a finally), or bound the observation with a deadline derived from operationDeadlineNanos() rather than a single lease, so a slow-but-successful observation can still be committed.
One premise this fix must not break: The renewal must not outlive the reconciliation deadline: operationDeadlineNanos() is leaseNanos * 4 (RuntimeBrokerService:1547-1552) and §1 of the new design doc documents "four 30-second operation leases" as the deadline, so an unbounded renewal would break the documented bound. ownsOperation (1541-1545) checks owner and generation but not lease liveness, so the fix cannot rely on it to detect an expired claim.
Fix acceptance — A DurableRuntimeRecoveryTest arm whose provisioner reconcile completes after one lease period but within the reconciliation deadline, asserting the binding reaches RELEASED with both evidence records persisted and the execution ABANDONED; without the renewal that arm gets 503 runtime_broker_runtime_lost with null evidence. Please prove it by mutation: apply the fix, then revert it and confirm that test goes red.
中文说明
[Critical] R1-3:新的取证式 reconcile 调用用 orTimeout(operationLeaseDuration) 限界,但这个计时器是在 claimOperation 与第一次 recoverLost 事务之后才武装的,而 reclaimLostBindingNow 全程不续租——两个兄弟路径都续(provisionBinding:920-921、startReconciliation:1191-1193)。因此只要观测耗时超过租约余量,超时就必然发生在租约已过期之后:compareAndSet 因死租约返回 null,第二次 recoverLost 在 hasLiveOperationAt 处返回 null,观测结果连同证据被整体丢弃,调用方拿到 503 runtime_broker_runtime_lost。.exceptionally(error -> null) 还把 TimeoutException 与“确实没有停写证明”折叠成同一个 null,两者无从区分。这在生产配置下就可达,不只是短租约测试的问题:EmbeddedRuntimeBroker.LEASE 对 operation 与 dispatch 都是 30s,而随包的观测路径 attest 超时同样是 30s(LocalProcessRuntimeProvisioner.READY_TIMEOUT、HttpRuntimeTransport:46、ManagedAgentProperties 的 requestTimeout),所以一次卡住的 worker attestation 就能吃光整个未续租的租约,该代次的恢复永远推进不了。README:10-12 那句常驻承诺“renews them while external work is in flight”也被这个新调用点证伪。
修复建议:像 provisionBinding 与 startReconciliation 那样在外部观测期间续租(BindingRenewal renewal = new BindingRenewal(claimed); renewal.start(); 包住 reconcile 这一段,并在 finally 中停止),或者把观测的限界改成由 operationDeadlineNanos() 推导而不是单个租约,让一次慢但成功的观测仍然能提交。
修复不得违反的既有事实:续租不得越过对账截止期:operationDeadlineNanos() 是 leaseNanos * 4(RuntimeBrokerService:1547-1552),新设计文档 §1 也把“四个 30 秒操作租约”写成截止期,所以无界续租会破坏这个成文上界。另外 ownsOperation(1541-1545)只校验 owner 与 generation、不校验租约是否存活,修复不能依赖它来发现过期认领。
修复验收:上方 “Fix acceptance” 一句点名的测试即验收标准(测试名与标识符在两种语言中逐字保留);请用变异验证——先应用修复,再回退它,确认该测试变红。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)
| recovered.getLease())).toCompletableFuture() | ||
| .orTimeout(operationLeaseDuration.toMillis(), TimeUnit.MILLISECONDS) | ||
| .exceptionally(error -> null); | ||
| return observation.thenApply(observed -> { |
There was a problem hiding this comment.
[Critical] R1-4: [fails-closed] [new-surface] The post-observation recovery work runs blocking JDBC on the JVM's single shared CompletableFuture Delayer thread
reclaimLostBindingNow chains orTimeout(...).exceptionally(error -> null) into a plain thenApply with no executor, so when the timeout fires the continuation runs inline on the JDK's process-wide single-threaded CompletableFutureDelayScheduler. That continuation performs four blocking repository calls — findById (which AES-GCM-decrypts the seed and lease credential), compareAndSet, recoverLost (a transaction that opens with the tenant placement-guard row lock) and releaseOperationQuietly in whenComplete. While any of them blocks, no other orTimeout in the JVM can fire: HttpRuntimeTransport's request deadlines at :158 and :789 stop working, and that file's own comment (1103-1106) says a body that stalls after the headers is bounded only by that stage deadline; RuntimeBrokerService:512's lookup.orTimeout stops working too. In-flight requests then run past operationLeaseDuration/dispatchLeaseDuration, their leases expire, and another Broker can claim the same operation — the double-write this design exists to prevent. A host crash fences every binding on that host at once, so several reclamations can time out together and serialize on the one Delayer thread. On the non-timeout path the same JDBC runs on LocalProcessRuntimeProvisioner's cached thread pool, coupling database latency into provisioning.
Witness:
probe with in-memory repositories behind a recording proxy, a provisioner whose reconcile never completes, operationLeaseDuration=300ms: PR as written -> `OBSERVED recovery-thread=CompletableFutureDelayScheduler entered-blocking-repository-call=true unrelated-orTimeout(100ms)-fired-while-recovery-blocked=false`, then `fired-after-unblock=true`; with thenApply -> thenApplyAsync (one-word edit) -> `recovery-thread=ForkJoinPool.commonPool-worker-1 ... unrelated-orTimeout(100ms)-fired-while-recovery-blocked=true`, same 169-test selection green.
Suggested fix: Move every post-observation repository call off the completing thread onto an executor the service owns and closes: observation.thenApplyAsync(observed -> {...}, recoveryExecutor) and .whenCompleteAsync((ignored, error) -> releaseOperationQuietly(...), recoveryExecutor). Apply the same to the pre-existing lookup continuation at :512/:515.
One premise this fix must not break: Do not reuse scheduler as the recovery executor: it carries DispatchRenewal.start()'s scheduleWithFixedDelay renewals (RuntimeBrokerService:2588), BindingRenewal.renew() (:2544-2556) and the provisioning deadlineTask (:1013), so blocking JDBC there would stop lease renewal — a worse failure than this one. Do not reuse the provisioner's newCachedThreadPool either.
Fix acceptance — A DurableRuntimeRecoveryTest arm whose provisioner reconcile never completes, recording Thread.currentThread() inside a stub binding repository's findById and asserting it is the recovery executor's thread, plus a positive control asserting an unrelated orTimeout(50ms) still fires while recovery is blocked. Reverting to thenApply reddens both. Please prove it by mutation: apply the fix, then revert it and confirm that test goes red.
中文说明
[Critical] R1-4:reclaimLostBindingNow 把 orTimeout(...).exceptionally(...) 之后的续接写成不带 executor 的 thenApply,因此超时触发时这段代码内联运行在 JDK 全进程共享的单线程 CompletableFutureDelayScheduler 上。而这段续接做的是四次阻塞式仓储调用——findById(会 AES-GCM 解密 seed 与 lease 凭据)、compareAndSet、recoverLost(事务第一句就是租户级 placement guard 行锁)、以及 whenComplete 里的 releaseOperationQuietly。只要其中任何一步阻塞,JVM 内所有其它 orTimeout 都不会触发:HttpRuntimeTransport:158/:789 的请求截止期失效(该文件 1103-1106 行的注释明确写着“响应头之后卡住的 body 由 stage deadline 限界,而不是 HttpRequest.timeout”),RuntimeBrokerService:512 的 lookup.orTimeout 同样失效。于是在飞请求会越过 operationLeaseDuration/dispatchLeaseDuration 继续等待,租约过期,另一个 Broker 就能认领同一个 operation——正是本设计要防止的双写。宿主崩溃会一次性把该宿主上所有绑定打成 LOST,多个回收同时超时就会在这一个线程上串行,放大倍数等于并发回收数。非超时路径同样有误:此时这段 JDBC 跑在 LocalProcessRuntimeProvisioner 的 cached 线程池上,把数据库时延耦合进 provisioning。
修复建议:把观测之后的所有仓储调用从 completing thread 移走,交给 service 自己拥有并在 close() 中关闭的有界 IO executor:observation.thenApplyAsync(observed -> {...}, recoveryExecutor),并把 .whenComplete(...) 改成 .whenCompleteAsync(..., recoveryExecutor),让 releaseOperation 也不落在 Delayer 上。:512/:515 那个既有实例应一并修正。
修复不得违反的既有事实:不要把 scheduler 当作 recoveryExecutor 复用:它承载着 DispatchRenewal.start() 的 scheduleWithFixedDelay 续租(RuntimeBrokerService:2588)、BindingRenewal.renew()(:2544-2556)以及 provisioning 的 deadlineTask = scheduler.schedule(...)(:1013),在它上面跑阻塞 JDBC 会直接停掉续租,比本缺陷更严重。也不要复用 provisioner 的 newCachedThreadPool。
修复验收:上方 “Fix acceptance” 一句点名的测试即验收标准(测试名与标识符在两种语言中逐字保留);请用变异验证——先应用修复,再回退它,确认该测试变红。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)
| @@ -1,3 +1,8 @@ | |||
| CREATE TABLE IF NOT EXISTS qwen_runtime_placement_guard ( | |||
There was a problem hiding this comment.
[Critical] R1-5: [certifies-falsely] [regression] README still documents four private Broker tables; this change adds a fifth that the embedding service must own
JdbcRuntimeBrokerSchema.initialize executes schema.sql verbatim, and schema.sql now holds five CREATE TABLE statements (base held four) because this change adds qwen_runtime_placement_guard; V16__runtime_loss_evidence.sql adds it to Flyway too. The untouched sentence at packages/sdk-java/runtime-broker/README.md:43 still says initialize "installs the four private Broker tables". README frames the embedding service as owning the connection pool and schema lifecycle, so an integrator writing GRANTs, backup/restore verification, a count-keyed schema-drift assertion or a teardown against the documented four-table contract misses the guard table. lockPlacementDomain writes that table on every findOrCreate, compareAndSet and recoverLost, so the omission surfaces as a runtime SQL permission or missing-table error on the placement path rather than a clean startup failure.
Witness:
base/HEAD count over the SQL resource initialize actually executes: `BASE (df9374d435) schema.sql CREATE TABLE count: 4` / `HEAD schema.sql CREATE TABLE count: 5` (new: qwen_runtime_placement_guard at schema.sql:1) / `BASE README:43 == HEAD README:43` (byte-identical, still "the four private").
Suggested fix: Change README:43 to "installs the five private Broker tables", or list the five names including qwen_runtime_placement_guard.
中文说明
[Critical] R1-5:JdbcRuntimeBrokerSchema.initialize 逐条执行 schema.sql,而 schema.sql 现在有五条 CREATE TABLE(基线是四条),因为本次改动新增了 qwen_runtime_placement_guard;V16__runtime_loss_evidence.sql 也把它加进了 Flyway。但未被改动的 packages/sdk-java/runtime-broker/README.md:43 仍然写着 initialize “installs the four private Broker tables”。README 自己把连接池与 schema 生命周期归给嵌入方,所以按这份四表契约去写 GRANT、备份/恢复校验、按表数量断言 schema 漂移或 teardown 的集成方会漏掉 guard 表;而 lockPlacementDomain 在每一次 findOrCreate、compareAndSet、recoverLost 上都要写这张表,于是漏掉的后果不是启动检查干净地失败,而是运行期在放置路径上拿到 SQL 权限或缺表错误。
修复建议:把 README:43 改成 “installs the five private Broker tables”,或直接列出五张表名(含 qwen_runtime_placement_guard)。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)
| }); | ||
| var fence = pool.submit(() -> { start.await(); return raced.lose(true); }); | ||
| start.countDown(); | ||
| boolean admitted = admission.get(10, TimeUnit.SECONDS); |
There was a problem hiding this comment.
[Suggestion] R1-46: The only concurrency probe pairing admission against loss recovery joins both racers before calling recoverLost, so recovery never overlaps an in-flight admission
The fence task is only start.await(); return raced.lose(true); — recoverLost is called at :133 after both Future.get() calls return, so recovery always runs on a quiescent ledger in a single pass. The invariant the loop is meant to prove (an admission either is refused, or the session and execution it wrote are reclaimed before the binding reaches RELEASED) is guarded in production by admitSession's binding-row FOR UPDATE and by the re-checks after abandonByBinding and releaseLost; the dangerous interleaving is an admission landing after the re-check scan and before the release commits, and this harness structurally excludes it. Removing admitSession's binding-row lock, or moving the hasActiveByBinding re-check before abandonByBinding, would let a straggler admission survive a RELEASED binding with generation 2 already provisioned — an unsettled execution on an old writer domain and two contradictory client answers — while this loop stays green in all five of its arms.
Witness:
probe counting overlap between admission and recovery: `shipped-harness memory arm: admitCalls=253 recoverLostCalls=27 overlappingPairs=0`; shape A/B over 2000 iterations each — `shipped shape (join, then recoverLost): overlappingPairs=0 admissionsAccepted=1694` versus `restructured shape (recoverLost inside the fence task): overlappingPairs=9 admissionsAccepted=1985`; and mutation M4 (admitSession's binding-row FOR UPDATE changed to a non-locking read) leaves the whole JDBC contract green, `Tests run: 1, Failures: 0, Errors: 0 / BUILD SUCCESS`.
Suggested fix: Make recovery overlap admission: move recoverLost into the fence task and turn it into a bounded retry until it returns RELEASED, join the admission thread afterwards, then assert the invariant (for every admitted == true, its execution is ABANDONED and the active-session count is 0). Run the trial in a bounded loop (e.g. 50 iterations), otherwise the window is rarely hit. The dead GatedSessionRepository double in DurableRuntimeRecoveryTest:1323 is already written and is the only repository-level latch in the tree that can open a window mid-recovery — reuse it rather than deleting it.
One premise this fix must not break: The in-memory arm cannot produce a real interleaving — InMemoryRuntimeBindingRepository.recoverLost (:47) and admitSession (:92) share one monitor — so the overlap assertion is JDBC-only. Fixture's operation lease is claimed once for five minutes and never renewed while recoverLost requires hasLiveOperationAt (InMemoryRuntimeBindingRepository:56, JdbcRuntimeBindingRepository:118-121, the latter throwing 409 on expiry), so any retry loop must stay inside that window.
Fix acceptance — The new overlapping arm must go red on the JDBC side when admitSession's binding-row FOR UPDATE is removed, or when the post-abandon hasActiveByBinding re-check is removed; both mutations are green against the contract today. Please prove it by mutation: apply the fix, then revert it and confirm that test goes red.
中文说明
[Suggestion] R1-46:围栏任务只有 start.await(); return raced.lose(true);,recoverLost 在 :133、即两个 Future.get() 都返回之后才执行,所以回收总是在静止账本上单趟完成。这个循环要证明的不变量是“准入要么被拒,要么它写下的 session/execution 在绑定变 RELEASED 之前一定被回收”,生产侧靠 admitSession 的 binding 行 FOR UPDATE 与 abandonByBinding/releaseLost 之后的复检共同保证;而危险的交错是“准入落在复检扫描之后、release 提交之前”,本 harness 在结构上排除了它。删掉 admitSession 的 binding 行锁,或把 hasActiveByBinding 复检挪到 abandonByBinding 之前,都能让一个 straggler 准入在绑定已 RELEASED、findOrCreate 已开出第 2 代之后仍存活——旧写入域上留下未终结的执行、客户端拿到两个互相矛盾的答案——而这个循环的五个臂全绿。
修复建议:让回收与准入真正重叠:把 recoverLost 移入围栏任务,并改成有界重试直到返回 RELEASED,之后再 join 准入线程并断言不变量(凡 admitted == true 者,其执行必须为 ABANDONED、活跃会话计数必须为 0)。把单次试验改成有界循环(例如 50 次),否则窗口命中率不足以构成守卫。DurableRuntimeRecoveryTest:1323 那个已成死代码的 GatedSessionRepository 正是全树唯一能在回收事务中途开窗口的仓储级 latch——与其删掉它,不如复用。
修复不得违反的既有事实:内存臂不可能产生真实交错——InMemoryRuntimeBindingRepository.recoverLost(:47)与 admitSession(:92)共用同一监视器——所以重叠断言只在 JDBC 臂有效。Fixture 的 operation lease 在构造时以 claimOperation(..., Duration.ofMinutes(5)) 取得且从不续租,而 recoverLost 要求 hasLiveOperationAt(InMemoryRuntimeBindingRepository:56、JdbcRuntimeBindingRepository:118-121,后者超时抛 409),所以任何重试循环都必须留在这 5 分钟窗口内。
修复验收:上方 “Fix acceptance” 一句点名的测试即验收标准(测试名与标识符在两种语言中逐字保留);请用变异验证——先应用修复,再回退它,确认该测试变红。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)
| } | ||
| } | ||
| Fixture bounded = new Fixture(bindings, sessions, executions, prefix + "-batch"); | ||
| for (int index = 0; index < 103; index++) { |
There was a problem hiding this comment.
[Suggestion] R1-47: The execution-drain batch bound is pinned only to the range [52, 102], while the session bound in the same critical section is pinned to exactly 100
The bounded fixture prepares 103 executions and asserts only that the first recoverLost returns LOST with active executions remaining and the second returns RELEASED. For a bound B those two assertions require B <= 102 and B >= 52, so changing LIMIT 100 in JdbcToolExecutionRepository:42 or .limit(100) in InMemoryToolExecutionRepository:35 to 60 or 102 keeps all three backends green. The constant is not academic: the drain runs inside the tenant placement-guard row lock that this PR also added to every compareAndSet (measured elsewhere as blocking an unrelated same-tenant transition for over three seconds), and each drained row costs a full-column SELECT plus its own UPDATE round trip. The bound therefore sets both how long one recovery freezes the tenant and how many recoverLost calls an operator needs to clear a backlog — lowering it to 60 turns a 103-row backlog from "2 calls, 3 left after the first" into "2 calls, 43 left", and a 600-row backlog from 6 calls into 10, with nothing turning red. The session half of the same critical section does not have this problem: assertEquals(3, countActiveByBinding(...)) at :163 pins it to exactly 100.
Witness:
the two execution-side literals were parameterized and the boundaries measured, everything else unchanged: `B=100 GREEN/GREEN, B=60 GREEN/GREEN, B=52 GREEN/GREEN, B=51 RED (expected: <RELEASED> but was: <LOST>), B=102 GREEN/GREEN, B=103 RED (expected: <LOST> but was: <RELEASED>)` on both the memory and H2 backends. Control arm for the session bound: `sessionBatch=100 GREEN/GREEN, 99 RED (expected: <3> but was: <4>), 101 RED (expected: <3> but was: <2>)`.
Suggested fix: Replace the one-sided 103 fixture with a two-sided squeeze: one fixture with exactly 100 executions asserting the FIRST recoverLost reaches RELEASED (pins B >= 100), and one with 101 asserting the first recoverLost is still LOST (pins B <= 100). Both use lose(true) and otherwise keep the existing :147-152 shape.
One premise this fix must not break: The four bound literals are independent with no shared constant: JdbcToolExecutionRepository:42 LIMIT 100, InMemoryToolExecutionRepository:35 .limit(100), JdbcRuntimeSessionRepository:38 LIMIT 100, InMemoryRuntimeSessionRepository:20 .limit(100). Because verify is shared by the in-memory runner and the JDBC/H2/Flyway runners, a new fixture pins both execution-side literals only while they are equal; if one side changes alone, the other backend reddens rather than the same assertion.
Fix acceptance — Those two arms: B = 101 reddens the 100-row fixture's RELEASED assertion and B = 99 reddens the 101-row fixture's LOST assertion; today both arms are absent and every value in [52, 102] is green. Please prove it by mutation: apply the fix, then revert it and confirm that test goes red.
中文说明
[Suggestion] R1-47:bounded fixture 准备 103 条执行后,只断言第一次 recoverLost 返回 LOST 且仍有活跃执行、第二次返回 RELEASED。设上界为 B,这两条断言只要求 B <= 102 与 B >= 52,所以把 JdbcToolExecutionRepository:42 的 LIMIT 100 或 InMemoryToolExecutionRepository:35 的 .limit(100) 改成 60 或 102,三套后端全部保持绿色。这个常数不是学术性的:整段排空运行在本 PR 同时加进每一次 compareAndSet 的租户级 placement-guard 行锁之内(另一条发现实测到无关的同租户 CAS 在持锁期间阻塞超过 3 秒),而每条被排空的行还要付一次全列 SELECT 加一次独立 UPDATE 往返。所以上界同时决定“单次恢复把整租户冻住多久”和“运维要调多少次 recoverLost 才能排空积压”:降到 60,103 条积压从“2 次调用、首轮余 3 条”变成“2 次调用、首轮余 43 条”,600 条积压所需调用次数从 6 次涨到 10 次,而没有任何测试变红。同一临界区的会话那一半没有这个问题——:163 的 assertEquals(3, ...) 把它钉成恰好 100。
修复建议:把单侧的 103 fixture 换成两侧夹逼:一个恰好 100 条执行的 fixture,断言第一次 recoverLost 就到达 RELEASED(钉 B >= 100);一个 101 条的 fixture,断言第一次 recoverLost 仍是 LOST(钉 B <= 100)。两者都用 lose(true),其余断言沿用现有 :147-152 的形状。
修复不得违反的既有事实:四个上界字面量彼此独立、没有共享常量:JdbcToolExecutionRepository:42 的 LIMIT 100、InMemoryToolExecutionRepository:35 的 .limit(100)、JdbcRuntimeSessionRepository:38 的 LIMIT 100、InMemoryRuntimeSessionRepository:20 的 .limit(100)。因为 verify 由内存与 JDBC/H2/Flyway 两侧共用,新 fixture 只在两个执行侧字面量取值相等时才同时钉住二者;若将来只改一侧,变红的是另一个后端而不是同一条断言。
修复验收:上方 “Fix acceptance” 一句点名的测试即验收标准(测试名与标识符在两种语言中逐字保留);请用变异验证——先应用修复,再回退它,确认该测试变红。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)
| ToolExecutionRecord otherCall = other.prepare("call"); | ||
| RuntimeBindingRecord scopedLost = scoped.lose(true); | ||
| assertEquals(RuntimeBindingRecord.State.RELEASED, bindings.recoverLost(sessions, executions, scopedLost).getState()); | ||
| assertFalse(executions.hasActiveByRuntimeSession(scoped.binding.getBindingId(), 1, "shared-id")); |
There was a problem hiding this comment.
[Suggestion] R1-48: All three assertions on the new scoped hasActiveByRuntimeSession are insensitive to its runtimeSessionId predicate, so the method can silently degrade to hasActiveByBinding
The three-argument overload exists to stop one session's activity from blocking another's release, and its only three assertions (:90, :172, :173) are each satisfied regardless of the session predicate: at :90 the binding's only execution is already terminal and the late prepare was refused; at :172 the scoped binding has been drained by recoverLost (the same call asserts RELEASED, whose precondition is an empty active scan); at :173 the other binding's single PREPARED execution belongs to the very "shared-id" being queried. The service layer does not supply the missing pin either — all three runtime_session_busy assertions use a single session whose own execution is the active one. Deleting the predicate from both implementations keeps the contract green on the memory, H2 and Flyway backends. The cost lands on the user-visible release busy check (RuntimeBrokerService:677-682) and its negated fence (:568): with the predicate lost or its parameters transposed, any other harness session with a tool call in flight on the same runtime generation makes this session's release fail with a non-retryable 409 runtime_session_busy. One binding carrying many sessions is the norm, not a corner — this contract's own sessionBatch fixture puts 103 sessions on one binding.
Witness:
three-arm probe: mutation deleting the session predicate from BOTH implementations (InMemoryToolExecutionRepository:246 and JdbcToolExecutionRepository:376 plus its :380 bind and Java-side re-check) leaves `RESULT contract-memory: PASS` and `RESULT contract-h2: PASS`; adding the suggested idle-session discrimination arm turns both red at `RuntimeRecoveryContract.java:181 (assertFalse)`; restoring the predicate with the new assertion kept gives PASS on both backends.
Suggested fix: Add an idle session to the other binding so the pair discriminates: admit a second RuntimeSessionRecord with a distinct runtimeSessionId on other.binding, then assertFalse(executions.hasActiveByRuntimeSession(other.binding.getBindingId(), 1, "idle-session")); right after the existing assertTrue for "shared-id". The two lines together are what pin the predicate.
One premise this fix must not break: The discriminating arm must hang off a binding that still has an active execution — other's is pinned as PREPARED at :174 — and admitSession requires the binding to be READY (InMemoryRuntimeBindingRepository:95), so the new session can only be added on the other side, since scoped is the fenced one.
Fix acceptance — The new assertFalse: deleting AND runtime_session_key = ? from the JDBC implementation or the runtimeSessionId equality from the in-memory one must redden it on the corresponding backend; it stays green on the unmodified tree because other's active execution belongs to "shared-id". Please prove it by mutation: apply the fix, then revert it and confirm that test goes red.
中文说明
[Suggestion] R1-48:三参 hasActiveByRuntimeSession 存在的意义就是阻止一个会话的活跃执行阻塞另一个会话的释放,而它仅有的三处断言(:90、:172、:173)对会话谓词全部不敏感::90 时该绑定的唯一执行已终态且迟到的 prepare 被拒;:172 的 scoped 绑定已被 recoverLost 排空(同一次调用断言 RELEASED,其放行前置就是活跃扫描为空);:173 的 other 绑定恰有一条属于被查询的 "shared-id" 的 PREPARED 执行。服务层也补不上这颗钉子——三处 runtime_session_busy 断言都只用单个会话,活跃执行就属于它自己。从两个实现里删掉该谓词,契约在 memory、H2、Flyway 三套后端上都保持绿色。代价落在用户可见的 release 忙判定(RuntimeBrokerService:677-682)与其取反围栏(:568)上:谓词丢失或参数绑错时,同一 runtime generation 上任意另一个 harness 会话有在执行的工具调用,就会让本会话的 release 收到不可重试的 409 runtime_session_busy。一个绑定承载多个会话是常态——本契约自己的 sessionBatch fixture 就在同一个绑定上放了 103 个会话。
修复建议:给 other 绑定再加一个空闲会话,构成判别对:用不同 runtimeSessionId admit 第二条 RuntimeSessionRecord,然后在 :173 的 assertTrue 之后加 assertFalse(executions.hasActiveByRuntimeSession(other.binding.getBindingId(), 1, "idle-session"));。两行合起来才钉住会话谓词。
修复不得违反的既有事实:判别臂必须挂在仍有活跃执行的绑定上——other 的这条执行由 :174 钉为 PREPARED;而 admitSession 要求绑定处于 READY(InMemoryRuntimeBindingRepository:95 的 RuntimeAdmission.requireReady),被围栏的是 scoped,所以新会话只能加在 other 一侧。
修复验收:上方 “Fix acceptance” 一句点名的测试即验收标准(测试名与标识符在两种语言中逐字保留);请用变异验证——先应用修复,再回退它,确认该测试变红。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)
| assertThrows(RuntimeBrokerException.class, () -> bindings.findOrCreate( | ||
| new RuntimeProvisionRequest(changed, null, "other-provisioner"))); | ||
| RuntimeRecoveryEvidence foreign = evidence(bounded.binding, RuntimeRecoveryEvidence.Fact.JOURNAL_LOST); | ||
| assertThrows(IllegalArgumentException.class, () -> bare.withRecoveryEvidence(foreign, null, Instant.now())); |
There was a problem hiding this comment.
[Suggestion] R1-49: The contract's only evidence-forgery probe never reaches requireSafeReplacement, so two of that new method's three clauses are unpinned
bare.withRecoveryEvidence(foreign, null, now) is a record-layer call whose first act is requireEvidence(loss, JOURNAL_LOST); foreign is built from another binding, so matches() fails and it throws "Recovery evidence identity differs" before any repository call. requireSafeReplacement — reached only through compareAndSet in both implementations — therefore never runs, and its two new clauses ("Recovery evidence cannot be overwritten" and "Unproved runtime loss cannot free a placement", both added by this diff) have no test that can fail when they are deleted. The only assertion in the tree that a binding becomes FAILED goes through PROVISIONING -> FAILED, which the new clause does not cover. The consequence: the premise that makes this PR's own seeded-FAILED placement rule safe — an unproved READY binding must not be laundered into FAILED, which isActive() treats as released so findOrCreate would provision a second runtime beside a live one — is unguarded, and this diff already reduced failBinding's call sites from four to one, so a future re-introduction on READY or DRAINING would ship green. Separately, the unproved fixture makes only negative assertions about the LOST-without-evidence state and never exercises the single transition out of it, so narrowing that exit would wedge tenants silently.
Witness:
mutant/control pair over the same 188-test population and both implementations: MUTANT-R23 (the two clauses removed, the third left in place) -> GREEN `Tests run: 188, Failures: 0, Errors: 0 / BUILD SUCCESS`; CONTROL (requireSafeReplacement reduced to a no-op) -> RED on both sides, `RuntimeRecoveryTest.memoryRepositoryHonorsRecoveryContract:17 Expected java.lang.IllegalArgumentException to be thrown, but nothing was thrown` and `JdbcRepositoryTest.repositoriesPreserveTheirContractsAcrossInstances:12` likewise. The control is what makes the green meaningful: the method IS exercised by the contract on both implementations, and exactly one of its three clauses is pinned.
Suggested fix: Add a repository-layer probe beside the unproved fixture: assertThrows(IllegalArgumentException.class, () -> bindings.compareAndSet(unproved.binding, unproved.binding.withState(RuntimeBindingRecord.State.FAILED, null, Instant.now()))) (READY -> FAILED on a record carrying no evidence yet). For the overwrite clause, either build a same-identity record with a different lossEvidence through the package-private constructor and assert compareAndSet refuses it, or delete the clause as unreachable. Complete the unproved fixture with its positive control: attach matching loss evidence, then stop evidence, then assert recovery reaches RELEASED and the tenant can place again.
One premise this fix must not break: The probe must use a record that carries no loss evidence yet: RuntimeBindingRecord.java:152-154 (if (lossEvidence != null && state != State.LOST && state != State.RELEASED) throw new IllegalArgumentException("Loss evidence requires a fenced binding")) would otherwise throw first and leave requireSafeReplacement uncovered — that is precisely what the contract's :53-54 withState(READY, ...) assertion hits. The positive control's evidence must be built from bare itself, since :198-199 already asserts that evidence derived from another binding is rejected.
Fix acceptance — The new READY->FAILED assertThrows in RuntimeRecoveryContract.verify, which runs on both implementations (memory and H2/Flyway/MySQL). Mutation check: removing requireSafeReplacement's FAILED clause must turn it red. Please prove it by mutation: apply the fix, then revert it and confirm that test goes red.
中文说明
[Suggestion] R1-49:bare.withRecoveryEvidence(foreign, null, now) 是记录层调用,它的第一步就是 requireEvidence(loss, JOURNAL_LOST),而 foreign 由另一个绑定构造,matches() 为假,于是抛 “Recovery evidence identity differs”,根本没有发生仓储调用;requireSafeReplacement(只经两个实现的 compareAndSet 到达)因此从未被执行,它新增的两条子句(“Recovery evidence cannot be overwritten”、“Unproved runtime loss cannot free a placement”)删掉也不会让任何测试变红。全树唯一断言绑定变成 FAILED 的用例走的是 PROVISIONING -> FAILED,正好在这条子句覆盖范围之外。后果是:让本 PR 自己的 seeded-FAILED 放置规则成立的前提——未被举证停写的 READY 绑定不能被洗成 FAILED(isActive() 把 FAILED 视为已释放,于是 findOrCreate 会为同一 slot 再拉起一个仍然活着的 runtime)——完全没有测试;而本 diff 已把 failBinding 的调用点从四处删到一处,日后在 READY/DRAINING 上重新引入一次调用就会静默放行这条绕过路径。另外,unproved fixture 对 LOST-无证据状态只做负向断言,从未走过离开该状态的那一次转换,所以收窄这个唯一出口会让租户被静默钉死。
修复建议:在 unproved 段补一条仓储层探针:assertThrows(IllegalArgumentException.class, () -> bindings.compareAndSet(unproved.binding, unproved.binding.withState(RuntimeBindingRecord.State.FAILED, null, Instant.now())))(READY -> FAILED,且该记录尚未携带任何证据)。证据覆写子句要么用同包包级构造器造一个 lossEvidence 不同、其余身份相同的替换记录经 compareAndSet 断言被拒,要么按“不为不可能的场景写错误处理”删掉它。同时给 unproved fixture 补上正向对照:附着匹配的丢失证据、再附着停写证据,然后断言恢复到达 RELEASED 且该租户可以重新放置。
修复不得违反的既有事实:探针必须用尚未携带丢失证据的记录:RuntimeBindingRecord.java:152-154(if (lossEvidence != null && state != State.LOST && state != State.RELEASED) throw new IllegalArgumentException("Loss evidence requires a fenced binding"))会先抛,使 requireSafeReplacement 依旧未被覆盖——这正是契约 :53-54 那条 withState(READY, ...) 断言实际命中的位置。正向对照的证据必须由 bare 自己派生,因为 :198-199 已断言来自另一个绑定的证据会被拒。
修复验收:上方 “Fix acceptance” 一句点名的测试即验收标准(测试名与标识符在两种语言中逐字保留);请用变异验证——先应用修复,再回退它,确认该测试变红。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)
| RuntimeBindingRecord proof = bindings.compareAndSet(lost, lost.withRecoveryEvidence(null, | ||
| RuntimeRecoveryContract.evidence(lost, RuntimeRecoveryEvidence.Fact.WRITERS_STOPPED), | ||
| java.time.Instant.now())); | ||
| bindings.recoverLost(sessions, executions, proof); |
There was a problem hiding this comment.
[Suggestion] R1-50: The released=true arm discards recoverLost's return value, so the RELEASED transition the test name promises is never asserted
The executions were already abandoned by the recoverLost call outside the loop, so this call's only remaining job is releasing the session and writing RELEASED. InMemoryRuntimeBindingRepository.recoverLost returns the UNCHANGED LOST record — no throw, no null — when active executions or sessions remain or the stop evidence is not accepted (:68). If it silently no-ops, the released=true iteration becomes observationally identical to released=false and every later assertion still passes, because the HTTP reads and the cancel/prepare/create/reconcile calls all short-circuit on the terminal record and none reads binding or session state. The honest cost is narrower than first filed: the RELEASED transition IS pinned elsewhere (RuntimeRecoveryContract:97-98, :133, :152, :164, :171, :187 and DurableRuntimeRecoveryTest:1203), so what is wrong is that this test's name — terminalHttpReadsUseSavedOwnershipAfterRestartAndRelease — claims coverage it does not provide.
Witness:
mutation inserting `if (Boolean.TRUE) { return current; }` after abandonByBinding and before releaseLost: baseline `-Dtest=RuntimeRecoveryTest -> Tests run: 4, Failures: 0`; MUTANT `Tests run: 4, Failures: 1` with the only failure `memoryRepositoryHonorsRecoveryContract (expected: <RELEASED> but was: <LOST> at RuntimeRecoveryContract.java:98)` and surefire showing both terminalHttpReads parameterizations as self-closed passing testcases; MUTANT+FIX (capture the result and assert binding and session RELEASED) `Tests run: 4, Failures: 3` with both parameterizations failing `expected: <RELEASED> but was: <LOST> at RuntimeRecoveryTest.java:99`; FIX alone `Tests run: 4, Failures: 0`.
Suggested fix: Capture the result and assert the release happened: RuntimeBindingRecord released = bindings.recoverLost(sessions, executions, proof); assertEquals(State.RELEASED, released.getState()); assertEquals(RuntimeSessionRecord.State.RELEASED, sessions.findById(lost.getRequest().getScope(), fixture.session.getRuntimeSessionId()).getState());
One premise this fix must not break: InMemoryRuntimeBindingRepository.recoverLost's release preconditions at :68 (!current.hasStoppedWriters() || executions.hasActiveByBinding(...) -> return current), :73 (releaseLost) and :81 (write State.RELEASED): the new assertions depend on stop evidence being present and on no active execution or session for that binding.
Fix acceptance — The new RELEASED assertions in both deferred parameterizations. Mutation check: replacing recoverLost's RELEASED write with return current; must turn them red. Please prove it by mutation: apply the fix, then revert it and confirm that test goes red.
中文说明
[Suggestion] R1-50:循环外那次 recoverLost 已经放弃了全部执行,所以循环内这次调用只剩“释放会话 + 写 RELEASED”一件事。InMemoryRuntimeBindingRepository.recoverLost 在仍有活跃执行/会话或停写证据未被接受时原样返回 LOST 记录(不抛异常、不返回 null,:68),此时 released=true 这一轮与 released=false 在可观测行为上完全等价,而循环体内后续每条断言都照样通过——HTTP 读取与 cancel/prepare/create/reconcile 全都在终态记录上短路,没有一个读绑定或会话状态。需要如实收窄代价:RELEASED 转换在别处是被钉住的(RuntimeRecoveryContract:97-98、:133、:152、:164、:171、:187 以及 DurableRuntimeRecoveryTest:1203),所以真正的问题是这个名字——terminalHttpReadsUseSavedOwnershipAfterRestartAndRelease——承诺了它并不提供的覆盖。
修复建议:接住返回值并断言释放确实发生:RuntimeBindingRecord released = bindings.recoverLost(sessions, executions, proof); assertEquals(State.RELEASED, released.getState()); assertEquals(RuntimeSessionRecord.State.RELEASED, sessions.findById(lost.getRequest().getScope(), fixture.session.getRuntimeSessionId()).getState());
修复不得违反的既有事实:InMemoryRuntimeBindingRepository.recoverLost 的释放前置条件在 :68(!current.hasStoppedWriters() || executions.hasActiveByBinding(...) -> return current)、:73(releaseLost)与 :81(写 State.RELEASED):新断言依赖停写证据存在、且该绑定下没有活跃执行与会话。
修复验收:上方 “Fix acceptance” 一句点名的测试即验收标准(测试名与标识符在两种语言中逐字保留);请用变异验证——先应用修复,再回退它,确认该测试变红。
证据见上方 Witness 代码块:那是程序输出,按规则逐字保留、不翻译。
— kimi-k3 via Qwen Code /review (v0.24.7)












What this PR does
Implements W0e-1: executions whose original Runtime journal is authoritatively lost can reach
ABANDONED, an immutable terminal state with an unknown outcome. It retains the original request, idempotency receipt and execution identity without inventing a tool result or replaying work. Separate, generation-matched writer-stop evidence is required before releasing Runtime Sessions and retiring their placement.The change adds durable evidence, atomic admission fences, bounded recovery batches, original-ownership terminal reads, and terminal-uncertainty metadata for private clients. It also protects ordinary release against cached state and late deactivation responses crossing the loss fence. Additive SQL migrations preserve existing receipts. The full W0e design and follow-up boundaries are synchronized in English and 简体中文.
Why it's needed
Once the original journal is permanently unavailable, its in-flight executions cannot obtain a result and currently pin their generation indefinitely. Ending result lookup must not imply that files are safe to reuse: descendants can outlive a worker, and process exit alone does not establish that the old writer domain stopped. This slice separates the two decisions before adding production restart adoption.
Reviewer Test Plan
How to verify
Evidence (Before & After)
Before: the existing real-process reproduction of #12670 retained an EXECUTING record behind a LOST generation and returned IN_FLIGHT indefinitely.
After: the real-process loss gate returns ABANDONED with no result and no replay. It still refuses release or replacement because the process test does not prove every old writer stopped. Deterministic evidence fixtures separately verify conditional cleanup. Detailed test results and audit findings are posted in the E2E report comment.
Tested on
Environment (optional)
Java 21, Node.js, real Broker JVMs and worker processes, durable H2 and an isolated local MySQL instance. No model invocation or host reboot is required for these W0e-1 checks.
Risk & Scope
Linked Issues
Part of #12380. Implements the W0e-1 decision proposed for #12670; physical reuse stays gated by independent stop proof. Production adoption for #12766 follows separately. Coordinates with #12831 without duplicating Hosted tool orchestration.
中文说明
本 PR 做什么
实现 W0e-1:原 Runtime journal 被权威证据确认丢失后,执行可进入不可变的
ABANDONED终态,结果仍未知。保留原请求、幂等回执和执行身份,不伪造工具结果、不重放工作。释放 Runtime Session 和退役 placement 还需要独立且匹配原代数的停写证明。本次增加持久证据、原子准入屏障、有界恢复批次、按原归属读取终态,以及私有客户端的终态不确定性元数据。同时防止普通释放被本地缓存状态或跨越失联屏障的迟到停用响应绕过。增量 SQL 迁移保留现有回执。完整 W0e 设计和后续边界已同步到英文与中文。
为什么需要
原 journal 永久不可用后,在途执行无法再取得结果,当前会永久钉住该代数。结束结果查询不能代表文件可以安全复用:后代进程可能在 worker 退出后继续运行,根进程退出不能证明旧写入域已经停止。本切片在生产重启接管之前明确分离这两个决策。
评审者测试计划
如何验证
前后证据
此前:#12670 的现有真实进程复现中,LOST 代数后的执行停留在 EXECUTING,并持续返回 IN_FLIGHT。
此后:真实进程失联门禁返回 ABANDONED,无结果、不重放。由于该进程测试不能证明所有旧写入者都停止,释放和替换仍被拒绝。确定性证据 fixture 单独验证条件清理。详细测试结果与审计发现记录在 E2E 报告评论中。
测试平台
环境
Java 21、Node.js、真实 Broker JVM 与 worker 进程、持久化 H2 和隔离的本地 MySQL 实例。这些 W0e-1 检查不需要模型调用或宿主机重启。
风险与范围
关联 issue
属于 #12380,实现针对 #12670 提出的 W0e-1 决策;物理复用继续由独立停写证明保护。#12766 的生产接管另行交付。与 #12831 协调,不重复 Hosted 工具编排。