Summary
If the Managed Session Store (hosted by Spring) is unreachable while the Hosted Harness is writing a running Turn, the Harness stops writing that Session for good. The Turn can then never complete and never be cancelled. A transient outage becomes a permanent wedge.
Observed by @wenshao during the Round 4 real-stack re-verification of #13163 (comment), who noted it "is a different failure from #13054 and deserves its own issue", and that it reproduces on main — #13163 neither introduces nor fixes it.
Observed behaviour
wenshao's rig, identical on both arms (the PR head and its merge base with main).
Repro: a bound Turn whose fake model streams a text delta every 200 ms for 40 s. Spring is restarted mid-stream. No grant changes.
- The Harness logs
session log writes stopped after an earlier failure: Managed Session Store request failed: fetch failed, then could not settle.
- 35 s after the stream ended, the Turn is still
RUNNING.
- The creator's cancel gets
202. 60 s later the Turn is only CANCELLING.
- The next submit gets
409 turn_active.
Mechanism
packages/core/src/managed-runtime/managed-session-authority.ts. Read statically from source — I did not execute this.
The write path latches on the first failure (line 2152-2160):
try {
await this.journal.appendTransaction(records);
this.lastRecordUuid = final.uuid;
} catch (cause) {
// Records may already be on disk, so the sequences this transaction
// claimed are spent whether or not the marker landed.
this.writeFailure =
cause instanceof Error ? cause : new Error(String(cause));
throw cause;
}
Every later write is then refused (line 1980-1982):
if (this.writeFailure !== undefined) {
...
`session log writes stopped after an earlier failure: ${this.writeFailure.message}`,
writeFailure is declared at line 378, read at 463 / 1980 / 1982, and assigned at exactly one place, line 2157. Nothing ever resets it to undefined. The latch is one-way for the lifetime of the Session authority instance.
The in-code comment shows the latch is deliberate: after a failed appendTransaction the sequences that transaction claimed are spent whether or not the marker landed, so reusing them could corrupt the journal. The gap is that the same permanent latch is applied to a transient cause — a store that is briefly unreachable during a restart — with no retry, no re-probe, and no path back once the store returns.
Impact
- The Turn never reaches a terminal state, so the Session stays
RUNNING and every later submit is refused 409 turn_active.
- Cancellation cannot rescue it: the cancel is accepted (
202) but the settlement write is refused by the same latch, so the Turn parks in CANCELLING.
- One Spring restart during an active Turn is enough. wenshao hit it by chance in a WebShell run where Spring began shutting down about 0.6 s after the Turn started.
What a fix has to reconcile
Recovery must not reuse the sequences the failed transaction already claimed. Directions worth weighing:
- re-probe the store and resume from a freshly claimed sequence range;
- persist or derive the spent-sequence high-water mark so the latch can be lifted safely;
- or escalate the Session to a terminal failed state instead of leaving it indefinitely
RUNNING, so a later submit is not refused turn_active forever.
Scope
中文说明
问题。 Hosted Harness 正在写一个运行中的 Turn 时,如果承载 Managed Session Store 的 Spring 不可达,Harness 会永久停止写这个 Session。此后该 Turn 既不能完成也不能被取消,一次瞬时抖动变成永久卡死。
现象(@wenshao 在真实栈上实测,PR head 与其 main merge base 两个 arm 表现一致):绑定的 Turn 用假模型每 200 ms 输出一个 text delta、持续 40 s,中途重启 Spring,不改任何授权。Harness 打出 session log writes stopped after an earlier failure: Managed Session Store request failed: fetch failed,随后 could not settle;流结束 35 s 后 Turn 仍是 RUNNING;创建者取消拿到 202,60 s 后 Turn 只到 CANCELLING;下一次 submit 拿到 409 turn_active。
机制(我静态读源码得到,未实际执行):packages/core/src/managed-runtime/managed-session-authority.ts 的写入路径在第一次失败时置位 writeFailure(2157 行)并重新抛出;之后所有写入在 1980-1982 行被直接拒绝。该字段声明于 378 行,读取于 463 / 1980 / 1982 行,只有 2157 行这一处赋值,没有任何地方把它重置为 undefined,所以在 Session authority 实例的整个生命周期里是单向闩锁。代码注释说明闩锁是刻意的:appendTransaction 失败后,那次事务已经占用的 sequence 无论 marker 是否落盘都算用掉了,重用可能损坏 journal。缺口在于同一个永久闩锁被套用在瞬时成因上——重启期间短暂不可达的 store——既没有重试,也没有恢复后重新探测的通路。
影响。 Turn 永远到不了终止态,Session 一直停在 RUNNING,之后每次 submit 都被 409 turn_active 拒绝;取消也救不回来:取消被接受(202),但结算写入被同一个闩锁拒绝,Turn 卡在 CANCELLING。一次 Spring 重启就足够触发。
修复要兼顾的约束。 恢复时不得重用失败事务已经占用的 sequence。可考虑的方向:重新探测 store 并从新申请的 sequence 区间继续;持久化或推导出已用 sequence 的高水位,使闩锁可以被安全解除;或者把 Session 直接推进到终止失败态,而不是无限期停在 RUNNING。
范围。 在 main 和 #13163 的 30f092d0 上都能复现,不是该 PR 引入的回归;与 #13054(继承的 recovery-blocked Turn)是不同故障。
Summary
If the Managed Session Store (hosted by Spring) is unreachable while the Hosted Harness is writing a running Turn, the Harness stops writing that Session for good. The Turn can then never complete and never be cancelled. A transient outage becomes a permanent wedge.
Observed by @wenshao during the Round 4 real-stack re-verification of #13163 (comment), who noted it "is a different failure from #13054 and deserves its own issue", and that it reproduces on
main— #13163 neither introduces nor fixes it.Observed behaviour
wenshao's rig, identical on both arms (the PR head and its merge base with
main).Repro: a bound Turn whose fake model streams a text delta every 200 ms for 40 s. Spring is restarted mid-stream. No grant changes.
session log writes stopped after an earlier failure: Managed Session Store request failed: fetch failed, thencould not settle.RUNNING.202. 60 s later the Turn is onlyCANCELLING.409 turn_active.Mechanism
packages/core/src/managed-runtime/managed-session-authority.ts. Read statically from source — I did not execute this.The write path latches on the first failure (line 2152-2160):
Every later write is then refused (line 1980-1982):
writeFailureis declared at line 378, read at 463 / 1980 / 1982, and assigned at exactly one place, line 2157. Nothing ever resets it toundefined. The latch is one-way for the lifetime of the Session authority instance.The in-code comment shows the latch is deliberate: after a failed
appendTransactionthe sequences that transaction claimed are spent whether or not the marker landed, so reusing them could corrupt the journal. The gap is that the same permanent latch is applied to a transient cause — a store that is briefly unreachable during a restart — with no retry, no re-probe, and no path back once the store returns.Impact
RUNNINGand every later submit is refused409 turn_active.202) but the settlement write is refused by the same latch, so the Turn parks inCANCELLING.What a fix has to reconcile
Recovery must not reuse the sequences the failed transaction already claimed. Directions worth weighing:
RUNNING, so a later submit is not refusedturn_activeforever.Scope
mainand on fix(managed-agent): stop a bound Turn under refused authorization #13163 at30f092d0. Not a regression from that PR.中文说明
问题。 Hosted Harness 正在写一个运行中的 Turn 时,如果承载 Managed Session Store 的 Spring 不可达,Harness 会永久停止写这个 Session。此后该 Turn 既不能完成也不能被取消,一次瞬时抖动变成永久卡死。
现象(@wenshao 在真实栈上实测,PR head 与其 main merge base 两个 arm 表现一致):绑定的 Turn 用假模型每 200 ms 输出一个 text delta、持续 40 s,中途重启 Spring,不改任何授权。Harness 打出
session log writes stopped after an earlier failure: Managed Session Store request failed: fetch failed,随后could not settle;流结束 35 s 后 Turn 仍是RUNNING;创建者取消拿到202,60 s 后 Turn 只到CANCELLING;下一次 submit 拿到409 turn_active。机制(我静态读源码得到,未实际执行):
packages/core/src/managed-runtime/managed-session-authority.ts的写入路径在第一次失败时置位writeFailure(2157 行)并重新抛出;之后所有写入在 1980-1982 行被直接拒绝。该字段声明于 378 行,读取于 463 / 1980 / 1982 行,只有 2157 行这一处赋值,没有任何地方把它重置为undefined,所以在 Session authority 实例的整个生命周期里是单向闩锁。代码注释说明闩锁是刻意的:appendTransaction失败后,那次事务已经占用的 sequence 无论 marker 是否落盘都算用掉了,重用可能损坏 journal。缺口在于同一个永久闩锁被套用在瞬时成因上——重启期间短暂不可达的 store——既没有重试,也没有恢复后重新探测的通路。影响。 Turn 永远到不了终止态,Session 一直停在
RUNNING,之后每次 submit 都被409 turn_active拒绝;取消也救不回来:取消被接受(202),但结算写入被同一个闩锁拒绝,Turn 卡在CANCELLING。一次 Spring 重启就足够触发。修复要兼顾的约束。 恢复时不得重用失败事务已经占用的 sequence。可考虑的方向:重新探测 store 并从新申请的 sequence 区间继续;持久化或推导出已用 sequence 的高水位,使闩锁可以被安全解除;或者把 Session 直接推进到终止失败态,而不是无限期停在
RUNNING。范围。 在
main和 #13163 的30f092d0上都能复现,不是该 PR 引入的回归;与 #13054(继承的 recovery-blocked Turn)是不同故障。