You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
fix(managed-agent): residual admission gap-lock deadlock on the shared command index (mutation-admission siblings) #13374
PR #13365 removed the deposition-critical gap-lock shape behind #13333 (insertTurnCommand's FOR UPDATE replay probe) and proved the deterministic N−1 admission deadlock gone. One narrower window of the same InnoDB gap-lock family remains open on the tenant-shared managed_agent_command index, verified by code inspection during that fix's review and disclosed in the PR:
Mutation-admission siblings keep the old locking-probe shape.beginSessionMutation and unarchiveWorkspaceSession still probe the idempotency command row with a FOR UPDATE read that — when the row does not exist — next-key-locks a gap of the tenant-shared index, and their service callers (renameSession, unarchive) have no DuplicateKeyException replay fallback. Two concurrent same-tenant mutations on different sessions can therefore deadlock the same way (one 500 after InnoDB kills the loser), and two same-key concurrent mutations surface a raw 500 instead of an idempotent replay.
Creation scope catch path, same-key triplicate race.CLOSED (2026-10-04, PR fix(managed-agent): stop gap-locked replay probes deadlocking same-tenant admission bursts #13365 head 1f2ffe30a5). The requireCreationScope rewrite that created this window (plain INSERT + DuplicateKeyException catch + FOR UPDATE re-read) deadlocked N−2 same-key racers because each failed duplicate check retains a shared lock on the committed row and the upgrades cross-block — measured on a JDBC probe against MySQL 8.0.46 with 8 barrier-released racers (7/8, 7/8, 6/8 losers killed), while the base ON DUPLICATE KEY UPDATE sequence queues the losers on its exclusive lock (0/8 every round). The hunk was reverted to the base sequence exactly; the window no longer exists. Context in this comment and in the PR delta comments.
What did you expect to happen?
Concurrent same-tenant admissions of every kind — turns, session creates, and lifecycle mutations — must never deadlock or surface lock errors; same-key races must resolve to a deterministic replay or 409 idempotency_conflict, never a 500. For window 1 that means routing the mutation siblings through a service-side replay that preserves beginSessionMutation's FAILED→PENDING resurrection semantics in its existing-row branch.
Client information
Code inspection at main after PR #13365 (ManagedAgentStore.javabeginSessionMutation / unarchiveWorkspaceSession; ManagedAgentService.java rename/unarchive admission paths). The window-1 deadlock is reachable with two concurrent same-tenant mutations, no heavy load required.
A fix for window 1 must keep beginSessionMutation's FAILED→PENDING resurrection (ManagedAgentStore.java existing-row branch) reachable — a naive DuplicateKeyException → replayCommand fallback would drop it — and should extend the burst IT with a mutation-burst round rather than inventing a new harness.
Watch-out: window 2 is a template for what NOT to do — replacing an ON DUPLICATE KEY UPDATE on an existing row with plain INSERT + a FOR UPDATE re-read converts queued losers (exclusive upsert lock) into cross-blocking losers (retained shared lock + lock upgrade). Don't reintroduce that shape.
What happened?
PR #13365 removed the deposition-critical gap-lock shape behind #13333 (
insertTurnCommand'sFOR UPDATEreplay probe) and proved the deterministic N−1 admission deadlock gone. One narrower window of the same InnoDB gap-lock family remains open on the tenant-sharedmanaged_agent_commandindex, verified by code inspection during that fix's review and disclosed in the PR:Mutation-admission siblings keep the old locking-probe shape.
beginSessionMutationandunarchiveWorkspaceSessionstill probe the idempotency command row with aFOR UPDATEread that — when the row does not exist — next-key-locks a gap of the tenant-shared index, and their service callers (renameSession,unarchive) have noDuplicateKeyExceptionreplay fallback. Two concurrent same-tenant mutations on different sessions can therefore deadlock the same way (one 500 after InnoDB kills the loser), and two same-key concurrent mutations surface a raw 500 instead of an idempotent replay.Creation scope catch path, same-key triplicate race.CLOSED (2026-10-04, PR fix(managed-agent): stop gap-locked replay probes deadlocking same-tenant admission bursts #13365 head1f2ffe30a5). TherequireCreationScoperewrite that created this window (plainINSERT+DuplicateKeyExceptioncatch +FOR UPDATEre-read) deadlocked N−2 same-key racers because each failed duplicate check retains a shared lock on the committed row and the upgrades cross-block — measured on a JDBC probe against MySQL 8.0.46 with 8 barrier-released racers (7/8, 7/8, 6/8 losers killed), while the baseON DUPLICATE KEY UPDATEsequence queues the losers on its exclusive lock (0/8 every round). The hunk was reverted to the base sequence exactly; the window no longer exists. Context in this comment and in the PR delta comments.What did you expect to happen?
Concurrent same-tenant admissions of every kind — turns, session creates, and lifecycle mutations — must never deadlock or surface lock errors; same-key races must resolve to a deterministic replay or 409
idempotency_conflict, never a 500. For window 1 that means routing the mutation siblings through a service-side replay that preservesbeginSessionMutation'sFAILED→PENDINGresurrection semantics in its existing-row branch.Client information
Code inspection at
mainafter PR #13365 (ManagedAgentStore.javabeginSessionMutation/unarchiveWorkspaceSession;ManagedAgentService.javarename/unarchive admission paths). The window-1 deadlock is reachable with two concurrent same-tenant mutations, no heavy load required.Anything else we need to know?
HostedConcurrentTurnBurstMySqlIT.beginSessionMutation'sFAILED→PENDINGresurrection (ManagedAgentStore.javaexisting-row branch) reachable — a naiveDuplicateKeyException→replayCommandfallback would drop it — and should extend the burst IT with a mutation-burst round rather than inventing a new harness.ON DUPLICATE KEY UPDATEon an existing row with plainINSERT+ aFOR UPDATEre-read converts queued losers (exclusive upsert lock) into cross-blocking losers (retained shared lock + lock upgrade). Don't reintroduce that shape.中文
发生了什么?
PR #13365 移除了 #13333 的关键 gap-lock 形态(
insertTurnCommand的FOR UPDATE回放探针),并验证确定性 N−1 准入死锁已消除。租户共享的managed_agent_command索引上仍有一个更窄的同族锁窗口(修复评审期经代码核查并在 PR 中披露):mutation 准入兄弟路径仍是旧锁探针形态。
beginSessionMutation与unarchiveWorkspaceSession仍用FOR UPDATE读探测幂等命令行——行不存在时会对租户共享索引的间隙加 next-key 锁,且其 service 调用方(renameSession、unarchive)没有DuplicateKeyException回放兜底。两个同租户不同会话的并发 mutation 可按同一机制死锁(InnoDB 判死一个,客户端收 500);同键并发 mutation 则直接表面为裸 500,而非幂等回放。创建作用域 catch 路径的同键三连竞态。已关闭(2026-10-04,PR fix(managed-agent): stop gap-locked replay probes deadlocking same-tenant admission bursts #13365 head1f2ffe30a5)。 产生该窗口的requireCreationScope重写(普通 INSERT + 捕获重复键 +FOR UPDATE复查)会让同键竞态死锁 N−2 个输家——每个输家的失败重复检查对已提交行保留共享锁,升级交叉阻塞(JDBC 探针在 MySQL 8.0.46 上以 8 个屏障齐发竞者实测 7/8、7/8、6/8 个被判死);而基线ON DUPLICATE KEY UPDATE序列让输家在独占锁上排队(每轮 0/8 死锁)。该 hunk 已逐字节回退至基线,窗口不复存在。背景见本评论与 PR 的增量评论。期望行为?
任何形式(Turn、Session 创建、生命周期 mutation)的同租户并发准入都不得死锁或表面锁错误;同键竞态必须确定性结算为回放或 409
idempotency_conflict,绝不返回 500。对窗口 1 即:把 mutation 兄弟路径接入保留beginSessionMutation既有行分支FAILED→PENDING复活语义的 service 侧回放。客户端信息
PR #13365 之后
main上的代码核查(ManagedAgentStore.java的beginSessionMutation/unarchiveWorkspaceSession;ManagedAgentService.java的 rename/unarchive 准入路径)。窗口 1 在中等并发下两个并发同租户 mutation 即可达,无需高负载。其他需要知道的事?
HostedConcurrentTurnBurstMySqlIT。beginSessionMutation既有行分支的FAILED→PENDING复活可达(朴素的DuplicateKeyException→replayCommand会把它丢掉),并应给突发 IT 增加 mutation-burst 轮次,而不是另造 harness。ON DUPLICATE KEY UPDATE换成普通INSERT+FOR UPDATE复查,会把排队的输家(upsert 独占锁)变成互相阻塞的输家(保留共享锁 + 锁升级)。不要再引入该形态。