fix(managed-agent): ≥8 concurrent Turns stall after the model answers on modest hardware (lock convoy in the store path)
What happened?
Found by bisection with the new --list-pagination mode of
scripts/run-managed-agent-server-e2e.ts (full packaged stack), varying the
number of Sessions created at once on the same tenant:
| Concurrent Turns |
macOS arm64 (M-series) |
Linux aarch64 (Orange Pi 5, 8 cores) |
| 4 |
all complete |
all complete |
| 8 |
all complete |
all stall: 8/8 RUNNING, zero events after turn.started for 120+s |
| 12 |
all complete (~35 s) |
all stall identically |
The stall shape on the rig, twice deterministically:
- all 8 Sessions admitted, all 8 Turns reach
turn.started;
- all 8 model requests reach the (fake) provider and are answered instantly —
so the model call is not the bottleneck;
- zero
message.delta/completion commits follow: the pipeline between
"model answered" and "journal commit" produces nothing for 120+ s;
- the coordinator logs
Managed Turn lease renewal failed … failure=CannotAcquireLockException on the -dispatch-lease thread (3×).
So the Turns don't fail and don't finish — they wedge the same way
#13322 records for a held model stream, here triggered purely by concurrency.
What did you expect to happen?
N concurrent Turns may be slower, but they must keep making progress and
settle; a lock in the store/dispatch path must not convoy every writer
behind one another until the InnoDB lock wait times out. If the intended
shape is a concurrency cap, the admission should queue Turns explicitly
instead of starting all of them into a stall.
Client information
Packaged-stack E2E runner (--list-pagination, twelve Sessions), repo HEAD
1cf80d46d0. Rig: Orange Pi 5 (8× Cortex-A76/A55, 16 GB), Ubuntu 22.04,
MySQL 8.0.45, JDK 21. Compare: the same 12-Turn burst passes on macOS arm64
in ~35 s.
Anything else we need to know?
Parent: #12380 / Stage G (#12952). Same-wedge family: #13322 (held stream),
#13327 (coordinator-only crash). The CannotAcquireLockException on dispatch
lease renewal is the only diagnostic surfaced; it may be a symptom (a Turn's
commit transaction holding the row the renewal wants) rather than the cause.
中文
发生了什么?
用 --list-pagination 模式的并发数二分定位(同一租户、整栈):
| 并发 Turn 数 |
macOS arm64 |
Linux aarch64(Orange Pi 5,8 核) |
| 4 |
全部完成 |
全部完成 |
| 8 |
全部完成 |
全部卡死:8/8 RUNNING,turn.started 后 120+ 秒零事件 |
| 12 |
全部完成(约 35s) |
同样全部卡死(两次确定性复现) |
卡死形态:8 个 Session 全部准入、8 个 Turn 全部 turn.started;8 个模型请求都到达了
(fake)provider 并被即时应答——模型调用不是瓶颈;但之后没有任何
message.delta/完成落账,120+ 秒管线零产出;协调器在 -dispatch-lease 线程报
CannotAcquireLockException(租约续约失败,3 次)。Turn 不失败也不完成——
与 #13322 的挂起同族,只是触发源是并发。
期望行为
N 个并发 Turn 可以变慢,但必须持续推进并结算;store/dispatch 路径上的锁不能把
所有写者堵到 InnoDB 锁等待超时。若设计意图是并发上限,应在准入处显式排队,
而不是把所有 Turn 放进 stall。
fix(managed-agent): ≥8 concurrent Turns stall after the model answers on modest hardware (lock convoy in the store path)
What happened?
Found by bisection with the new
--list-paginationmode ofscripts/run-managed-agent-server-e2e.ts(full packaged stack), varying thenumber of Sessions created at once on the same tenant:
turn.startedfor 120+sThe stall shape on the rig, twice deterministically:
turn.started;so the model call is not the bottleneck;
message.delta/completion commits follow: the pipeline between"model answered" and "journal commit" produces nothing for 120+ s;
Managed Turn lease renewal failed … failure=CannotAcquireLockExceptionon the-dispatch-leasethread (3×).So the Turns don't fail and don't finish — they wedge the same way
#13322 records for a held model stream, here triggered purely by concurrency.
What did you expect to happen?
N concurrent Turns may be slower, but they must keep making progress and
settle; a lock in the store/dispatch path must not convoy every writer
behind one another until the InnoDB lock wait times out. If the intended
shape is a concurrency cap, the admission should queue Turns explicitly
instead of starting all of them into a stall.
Client information
Packaged-stack E2E runner (
--list-pagination, twelve Sessions), repo HEAD1cf80d46d0. Rig: Orange Pi 5 (8× Cortex-A76/A55, 16 GB), Ubuntu 22.04,MySQL 8.0.45, JDK 21. Compare: the same 12-Turn burst passes on macOS arm64
in ~35 s.
Anything else we need to know?
Parent: #12380 / Stage G (#12952). Same-wedge family: #13322 (held stream),
#13327 (coordinator-only crash). The CannotAcquireLockException on dispatch
lease renewal is the only diagnostic surfaced; it may be a symptom (a Turn's
commit transaction holding the row the renewal wants) rather than the cause.
中文
发生了什么?
用
--list-pagination模式的并发数二分定位(同一租户、整栈):卡死形态:8 个 Session 全部准入、8 个 Turn 全部 turn.started;8 个模型请求都到达了
(fake)provider 并被即时应答——模型调用不是瓶颈;但之后没有任何
message.delta/完成落账,120+ 秒管线零产出;协调器在 -dispatch-lease 线程报
CannotAcquireLockException(租约续约失败,3 次)。Turn 不失败也不完成——与 #13322 的挂起同族,只是触发源是并发。
期望行为
N 个并发 Turn 可以变慢,但必须持续推进并结算;store/dispatch 路径上的锁不能把
所有写者堵到 InnoDB 锁等待超时。若设计意图是并发上限,应在准入处显式排队,
而不是把所有 Turn 放进 stall。