Skip to content

fix(managed-agent): ≥8 concurrent Turns stall after the model answers on modest hardware (lock convoy in the store path) #13333

Description

@wenshao

fix(managed-agent): ≥8 concurrent Turns stall after the model answers on modest hardware (lock convoy in the store path)

What happened?

Found by bisection with the new --list-pagination mode of
scripts/run-managed-agent-server-e2e.ts (full packaged stack), varying the
number of Sessions created at once on the same tenant:

Concurrent Turns macOS arm64 (M-series) Linux aarch64 (Orange Pi 5, 8 cores)
4 all complete all complete
8 all complete all stall: 8/8 RUNNING, zero events after turn.started for 120+s
12 all complete (~35 s) all stall identically

The stall shape on the rig, twice deterministically:

  • all 8 Sessions admitted, all 8 Turns reach turn.started;
  • all 8 model requests reach the (fake) provider and are answered instantly —
    so the model call is not the bottleneck;
  • zero message.delta/completion commits follow: the pipeline between
    "model answered" and "journal commit" produces nothing for 120+ s;
  • the coordinator logs Managed Turn lease renewal failed … failure=CannotAcquireLockException on the -dispatch-lease thread (3×).

So the Turns don't fail and don't finish — they wedge the same way
#13322 records for a held model stream, here triggered purely by concurrency.

What did you expect to happen?

N concurrent Turns may be slower, but they must keep making progress and
settle; a lock in the store/dispatch path must not convoy every writer
behind one another until the InnoDB lock wait times out. If the intended
shape is a concurrency cap, the admission should queue Turns explicitly
instead of starting all of them into a stall.

Client information

Packaged-stack E2E runner (--list-pagination, twelve Sessions), repo HEAD
1cf80d46d0. Rig: Orange Pi 5 (8× Cortex-A76/A55, 16 GB), Ubuntu 22.04,
MySQL 8.0.45, JDK 21. Compare: the same 12-Turn burst passes on macOS arm64
in ~35 s.

Anything else we need to know?

Parent: #12380 / Stage G (#12952). Same-wedge family: #13322 (held stream),
#13327 (coordinator-only crash). The CannotAcquireLockException on dispatch
lease renewal is the only diagnostic surfaced; it may be a symptom (a Turn's
commit transaction holding the row the renewal wants) rather than the cause.

中文

发生了什么?

用 --list-pagination 模式的并发数二分定位(同一租户、整栈):

并发 Turn 数 macOS arm64 Linux aarch64(Orange Pi 5,8 核)
4 全部完成 全部完成
8 全部完成 全部卡死:8/8 RUNNING,turn.started 后 120+ 秒零事件
12 全部完成(约 35s) 同样全部卡死(两次确定性复现)

卡死形态:8 个 Session 全部准入、8 个 Turn 全部 turn.started;8 个模型请求都到达了
(fake)provider 并被即时应答——模型调用不是瓶颈;但之后没有任何
message.delta/完成落账,120+ 秒管线零产出;协调器在 -dispatch-lease 线程报
CannotAcquireLockException(租约续约失败,3 次)。Turn 不失败也不完成——
与 #13322 的挂起同族,只是触发源是并发。

期望行为

N 个并发 Turn 可以变慢,但必须持续推进并结算;store/dispatch 路径上的锁不能把
所有写者堵到 InnoDB 锁等待超时。若设计意图是并发上限,应在准入处显式排队,
而不是把所有 Turn 放进 stall。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions