Skip to content

fix(managed-agent): bounded growth for managed session stores and panel projection #13184

Description

@wenshao

What happened?

"Unbounded in the time dimension" is the default posture across the managed-agent persistence and UI layers — everything only grows. Audit-verified against main (a7deb01):

TS core (packages/core/src/managed-runtime):

  1. The session authority keeps every event of a session in process memory for its lifetime (managed-session-authority.ts:346, :361, :1910 — push-only), and the read paths replay everything: the local JSONL journal reads the whole file per open (local-jsonl-managed-session-journal-store.ts:70), the HTTP store pages the whole journal from revision 0 and concatenates (http-managed-session-store.ts:436-498), and each transcript-chain rebuild re-reads the full transcript plus every referenced resource body (chatRecordingService.ts:1959 → config.ts:445).
  2. The activation store and inbox journals never compact: released activations stay in a permanent in-memory map and an append-only JSONL journal that a restart replays in full (managed-activation-store.ts:265, :528, :288-320; managed-session-inbox.ts:385). Lease renewals append a durable journal record roughly every 100 seconds per open session for as long as the process lives (managed-session-assembly.ts:33-46).
  3. Process-level maps never evict: dispatch gates keep settled/handed-off entries forever (managed-runtime-dispatch-gate.ts:31-46, :63), the operation grant gate documents entries staying until process end (managed-operation-grant-gate.ts:93-101), shell-result prepared entries accumulate (managed-shell-result-session.ts:66).
  4. Hardening is uneven between sibling stores: the tool-result store has pending+link durability, directory symlink refusal, O_NOFOLLOW opens, and 0700 directories, while the adjacent session-resource store has no size cap and weaker directory checks (managed-session-resources.ts:64-107 vs local-managed-tool-result-store.ts:107-114). Hot read paths verify by re-scanning instead of using an index (resource-tool-result-store.ts:264-291 full page reads for any range; managed-shell-result-session.ts:132-150, :219-226, :360-388 repeated O(history) scans and re-hashing per shell call).

Web Shell (packages/web-shell/client/components/managed):

  1. Message projection re-reduces the entire event array on every delta (managedEventsToMessages called per state change), concatenating the full assistant text per delta; the actions hook re-scans the whole event array per event; loaded older pages accumulate without bound (the merge side was optimized in fix(managed-agent): harden the managed panel failure lifecycle and worker path containment #13179 — this remaining piece is the dominant per-event cost).
  2. A single corrupt persisted SSE frame wedges a session panel forever: JSON.parse failure propagates out of the frame loop (java-managed-agent-client.ts:448), resubscribing from the same cursor replays the same frame every 3 seconds.
  3. A stream-gap resync replaces the event array wholesale (use-managed-session.ts gap branch), silently discarding history pages the user had loaded and scrolled into.

What did you expect to happen?

Bounded growth as a design default: lightweight indices plus checkpoint/compaction windows instead of full-event retention, snapshot+truncate for append-only journals, TTL/LRU eviction for settled gate/grant/prepared entries, a stop condition for renewal loops on persistent failure or write-stop, one shared set of hardened directory primitives for all durable stores, size caps at every publish boundary, indexed lookups over kind/executionCallId instead of full scans, and an incremental UI projector plus gap handling that preserves loaded history and skips corrupt frames with a bounded failure budget.

Client information

N/A — code audit finding.

Anything else we need to know?

None of these bite in short demo sessions; they are the difference between the feature surviving long-running hosted tenants and not. Fix order suggestion: (1) and (4) first — they shape the disk and memory envelope of every session; (5)-(7) are independent and web-side only.

中文

发生了什么?

「时间维度无界」是 managed-agent 持久层与 UI 层的默认姿态——一切只增不减。全部对照 main(a7deb01bcb)审计核实:

TS core(packages/core/src/managed-runtime):

  1. 会话 authority 把会话全部事件常驻进程内存(managed-session-authority.ts:346, :361, :1910 只增不删);读路径整体重放:本地 JSONL journal 每次 open 整文件重读(local-jsonl-managed-session-journal-store.ts:70),HTTP 店从 revision 0 分页全量拉取拼接(http-managed-session-store.ts:436-498),transcript 链重建每次重读全部 transcript 与所有引用资源体(chatRecordingService.ts:1959 → config.ts:445)。
  2. activation store 与 inbox 永不压缩:已释放激活永久驻留内存 Map 和 append-only JSONL journal,重启全量回放(managed-activation-store.ts:265, :528, :288-320;managed-session-inbox.ts:385)。续约定时器给每个打开的会话约每 100 秒写一条持久 journal 记录,随进程寿命无限累积(managed-session-assembly.ts:33-46)。
  3. 进程级 Map 群不淘汰:dispatch gate 的 settled/handed_off 条目永久留存(managed-runtime-dispatch-gate.ts:31-46, :63)、operation grant gate 注释自述条目驻留到进程结束(managed-operation-grant-gate.ts:93-101)、shell prepared 条目累积(managed-shell-result-session.ts:66)。
  4. 姊妹存储的加固不均:tool-result 店有 pending+link 耐久、目录拒 symlink、O_NOFOLLOW、0700 目录,相邻的会话资源店无 size 上限、目录检查更弱(managed-session-resources.ts:64-107 对 local-managed-tool-result-store.ts:107-114)。热读路径靠全量重扫而非索引(resource-tool-result-store.ts:264-291 任意 range 全页读;managed-shell-result-session.ts:132-150, :219-226, :360-388 每次 shell 调用约三遍 O(历史) 扫描与重哈希)。

Web Shell(packages/web-shell/client/components/managed):

  1. 消息投影对每条 delta 全量重归约整个事件数组(每次 state 变化都调 managedEventsToMessages),assistant 文本每条 delta 全量重拷;actions hook 对每事件全扫一次事件数组;翻出的历史页无上限累积(合并侧已在 fix(managed-agent): harden the managed panel failure lifecycle and worker path containment #13179 优化,这部分是每事件成本的主体)。
  2. 一帧损坏的持久化 SSE 帧会让会话面板永久停摆:帧循环里 JSON.parse 失败直接上抛(java-managed-agent-client.ts:448),用同一游标重订阅后每 3 秒重放同一坏帧。
  3. stream_gap 重同步整体替换事件数组(use-managed-session.ts gap 分支),静默丢弃用户已翻出并正在阅读的历史页。

期望行为

有界增长成为默认设计:轻量索引 + checkpoint/压缩窗口替代全事件驻留;append-only journal 做 snapshot+truncate;settled 的 gate/grant/prepared 条目 TTL/LRU 淘汰;续约循环在持续失败或写入停止时停表;所有持久存储共享一套加固目录原语;publish 边界统一 size 上限;按 kind/executionCallId 建索引替代全表扫描;UI 投影增量化;gap 处理保留已加载历史;坏帧跳过并受失败预算约束。

其他

短 demo 会话里这些都无感;它们决定的是长托管租户形态能否存活。建议修复顺序:(1)(4) 先行——它们划定每个会话的磁盘与内存包络;(5)-(7) 相互独立,纯前端。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions