Skip to content

Track structured Auto Memory rollout readiness on main #12947

Description

@yiliang114

Goal

Track the remaining correctness, effectiveness, and validation work for structured Auto Memory on main before considering broader rollout.

This is the memory-specific closeout tracker for #10151 and a child track of the broader token-governance umbrella #12028. Structured recall remains intentionally opt-in and disabled by default while this work is incomplete. Completing this checklist enables a separate default-on decision; it does not require or pre-commit that decision.

Baseline reviewed for this tracker: main at 92f4d4f6455bf02a41e0d856afe63b79108f936f.

Landed foundation

Remaining correctness work

Remaining hardening work

Effectiveness validation

  • Close Validate the recall initial-turn budget against production delivery telemetry #8998 with a controlled selector ablation. Measure unique selector deliveries, overlap with deterministic fast recall, already_delivered, no_safe_delivery_point, answer quality, and the extra selector/model requests. Do not use Validate the recall initial-turn budget against production delivery telemetry #8998 as the umbrella tracker.
  • Rerun the matched end-to-end matrix after the correctness fixes land: legacy-only corpus, fully structured corpus, metadata-only answer, body fetch, tool-using interaction, unrelated query, and compression/refetch.
  • Report main-model uncached input, cache-read input, output, request count, selector usage, extraction usage, Dream usage, and migration usage separately. Do not present a lower whole-process total as an optimization when a background task was skipped.
  • Recheck tail latency. The earlier 36-case result improved the median but regressed P95 (191,893 ms to 276,311 ms); the final main implementation needs a matched rerun before rollout.
  • Confirm answer and recall quality do not regress while token reduction remains meaningful on the final implementation. The corrected three-case matched smoke reduced main-model gross input from 73,732 to 62,571 tokens (-15.14%) with 3/3 answers passing. The observed whole-process -27.20% is not a stable claim because extractor request counts differed between arms; this remains directional evidence, not the final acceptance result.

Rollout gate

Keep memory.enableStructuredRecall opt-in until:

  • the correctness items above are fixed;
  • legacy and structured memory paths both pass the end-to-end matrix;
  • the measured token result separates foreground savings from background cost;
  • selector contribution and tail latency have a clear disposition; and
  • no tool-using interaction silently skips required memory lifecycle work.

Default-on remains a product and rollout decision after these gates pass, not an automatic consequence of closing this tracker.

Relationship to other issues

中文说明

目标

这是一份针对当前 main 的结构化 Auto Memory 收口清单,也是整体 token 治理总 issue #12028 下的 memory 子轨道。核心能力已经合入,但正确性、收益与完整链路尚未全部验收,因此 memory.enableStructuredRecall 继续默认关闭。完成本清单只代表具备讨论默认开启的条件,并不代表必须强制开启。

已合入

剩余正确性问题

剩余加固工作

效果验收

  • 在 Validate the recall initial-turn budget against production delivery telemetry #8998 完成 selector controlled ablation,区分真正交付、与 fast recall 重复、already_delivered 和 no_safe_delivery_point。
  • 正确性修复后,重新跑旧记忆、纯新记忆、metadata-only、正文 fetch、工具型交互、无关问题、压缩后重新 fetch 的完整矩阵。
  • 主模型、selector、extraction、Dream、migration 分别记账,不能把后台任务没有执行造成的低 token 当成优化收益。
  • 复查 P95 尾延迟;此前 36 条评测从 191,893 ms 上升到 276,311 ms。
  • 在最终 main 上确认答案与召回质量不退化,并验证 token 下降仍然有意义。修正后的 3 组配对冒烟中,主模型 gross input 从 73,732 降到 62,571(-15.14%),3/3 答案通过;全进程观察值 -27.20% 因两侧 extractor 请求数不等,不能作为稳定收益,只能算方向性证据。

Issue 分工

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions