Skip to content

RFC: Reliable auto-memory recall — timing, quality, and telemetry #7040

Description

@jifeng

Status

PR Scope State
1 Recall delivery telemetry Merged — #7393
2 Bounded initial-turn recall + deterministic fast path In review — #8716
3 Recall precision and multilingual evaluation In review — #8716

PR 2's design was amended after measurement: the two-result Fast/Refined
architecture originally specified here is replaced by a minimal deterministic
fast path. This body describes the current normative design. The decision
record for the change, with the measurements behind it, is in
this comment.

Normative design documents, committed in #8716:
docs/design/2026-08-08-native-memory-recall-reliability.md and
docs/design/2026-08-09-bounded-memory-recall-candidates.md.

Decision and scope

After feedback from the Core memory maintainer, this RFC is intentionally narrowed. Qwen Code Core should improve the recall path that benefits every user without becoming an enterprise memory-governance platform.

The Core roadmap contains three independently reviewable PRs:

  1. Add recall-delivery telemetry.
  2. Add bounded initial-turn recall with a deterministic fast path.
  3. Improve multilingual recall precision and add a small labeled regression set.

The previous 13-PR governance roadmap is superseded. Candidate staging, approval inboxes, enterprise audit logs, organization identity, policy-based secret governance, provenance schemas, expiry, supersession, and conflict workflows are no longer proposed for Core.

Why these Core changes are needed

  • Memory reaches the first prompt only by luck, and a bounded wait alone does not fix it. Recall starts on UserQuery, but the initial request performed a zero-wait poll. Adding a fixed initial budget was the first attempt at a fix and is necessary but not sufficient: recall awaits the model selector, a network side query with a 30 s abort ceiling, so the budget is dominated by round-trip time rather than the incidental scheduler timing it was sized for. It therefore expires on the common path.
  • Tool-free turns then get nothing at all. On budget expiry, delivery falls through to the ToolResult point. A turn that makes no tool call never reaches one, so the result is discarded as no_safe_delivery_point. These are exactly the short, context-answered questions where user-level memory matters most.
  • The local heuristic was not an actual fast path. It ran only after the model selector failed or timed out, so it could not help the initial turn — and it was ASCII-oriented, so it was weak for CJK queries.
  • No-lexical-match documents could still score. Any non-empty body earned a positive score, so unrelated documents could be selected.
  • Topic documents were truncated to the 200 most recent per scope before relevance selection. The scanner already reads and parses the files, so the cap created deterministic misses without avoiding most I/O. Note the fix is a change of truncation key rather than a lifted ceiling — see the candidate-set note under PR 3.

Telemetry before PR 1 recorded selection but could not prove that selected memory was delivered to the main model.

Core implementation plan

PR 1: Recall delivery telemetry

Keep the existing selection event and add a low-cardinality delivery event that records:

  • phase, delivery point, a fixed discard reason, strategy, document count, and latency.

Do not record query text or hashes, memory content, file or project paths, session or message identifiers, model reasoning, secrets, or raw error text.

phase and strategy are orthogonal and neither subsumes the other.

  • phase is the delivery stage: fast = the deterministic result injected when the initial budget expired; refined = the model-selected result.
  • strategy is the selection method: none / heuristic / model.
  • A fast delivery is always heuristic. A refined delivery is model normally, and heuristic when the selector failed and the fallback ran. Reading delivery stage off strategy alone would merge "the deterministic result arrived first" with "the selector broke".
  • delivery_point: initial / tool_result / discarded.
  • discard_reason: no_safe_delivery_point / new_query / reset / abort / shutdown / no_relevant_results / already_delivered.

PR 2: Bounded initial-turn recall with a deterministic fast path

One recall lifecycle and model-primary selection, with a single deterministic delivery stage in front of it:

  • Give user-query recall a small internal ceiling, determined by benchmark rather than exposed as public configuration. The wait ends on whichever comes first: recall settling, the deterministic result being published, cancellation, or the ceiling. Benchmarks showed the ceiling was otherwise a fixed per-turn tax: the deterministic result is published once the memory tree has been scanned, which is tens of milliseconds for an ordinary tree, while the selector round trip is assumed to miss the budget entirely.
  • A result settling inside the budget is delivered in the initial prompt.
  • On budget expiry, deliver the deterministic result rather than nothing. selectModelCandidateDocuments already computes lexically ranked, active-tool-filtered candidates in order to build the model manifest; the fast result reuses them, so it costs no extra scan or I/O. Cap it at two documents (MAX_FAST_RECALL_DOCS), well below the five-document prompt limit, because it carries no model judgement.
  • Leave recall pending after a fast delivery so the model-selected result still lands at the existing same-query ToolResult delivery point.
  • Exclude documents the fast phase already delivered from that later delivery and rebuild the prompt from the remainder. Both results come from one scan, so the selector never saw the fast documents as excluded and can legitimately re-select them. When nothing remains, record already_delivered.
  • Do not abort recall merely because the budget expired.
  • A new Query, Reset, Abort, or Shutdown cancels the recall plan. A cancelled turn delivers no fast result. Results must never cross query boundaries or inject the same document twice. The prompt obeys both document-count and aggregate-size budgets.
  • MAX_RELEVANT_DOCS = 5 bounds one prompt, not one turn. A fast delivery of two documents plus a refined delivery of five disjoint ones puts seven in front of the model; deduplication removes repeats, not the sum. This is the direct consequence of dropping combined fast/refined budget accounting, which this RFC originally specified as a fill-to-five limit across both phases. Both prompts stay individually bounded and each body is still truncated, so the worst case is bounded and small — it is simply not five. If a hard aggregate ceiling is ever wanted, the cheap version is passing limit - fastDeliveredPaths.size as the refined limit rather than reintroducing a second budget.

Explicitly not adopted: the two-result shared-scan Fast/Refined architecture, i.e. a second selection pathway with its own scan plumbing, its own budget accounting, and cross-phase fill-to-five logic. The delivery guarantee is worth having; that machinery is not, and cross-phase bookkeeping is the source of the duplicate-injection bug class this RFC already warned about. One callback plus one exclusion set achieves the same guarantee.

No public configuration is added.

PR 3: Recall precision and multilingual evaluation

Improve the deterministic scorer — which now serves both the fast path and the selector-failure fallback — without introducing a persistent index or new retrieval dependency:

  • NFKC normalization; whole-run tokens for non-CJK letters, marks, and digits of at least three characters (\p{L}-based rather than [a-z0-9], so Cyrillic, Greek, Arabic, and accented Latin produce tokens instead of none, with CJK excluded per character so a Latin-initial run cannot swallow the CJK that follows it); Unicode code-point bigrams for Han, Hiragana, Katakana, and Hangul runs; no broad match from a single CJK character; no score without a lexical match; title and description matches weighted above body matches; type boosts applied only after a real match (no scope boost is implemented — scope precedence is carried by ordering, not by score); score ties broken by recency and then by input order, never by document type, since an alphabetical type comparison ranks user last and the two-document fast result would drop it; bounded query tokens retaining both query edges; scoring limited to the surfaced body window; existing active-tool noise filtering retained.

Replace the per-scope 200-document recency truncation before recall ranking with a global, query-aware candidate set (non-recall callers keep the capped scanner). To keep the model manifest bounded, run lightweight local ranking over all parsed topics, select a limited candidate set, and reserve part of that set for recent documents so semantic-but-not-lexical matches still have a chance. Candidate count and manifest bytes remain bounded internal constants.

Read the effect per pool size; it is not a uniform widening. At or under 200 documents nothing was excluded by count under either design, but the 25,000-byte manifest budget is a ceiling the old path lacked. Between 200 and 400 with neither scope over 200, the old path sent every document to the selector and the new one sends at most 200, so fewer reach the model — the survivors are chosen by relevance plus a recency reserve rather than by recency alone, which is the intended trade, but the raw count goes down. Only a scope over 200 is the case this is for, where an old but lexically matching document was permanently invisible. The manifest budget also packs rather than prefixes: a document whose line does not fit is skipped and later, shorter lines are still considered.

Maintain a labeled fixture covering Chinese, English, Japanese, Korean, mixed queries, NFKC, body-only matches, semantic/no-lexical cases, alphabetic scripts outside ASCII and CJK, and no-result cases. Report the corpus size and the Recall@5 a query-blind random scorer would achieve on it, so a small corpus cannot flatter the headline. The more-than-200-topics case is covered by memoryLifecycle.integration.test.ts against a real temporary memory tree, and active-tool noise by recall.test.ts, since neither is reachable from the pure scoring function the fixture evaluates.

Track Recall@5, top-1 accuracy, no-result precision, no-result recall, candidate-reduction ratio, and scan/fast/refined latency. "Fast-path precision" now means the precision of the deterministic delivery rather than of a separate Fast selector. BM25, persistent catalogs, vector databases, language-specific tokenizers, and new dependencies require separate evidence and a separate proposal.

Enterprise memory extension boundary

Enterprise review, identity, audit, retention, and secret-policy requirements should be implemented as an optional Memory Extension with its own storage and write path, rather than wrapping or taking over Core managed memory.

The existing extension system is sufficient for a first implementation: UserPromptSubmit / SessionStart hooks can inject additional context; Stop / SessionEnd hooks can enqueue asynchronous extraction; extension-owned MCP write tools can enforce organization secret and compliance policies; optional PreToolUse hooks can inspect generic tool calls, without assuming they intercept Core-internal writes; MCP tools can expose Search, Propose, Approve, Reject, and Audit; the enterprise service or extension owns review UI, identity, authorization, retention, and audit logs.

This RFC does not propose a dedicated Core memory-plugin API. If a real extension implementation identifies a missing generic capability, that API gap should be proposed independently and kept minimal.

Explicit non-goals

  • Candidate staging or Memory Inbox in Core.
  • A second scan, second selector, or separate fast/refined budget accounting.
  • Enterprise identity, tenant, audit, or policy engines in Core.
  • Memory Schema v2 or bulk migration.
  • Expiry, supersession, and conflict workflows.
  • A unified rewrite of existing secret guards.
  • BM25, vector databases, or a persistent catalog.
  • Public retrieval-mode or enterprise-policy configuration.
  • Changes to current Remember, Forget, Extraction, DREAM, Team Memory, or write semantics.

Verification and rollout

  1. Land delivery telemetry first. — Done (feat(core): add memory recall delivery telemetry #7393).
  2. Land bounded initial recall with internal thresholds and focused lifecycle tests. — Done in fix(memory): improve recall reliability and candidate coverage #8716, including the deterministic fast path and delivery-stage tests for the tool-free turn, later ToolResult delivery, fast/refined dedupe, cancellation inside the initial window, and no cross-query leakage.
  3. Land multilingual precision changes only after the labeled fixture shows no regression in English Recall@5 or no-result precision. — Gate measured and satisfied, see below.
  4. Keep model manifests, document count, aggregate prompt size, and initial-turn latency bounded. — Done, with a measured caveat. Deterministic scoring costs p50 0.036 ms / p95 0.053 ms, but scoring was never what decided initial-turn latency: the memory-tree scan is, and recall-scan-latency.test.ts measures it at ~29 ms (200 topics), ~70 ms (500), ~130 ms (1000). The initial wait is therefore a ceiling that ends on the deterministic result; past ~1000 topics in one scope the scan alone exceeds it and the turn delivers nothing, which is recorded as a limitation rather than fixed.
  5. Keep enterprise extension work in a separate project so it does not block Core recall improvements. — Held; continues under proposal(memory): Define an enterprise external-memory integration profile #7449 / proposal: Add a direct external context provider profile #7585.

Each PR must independently pass focused package tests, build, typecheck, self-audit, and review.

Measured results

Recall quality (packages/core/src/memory/recall-eval.test.ts, 51-case / 25-document labeled corpus scored by both the shipped scorer and a frozen copy of the pre-change scorer, so "no regression" is reproducible rather than asserted). A query-blind scorer returning 5 random documents scores 20.0% Recall@5 on this pool, which is the floor the measured numbers should be read against:

slice metric before after
overall (n=51) Recall@5 45.2% 92.9%
overall (n=51) top-1 accuracy 46.2% 97.4%
overall (n=51) no-result precision 16.0% 75.0%
overall (n=51) no-result recall 44.4% 100.0%
english (n=11) Recall@5 100.0% 100.0%
english (n=11) top-1 accuracy 100.0% 100.0%
cjk (n=17) Recall@5 0.0% 100.0%
cjk (n=17) top-1 accuracy 0.0% 100.0%
mixed (n=5) Recall@5 80.0% 100.0%
mixed (n=5) top-1 accuracy 60.0% 100.0%
other-script (n=3) Recall@5 33.3% 100.0%
other-script (n=3) top-1 accuracy 33.3% 100.0%
semantic-no-lexical (n=3) Recall@5 0.0% 0.0%

No-result precision measured over a no-result-only slice is 100% for any scorer that ever stays silent, so no-result recall — of genuinely unanswerable queries, how many got an empty answer — is the metric that detects the old scorer's actual failure.

The overall row is held below 100% by the semantic-no-lexical slice this RFC asked the fixture to cover: answerable queries sharing no token with their document. An earlier revision of the corpus labeled a no-result case with that category, so nothing measured the cost of "no score without a lexical match". With three genuine cases added, both the shipped scorer and the frozen one score 0% on the slice — the rule did not create the gap, but it does keep the deterministic path silent there. The slice is asserted separately and excluded from the quality floor; the overall no-result precision figure moves for the same reason, since staying silent on an answerable query counts against it.

Delivery (packages/core/src/memory/recall-delivery-eval.test.ts). Selector latency is modelled, not measured — a network round trip cannot be timed in a unit test — so results are reported per scenario; the structural claim holds for every scenario above the budget:

selector latency metric before after
inside budget first-turn delivery (tool-free) 92.9% 92.9%
above budget first-turn delivery (tool-free) 0.0% 92.9%
above budget delivered at all (tool-free) 0.0% 92.9%
above budget delivered at all (tool-using) 92.9% 92.9%
above budget duplicate delivery 0.0% 0.0%

The residual 7.1% in every row is the semantic-no-lexical slice, asserted separately at 0%: the fast path closes the timing gap, not the matching gap. The stand-in for the model selector is the deterministic top-5, so the simulation cannot answer those queries either — which is why the inside-budget row is 92.9% rather than 100% for both designs.

Tool-using turns were never broken — they receive memory one request later — so the fix is targeted at the tool-free gap. Fast/refined overlap occurs in 92.9% of above-budget cases by construction (the harness stands in for the model's choice with the deterministic top-5, which always contains the fast top-2), and dedupe still yields 0% duplicate delivery. That last figure is a property of the simulation, which does not drive tryConsumeMemoryPrefetch; the shipped dedupe is covered by client.test.ts, verified to fail when the exclusion filter is removed.

Known limitations

  • Selector latency is modelled rather than measured.
  • The fast result carries no model judgement. Two documents bounds the cost of being wrong, but on a tool-free turn where the selector never lands, a mis-ranked fast document is what the model sees.
  • Scoring is substring-based, so a query token can match inside a longer word ("owner" inside "ownership"). The evaluation corpus records one such case rather than hiding it.
  • A query sharing no token with its document produces no deterministic result, so a tool-free turn asking it still ends with nothing delivered. Only the model selector can serve those, and on a tool-free turn it never lands.
  • Past roughly a thousand topics in one scope, the memory-tree scan alone exceeds the initial ceiling, so the turn spends the whole budget and still delivers nothing — worse than the zero-wait behaviour it replaced. Ending the wait on the deterministic result bounds this rather than removing it; a persistent catalog is the actual fix and stays out of scope.
  • Scripts written without word separators and outside the CJK set (Thai, Khmer, Lao) now produce a token where they produced none, but the token is the whole run. That is not segmentation.

The previous broad design at b6dec4a is retained for history but is no longer normative.

详细中文方案(评审与合并参考)

状态

PR 范围 状态
1 Recall Delivery Telemetry 已合入 — #7393
2 有界首轮 Recall + 确定性 Fast Path 评审中 — #8716
3 召回精度与多语言评估 评审中 — #8716

PR 2 的设计在实测后已修订:原先规定的双结果 Fast/Refined 架构,替换为最小化的确定性 Fast Path。本文描述的是当前规范设计;变更的决策记录与实测依据见该评论。

规范设计文档已随 #8716 提交:docs/design/2026-08-08-native-memory-recall-reliability.md 与 docs/design/2026-08-09-bounded-memory-recall-candidates.md。

决策与范围

根据 Core Memory 负责同学的反馈,本 RFC 明确收缩范围:Qwen Code Core 只解决所有用户都会遇到的召回时机、召回精度和可观测性问题,不在 Core 内建设企业 Memory 治理平台。

Core 路线图为 3 个可独立评审和回滚的 PR:

  1. Recall Delivery Telemetry;
  2. 有界的首轮 Recall(含确定性 Fast Path);
  3. 多语言召回精度和小规模标注回归集。

原方案中的 Candidate Staging、Memory Inbox、Apply Journal、企业审计日志、组织身份、企业秘密策略、Provenance Schema、Expiry、Supersession 和 Conflict Workflow 不再进入 Core 路线图。

现有 Remember、Forget、Extraction、DREAM、User/Project/Team Scope 和 Team Secret Guard 保持现状。

当前 Core 问题

1. 首轮召回依赖偶然时序,且仅靠有界等待无法解决

Client 在 UserQuery 后启动 Recall Prefetch,首轮请求前原本只做零等待检查。加入固定初始预算是第一步,但并不充分:当存在 Config 时(常规情况)Recall 会等待 Model Selector,而它是一次网络 Side Query,中止上限为 30 秒。因此该预算由网络往返时间主导,而不是它原本针对的偶然调度抖动,常规路径上必然超时。

2. 超时后,无工具调用的轮次完全拿不到 Memory

预算超时后投递会退到 ToolResult 时机。没有工具调用的轮次永远不会到达该时机,结果被记为 no_safe_delivery_point 丢弃。而这类"短问答、靠上下文回答"的轮次,恰恰是用户级 Memory 最重要的场景。

3. Heuristic 不是实际 Fast Path

它只在 Model Selector 失败或超时后执行,无法解决首轮时机问题;且以 ASCII 为主,对 CJK Query 效果很差。

4. 无匹配文档也可能得到分数

只要正文非空就加分,没有任何词面匹配的文档仍可能进入结果。新方案要求没有 Lexical Match 时分数必须为 0。

5. 按 scope 各留最近 200 篇的截断发生在相关性判断之前

Scanner 已经读取并解析 Markdown 才截断,造成漏召回却没有节省主要读取成本。注意修复方式是换截断依据而非抬高上限,效果分档见 PR 3。

PR 1:Recall Delivery Telemetry

保留现有 Selection Event,新增低基数 Delivery Event,记录 phase、delivery point、固定枚举的丢弃原因、strategy、文档数量和耗时。Telemetry 禁止记录 Query 或 Hash、Memory 正文、文件路径、Project Path、Session/Message ID、模型 Reasoning、Secret 或原始错误正文。

phase 与 strategy 正交,互不替代:

  • phase 是投递阶段:fast = 预算超时后注入的确定性结果;refined = Model Selector 选出的结果。
  • strategy 是选择方式:none / heuristic / model。
  • fast 投递必然是 heuristic;refined 投递常规为 model,在 Selector 失败走 Fallback 时为 heuristic。仅凭 strategy 判断阶段,会把"确定性结果先到"与"Selector 故障"混为一谈。
  • delivery_point:initial / tool_result / discarded。
  • discard_reason:no_safe_delivery_point / new_query / reset / abort / shutdown / no_relevant_results / already_delivered。

PR 2:有界的首轮 Recall(含确定性 Fast Path)

保持单一 Recall 生命周期与 Model 优先选择,在其前面增加一个确定性投递阶段:

  • 首轮 UserQuery 前等待一个很小的内部上限,数值由 Benchmark 决定,不作为公开配置。等待以先到者为准结束:Recall 完成、确定性结果发布、取消,或到达上限。实测表明否则这个上限就是每轮固定的税:确定性结果在扫完 Memory 树后就会发布(普通树是几十毫秒),而 Selector 往返按本设计的假定根本赶不上预算。
  • 预算内完成的结果直接进入首轮 Prompt。
  • 预算超时时投递确定性结果,而不是什么都不投。 selectModelCandidateDocuments 为构建 Model Manifest 本就计算了按词面排序、且已过滤 Active Tool 噪声的候选;Fast 结果直接复用,不产生额外扫描或 I/O。上限两个文档(MAX_FAST_RECALL_DOCS),远低于五个文档的 Prompt 上限,因为它没有模型判断背书。
  • Fast 投递后 Recall 继续保留,Model 选出的结果仍在同 Query 的 ToolResult 时机投递。
  • 该次投递需排除 Fast 阶段已投递的文档,并按剩余文档重建 Prompt。两个结果来自同一次扫描,Selector 并未把 Fast 文档视为已排除,重复选中是正常的。若无剩余,记录 already_delivered。
  • 不因预算超时而中止 Recall。
  • 新 Query、Reset、Abort、Shutdown 取消 Recall Plan;被取消的轮次不投递 Fast 结果。结果不得跨 Query 投递或重复注入同一文档,Prompt 同时遵守文档数量与聚合大小预算。
  • MAX_RELEVANT_DOCS = 5 限制的是单次注入,不是单轮总量。 Fast 投递 2 篇加上 refined 投递 5 篇不重叠文档时,本轮进入模型的是 7 篇;去重只消除重复,不压缩总和。这是放弃跨阶段预算核算的直接结果——本 RFC 原方案规定的是两阶段合计五篇。两次 Prompt 各自有界、每篇 body 仍截断,因此最坏情况有界且不大,只是不等于 5。若确需硬性总量上限,最省事的做法是把 refined 的 limit 传成 limit - fastDeliveredPaths.size,而不是重新引入第二套预算。

明确不采用:双结果共享扫描的 Fast/Refined 架构,即带独立扫描链路、独立预算核算和跨阶段补齐到五篇逻辑的第二条选择通路。投递保证值得保留,这套机制不值得——跨阶段簿记正是本 RFC 早已警示的重复注入缺陷来源。一个回调加一个排除集合即可获得同样保证。

不增加公开配置。

PR 3:召回精度与多语言评估

确定性评分器(现同时服务 Fast Path 与 Selector 失败 Fallback)改进:NFKC 规范化;至少 3 个字符的非 CJK 字母/组合符/数字整串 Token(基于 \p{L} 而非 [a-z0-9],因此西里尔、希腊、阿拉伯和带重音拉丁文都能产生 Token;CJK 逐字符排除,避免拉丁开头的串吞掉后面的 CJK);Han、Hiragana、Katakana、Hangul 连续片段的 Unicode Code Point Bigram;单个 CJK 字符不做宽泛匹配;没有 Lexical Match 时得分为 0;Title/Description 匹配权重高于 Body;Type 仅在真实匹配后加权(Scope 没有单独加权,Scope 优先级由排序而非分数承载);同分时按 mtime 降序再按输入顺序,不按文档 Type——Type 字典序会把 user 排到最后,而两篇的 Fast 结果会把它整个丢掉;Query Token 有界且保留首尾;只对可进入 Prompt 的正文窗口评分;保留 Active Tool 降噪。

Recall 路径把「按 scope 各留最近 200 篇」的截断,换成全局、感知 Query 的候选集(其他调用方保留原有 capped scanner)。为避免 Manifest 无限增长:对全部 Topic 轻量本地评分 → 选择有限 Model Candidate Set → 优先高分词面匹配 → 保留部分近期文档额度 → Selector 只对有限集合精排。候选数量与 Manifest Bytes 均为有界内部常量。

效果要分档看,不是一律放宽。 总量 ≤200 时两种设计都不按数量丢弃,但 25,000 字节的 Manifest 上限是旧路径没有的。总量在 200–400 且单个 Scope 不超 200 时,旧路径会把全部文档送给 Selector,新路径最多送 200 篇,候选变少——留下来的是按相关性 + recency reserve 选的而不是纯按时间,这是有意的取舍,但绝对数量确实下降。只有单个 Scope 超过 200 时才是这次真正要解决的场景:老而词法命中的文档原本永久不可见。Manifest 预算是「装箱」而非「前缀截断」:某篇的行装不下会跳过,后面更短的行仍会被考虑。

标注回归集覆盖中文、英文、日文、韩文、混合 Query、NFKC、Body-only、同义无词面、ASCII 与 CJK 之外的字母文字,以及无相关 Memory。同时报告语料规模和「不看 Query 的随机 Scorer」在该语料上的 Recall@5,避免小语料美化结论。超过 200 个 Topic 由 memoryLifecycle.integration.test.ts 在真实临时 Memory 树上覆盖,Active Tool Noise 由 recall.test.ts 覆盖——这两者都不是标注集所评估的纯打分函数能触达的。记录 Recall@5、Top-1 Accuracy、No-result Precision、No-result Recall、Candidate Reduction Ratio 以及 Scan/Fast/Refined Latency。其中 "Fast-path Precision" 现指确定性投递的精度,而非独立 Fast Selector 的精度。

Enterprise Memory Extension 边界

企业审核、身份、审计、Retention 和 Secret Policy 由独立可选 Memory Extension 实现,拥有自己的存储和写入链路,不包装或接管 Core Managed Memory。现有 Extension 能力(Hook、MCP Write Tool、Extension Commands/Settings)足以支持第一版。本 RFC 不要求 Core 新增专用 Memory Plugin API。

明确不做

  • 不在 Core 引入 Candidate Staging 或 Memory Inbox;
  • 不引入第二次扫描、第二个 Selector 或独立的 Fast/Refined 预算核算;
  • 不在 Core 引入企业 Identity、Tenant、Audit 或 Policy Engine;
  • 不引入 Memory Schema v2 或批量迁移;
  • 不增加 Expiry、Supersession 或 Conflict Workflow;
  • 不统一重写现有 Secret Guard;
  • 不引入 BM25、Vector Database 或 Persistent Catalog;
  • 不增加公开 Retrieval Mode 或企业 Policy Config;
  • 不修改现有 Memory 写入语义。

验证与发布

  1. 先合入 Delivery Telemetry —— 已完成(feat(core): add memory recall delivery telemetry #7393);
  2. 再合入 Bounded Initial Recall 与确定性 Fast Path,含投递阶段测试 —— 已在 fix(memory): improve recall reliability and candidate coverage #8716 完成;
  3. 多语言 Precision 只有在标注集证明 English Recall@5 与 No-result Precision 不退化后才合入 —— 已实测并满足;
  4. Model Manifest、文档数量、聚合 Prompt 大小和首轮附加延迟始终有界 —— 已完成,但实测发现一个前提。确定性评分耗时 p50 0.036 ms / p95 0.053 ms,但决定首轮延迟的从来不是评分,而是 Memory 树扫描:recall-scan-latency.test.ts 实测 200 篇约 29 ms、500 篇约 70 ms、1000 篇约 130 ms。因此首轮等待改为「上限 + 确定性结果就绪即结束」;单个 Scope 超过约 1000 篇时扫描本身就超上限、该轮什么都投不到,作为已知限制记录而非修复;
  5. Enterprise Extension 单独立项,不阻塞 Core Recall 改进 —— 保持,见 proposal(memory): Define an enterprise external-memory integration profile #7449 / proposal: Add a direct external context provider profile #7585。

实测结果(英文/CJK/混合分片、首轮与无工具轮投递率、去重、延迟)见上文英文小节的两张表格。

已知限制

  • Selector 延迟为建模值而非实测值;
  • Fast 结果没有模型判断背书,两篇文档限制了出错代价,但在 Selector 始终未返回的无工具轮次中,模型看到的就是它;
  • 评分基于子串匹配,Query Token 可能匹配到更长单词内部(如 "owner" 命中 "ownership")。评估集保留了一个此类用例而非回避它;
  • 单个 Scope 超过约 1000 篇 Topic 时,Memory 树扫描本身就超出首轮上限,该轮付满预算且什么都投不到——比它替换掉的零等待更差。确定性结果就绪即结束等待只能限制这种情况,消除不了它;真正的解法是持久化 Catalog,仍在范围外;
  • 与文档没有任何词面重叠的 Query 产生不了确定性结果,这类 Query 在无工具轮次仍然拿不到 Memory。只有 Model Selector 能覆盖它们,而无工具轮次等不到 Selector。语料的 semantic-no-lexical 分片直接测量这一点:新旧两个 Scorer 在该分片都是 0%,因此"要求词面匹配"没有制造这个缺口,但确实让 Fast Path 在这里保持沉默——上文 92.9% 而非 100% 的残差正是它;
  • Thai/Khmer/Lao 这类无分词符又不在 CJK 集合内的文字,现在会产生 Token(之前完全没有),但整段是一个 Token,这不是分词。

旧的 b6dec4a 宽方案 仅保留为历史记录,不再是规范方案。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions