You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
RFC: Reliable auto-memory recall — timing, quality, and telemetry #7040
PR 2's design was amended after measurement: the two-result Fast/Refined
architecture originally specified here is replaced by a minimal deterministic
fast path. This body describes the current normative design. The decision
record for the change, with the measurements behind it, is in this comment.
Normative design documents, committed in #8716: docs/design/2026-08-08-native-memory-recall-reliability.md and docs/design/2026-08-09-bounded-memory-recall-candidates.md.
Decision and scope
After feedback from the Core memory maintainer, this RFC is intentionally narrowed. Qwen Code Core should improve the recall path that benefits every user without becoming an enterprise memory-governance platform.
The Core roadmap contains three independently reviewable PRs:
Add recall-delivery telemetry.
Add bounded initial-turn recall with a deterministic fast path.
Improve multilingual recall precision and add a small labeled regression set.
The previous 13-PR governance roadmap is superseded. Candidate staging, approval inboxes, enterprise audit logs, organization identity, policy-based secret governance, provenance schemas, expiry, supersession, and conflict workflows are no longer proposed for Core.
Why these Core changes are needed
Memory reaches the first prompt only by luck, and a bounded wait alone does not fix it. Recall starts on UserQuery, but the initial request performed a zero-wait poll. Adding a fixed initial budget was the first attempt at a fix and is necessary but not sufficient: recall awaits the model selector, a network side query with a 30 s abort ceiling, so the budget is dominated by round-trip time rather than the incidental scheduler timing it was sized for. It therefore expires on the common path.
Tool-free turns then get nothing at all. On budget expiry, delivery falls through to the ToolResult point. A turn that makes no tool call never reaches one, so the result is discarded as no_safe_delivery_point. These are exactly the short, context-answered questions where user-level memory matters most.
The local heuristic was not an actual fast path. It ran only after the model selector failed or timed out, so it could not help the initial turn — and it was ASCII-oriented, so it was weak for CJK queries.
No-lexical-match documents could still score. Any non-empty body earned a positive score, so unrelated documents could be selected.
Topic documents were truncated to the 200 most recent per scope before relevance selection. The scanner already reads and parses the files, so the cap created deterministic misses without avoiding most I/O. Note the fix is a change of truncation key rather than a lifted ceiling — see the candidate-set note under PR 3.
Telemetry before PR 1 recorded selection but could not prove that selected memory was delivered to the main model.
Core implementation plan
PR 1: Recall delivery telemetry
Keep the existing selection event and add a low-cardinality delivery event that records:
phase, delivery point, a fixed discard reason, strategy, document count, and latency.
Do not record query text or hashes, memory content, file or project paths, session or message identifiers, model reasoning, secrets, or raw error text.
phase and strategy are orthogonal and neither subsumes the other.
phase is the delivery stage: fast = the deterministic result injected when the initial budget expired; refined = the model-selected result.
strategy is the selection method: none / heuristic / model.
A fast delivery is always heuristic. A refined delivery is model normally, and heuristic when the selector failed and the fallback ran. Reading delivery stage off strategy alone would merge "the deterministic result arrived first" with "the selector broke".
PR 2: Bounded initial-turn recall with a deterministic fast path
One recall lifecycle and model-primary selection, with a single deterministic delivery stage in front of it:
Give user-query recall a small internal ceiling, determined by benchmark rather than exposed as public configuration. The wait ends on whichever comes first: recall settling, the deterministic result being published, cancellation, or the ceiling. Benchmarks showed the ceiling was otherwise a fixed per-turn tax: the deterministic result is published once the memory tree has been scanned, which is tens of milliseconds for an ordinary tree, while the selector round trip is assumed to miss the budget entirely.
A result settling inside the budget is delivered in the initial prompt.
On budget expiry, deliver the deterministic result rather than nothing.selectModelCandidateDocuments already computes lexically ranked, active-tool-filtered candidates in order to build the model manifest; the fast result reuses them, so it costs no extra scan or I/O. Cap it at two documents (MAX_FAST_RECALL_DOCS), well below the five-document prompt limit, because it carries no model judgement.
Leave recall pending after a fast delivery so the model-selected result still lands at the existing same-query ToolResult delivery point.
Exclude documents the fast phase already delivered from that later delivery and rebuild the prompt from the remainder. Both results come from one scan, so the selector never saw the fast documents as excluded and can legitimately re-select them. When nothing remains, record already_delivered.
Do not abort recall merely because the budget expired.
A new Query, Reset, Abort, or Shutdown cancels the recall plan. A cancelled turn delivers no fast result. Results must never cross query boundaries or inject the same document twice. The prompt obeys both document-count and aggregate-size budgets.
MAX_RELEVANT_DOCS = 5 bounds one prompt, not one turn. A fast delivery of two documents plus a refined delivery of five disjoint ones puts seven in front of the model; deduplication removes repeats, not the sum. This is the direct consequence of dropping combined fast/refined budget accounting, which this RFC originally specified as a fill-to-five limit across both phases. Both prompts stay individually bounded and each body is still truncated, so the worst case is bounded and small — it is simply not five. If a hard aggregate ceiling is ever wanted, the cheap version is passing limit - fastDeliveredPaths.size as the refined limit rather than reintroducing a second budget.
Explicitly not adopted: the two-result shared-scan Fast/Refined architecture, i.e. a second selection pathway with its own scan plumbing, its own budget accounting, and cross-phase fill-to-five logic. The delivery guarantee is worth having; that machinery is not, and cross-phase bookkeeping is the source of the duplicate-injection bug class this RFC already warned about. One callback plus one exclusion set achieves the same guarantee.
No public configuration is added.
PR 3: Recall precision and multilingual evaluation
Improve the deterministic scorer — which now serves both the fast path and the selector-failure fallback — without introducing a persistent index or new retrieval dependency:
NFKC normalization; whole-run tokens for non-CJK letters, marks, and digits of at least three characters (\p{L}-based rather than [a-z0-9], so Cyrillic, Greek, Arabic, and accented Latin produce tokens instead of none, with CJK excluded per character so a Latin-initial run cannot swallow the CJK that follows it); Unicode code-point bigrams for Han, Hiragana, Katakana, and Hangul runs; no broad match from a single CJK character; no score without a lexical match; title and description matches weighted above body matches; type boosts applied only after a real match (no scope boost is implemented — scope precedence is carried by ordering, not by score); score ties broken by recency and then by input order, never by document type, since an alphabetical type comparison ranks user last and the two-document fast result would drop it; bounded query tokens retaining both query edges; scoring limited to the surfaced body window; existing active-tool noise filtering retained.
Replace the per-scope 200-document recency truncation before recall ranking with a global, query-aware candidate set (non-recall callers keep the capped scanner). To keep the model manifest bounded, run lightweight local ranking over all parsed topics, select a limited candidate set, and reserve part of that set for recent documents so semantic-but-not-lexical matches still have a chance. Candidate count and manifest bytes remain bounded internal constants.
Read the effect per pool size; it is not a uniform widening. At or under 200 documents nothing was excluded by count under either design, but the 25,000-byte manifest budget is a ceiling the old path lacked. Between 200 and 400 with neither scope over 200, the old path sent every document to the selector and the new one sends at most 200, so fewer reach the model — the survivors are chosen by relevance plus a recency reserve rather than by recency alone, which is the intended trade, but the raw count goes down. Only a scope over 200 is the case this is for, where an old but lexically matching document was permanently invisible. The manifest budget also packs rather than prefixes: a document whose line does not fit is skipped and later, shorter lines are still considered.
Maintain a labeled fixture covering Chinese, English, Japanese, Korean, mixed queries, NFKC, body-only matches, semantic/no-lexical cases, alphabetic scripts outside ASCII and CJK, and no-result cases. Report the corpus size and the Recall@5 a query-blind random scorer would achieve on it, so a small corpus cannot flatter the headline. The more-than-200-topics case is covered by memoryLifecycle.integration.test.ts against a real temporary memory tree, and active-tool noise by recall.test.ts, since neither is reachable from the pure scoring function the fixture evaluates.
Track Recall@5, top-1 accuracy, no-result precision, no-result recall, candidate-reduction ratio, and scan/fast/refined latency. "Fast-path precision" now means the precision of the deterministic delivery rather than of a separate Fast selector. BM25, persistent catalogs, vector databases, language-specific tokenizers, and new dependencies require separate evidence and a separate proposal.
Enterprise memory extension boundary
Enterprise review, identity, audit, retention, and secret-policy requirements should be implemented as an optional Memory Extension with its own storage and write path, rather than wrapping or taking over Core managed memory.
The existing extension system is sufficient for a first implementation: UserPromptSubmit / SessionStart hooks can inject additional context; Stop / SessionEnd hooks can enqueue asynchronous extraction; extension-owned MCP write tools can enforce organization secret and compliance policies; optional PreToolUse hooks can inspect generic tool calls, without assuming they intercept Core-internal writes; MCP tools can expose Search, Propose, Approve, Reject, and Audit; the enterprise service or extension owns review UI, identity, authorization, retention, and audit logs.
This RFC does not propose a dedicated Core memory-plugin API. If a real extension implementation identifies a missing generic capability, that API gap should be proposed independently and kept minimal.
Explicit non-goals
Candidate staging or Memory Inbox in Core.
A second scan, second selector, or separate fast/refined budget accounting.
Enterprise identity, tenant, audit, or policy engines in Core.
Memory Schema v2 or bulk migration.
Expiry, supersession, and conflict workflows.
A unified rewrite of existing secret guards.
BM25, vector databases, or a persistent catalog.
Public retrieval-mode or enterprise-policy configuration.
Changes to current Remember, Forget, Extraction, DREAM, Team Memory, or write semantics.
Land bounded initial recall with internal thresholds and focused lifecycle tests. — Done in fix(memory): improve recall reliability and candidate coverage #8716, including the deterministic fast path and delivery-stage tests for the tool-free turn, later ToolResult delivery, fast/refined dedupe, cancellation inside the initial window, and no cross-query leakage.
Land multilingual precision changes only after the labeled fixture shows no regression in English Recall@5 or no-result precision. — Gate measured and satisfied, see below.
Keep model manifests, document count, aggregate prompt size, and initial-turn latency bounded. — Done, with a measured caveat. Deterministic scoring costs p50 0.036 ms / p95 0.053 ms, but scoring was never what decided initial-turn latency: the memory-tree scan is, and recall-scan-latency.test.ts measures it at ~29 ms (200 topics), ~70 ms (500), ~130 ms (1000). The initial wait is therefore a ceiling that ends on the deterministic result; past ~1000 topics in one scope the scan alone exceeds it and the turn delivers nothing, which is recorded as a limitation rather than fixed.
Each PR must independently pass focused package tests, build, typecheck, self-audit, and review.
Measured results
Recall quality (packages/core/src/memory/recall-eval.test.ts, 51-case / 25-document labeled corpus scored by both the shipped scorer and a frozen copy of the pre-change scorer, so "no regression" is reproducible rather than asserted). A query-blind scorer returning 5 random documents scores 20.0% Recall@5 on this pool, which is the floor the measured numbers should be read against:
slice
metric
before
after
overall (n=51)
Recall@5
45.2%
92.9%
overall (n=51)
top-1 accuracy
46.2%
97.4%
overall (n=51)
no-result precision
16.0%
75.0%
overall (n=51)
no-result recall
44.4%
100.0%
english (n=11)
Recall@5
100.0%
100.0%
english (n=11)
top-1 accuracy
100.0%
100.0%
cjk (n=17)
Recall@5
0.0%
100.0%
cjk (n=17)
top-1 accuracy
0.0%
100.0%
mixed (n=5)
Recall@5
80.0%
100.0%
mixed (n=5)
top-1 accuracy
60.0%
100.0%
other-script (n=3)
Recall@5
33.3%
100.0%
other-script (n=3)
top-1 accuracy
33.3%
100.0%
semantic-no-lexical (n=3)
Recall@5
0.0%
0.0%
No-result precision measured over a no-result-only slice is 100% for any scorer that ever stays silent, so no-result recall — of genuinely unanswerable queries, how many got an empty answer — is the metric that detects the old scorer's actual failure.
The overall row is held below 100% by the semantic-no-lexical slice this RFC asked the fixture to cover: answerable queries sharing no token with their document. An earlier revision of the corpus labeled a no-result case with that category, so nothing measured the cost of "no score without a lexical match". With three genuine cases added, both the shipped scorer and the frozen one score 0% on the slice — the rule did not create the gap, but it does keep the deterministic path silent there. The slice is asserted separately and excluded from the quality floor; the overall no-result precision figure moves for the same reason, since staying silent on an answerable query counts against it.
Delivery (packages/core/src/memory/recall-delivery-eval.test.ts). Selector latency is modelled, not measured — a network round trip cannot be timed in a unit test — so results are reported per scenario; the structural claim holds for every scenario above the budget:
selector latency
metric
before
after
inside budget
first-turn delivery (tool-free)
92.9%
92.9%
above budget
first-turn delivery (tool-free)
0.0%
92.9%
above budget
delivered at all (tool-free)
0.0%
92.9%
above budget
delivered at all (tool-using)
92.9%
92.9%
above budget
duplicate delivery
0.0%
0.0%
The residual 7.1% in every row is the semantic-no-lexical slice, asserted separately at 0%: the fast path closes the timing gap, not the matching gap. The stand-in for the model selector is the deterministic top-5, so the simulation cannot answer those queries either — which is why the inside-budget row is 92.9% rather than 100% for both designs.
Tool-using turns were never broken — they receive memory one request later — so the fix is targeted at the tool-free gap. Fast/refined overlap occurs in 92.9% of above-budget cases by construction (the harness stands in for the model's choice with the deterministic top-5, which always contains the fast top-2), and dedupe still yields 0% duplicate delivery. That last figure is a property of the simulation, which does not drive tryConsumeMemoryPrefetch; the shipped dedupe is covered by client.test.ts, verified to fail when the exclusion filter is removed.
Known limitations
Selector latency is modelled rather than measured.
The fast result carries no model judgement. Two documents bounds the cost of being wrong, but on a tool-free turn where the selector never lands, a mis-ranked fast document is what the model sees.
Scoring is substring-based, so a query token can match inside a longer word ("owner" inside "ownership"). The evaluation corpus records one such case rather than hiding it.
A query sharing no token with its document produces no deterministic result, so a tool-free turn asking it still ends with nothing delivered. Only the model selector can serve those, and on a tool-free turn it never lands.
Past roughly a thousand topics in one scope, the memory-tree scan alone exceeds the initial ceiling, so the turn spends the whole budget and still delivers nothing — worse than the zero-wait behaviour it replaced. Ending the wait on the deterministic result bounds this rather than removing it; a persistent catalog is the actual fix and stays out of scope.
Scripts written without word separators and outside the CJK set (Thai, Khmer, Lao) now produce a token where they produced none, but the token is the whole run. That is not segmentation.
The previous broad design at b6dec4a is retained for history but is no longer normative.
Status
PR 2's design was amended after measurement: the two-result Fast/Refined
architecture originally specified here is replaced by a minimal deterministic
fast path. This body describes the current normative design. The decision
record for the change, with the measurements behind it, is in
this comment.
Normative design documents, committed in #8716:
docs/design/2026-08-08-native-memory-recall-reliability.mdanddocs/design/2026-08-09-bounded-memory-recall-candidates.md.Decision and scope
After feedback from the Core memory maintainer, this RFC is intentionally narrowed. Qwen Code Core should improve the recall path that benefits every user without becoming an enterprise memory-governance platform.
The Core roadmap contains three independently reviewable PRs:
The previous 13-PR governance roadmap is superseded. Candidate staging, approval inboxes, enterprise audit logs, organization identity, policy-based secret governance, provenance schemas, expiry, supersession, and conflict workflows are no longer proposed for Core.
Why these Core changes are needed
no_safe_delivery_point. These are exactly the short, context-answered questions where user-level memory matters most.Telemetry before PR 1 recorded selection but could not prove that selected memory was delivered to the main model.
Core implementation plan
PR 1: Recall delivery telemetry
Keep the existing selection event and add a low-cardinality delivery event that records:
Do not record query text or hashes, memory content, file or project paths, session or message identifiers, model reasoning, secrets, or raw error text.
phaseandstrategyare orthogonal and neither subsumes the other.phaseis the delivery stage:fast= the deterministic result injected when the initial budget expired;refined= the model-selected result.strategyis the selection method:none/heuristic/model.fastdelivery is alwaysheuristic. Arefineddelivery ismodelnormally, andheuristicwhen the selector failed and the fallback ran. Reading delivery stage offstrategyalone would merge "the deterministic result arrived first" with "the selector broke".delivery_point:initial/tool_result/discarded.discard_reason:no_safe_delivery_point/new_query/reset/abort/shutdown/no_relevant_results/already_delivered.PR 2: Bounded initial-turn recall with a deterministic fast path
One recall lifecycle and model-primary selection, with a single deterministic delivery stage in front of it:
selectModelCandidateDocumentsalready computes lexically ranked, active-tool-filtered candidates in order to build the model manifest; the fast result reuses them, so it costs no extra scan or I/O. Cap it at two documents (MAX_FAST_RECALL_DOCS), well below the five-document prompt limit, because it carries no model judgement.already_delivered.MAX_RELEVANT_DOCS = 5bounds one prompt, not one turn. A fast delivery of two documents plus a refined delivery of five disjoint ones puts seven in front of the model; deduplication removes repeats, not the sum. This is the direct consequence of dropping combined fast/refined budget accounting, which this RFC originally specified as a fill-to-five limit across both phases. Both prompts stay individually bounded and each body is still truncated, so the worst case is bounded and small — it is simply not five. If a hard aggregate ceiling is ever wanted, the cheap version is passinglimit - fastDeliveredPaths.sizeas the refined limit rather than reintroducing a second budget.Explicitly not adopted: the two-result shared-scan Fast/Refined architecture, i.e. a second selection pathway with its own scan plumbing, its own budget accounting, and cross-phase fill-to-five logic. The delivery guarantee is worth having; that machinery is not, and cross-phase bookkeeping is the source of the duplicate-injection bug class this RFC already warned about. One callback plus one exclusion set achieves the same guarantee.
No public configuration is added.
PR 3: Recall precision and multilingual evaluation
Improve the deterministic scorer — which now serves both the fast path and the selector-failure fallback — without introducing a persistent index or new retrieval dependency:
\p{L}-based rather than[a-z0-9], so Cyrillic, Greek, Arabic, and accented Latin produce tokens instead of none, with CJK excluded per character so a Latin-initial run cannot swallow the CJK that follows it); Unicode code-point bigrams for Han, Hiragana, Katakana, and Hangul runs; no broad match from a single CJK character; no score without a lexical match; title and description matches weighted above body matches; type boosts applied only after a real match (no scope boost is implemented — scope precedence is carried by ordering, not by score); score ties broken by recency and then by input order, never by document type, since an alphabetical type comparison ranksuserlast and the two-document fast result would drop it; bounded query tokens retaining both query edges; scoring limited to the surfaced body window; existing active-tool noise filtering retained.Replace the per-scope 200-document recency truncation before recall ranking with a global, query-aware candidate set (non-recall callers keep the capped scanner). To keep the model manifest bounded, run lightweight local ranking over all parsed topics, select a limited candidate set, and reserve part of that set for recent documents so semantic-but-not-lexical matches still have a chance. Candidate count and manifest bytes remain bounded internal constants.
Read the effect per pool size; it is not a uniform widening. At or under 200 documents nothing was excluded by count under either design, but the 25,000-byte manifest budget is a ceiling the old path lacked. Between 200 and 400 with neither scope over 200, the old path sent every document to the selector and the new one sends at most 200, so fewer reach the model — the survivors are chosen by relevance plus a recency reserve rather than by recency alone, which is the intended trade, but the raw count goes down. Only a scope over 200 is the case this is for, where an old but lexically matching document was permanently invisible. The manifest budget also packs rather than prefixes: a document whose line does not fit is skipped and later, shorter lines are still considered.
Maintain a labeled fixture covering Chinese, English, Japanese, Korean, mixed queries, NFKC, body-only matches, semantic/no-lexical cases, alphabetic scripts outside ASCII and CJK, and no-result cases. Report the corpus size and the Recall@5 a query-blind random scorer would achieve on it, so a small corpus cannot flatter the headline. The more-than-200-topics case is covered by
memoryLifecycle.integration.test.tsagainst a real temporary memory tree, and active-tool noise byrecall.test.ts, since neither is reachable from the pure scoring function the fixture evaluates.Track Recall@5, top-1 accuracy, no-result precision, no-result recall, candidate-reduction ratio, and scan/fast/refined latency. "Fast-path precision" now means the precision of the deterministic delivery rather than of a separate Fast selector. BM25, persistent catalogs, vector databases, language-specific tokenizers, and new dependencies require separate evidence and a separate proposal.
Enterprise memory extension boundary
Enterprise review, identity, audit, retention, and secret-policy requirements should be implemented as an optional Memory Extension with its own storage and write path, rather than wrapping or taking over Core managed memory.
The existing extension system is sufficient for a first implementation: UserPromptSubmit / SessionStart hooks can inject additional context; Stop / SessionEnd hooks can enqueue asynchronous extraction; extension-owned MCP write tools can enforce organization secret and compliance policies; optional PreToolUse hooks can inspect generic tool calls, without assuming they intercept Core-internal writes; MCP tools can expose Search, Propose, Approve, Reject, and Audit; the enterprise service or extension owns review UI, identity, authorization, retention, and audit logs.
This RFC does not propose a dedicated Core memory-plugin API. If a real extension implementation identifies a missing generic capability, that API gap should be proposed independently and kept minimal.
Explicit non-goals
Verification and rollout
recall-scan-latency.test.tsmeasures it at ~29 ms (200 topics), ~70 ms (500), ~130 ms (1000). The initial wait is therefore a ceiling that ends on the deterministic result; past ~1000 topics in one scope the scan alone exceeds it and the turn delivers nothing, which is recorded as a limitation rather than fixed.Each PR must independently pass focused package tests, build, typecheck, self-audit, and review.
Measured results
Recall quality (
packages/core/src/memory/recall-eval.test.ts, 51-case / 25-document labeled corpus scored by both the shipped scorer and a frozen copy of the pre-change scorer, so "no regression" is reproducible rather than asserted). A query-blind scorer returning 5 random documents scores 20.0% Recall@5 on this pool, which is the floor the measured numbers should be read against:No-result precision measured over a no-result-only slice is 100% for any scorer that ever stays silent, so no-result recall — of genuinely unanswerable queries, how many got an empty answer — is the metric that detects the old scorer's actual failure.
The overall row is held below 100% by the
semantic-no-lexicalslice this RFC asked the fixture to cover: answerable queries sharing no token with their document. An earlier revision of the corpus labeled a no-result case with that category, so nothing measured the cost of "no score without a lexical match". With three genuine cases added, both the shipped scorer and the frozen one score 0% on the slice — the rule did not create the gap, but it does keep the deterministic path silent there. The slice is asserted separately and excluded from the quality floor; the overall no-result precision figure moves for the same reason, since staying silent on an answerable query counts against it.Delivery (
packages/core/src/memory/recall-delivery-eval.test.ts). Selector latency is modelled, not measured — a network round trip cannot be timed in a unit test — so results are reported per scenario; the structural claim holds for every scenario above the budget:The residual 7.1% in every row is the
semantic-no-lexicalslice, asserted separately at 0%: the fast path closes the timing gap, not the matching gap. The stand-in for the model selector is the deterministic top-5, so the simulation cannot answer those queries either — which is why the inside-budget row is 92.9% rather than 100% for both designs.Tool-using turns were never broken — they receive memory one request later — so the fix is targeted at the tool-free gap. Fast/refined overlap occurs in 92.9% of above-budget cases by construction (the harness stands in for the model's choice with the deterministic top-5, which always contains the fast top-2), and dedupe still yields 0% duplicate delivery. That last figure is a property of the simulation, which does not drive
tryConsumeMemoryPrefetch; the shipped dedupe is covered byclient.test.ts, verified to fail when the exclusion filter is removed.Known limitations
The previous broad design at b6dec4a is retained for history but is no longer normative.
详细中文方案(评审与合并参考)
状态
PR 2 的设计在实测后已修订:原先规定的双结果 Fast/Refined 架构,替换为最小化的确定性 Fast Path。本文描述的是当前规范设计;变更的决策记录与实测依据见该评论。
规范设计文档已随 #8716 提交:
docs/design/2026-08-08-native-memory-recall-reliability.md与docs/design/2026-08-09-bounded-memory-recall-candidates.md。决策与范围
根据 Core Memory 负责同学的反馈,本 RFC 明确收缩范围:Qwen Code Core 只解决所有用户都会遇到的召回时机、召回精度和可观测性问题,不在 Core 内建设企业 Memory 治理平台。
Core 路线图为 3 个可独立评审和回滚的 PR:
原方案中的 Candidate Staging、Memory Inbox、Apply Journal、企业审计日志、组织身份、企业秘密策略、Provenance Schema、Expiry、Supersession 和 Conflict Workflow 不再进入 Core 路线图。
现有 Remember、Forget、Extraction、DREAM、User/Project/Team Scope 和 Team Secret Guard 保持现状。
当前 Core 问题
1. 首轮召回依赖偶然时序,且仅靠有界等待无法解决
Client 在 UserQuery 后启动 Recall Prefetch,首轮请求前原本只做零等待检查。加入固定初始预算是第一步,但并不充分:当存在 Config 时(常规情况)Recall 会等待 Model Selector,而它是一次网络 Side Query,中止上限为 30 秒。因此该预算由网络往返时间主导,而不是它原本针对的偶然调度抖动,常规路径上必然超时。
2. 超时后,无工具调用的轮次完全拿不到 Memory
预算超时后投递会退到 ToolResult 时机。没有工具调用的轮次永远不会到达该时机,结果被记为
no_safe_delivery_point丢弃。而这类"短问答、靠上下文回答"的轮次,恰恰是用户级 Memory 最重要的场景。3. Heuristic 不是实际 Fast Path
它只在 Model Selector 失败或超时后执行,无法解决首轮时机问题;且以 ASCII 为主,对 CJK Query 效果很差。
4. 无匹配文档也可能得到分数
只要正文非空就加分,没有任何词面匹配的文档仍可能进入结果。新方案要求没有 Lexical Match 时分数必须为 0。
5. 按 scope 各留最近 200 篇的截断发生在相关性判断之前
Scanner 已经读取并解析 Markdown 才截断,造成漏召回却没有节省主要读取成本。注意修复方式是换截断依据而非抬高上限,效果分档见 PR 3。
PR 1:Recall Delivery Telemetry
保留现有 Selection Event,新增低基数 Delivery Event,记录 phase、delivery point、固定枚举的丢弃原因、strategy、文档数量和耗时。Telemetry 禁止记录 Query 或 Hash、Memory 正文、文件路径、Project Path、Session/Message ID、模型 Reasoning、Secret 或原始错误正文。
phase与strategy正交,互不替代:phase是投递阶段:fast= 预算超时后注入的确定性结果;refined= Model Selector 选出的结果。strategy是选择方式:none/heuristic/model。fast投递必然是heuristic;refined投递常规为model,在 Selector 失败走 Fallback 时为heuristic。仅凭strategy判断阶段,会把"确定性结果先到"与"Selector 故障"混为一谈。delivery_point:initial/tool_result/discarded。discard_reason:no_safe_delivery_point/new_query/reset/abort/shutdown/no_relevant_results/already_delivered。PR 2:有界的首轮 Recall(含确定性 Fast Path)
保持单一 Recall 生命周期与 Model 优先选择,在其前面增加一个确定性投递阶段:
selectModelCandidateDocuments为构建 Model Manifest 本就计算了按词面排序、且已过滤 Active Tool 噪声的候选;Fast 结果直接复用,不产生额外扫描或 I/O。上限两个文档(MAX_FAST_RECALL_DOCS),远低于五个文档的 Prompt 上限,因为它没有模型判断背书。already_delivered。MAX_RELEVANT_DOCS = 5限制的是单次注入,不是单轮总量。 Fast 投递 2 篇加上 refined 投递 5 篇不重叠文档时,本轮进入模型的是 7 篇;去重只消除重复,不压缩总和。这是放弃跨阶段预算核算的直接结果——本 RFC 原方案规定的是两阶段合计五篇。两次 Prompt 各自有界、每篇 body 仍截断,因此最坏情况有界且不大,只是不等于 5。若确需硬性总量上限,最省事的做法是把 refined 的 limit 传成limit - fastDeliveredPaths.size,而不是重新引入第二套预算。明确不采用:双结果共享扫描的 Fast/Refined 架构,即带独立扫描链路、独立预算核算和跨阶段补齐到五篇逻辑的第二条选择通路。投递保证值得保留,这套机制不值得——跨阶段簿记正是本 RFC 早已警示的重复注入缺陷来源。一个回调加一个排除集合即可获得同样保证。
不增加公开配置。
PR 3:召回精度与多语言评估
确定性评分器(现同时服务 Fast Path 与 Selector 失败 Fallback)改进:NFKC 规范化;至少 3 个字符的非 CJK 字母/组合符/数字整串 Token(基于
\p{L}而非[a-z0-9],因此西里尔、希腊、阿拉伯和带重音拉丁文都能产生 Token;CJK 逐字符排除,避免拉丁开头的串吞掉后面的 CJK);Han、Hiragana、Katakana、Hangul 连续片段的 Unicode Code Point Bigram;单个 CJK 字符不做宽泛匹配;没有 Lexical Match 时得分为 0;Title/Description 匹配权重高于 Body;Type 仅在真实匹配后加权(Scope 没有单独加权,Scope 优先级由排序而非分数承载);同分时按 mtime 降序再按输入顺序,不按文档 Type——Type 字典序会把user排到最后,而两篇的 Fast 结果会把它整个丢掉;Query Token 有界且保留首尾;只对可进入 Prompt 的正文窗口评分;保留 Active Tool 降噪。Recall 路径把「按 scope 各留最近 200 篇」的截断,换成全局、感知 Query 的候选集(其他调用方保留原有 capped scanner)。为避免 Manifest 无限增长:对全部 Topic 轻量本地评分 → 选择有限 Model Candidate Set → 优先高分词面匹配 → 保留部分近期文档额度 → Selector 只对有限集合精排。候选数量与 Manifest Bytes 均为有界内部常量。
效果要分档看,不是一律放宽。 总量 ≤200 时两种设计都不按数量丢弃,但 25,000 字节的 Manifest 上限是旧路径没有的。总量在 200–400 且单个 Scope 不超 200 时,旧路径会把全部文档送给 Selector,新路径最多送 200 篇,候选变少——留下来的是按相关性 + recency reserve 选的而不是纯按时间,这是有意的取舍,但绝对数量确实下降。只有单个 Scope 超过 200 时才是这次真正要解决的场景:老而词法命中的文档原本永久不可见。Manifest 预算是「装箱」而非「前缀截断」:某篇的行装不下会跳过,后面更短的行仍会被考虑。
标注回归集覆盖中文、英文、日文、韩文、混合 Query、NFKC、Body-only、同义无词面、ASCII 与 CJK 之外的字母文字,以及无相关 Memory。同时报告语料规模和「不看 Query 的随机 Scorer」在该语料上的 Recall@5,避免小语料美化结论。超过 200 个 Topic 由
memoryLifecycle.integration.test.ts在真实临时 Memory 树上覆盖,Active Tool Noise 由recall.test.ts覆盖——这两者都不是标注集所评估的纯打分函数能触达的。记录 Recall@5、Top-1 Accuracy、No-result Precision、No-result Recall、Candidate Reduction Ratio 以及 Scan/Fast/Refined Latency。其中 "Fast-path Precision" 现指确定性投递的精度,而非独立 Fast Selector 的精度。Enterprise Memory Extension 边界
企业审核、身份、审计、Retention 和 Secret Policy 由独立可选 Memory Extension 实现,拥有自己的存储和写入链路,不包装或接管 Core Managed Memory。现有 Extension 能力(Hook、MCP Write Tool、Extension Commands/Settings)足以支持第一版。本 RFC 不要求 Core 新增专用 Memory Plugin API。
明确不做
验证与发布
recall-scan-latency.test.ts实测 200 篇约 29 ms、500 篇约 70 ms、1000 篇约 130 ms。因此首轮等待改为「上限 + 确定性结果就绪即结束」;单个 Scope 超过约 1000 篇时扫描本身就超上限、该轮什么都投不到,作为已知限制记录而非修复;实测结果(英文/CJK/混合分片、首轮与无工具轮投递率、去重、延迟)见上文英文小节的两张表格。
已知限制
semantic-no-lexical分片直接测量这一点:新旧两个 Scorer 在该分片都是 0%,因此"要求词面匹配"没有制造这个缺口,但确实让 Fast Path 在这里保持沉默——上文 92.9% 而非 100% 的残差正是它;旧的 b6dec4a 宽方案 仅保留为历史记录,不再是规范方案。