Skip to content

Cancellation provenance: five decisions deferred out of PR 13436 #13502

Description

@yiliang114

Active client-close follow-up — 2026-10-08

The active/admission-waiting client Close path is implemented and verified in the still-unmerged PR at 82129f6aad3964dab991cc8fbb69108e081791d8. Local checks and actual native tmux pair show REST Close persists user intent and cold Continue refuses it without a new provider request; runtime disposal remains recoverable. Closing a session that was already idle after an earlier interruption still needs a durable closed marker. Legacy cancellation provenance and unmarked cron/steering ownership remain open. This is a partial implementation update, not acceptance of those limitations or full #6710 closure.

中文:active/等待准入的主动 Close 已实现并验证,原生重启对照拒绝续跑且不新增模型请求;runtime disposal 仍可恢复。先前中断后已空闲再 Close 的持久化关闭标记、旧取消记录来源及无标识 cron/steering 归属仍开放,不宣称完整关闭 #6710 或接受兼容损失。


Why this issue exists

This issue tracks the remaining cancellation-provenance decisions. PR #13436 is open and unmerged, currently at 82129f6aad3964dab991cc8fbb69108e081791d8. The five sections below retain the historical findings from 59937eb561; their old line numbers/signatures and claim that work had shipped are not current implementation or release facts. The design describes known limitations pending a maintainer ruling, not accepted compatibility losses.

Current disposition

Item Current status
1 — legacy cancelledAt Open. No durable provenance/version field or accepted compatibility ruling.
2 — unmarked automatic boundaries Partial: exact Todo Stop Guard/task-notification tails now preserve ownership. Unmarked cron/steering boundaries remain open.
3 — resumed-attempt owner Known guard/notification retry cases now retain the original ID in both attempt recording and send. General unknown-owner/pre-send acceptance remains open; this is not a claim that every producer was replaced by recovery-plan ownership.
4 — channel cancellation metadata Implemented in the unmerged PR: interface, deadline/teardown callers, SDK reason preservation/coalescing and later user upgrade have focused checks. Pending merge/review, not main/release completion.
5 — Goal pause reason Open. No infrastructure-interrupt protocol member is introduced by the latest follow-up.

Local checks, actual native daemon/ACP tmux pair and byte identity. Six concrete follow-up Criticals are resolved; the unmarked-tail Critical remains open for the residual boundary, as does legacy provenance. This issue stays open.

Test pins for the same area are tracked separately in #13478 and are not duplicated here.

1. cancelledAt is read as user intent, but pre-#13436 builds stamped it for every controlled abort

isLastApiPromptCancelled returns true on state === 'cancelled' && cancelledAt !== undefined (packages/core/src/services/session-api-history.ts:156). At the merge base, recordAdmissionCancellation stamped cancelledAt unconditionally and every non-dispose abort resolved to USER_CANCEL_ABORT_REASON, so a prompt deadline, an HTTP response-close or a superseding prompt all persisted turn_result{state:'cancelled', cancelledAt}.

A user upgrading from a pre-#13436 build therefore has those turns read as explicitly cancelled: buildSessionRecoveryPlanFromApiHistory certifies {kind:'clean', canContinue:false} and recovery is permanently refused, with no version gate able to distinguish the record from one this change wrote.

Needs a new provenance field on the persisted turn_result, written only on the user-cancel branch, threaded Session.ts -> chatRecordingService -> turn-result-record.ts -> session-api-history.ts (four files, two packages). It must stay inside TURN_RESULT_IDENTIFIER_MAX_CHARS, whose validator makes recordTurnResult silently drop over-contract payloads.

2. The unmarked-entry rule is fail-open; which way it should break is a product ruling

getLastApiHistoryPromptId stops at the last user-shaped entry and returns that entry's mark (session-api-history.ts:96-110), so an unmarked tail yields undefined and isLastApiPromptCancelled bails at :124-125. Four producers are confirmed:

  • a reminder-bearing notification — only a bare <task-notification> envelope is trimmed, and the live continuedPromptId call site omits ACP_API_USER_PROMPT_OPTIONS
  • a cron or loop echo — isSystemNotificationRecord excludes 'cron' and createNotificationRecord never sets promptId
  • a pushed mid_turn_user_message — appendApiHistoryRecord marks only type === 'user' && !subtype
  • a steering entry pushed after a dropped code_mode_tool_result (session-api-history.ts:305), which makes the merge branch see a model previous and push unmarked

When one lands after a cancelled turn, recovery stays available and a restart can re-offer work the user explicitly stopped — the behaviour #6710 was closed against.

Both candidate fixes were measured and both collide with the documented rule "a newer unmarked user entry prevents an older cancellation from being reused": the mark-based variant breaks 2 tests #13436 added, the hint-based variant breaks 1. Any fix must skip the unmarked entry rather than stamp it, because findApiHistoryPromptIndex returns -1 when one promptId marks two entries. The ruling needed: re-open #6710 for these entrances, or accept re-executing stopped work.

3. Producer-side owner resolution (closes two findings at once)

recordTurnAttempt(promptId: string | undefined, ...) accepts undefined and durably appends an unlinkable turn_attempt (packages/core/src/services/chatRecordingService.ts:3419-3421), while the write guard checks continuesCurrentWorkChain && daemonPromptId && !reattemptRecorded and never continuedPromptId (Session.ts:6907-6911). The owner scan then adopts any later attempt hint with no identity check (session-api-history.ts:133) and rejects on owner.promptId !== promptId (:138), so one ownerless record flips the classification permanently.

The same producer-side gap is what leaves a Retry/Continue cancelled before the model send with a cancelled turn_result and no turn_attempt. Resolving the owner from the accepted recovery plan instead of re-deriving it from live history closes both.

The reader-side alternative is contested: the design doc states "a binding with unknown identity keeps recovery available", so a reader guard that skips unknown-identity bindings needs that sentence overturned first.

4. Two ChannelAgentBridge cancels still cannot express provenance

ChannelAgentBridge.cancelSession(sessionId: string) (packages/channels/base/src/ChannelAgentBridge.ts:292) takes no metadata, so two infrastructure cancels reach the absent-key arm and are stamped user intent:

  • ChannelBase.ts:2800 — cancelTimedOutLoopPrompt, reached from the loop-deadline race
  • DaemonChannelBridge.ts:925 — void session.cancel() inside stop() teardown

This contradicts the design doc's own rule that deadlines and session teardown carry interruption. Needs the interface widened plus both implementations and both call sites.

The third site named in the same finding — session-dispatch-port.ts over-budget enforcement — was handled by #13436 at 59937eb561; its bridge is the narrower AgentSessionBridge pick, so it needed no widening.

5. Goal pause reasons have no infrastructure-interrupt member

Four GOAL_PAUSE_REASON_USER_INTERRUPT producers exist in Session.ts; #13436 folded INTERRUPTED_PROMPT_ABORT_REASON into two of them, so a Goal turn killed by infrastructure is journalled, carded and emitted as "Interrupted by the user. Run /goal resume to continue." The other two producers read no reason ladder at all.

goal-protocol.ts carries nine pause-reason constants and none means infrastructure interrupt. Needs a new member plus a shared mapping helper across all four producers.

中文说明

本 issue 跟进剩余取消来源决策。#13436 开放且未合并,当前为 2ba5cbfe0a883dd1a5b779bb530dd5bb3a5ea706。下列五项保留 59937eb561 的历史发现,旧行号/签名及“已经交付”的表述不代表当前实现或发布;设计现已明确为“已知限制、待维护者裁决”,没有视为已接受的兼容损失。

当前处置

项目 当前状态
1 — 旧 cancelledAt 开放,未新增持久化来源/版本字段,也没有接受该兼容行为的裁决。
2 — 无标识自动边界 部分修复:确切 Todo Stop Guard/task-notification 尾项保持归属;未标识 cron/steering 仍开放。
3 — 重试归属 已知 guard/notification 重试在 attempt 记录和发送中保留原 ID;一般未知归属/发送前验收仍开放,没有声称所有产生源已改为 recovery-plan owner。
4 — channel 元信息 已在未合并 PR 实现接口、超时/退出调用方、SDK 保留来源/合并与后续 user 升级,并有定向验证;仍待合并/评审,不代表 main 或发布已修。
5 — Goal pause reason 开放,最新跟进未引入基础设施中断协议成员。

本地检查、真实 daemon/ACP tmux 双组验证与字节身份。六条具体跟进 Critical 已解决,未标识尾项及旧转录来源仍开放。本 issue 保持开放。

同区域的测试补钉由 #13478 单独跟踪,此处不重复。

  1. cancelledAt 被当作用户意图:session-api-history.ts:156 判 state === 'cancelled' && cancelledAt !== undefined;merge base 上该字段对所有受控 admission abort 无条件写入,因此旧版本记录的超时/响应关闭/被后续输入取代的轮次,升级后会被永久判为「主动取消」而拒绝恢复。修法是新增持久化来源字段,跨 cli→core 四个文件,且必须留在 TURN_RESULT_IDENTIFIER_MAX_CHARS 契约内(超限会被静默丢弃)。

  2. 「无标识用户输入」规则 fail-open:四类产生源已确认(带 reminder 的通知、cron/loop 回显、被推入的轮次中途插话、丢弃 code_mode_tool_result 后推入的插话)。两个候选修法都与已记录规则冲突(分别挂 2 条/1 条新测试),需裁决方向:为这些入口重开 fix(acp): distinguish user-cancelled turns from unexpected interruption after restore #6710,还是接受重跑用户已停止的工作。

  3. producer 侧归属解析:recordTurnAttempt 接受 promptId: undefined 并落盘不可归属的 turn_attempt,owner 扫描无身份检查地采纳后置 attempt,一条无主记录即永久翻转分类。改为「从已接受的恢复计划解析 owner」可同时关闭它与「模型发送前取消的重试」那条。reader 侧修法与设计文档「归属未知时保持可恢复」冲突,需先推翻该句。

  4. 两个 ChannelAgentBridge 取消无法表达来源:接口是单参签名,loop 超时与 stop() 退出两处基础设施取消被打成用户意图,违反设计文档自己的规则。需拓宽接口 + 两个实现 + 两个调用点。(同一发现里的第三处 session-dispatch-port.ts 超预算路径已在 59937eb561 修好,其 bridge 是更窄的 AgentSessionBridge pick,无需拓宽。)

  5. Goal pause reason 缺基础设施中断成员:四个 GOAL_PAUSE_REASON_USER_INTERRUPT 产生源中两个已折入新 abort 原因,导致基础设施中断被记录并显示为「Interrupted by the user」;另两个根本不读原因阶梯。需新增成员 + 四个产生源共用的映射helper。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions