Problem
Two real /goal-draft sessions on 2026-09-16 (one on deepseek-v4.1-flash writing the Goal architecture design doc, one on qwen3.8-max writing the dynamic-workflow developer guide) each finished the whole objective inside a single Goal turn of roughly 100 tool calls. What followed was the same in both:
- The evidence catalog overflowed. Every model step records one
assistant_output entry and one tool_result entry (goal-evidence.ts:1062-1081), so ~100 steps produce ~200 candidates for a catalog bounded at 100 entries / 24 000 bytes (goal-evidence.ts:32-33). get_goal reported truncated: true with 65 entries, 59 of them shortened.
update_goal status=complete was refused solely because the catalog was truncated (goal-tools.ts:359-376), and the proposal was discarded.
- The turn-end checkpoint then had to compress the overflowed window through a single-shot fast-model side query. On deepseek it returned invalid claims, then invalid JSON on batch 3/3; three stalls stopped the Goal as
usage_limited. On qwen the checkpoint succeeded, yet "32 claims while the window overflowed" still counted a stall (goal-checkpoint.ts:65-72) and cost an extra turn.
- After
/goal resume the evidence-limited branch repointed the cursor and dropped the checkpoint (goal-reducer.ts:237-252), so get_goal returned an empty catalog while the tool's nextAction still said to retry with the UUIDs it returns. The model recovered only by re-running the verification commands to mint fresh evidence.
The decisive fact: in both sessions the evidence the verifier finally accepted was the verification output re-run in the closing turn. None of the earlier catalog entries were needed. The deliverable was on disk and the check was re-runnable.
Why the mechanism, not the tuning
The evidence catalog, UUID citations, checkpoint compaction and the stall breaker are one layer of about 4 500 production lines. Since 2026-08-20, 10 of the 42 goal commits exist to keep that layer alive: #9975 → #9840 → #9835 → #9973 → #11304 → #11305 → #11365 → #11576 → #11578 → #11690 → #11693 → #11796. Each fixed the previous fix's edge. The two sessions above hit the layer after all of them landed.
Neither of the two comparable products carries this layer. One judges a goal condition with a separate evaluator that reads only the recent transcript and truncates the rest; the other has no second model at all and relies on a completion audit written into the continuation prompt. Both keep the same single-slot state machine, automatic continuation, and budget / no-progress stops that we have.
Proposal
Keep the independent verifier. Change what it sees: the current Goal turn's visible output and tool results, newest first, with the existing per-record (2 000 bytes) and per-request (256 000 bytes) limits, and an omittedEarlier count when the turn does not fit. Then delete the machinery that existed only to make older evidence citable:
- PR-1
feat(goal): judge terminal proposals from the current turn's evidence — verifier input becomes the current-turn window; update_goal.evidenceRefs becomes optional and ignored; the invalidEvidenceRefs, auto-citation and checkpointRequired gates go; get_goal stops returning the catalog; continuation prompt and /goal-draft skill say the verifier judges only this turn, and name the 1 500-character objective limit.
- PR-2
refactor(goal): remove evidence checkpoints and the stall breaker — delete goal-checkpoint.ts, goal-checkpoint-verifier.ts, the runtime checkpoint / batching / settle paths, the evidence-limited resume branch, the evidenceCheckpoint / checkpointStalls / lastCheckpointFailure record fields, the evidence_catalog / checkpoint_request limit kinds; deprecate model.goalCheckpointTimeoutSeconds.
- PR-3
refactor(goal): drop the evidence catalog and reference validation — collapse goal-evidence.ts to the record type, proof-kind mapping and the current-turn window builder.
- PR-4
chore(goal): stop carrying checkpoint state through the wire, UI and docs — SDK types, Web Shell mappers / gate / dialogs, CLI pill, docs.
Snapshot version stays v2; every removed field is optional, so old snapshots and old clients keep parsing. Budgets, pause reasons, the approval dialog and the legacy projection are untouched (the projection gets its own issue).
Out of scope
Removing the verifier itself. Folding usage_limited into paused is a possible follow-up after PR-4.
Problem
Two real
/goal-draftsessions on 2026-09-16 (one on deepseek-v4.1-flash writing the Goal architecture design doc, one on qwen3.8-max writing the dynamic-workflow developer guide) each finished the whole objective inside a single Goal turn of roughly 100 tool calls. What followed was the same in both:assistant_outputentry and onetool_resultentry (goal-evidence.ts:1062-1081), so ~100 steps produce ~200 candidates for a catalog bounded at 100 entries / 24 000 bytes (goal-evidence.ts:32-33).get_goalreportedtruncated: truewith 65 entries, 59 of them shortened.update_goal status=completewas refused solely because the catalog was truncated (goal-tools.ts:359-376), and the proposal was discarded.usage_limited. On qwen the checkpoint succeeded, yet "32 claims while the window overflowed" still counted a stall (goal-checkpoint.ts:65-72) and cost an extra turn./goal resumethe evidence-limited branch repointed the cursor and dropped the checkpoint (goal-reducer.ts:237-252), soget_goalreturned an empty catalog while the tool'snextActionstill said to retry with the UUIDs it returns. The model recovered only by re-running the verification commands to mint fresh evidence.The decisive fact: in both sessions the evidence the verifier finally accepted was the verification output re-run in the closing turn. None of the earlier catalog entries were needed. The deliverable was on disk and the check was re-runnable.
Why the mechanism, not the tuning
The evidence catalog, UUID citations, checkpoint compaction and the stall breaker are one layer of about 4 500 production lines. Since 2026-08-20, 10 of the 42 goal commits exist to keep that layer alive: #9975 → #9840 → #9835 → #9973 → #11304 → #11305 → #11365 → #11576 → #11578 → #11690 → #11693 → #11796. Each fixed the previous fix's edge. The two sessions above hit the layer after all of them landed.
Neither of the two comparable products carries this layer. One judges a goal condition with a separate evaluator that reads only the recent transcript and truncates the rest; the other has no second model at all and relies on a completion audit written into the continuation prompt. Both keep the same single-slot state machine, automatic continuation, and budget / no-progress stops that we have.
Proposal
Keep the independent verifier. Change what it sees: the current Goal turn's visible output and tool results, newest first, with the existing per-record (2 000 bytes) and per-request (256 000 bytes) limits, and an
omittedEarliercount when the turn does not fit. Then delete the machinery that existed only to make older evidence citable:feat(goal): judge terminal proposals from the current turn's evidence— verifier input becomes the current-turn window;update_goal.evidenceRefsbecomes optional and ignored; theinvalidEvidenceRefs, auto-citation andcheckpointRequiredgates go;get_goalstops returning the catalog; continuation prompt and/goal-draftskill say the verifier judges only this turn, and name the 1 500-character objective limit.refactor(goal): remove evidence checkpoints and the stall breaker— deletegoal-checkpoint.ts,goal-checkpoint-verifier.ts, the runtime checkpoint / batching / settle paths, the evidence-limited resume branch, theevidenceCheckpoint/checkpointStalls/lastCheckpointFailurerecord fields, theevidence_catalog/checkpoint_requestlimit kinds; deprecatemodel.goalCheckpointTimeoutSeconds.refactor(goal): drop the evidence catalog and reference validation— collapsegoal-evidence.tsto the record type, proof-kind mapping and the current-turn window builder.chore(goal): stop carrying checkpoint state through the wire, UI and docs— SDK types, Web Shell mappers / gate / dialogs, CLI pill, docs.Snapshot version stays v2; every removed field is optional, so old snapshots and old clients keep parsing. Budgets, pause reasons, the approval dialog and the legacy projection are untouched (the projection gets its own issue).
Out of scope
Removing the verifier itself. Folding
usage_limitedintopausedis a possible follow-up after PR-4.