Skip to content

Slim the Goal runtime: judge completion from the current turn's evidence and drop the evidence catalog and checkpoints #12053

Description

@qqqys

Problem

Two real /goal-draft sessions on 2026-09-16 (one on deepseek-v4.1-flash writing the Goal architecture design doc, one on qwen3.8-max writing the dynamic-workflow developer guide) each finished the whole objective inside a single Goal turn of roughly 100 tool calls. What followed was the same in both:

  1. The evidence catalog overflowed. Every model step records one assistant_output entry and one tool_result entry (goal-evidence.ts:1062-1081), so ~100 steps produce ~200 candidates for a catalog bounded at 100 entries / 24 000 bytes (goal-evidence.ts:32-33). get_goal reported truncated: true with 65 entries, 59 of them shortened.
  2. update_goal status=complete was refused solely because the catalog was truncated (goal-tools.ts:359-376), and the proposal was discarded.
  3. The turn-end checkpoint then had to compress the overflowed window through a single-shot fast-model side query. On deepseek it returned invalid claims, then invalid JSON on batch 3/3; three stalls stopped the Goal as usage_limited. On qwen the checkpoint succeeded, yet "32 claims while the window overflowed" still counted a stall (goal-checkpoint.ts:65-72) and cost an extra turn.
  4. After /goal resume the evidence-limited branch repointed the cursor and dropped the checkpoint (goal-reducer.ts:237-252), so get_goal returned an empty catalog while the tool's nextAction still said to retry with the UUIDs it returns. The model recovered only by re-running the verification commands to mint fresh evidence.

The decisive fact: in both sessions the evidence the verifier finally accepted was the verification output re-run in the closing turn. None of the earlier catalog entries were needed. The deliverable was on disk and the check was re-runnable.

Why the mechanism, not the tuning

The evidence catalog, UUID citations, checkpoint compaction and the stall breaker are one layer of about 4 500 production lines. Since 2026-08-20, 10 of the 42 goal commits exist to keep that layer alive: #9975 → #9840 → #9835 → #9973 → #11304 → #11305 → #11365 → #11576 → #11578 → #11690 → #11693 → #11796. Each fixed the previous fix's edge. The two sessions above hit the layer after all of them landed.

Neither of the two comparable products carries this layer. One judges a goal condition with a separate evaluator that reads only the recent transcript and truncates the rest; the other has no second model at all and relies on a completion audit written into the continuation prompt. Both keep the same single-slot state machine, automatic continuation, and budget / no-progress stops that we have.

Proposal

Keep the independent verifier. Change what it sees: the current Goal turn's visible output and tool results, newest first, with the existing per-record (2 000 bytes) and per-request (256 000 bytes) limits, and an omittedEarlier count when the turn does not fit. Then delete the machinery that existed only to make older evidence citable:

  • PR-1 feat(goal): judge terminal proposals from the current turn's evidence — verifier input becomes the current-turn window; update_goal.evidenceRefs becomes optional and ignored; the invalidEvidenceRefs, auto-citation and checkpointRequired gates go; get_goal stops returning the catalog; continuation prompt and /goal-draft skill say the verifier judges only this turn, and name the 1 500-character objective limit.
  • PR-2 refactor(goal): remove evidence checkpoints and the stall breaker — delete goal-checkpoint.ts, goal-checkpoint-verifier.ts, the runtime checkpoint / batching / settle paths, the evidence-limited resume branch, the evidenceCheckpoint / checkpointStalls / lastCheckpointFailure record fields, the evidence_catalog / checkpoint_request limit kinds; deprecate model.goalCheckpointTimeoutSeconds.
  • PR-3 refactor(goal): drop the evidence catalog and reference validation — collapse goal-evidence.ts to the record type, proof-kind mapping and the current-turn window builder.
  • PR-4 chore(goal): stop carrying checkpoint state through the wire, UI and docs — SDK types, Web Shell mappers / gate / dialogs, CLI pill, docs.

Snapshot version stays v2; every removed field is optional, so old snapshots and old clients keep parsing. Budgets, pause reasons, the approval dialog and the legacy projection are untouched (the projection gets its own issue).

Out of scope

Removing the verifier itself. Folding usage_limited into paused is a possible follow-up after PR-4.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions