Description
The bug: when a turn fails at the persistence layer (e.g. the session_message insert failure from #53146), opencode drains/retries the turn and re-executes its pending tool calls. A side-effectful call from the failed turn (in our case a bash tea issue create against a forge API) runs once per retry. The agent never receives any result between retries, so no agent-side protocol (search-before-create, verification) can intervene - the replay happens below the reasoning layer.
Blast radius (one incident, one host, one day): 35 duplicate forge issues from a single session, in two bursts at transport-retry cadence (~40-70 s apart), each burst bracketed by Failed to drain Session ERRORs in the server log.
Evidence (happy to share sanitized logs):
Trigger relationship to #53146: #53146 (independent per-process seq counters on a shared db) is one way to fail the turn; the replay is a distinct defect in the retry path. Fixing #53146 removes a common trigger but any other persistence failure would replay side effects again.
Expected behavior: a retried turn must not re-execute tool calls whose execution already started or completed (record intent before dispatch, or mark side-effectful calls as non-replayable and fail the turn explicitly). Failures should surface to the agent as a turn error it can observe, not be silently replayed.
Environment: opencode 2.0.22 (also observed on 1.18.34 fleet-side), Linux workstation, serve --service + TUI topology as in #53146.
Cross-ref: satware.ai/harness#1166 (RCA), satware.ai/harness#1168 (downstream duplication RCA + cleanup), #53146 (seq collision trigger).
Description
The bug: when a turn fails at the persistence layer (e.g. the
session_messageinsert failure from #53146), opencode drains/retries the turn and re-executes its pending tool calls. A side-effectful call from the failed turn (in our case a bashtea issue createagainst a forge API) runs once per retry. The agent never receives any result between retries, so no agent-side protocol (search-before-create, verification) can intervene - the replay happens below the reasoning layer.Blast radius (one incident, one host, one day): 35 duplicate forge issues from a single session, in two bursts at transport-retry cadence (~40-70 s apart), each burst bracketed by
Failed to drain SessionERRORs in the server log.Evidence (happy to share sanitized logs):
tea i createcommand line) spawned 41 times under one server run id, interleaved withFailed to drain Session ... insert into "session_message"errors.Trigger relationship to #53146: #53146 (independent per-process seq counters on a shared db) is one way to fail the turn; the replay is a distinct defect in the retry path. Fixing #53146 removes a common trigger but any other persistence failure would replay side effects again.
Expected behavior: a retried turn must not re-execute tool calls whose execution already started or completed (record intent before dispatch, or mark side-effectful calls as non-replayable and fail the turn explicitly). Failures should surface to the agent as a turn error it can observe, not be silently replayed.
Environment: opencode 2.0.22 (also observed on 1.18.34 fleet-side), Linux workstation,
serve --service+ TUI topology as in #53146.Cross-ref: satware.ai/harness#1166 (RCA), satware.ai/harness#1168 (downstream duplication RCA + cleanup), #53146 (seq collision trigger).