Repository navigation
core: 60m idle location eviction interrupts a running session and rejects pending questions #51343
Description
Activity
github-actions commented on Sep 25, 2026
Hello all,
I am having the same issue here. Yesterday I had opencode v1 + openchamber v1 running well with sessions of 3 to 12 hours (I have a very old server graphics card that take up to 3-4 hours to give the first token). Once I updated to opencode v2.0.16 + openchamber 2.0.1 I got the same issue. After 1 hour I got the message: Opencode failed to send message with error: Step interrupted
I tested several times with some parameters in the model (if possible to get the timeout) but it didn't work.
If the model answer in less than 60 minutes everything works correctly.
Regards.
Correction to the initial description: I first attributed this to a LayerMap idleTimeToLive: "60 minutes" in location-services.ts, but that was from the dev tree. On the released 2.0.8, location-services.ts builds the LayerMap with idleTimeToLive: Duration.infinity; the 60-minute mechanism is LocationActivity (packages/core/src/location-activity.ts), which renews only on durable session events carrying a location and, on expiry, interrupts active executions with reason: "inactivity". I have updated the description accordingly.
Still happens on 2.0.19, and seen here on 2.0.16 and 2.0.18 too (headless opencode serve, OpenChamber as the client). Five waits for an answer ended 60.2 to 60.6 minutes after the tool started, each on the same second as location services evicted for the session's directory: three on the built-in question tool and two on a plugin's review form (Plannotator's submit_plan). So forms created by plugins are cut off the same way. #50499 would cover all five.
Corroborating this on 2.0.18 (latest channel) — same failure, with some new data points. My topology differs from the original report: opencode serve --service with the TUI client on the same host (SQLite session storage), and the TUI process stayed alive with an ESTABLISHED socket for the entire window, so this is not a client disconnect.
Verified timeline (2026-10-01, UTC; from session storage + server log):
| Time (UTC) | Event |
|---|---|
| 10:58:49.998 | question tool part starts running, waiting for user input |
| 12:29:52.8–53.0 | snapshot spawns + location services evicted |
| 12:29:52.859 | pending question part completes with {"type":"aborted","message":"Tool execution interrupted"} — pending for 5462.9 s (~91 min) |
| 12:29:53 | session idle event {"outcome":"interrupted"}; time_idle/idle_outcome set |
| 13:46:49 | user returns and sends a message; only now does the model see the aborted tool result (77 min after the prompt died) |
Threshold: ~90 min, not 60 — worth reconciling. My prompt was pending 91.0 min before eviction. If part creation/ran counts as the last durable event, a 60-minute TTL should have fired ~31 min earlier, so either the TTL changed by 2.0.18 (looks like ~90 min now) or the renewal set is broader than described. For what it's worth, an occurrence in another session the previous day evicted only ~62 min after going pending, so my two data points straddle both 60 and ~90.
State of the pending question part in the session storage after eviction (on 2.0.18 it's the tool part, not the assistant step, that records the abort):
{
"type": "tool",
"name": "question",
"executed": false,
"state": {
"status": "error",
"error": { "type": "aborted", "message": "Tool execution interrupted" },
"time": {
"created": 1790852326718,
"ran": 1790852329998,
"completed": 1790857792859
}
}
}(created = 10:58:46.718Z, ran = 10:58:49.998Z, completed = 12:29:52.859Z.) No notification of any kind — the prompt is simply dead when the user returns.
Server log — note the location is re-booted 335 ms after eviction:
timestamp=2026-10-01T12:29:53.034Z level=INFO run=… message="location services evicted" directory=/…/workspace workspaceID=undefined role=server
timestamp=2026-10-01T12:29:53.369Z level=INFO run=… message="location services booted" directory=/…/workspace workspaceID=undefined durationMs=217 http.span=217 role=server
The abort timestamp (12:29:52.859) precedes the eviction log line by ~175 ms, consistent with the awaitSettlement: true sequencing below.
Eviction path in the shipped 2.0.18 bundle (reformatted, names shortened) — confirms the mechanism described in the original report is core bundle code, no plugin involved:
const now = currentTimeMillisUnsafe()
const expired = Array.from(locations.values()).filter((loc) => loc.expiresAt <= now)
if (expired.length === 0) return
for (const loc of expired) {
const runsInLocation = activeRuns.flatMap((run) =>
sameLocation(run.location, loc.ref) ? [run] : [],
)
yield* each(runsInLocation, (run) =>
sessions.interrupt(run.id, { reason: "inactivity", awaitSettlement: true }),
)
locations.delete(locationKey(loc.ref))
log("location services evicted", {
directory: loc.ref.directory,
workspaceID: loc.ref.workspaceID,
})
}Workaround: ask questions as plain assistant text instead of the interactive question tool — a completed turn survives inactivity and the conversation resumes when the user returns. As in the original report, I found no config, env var, or CLI flag to change the TTL or opt out of the interrupt.
Happy to provide the full session-storage dump or additional log context if useful.
Consolidating what the open PRs cover, since they patch the same eviction sweep:
| PR | Liveness signal it adds | Policy |
|---|---|---|
| #48730 | running terminals | renew while any exist |
| #50499 | pending forms + permissions | renew, whole location; explicit Stop still works |
| #51827 | live child processes | skip expired entries with live children |
| #51583 | per-session progress (model output, tool updates) | keeps a 1h bound for unanswered questions |
Two things worth deciding together:
- Policy for pending human waits. fix(core): preserve progressing sessions during location cleanup #51583's 1h bound for unanswered questions would still kill the headline case here — a question legitimately waiting on a human. The wait I measured before eviction was 5462.9s (~91 min; timeline in my earlier comment). fix(core): preserve pending human waits during idle cleanup #50499's shape (renew while a form/permission is pending, explicit Stop still interrupts) fits the reported failure better.
- These four patch the same ~5 lines with four different probes/registries. If you'd rather have one predicate, I'm happy to help consolidate (crediting everyone) — but that's a maintainer call, not something I'll do unilaterally.
Also worth knowing: branch dev no longer has location-activity.ts — it's back to the plain LayerMap idle TTL (location-services.ts, hardcoded "60 minutes") with no interrupt reason at all (SessionExecution.interrupt on dev can't even express one). If that's unintentional, the next v2→dev merge drops this whole fix. Happy to file a separate issue if useful.
Related reproduction from a different victim path — worth linking so the TTL is recognised as a broad defect, not just the shell-kill case.
On v2.0.22 a location was evicted at the 60-min idle TTL while a session in it was idle after dispatching a background: true shell job. The eviction didn't just interrupt a parked owner: it tore down the location-scoped plugin graph, which is where the background completion watcher is forked (ShellTool.notifyWhenDone, Effect.forkIn(scope, …), tool/plugin/shell.ts:154-186). The watcher fiber was interrupted before sessions.synthetic(...) ran, so the completion notification was never emitted at all (verified: no synthetic row in session_message for that shellID; session_inbox/session_pending empty) and the session was never resumed.
Agree with the config ask here: the TTL is hardcoded (location-activity.ts:25, options.timeToLive ?? "60 minutes") and no env/flag exists. A configurable or even merely higher default would also prevent this notification-loss case, so raising this alongside the parked-session case may help argue for the knob.
Two findings from T3 Code on 2.0.24, in case they help:
- Eviction also drops MCP servers added at runtime (
PUT /api/experimental/mcp/:server): they live in the Location's in-memory overrides. T3 registers one per thread and never re-adds it, so after an idle hour the thread silently has no T3 tools until T3 restarts (no catalog notice either). Verified in a VM withPOST /api/location/reload, which tears down the same way. - A plugin stopgap for parked asks has to go through the HTTP API:
ctx.session.update(...)from plugin code publishes without the Location envelope (the promise adapter's runtime has noLocation.Service), so it renews nothing.PATCH /api/session/:idagainst the server's own port does. Verified with a stepped clock: without it the parked question/permission was cancelled at +75 min; with a periodic PATCH it stayed pending through +100 min.
Description
A session that is still running is interrupted when its Location expires after 60 minutes without a durable session event. Closing the web UI/browser does not stop a run by itself, but a run that is parked (e.g. on a
question) or otherwise silent produces no durable events, so the Location's deadline passes and the active execution is interrupted.How it works on 2.0.8 (
packages/core/src/location-activity.ts):LocationActivitykeeps a per-Location 60-minute deadline and renews it only on durable session events that carry alocation(Schema.is(SessionEvent.Durable)in itsbus.listen). Plain HTTP requests, includingGET /api/session/{id}, do NOT renew it.execution.interrupt(session.id, { reason: "inactivity", awaitSettlement: true })) before invalidating it, which is why the run ends asStep interrupted. Pending questions/forms are rejected by the Location finalizers (packages/core/src/question.ts,location-lifecycle.ts).Observed on a run that was parked:
The same session's last durable events land on the same second:
Two problems:
reason: "inactivity") even when the run is only parked on user input or waiting on a subagent.LocationActivity(itslayer({ timeToLive })defaults to"60 minutes"and the node is built withlayer()); there is no config, env var, or CLI flag to change it, and no way to opt out of interrupting running work.Related: #48691 (terminals/shells killed by the same 60-minute location eviction), #44471 (pending question form lost when a location is evicted/reconnect).
Plugins
usage, ci-wake, harness-review (local, server-side). No desktop terminal/PTY involved.
OpenCode version
2.0.8 (latest, standalone
opencode serve)Steps to reproduce
opencode serveon an always-on host and open the web UI in a browser on another machine.questiontool, or wait on a stuck subagent).location services evictedfor the directory and the run ends withStep interrupted(idle outcome=interrupted).In the observed incident the eviction timestamp and the
idle outcome=interruptedevent fall on the same second, 60 minutes after the Location was last renewed by a durable event.Screenshot and/or share link
No response
Operating System
macOS 26.6.2 (Darwin 25.6.0, arm64) on the server; client is a MacBook browser
Terminal
n/a — headless
opencode serveplus the web UI in a browser