Description
I use opencode-go in a long-running coding agent workflow. With the same workflow, DeepSeek V4 Flash keeps stable and high cache reads across many turns. But after switching to GLM-5.1, cache reads randomly drop to 0, causing large unexpected cost spikes.
Observed behavior
In one session, DeepSeek V4 Flash ran from turn 1 to turn 116 without abnormal cacheRead=0 events after the initial request.
After switching to GLM-5.1, I saw two complete cache read failures within only 15 turns:
turn 123:
model: GLM-5.1
cacheRead: 0
input: 45918
turn 127:
model: GLM-5.1
cacheRead: 0
input: 63535
The surrounding turns had normal cache reads:
turn 122: cacheRead: 22272
turn 123: cacheRead: 0
turn 124: cacheRead: 35456
turn 126: cacheRead: 45888
turn 127: cacheRead: 0
turn 128: cacheRead: 57408
So GLM-5.1 does support cache sometimes, but cache reads randomly collapse to 0.
Prefix stability evidence
I added debug logging before the model call. Around both abnormal GLM-5.1 turns, the request prefix did not change:
sys_hash: 5d3ec78c
tools_hash: 6244fe09
prefix_4k_hash: 26ac1f5d
prefix_16k_hash: 30e260ca
prefix_64k_hash: d65f320d
The diff against the previous request also showed that only new items were appended:
turn 123:
previous input_items_count: 42
first_changed_input_index: 42
turn 127:
previous input_items_count: 57
first_changed_input_index: 57
This means the existing request prefix was unchanged. There was no mutation to the system prompt, tools, or earlier input items. From the client side, the prompt should have been cache-readable.
Request
Please investigate whether GLM-5.1 on opencode-go has unstable prompt cache routing or cache key/session handling.
Specifically:
- Are GLM-5.1 requests using sticky routing / session affinity for prompt cache?
- Is a stable cache key or session key passed to the upstream GLM API?
- Is GLM-5.1 routed differently from DeepSeek V4 Flash?
- Is
cacheRead: 0 expected even when the request prefix is unchanged?
- If opencode-go only forwards to the official GLM API, could this be escalated upstream with the evidence above?
This behavior makes GLM-5.1 unreliable for long-running coding agent workloads, because cache reads can randomly drop from tens of thousands of tokens to zero even when the request prefix is stable.
Plugins
None
OpenCode version
latest
Steps to reproduce
Use opencode-go with a long-running coding agent workflow through the OpenAI-compatible /v1/chat/completions API.
Screenshot and/or share link
Operating System
Windows 11
Terminal
Windows terminal
Description
I use opencode-go in a long-running coding agent workflow. With the same workflow, DeepSeek V4 Flash keeps stable and high cache reads across many turns. But after switching to GLM-5.1, cache reads randomly drop to 0, causing large unexpected cost spikes.
Observed behavior
In one session, DeepSeek V4 Flash ran from turn 1 to turn 116 without abnormal
cacheRead=0events after the initial request.After switching to GLM-5.1, I saw two complete cache read failures within only 15 turns:
The surrounding turns had normal cache reads:
So GLM-5.1 does support cache sometimes, but cache reads randomly collapse to 0.
Prefix stability evidence
I added debug logging before the model call. Around both abnormal GLM-5.1 turns, the request prefix did not change:
The diff against the previous request also showed that only new items were appended:
This means the existing request prefix was unchanged. There was no mutation to the system prompt, tools, or earlier input items. From the client side, the prompt should have been cache-readable.
Request
Please investigate whether GLM-5.1 on opencode-go has unstable prompt cache routing or cache key/session handling.
Specifically:
cacheRead: 0expected even when the request prefix is unchanged?This behavior makes GLM-5.1 unreliable for long-running coding agent workloads, because cache reads can randomly drop from tens of thousands of tokens to zero even when the request prefix is stable.
Plugins
None
OpenCode version
latest
Steps to reproduce
Use opencode-go with a long-running coding agent workflow through the OpenAI-compatible
/v1/chat/completionsAPI.Screenshot and/or share link
Operating System
Windows 11
Terminal
Windows terminal