Description
Summary
#39009 reported that the v2 OpenRouter route never sets cache_control for Anthropic models. It was closed by #39907. On 2.0.18, opencode run with Anthropic models on OpenRouter still reports zero cache reads and zero cache writes on every step, so every step bills the full prompt as fresh input.
The #39907 PR description says it uses a "stable session prompt_cache_key for default OpenRouter caching without emitting explicit markers". But OpenRouter only caches Anthropic models when explicit cache_control markers are present. So caching looks fixed only when a cache policy is set explicitly. The default case is still broken.
We confirmed three things:
The impact is large for headless, multi-step runs. In our case, a 2-step extraction run over a ~120k-token prompt cost $0.94. We estimate that caching would have cut that roughly in half.
Environment
- opencode version: 2.0.18 (
~/.opencode/bin/opencode, install/channel: latest). The comparison is with 1.18.32 (Homebrew).
- OS: Darwin 25.6.0 arm64 (macOS)
- Terminal: iTerm.app, TERM=xterm-256color
- Shell: /bin/zsh
- Provider:
openrouter (OpenRouter API key auth). OpenRouter routed every request in these tests to Amazon Bedrock (tool call IDs are toolu_bdrk_…) on both versions and with and without the workaround. So the upstream provider is not the difference.
- Active plugins: local files in
~/.config/opencode/plugins/: temp.ts (v2 Plugin.define), rtk.ts and herdr-agent-state.js. The last two fail to load on v2. None are referenced in opencode.json. The repro uses the default agent and no plugin tools.
Reproduction
In a directory that has a README.md and a pyproject.toml (any two small files will do):
- opencode 2.0.18, default config:
opencode run --format json --title cache-repro-v2 \
--model openrouter/anthropic/claude-sonnet-5 \
"Read README.md, then read pyproject.toml, then reply with just: ok"
Then run opencode session export <id> and look at messages[].tokens.cache on the assistant messages.
- Add this project
opencode.json and run the same command again:
{
"provider": { "openrouter": { "models": {
"anthropic/claude-sonnet-5": { "options": { "cache_control": { "type": "ephemeral" } } }
}}}
}
- Optional: run the same prompt on opencode 1.18.32 without the override, and look at
part.tokens.cache on the step-finish events.
Expected Behavior
With no extra config, step 1 writes the shared prefix (system prompt, tools, first user message) to the cache, and later steps read it. That is what 1.18.32 does by default:
| Step |
input |
cache write |
cache read |
cost |
| 1 |
2 |
12,044 |
0 |
$0.031 |
| 2 |
2 |
1,715 |
12,044 |
$0.008 |
| 3 |
2 |
218 |
13,759 |
$0.004 |
| 4 |
2 |
431 |
13,977 |
$0.004 |
Actual Behavior
On 2.0.18 with default config, every step reports cache: {read: 0, write: 0} and bills the whole prompt as input:
| Step |
input |
cache write |
cache read |
cost |
| 1 |
10,809 |
0 |
0 |
$0.0227 |
| 2 |
11,295 |
0 |
0 |
$0.0226 |
On 2.0.18 with the cache_control model option (step 2), caching works:
| Step |
input |
cache write |
cache read |
cost |
| 1 |
2 |
10,865 |
0 |
$0.0278 |
| 2 |
2 |
84 |
10,865 |
$0.0029 |
| 3 |
2 |
87 |
10,949 |
$0.0025 |
A cache write of 0 on step 1 without the override shows that no breakpoint reaches the provider. The model is not simply missing an existing cache.
Additional Context
Consistent. Every opencode run session with an Anthropic model on 2.0.18 without the override shows zero cache use, across 10 runs and every variant we tried:
Variant (2.0.18, opencode run) |
Model |
cache read |
Custom primary agent, large prompt attached with --file |
claude-sonnet-5 |
0 |
Custom primary agent, agent reads the file with the read tool (3 steps) |
claude-sonnet-5 |
0 |
Default agent, --standalone |
claude-sonnet-5 |
0 |
Default agent, --standalone |
claude-opus-5.5 |
0 |
Default agent, without --standalone |
claude-sonnet-5 |
0 |
Custom agent, variant #none |
claude-haiku-4.5 |
0 |
Interactive sessions are not affected. A long interactive session on 2.0.18 with openrouter/anthropic/claude-opus-5.5 shows 14.4M cache-read tokens. So the interactive path seems to supply a cache policy or hints that run does not. We have not confirmed this in the code.
Automatic-caching models are not affected. Models that OpenRouter caches automatically (for example moonshotai/kimi-k3) show cache reads under opencode run on 2.0.18. This fits the explanation that the Anthropic cache_control markers are what's missing.
Related:
Suggested fix: when an OpenRouter model is Anthropic-family (anthropic/*, ~anthropic/*), emit cache_control markers by default, or the top-level cache_control field. run should get the same behaviour as the interactive path.
Workaround: set "options": { "cache_control": { "type": "ephemeral" } } for each Anthropic model under provider.openrouter.models.
Plugins
Temp plugin
OpenCode version
2.0.18
Steps to reproduce
Run the same prompt in any directory that has a README.md and a pyproject.toml (any two small files will do):
- opencode 2.0.18:
~/.opencode/bin/opencode run --format json --title cache-repro-v2 \
--model openrouter/anthropic/claude-sonnet-5 \
"Read README.md, then read pyproject.toml, then reply with just: ok"
Then opencode session export <id> and look at messages[].tokens.cache on the assistant messages.
- opencode 1.18.32:
/opt/homebrew/bin/opencode run --format json \
--model openrouter/anthropic/claude-sonnet-5 \
"Read README.md, then read pyproject.toml, then reply with just: ok"
Look at part.tokens.cache on the step-finish events.
Screenshot and/or share link
No response
Operating System
Mac 25.6.0
Terminal
iTerm2
Description
Summary
#39009 reported that the v2 OpenRouter route never sets
cache_controlfor Anthropic models. It was closed by #39907. On 2.0.18,opencode runwith Anthropic models on OpenRouter still reports zero cache reads and zero cache writes on every step, so every step bills the full prompt as fresh input.The #39907 PR description says it uses a "stable session prompt_cache_key for default OpenRouter caching without emitting explicit markers". But OpenRouter only caches Anthropic models when explicit
cache_controlmarkers are present. So caching looks fixed only when a cache policy is set explicitly. The default case is still broken.We confirmed three things:
cache_controlis set by hand as a model option. This is the workaround from OpenRouter route never sets cache_control for Anthropic models #39009.The impact is large for headless, multi-step runs. In our case, a 2-step extraction run over a ~120k-token prompt cost $0.94. We estimate that caching would have cut that roughly in half.
Environment
~/.opencode/bin/opencode, install/channel: latest). The comparison is with 1.18.32 (Homebrew).openrouter(OpenRouter API key auth). OpenRouter routed every request in these tests to Amazon Bedrock (tool call IDs aretoolu_bdrk_…) on both versions and with and without the workaround. So the upstream provider is not the difference.~/.config/opencode/plugins/:temp.ts(v2Plugin.define),rtk.tsandherdr-agent-state.js. The last two fail to load on v2. None are referenced inopencode.json. The repro uses the default agent and no plugin tools.Reproduction
In a directory that has a
README.mdand apyproject.toml(any two small files will do):opencode run --format json --title cache-repro-v2 \ --model openrouter/anthropic/claude-sonnet-5 \ "Read README.md, then read pyproject.toml, then reply with just: ok"opencode session export <id>and look atmessages[].tokens.cacheon the assistant messages.opencode.jsonand run the same command again:{ "provider": { "openrouter": { "models": { "anthropic/claude-sonnet-5": { "options": { "cache_control": { "type": "ephemeral" } } } }}} }part.tokens.cacheon thestep-finishevents.Expected Behavior
With no extra config, step 1 writes the shared prefix (system prompt, tools, first user message) to the cache, and later steps read it. That is what 1.18.32 does by default:
Actual Behavior
On 2.0.18 with default config, every step reports
cache: {read: 0, write: 0}and bills the whole prompt as input:On 2.0.18 with the
cache_controlmodel option (step 2), caching works:A cache write of 0 on step 1 without the override shows that no breakpoint reaches the provider. The model is not simply missing an existing cache.
Additional Context
Consistent. Every
opencode runsession with an Anthropic model on 2.0.18 without the override shows zero cache use, across 10 runs and every variant we tried:opencode run)--filereadtool (3 steps)--standalone--standalone--standalone#noneInteractive sessions are not affected. A long interactive session on 2.0.18 with
openrouter/anthropic/claude-opus-5.5shows 14.4M cache-read tokens. So the interactive path seems to supply a cache policy or hints thatrundoes not. We have not confirmed this in the code.Automatic-caching models are not affected. Models that OpenRouter caches automatically (for example
moonshotai/kimi-k3) show cache reads underopencode runon 2.0.18. This fits the explanation that the Anthropiccache_controlmarkers are what's missing.Related:
prompt_cache_key"without emitting explicit markers". That is not enough for Anthropic models on OpenRouter.cache_controlfor Anthropic models.Suggested fix: when an OpenRouter model is Anthropic-family (
anthropic/*,~anthropic/*), emitcache_controlmarkers by default, or the top-levelcache_controlfield.runshould get the same behaviour as the interactive path.Workaround: set
"options": { "cache_control": { "type": "ephemeral" } }for each Anthropic model underprovider.openrouter.models.Plugins
Temp plugin
OpenCode version
2.0.18
Steps to reproduce
Run the same prompt in any directory that has a
README.mdand apyproject.toml(any two small files will do):opencode session export <id>and look atmessages[].tokens.cacheon the assistant messages./opt/homebrew/bin/opencode run --format json \ --model openrouter/anthropic/claude-sonnet-5 \ "Read README.md, then read pyproject.toml, then reply with just: ok"part.tokens.cacheon thestep-finishevents.Screenshot and/or share link
No response
Operating System
Mac 25.6.0
Terminal
iTerm2