Summary
glm-5.2 served through Zen (/zen/go/v1/chat/completions) is multi-sourced across upstream providers with no sticky routing, so byte-identical consecutive requests intermittently land on a provider with a cold prompt cache — a full prompt re-bill and a much slower prefill, for no client-side reason. The gateway already has the mechanism to prevent this (stickyProvider + the x-opencode-session sticky id); it just doesn't appear to be enabled for this model. Requesting it be enabled (at least "prefer") for glm-5.2 and other multi-sourced models.
Repro
10 byte-identical requests (same 19,370-token prompt, temperature: 0, stream: false, 3s apart, same client / API key / IP), non-streaming usage captured:
| call |
cached_tokens |
note |
| 1 |
0 |
cold (expected) |
| 2–5 |
19,328 |
warm |
| 6 |
0 |
full re-bill on identical bytes |
| 7 |
19,328 |
warm again |
| 8 |
10 |
effectively cold again |
| 9–10 |
19,369 |
warm — different cache granularity than calls 2–7 |
Two independent signals that these are different upstream providers, not cache eviction:
- the cached-token quantization flips between 19,328 (64-token-aligned) and 19,369 (exact),
- the
usage JSON shape itself flips between two schemas (one includes estimated_cost and reasoning_tokens, one doesn't) across calls in the same run.
Latency impact on this prompt size: ~6.3s cold vs ~1.6s warm per request. In an agent loop that re-sends a large prefix every turn, each reroute costs several seconds of prefill plus the full uncached token cost (~5.4× the cached rate at glm-5.2 pricing).
repro script (python, stdlib only)
import json, time, urllib.request
KEY = "sk-..." # zen api key
para = ("Section %d. The quick brown fox jumps over the lazy dog while the cache "
"measurement proceeds in a fully deterministic manner without any variation. ")
system = "You are a terse assistant. Reply with at most three words.\n\n" + "".join(para % i for i in range(1, 701))
msgs = [{"role": "system", "content": system}, {"role": "user", "content": "Reply with only: OK"}]
for i in range(10):
body = json.dumps({"model": "glm-5.2", "messages": msgs, "max_tokens": 8,
"temperature": 0, "stream": False}).encode()
req = urllib.request.Request("https://opencode.ai/zen/go/v1/chat/completions", data=body,
headers={"Content-Type": "application/json", "Authorization": f"Bearer {KEY}",
"User-Agent": "repro/1.0"})
with urllib.request.urlopen(req, timeout=120) as r:
u = json.load(r).get("usage", {})
print(i + 1, json.dumps(u))
time.sleep(3)
Where this comes from (from the gateway source in this repo)
packages/console/app/src/routes/zen/util/handler.ts:
stickyId = x-opencode-session header ?? workspaceID ?? ip
createStickyTracker(modelId, modelInfo.stickyProvider, stickyId) — the persisted session→provider pin only exists when the model's stickyProvider is set ("strict" | "prefer", stickyProviderTracker.ts). The code comments show it's enabled for codex models ("cannot change codex model providers mid-session").
- Without it,
selectProvider hashes the last 4 chars of stickyId modulo the filtered provider pool — and the pool composition changes with live budget/TPM/TPS state, so the same stickyId periodically resolves to a different provider → cold cache, even for a client sending identical bytes with a stable identity.
Ask
- Enable
stickyProvider (at least "prefer") for glm-5.2 — and consider it as the default for any multi-sourced model whose upstreams have prompt caches. The mechanism already exists and ships for codex models.
- Consider documenting
x-opencode-session for direct API users, so non-opencode clients (other agents/harnesses using the OpenAI-compatible endpoint) can supply a stable session id and benefit from sticky routing and per-session cache scoping.
Happy to provide more captures. Found while investigating prompt-cache behavior of agent harnesses driving Zen; we're a paying Go-bundle workspace.
Reported-by: @Peetiegonzalez
cc @Peetiegonzalez
Summary
glm-5.2served through Zen (/zen/go/v1/chat/completions) is multi-sourced across upstream providers with no sticky routing, so byte-identical consecutive requests intermittently land on a provider with a cold prompt cache — a full prompt re-bill and a much slower prefill, for no client-side reason. The gateway already has the mechanism to prevent this (stickyProvider+ thex-opencode-sessionsticky id); it just doesn't appear to be enabled for this model. Requesting it be enabled (at least"prefer") forglm-5.2and other multi-sourced models.Repro
10 byte-identical requests (same 19,370-token prompt,
temperature: 0,stream: false, 3s apart, same client / API key / IP), non-streamingusagecaptured:cached_tokensTwo independent signals that these are different upstream providers, not cache eviction:
usageJSON shape itself flips between two schemas (one includesestimated_costandreasoning_tokens, one doesn't) across calls in the same run.Latency impact on this prompt size: ~6.3s cold vs ~1.6s warm per request. In an agent loop that re-sends a large prefix every turn, each reroute costs several seconds of prefill plus the full uncached token cost (~5.4× the cached rate at glm-5.2 pricing).
repro script (python, stdlib only)
Where this comes from (from the gateway source in this repo)
packages/console/app/src/routes/zen/util/handler.ts:stickyId = x-opencode-session header ?? workspaceID ?? ipcreateStickyTracker(modelId, modelInfo.stickyProvider, stickyId)— the persisted session→provider pin only exists when the model'sstickyProvideris set ("strict" | "prefer",stickyProviderTracker.ts). The code comments show it's enabled for codex models ("cannot change codex model providers mid-session").selectProviderhashes the last 4 chars ofstickyIdmodulo the filtered provider pool — and the pool composition changes with live budget/TPM/TPS state, so the same stickyId periodically resolves to a different provider → cold cache, even for a client sending identical bytes with a stable identity.Ask
stickyProvider(at least"prefer") forglm-5.2— and consider it as the default for any multi-sourced model whose upstreams have prompt caches. The mechanism already exists and ships for codex models.x-opencode-sessionfor direct API users, so non-opencode clients (other agents/harnesses using the OpenAI-compatible endpoint) can supply a stable session id and benefit from sticky routing and per-session cache scoping.Happy to provide more captures. Found while investigating prompt-cache behavior of agent harnesses driving Zen; we're a paying Go-bundle workspace.
Reported-by: @Peetiegonzalez
cc @Peetiegonzalez