Skip to content

Zen: byte-identical requests to glm-5.2 intermittently reroute to a cold-cache provider — please enable stickyProvider for multi-sourced models #35402

Description

@Marvinthebored

Summary

glm-5.2 served through Zen (/zen/go/v1/chat/completions) is multi-sourced across upstream providers with no sticky routing, so byte-identical consecutive requests intermittently land on a provider with a cold prompt cache — a full prompt re-bill and a much slower prefill, for no client-side reason. The gateway already has the mechanism to prevent this (stickyProvider + the x-opencode-session sticky id); it just doesn't appear to be enabled for this model. Requesting it be enabled (at least "prefer") for glm-5.2 and other multi-sourced models.

Repro

10 byte-identical requests (same 19,370-token prompt, temperature: 0, stream: false, 3s apart, same client / API key / IP), non-streaming usage captured:

call cached_tokens note
1 0 cold (expected)
2–5 19,328 warm
6 0 full re-bill on identical bytes
7 19,328 warm again
8 10 effectively cold again
9–10 19,369 warm — different cache granularity than calls 2–7

Two independent signals that these are different upstream providers, not cache eviction:

  • the cached-token quantization flips between 19,328 (64-token-aligned) and 19,369 (exact),
  • the usage JSON shape itself flips between two schemas (one includes estimated_cost and reasoning_tokens, one doesn't) across calls in the same run.

Latency impact on this prompt size: ~6.3s cold vs ~1.6s warm per request. In an agent loop that re-sends a large prefix every turn, each reroute costs several seconds of prefill plus the full uncached token cost (~5.4× the cached rate at glm-5.2 pricing).

repro script (python, stdlib only)
import json, time, urllib.request

KEY = "sk-..."  # zen api key
para = ("Section %d. The quick brown fox jumps over the lazy dog while the cache "
        "measurement proceeds in a fully deterministic manner without any variation. ")
system = "You are a terse assistant. Reply with at most three words.\n\n" + "".join(para % i for i in range(1, 701))
msgs = [{"role": "system", "content": system}, {"role": "user", "content": "Reply with only: OK"}]

for i in range(10):
    body = json.dumps({"model": "glm-5.2", "messages": msgs, "max_tokens": 8,
                       "temperature": 0, "stream": False}).encode()
    req = urllib.request.Request("https://opencode.ai/zen/go/v1/chat/completions", data=body,
        headers={"Content-Type": "application/json", "Authorization": f"Bearer {KEY}",
                 "User-Agent": "repro/1.0"})
    with urllib.request.urlopen(req, timeout=120) as r:
        u = json.load(r).get("usage", {})
    print(i + 1, json.dumps(u))
    time.sleep(3)

Where this comes from (from the gateway source in this repo)

packages/console/app/src/routes/zen/util/handler.ts:

  • stickyId = x-opencode-session header ?? workspaceID ?? ip
  • createStickyTracker(modelId, modelInfo.stickyProvider, stickyId) — the persisted session→provider pin only exists when the model's stickyProvider is set ("strict" | "prefer", stickyProviderTracker.ts). The code comments show it's enabled for codex models ("cannot change codex model providers mid-session").
  • Without it, selectProvider hashes the last 4 chars of stickyId modulo the filtered provider pool — and the pool composition changes with live budget/TPM/TPS state, so the same stickyId periodically resolves to a different provider → cold cache, even for a client sending identical bytes with a stable identity.

Ask

  1. Enable stickyProvider (at least "prefer") for glm-5.2 — and consider it as the default for any multi-sourced model whose upstreams have prompt caches. The mechanism already exists and ships for codex models.
  2. Consider documenting x-opencode-session for direct API users, so non-opencode clients (other agents/harnesses using the OpenAI-compatible endpoint) can supply a stable session id and benefit from sticky routing and per-session cache scoping.

Happy to provide more captures. Found while investigating prompt-cache behavior of agent harnesses driving Zen; we're a paying Go-bundle workspace.

Reported-by: @Peetiegonzalez
cc @Peetiegonzalez

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions