Prompt caching affects latency, token cost, and local-model prefill time across long-running Qwen Code sessions. The work is currently spread across provider adapters, prompt construction, tool discovery, local KV-cache reuse, forks, and telemetry.
Goal: keep the reusable prompt prefix stable and observable so unchanged system instructions, tool schemas, and conversation history are not repeatedly reprocessed or rebilled.
This means:
- Semantically identical requests serialize cacheable prefixes deterministically.
- Dynamic state is appended after stable content instead of rewriting the front of the prompt.
- Side queries, deferred-tool discovery, forks, and resumed agents preserve cache reuse where their semantics allow it.
- Provider-specific cache controls, breakpoints, scopes, and retention settings work across supported endpoints.
- Local backends avoid unnecessary KV-cache invalidation and full re-prefill.
- Cache reads, writes, misses, and effective savings are visible in usage diagnostics.
How to Use This Tracker
This issue coordinates independently shippable work; it is not intended to become one large refactor.
- Own: take one scoped issue and deliver a focused PR with tests.
- Measure: include a before/after request-prefix comparison, provider usage counters, or local-backend prefill evidence.
- Triage: link new prompt-cache issues here and distinguish client-side prefix changes from provider-side cache behavior.
- Maintain: move completed work to the historical section and keep active PR status current.
Open Work
Stable prompt and tool prefixes
Local backends and compaction
Forks and cache sharing
Adjacent prompt composition
Completed and Historical Work
Prompt stability and tool discovery
Provider cache behavior
Local prefill and cache reuse
Forks, side queries, and resumed agents
Observability
Related Trackers
Scope Boundary
This tracker covers model prompt/prefix caching, provider prompt-cache controls, and local model KV/prefix reuse. It intentionally excludes unrelated caches such as file-search results, memory-recall storage, build artifacts, authentication state, and UI data caches.
Prompt caching affects latency, token cost, and local-model prefill time across long-running Qwen Code sessions. The work is currently spread across provider adapters, prompt construction, tool discovery, local KV-cache reuse, forks, and telemetry.
Goal: keep the reusable prompt prefix stable and observable so unchanged system instructions, tool schemas, and conversation history are not repeatedly reprocessed or rebilled.
This means:
How to Use This Tracker
This issue coordinates independently shippable work; it is not intended to become one large refactor.
Open Work
Stable prompt and tool prefixes
Local backends and compaction
Forks and cache sharing
Adjacent prompt composition
Completed and Historical Work
Prompt stability and tool discovery
tool_searchinvalidates LLM server KV-cache on every deferred-tool load #6265 — reduce KV-cache invalidation whentool_searchreveals deferred tools.Provider cache behavior
Local prefill and cache reuse
Forks, side queries, and resumed agents
/btwside queries.Observability
Related Trackers
Scope Boundary
This tracker covers model prompt/prefix caching, provider prompt-cache controls, and local model KV/prefix reuse. It intentionally excludes unrelated caches such as file-search results, memory-recall storage, build artifacts, authentication state, and UI data caches.