Skip to content

Better Prompt Caching #8277

Description

@DragonnZhang

Prompt caching affects latency, token cost, and local-model prefill time across long-running Qwen Code sessions. The work is currently spread across provider adapters, prompt construction, tool discovery, local KV-cache reuse, forks, and telemetry.

Goal: keep the reusable prompt prefix stable and observable so unchanged system instructions, tool schemas, and conversation history are not repeatedly reprocessed or rebilled.

This means:

  • Semantically identical requests serialize cacheable prefixes deterministically.
  • Dynamic state is appended after stable content instead of rewriting the front of the prompt.
  • Side queries, deferred-tool discovery, forks, and resumed agents preserve cache reuse where their semantics allow it.
  • Provider-specific cache controls, breakpoints, scopes, and retention settings work across supported endpoints.
  • Local backends avoid unnecessary KV-cache invalidation and full re-prefill.
  • Cache reads, writes, misses, and effective savings are visible in usage diagnostics.

How to Use This Tracker

This issue coordinates independently shippable work; it is not intended to become one large refactor.

  • Own: take one scoped issue and deliver a focused PR with tests.
  • Measure: include a before/after request-prefix comparison, provider usage counters, or local-backend prefill evidence.
  • Triage: link new prompt-cache issues here and distinguish client-side prefix changes from provider-side cache behavior.
  • Maintain: move completed work to the historical section and keep active PR status current.

Open Work

Stable prompt and tool prefixes

Local backends and compaction

Forks and cache sharing

Adjacent prompt composition

Completed and Historical Work

Prompt stability and tool discovery

Provider cache behavior

Local prefill and cache reuse

Forks, side queries, and resumed agents

Observability

Related Trackers

Scope Boundary

This tracker covers model prompt/prefix caching, provider prompt-cache controls, and local model KV/prefix reuse. It intentionally excludes unrelated caches such as file-search results, memory-recall storage, build artifacts, authentication state, and UI data caches.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions