Skip to content

Keep deferred tool discovery from invalidating prompt cache prefixes #6721

Description

@water-in-stone

What happened?

In the main session, deferred tools are discovered through tool_search. When the model searches for a hidden deferred tool, Qwen Code currently resolves the real tool schema, marks that tool as revealed, and calls GeminiClient.setTools() so the real tool declaration is added to the next API request.

That means the provider-facing tools / functionDeclarations block changes after discovery. For example:

Request 1 tools:
[read_file, edit, tool_search]

tool_search reveals cron_create
  -> revealDeferredTool("cron_create")
  -> setTools()

Request 2 tools:
[read_file, edit, tool_search, cron_create]

The model can call the discovered tool, but the request prefix has changed at the very front of the request. This can prevent provider prompt-cache reuse, and it can also hurt local prefix/KV reuse because unchanged system instruction and early history now sit behind a changed tools block.
Sorting tool declarations only fixes unstable ordering within the same tool set. It does not solve this case, because the actual tool set changes after a deferred tool is revealed.

What did you expect to happen?

Discovering a deferred tool should not mutate the provider-facing function declaration list for the main session.
Expected behavior:

Request 1 tools:
[read_file, edit, tool_search, deferred_tool_call, exit_plan_mode]

tool_search presents cron_create
  -> returns the current cron_create schema in tool result text
  -> records pending schema presentation metadata
  -> commits that presentation only after the result enters active history
  -> does not call setTools()
  -> does not add cron_create to provider tools

Request 2 tools:
[read_file, edit, tool_search, deferred_tool_call, exit_plan_mode]

Provider call:
deferred_tool_call({ name: "cron_create", arguments: {...} })

Internal scheduled call:
cron_create({...})

Provider response:
functionResponse({ name: "deferred_tool_call", id: originalCallId, ... })

This keeps the main-session functionDeclarations bytes stable after tool_search, while preserving the existing execution boundaries: target permission checks, schema validation, confirmation, hooks, telemetry, truncation, cancellation, and result recording should still run against the real target tool.
The proxy should not grant authorization by itself. It should only route a provider-visible stable wrapper call to a real deferred tool after the target schema has already been shown in the active model context and its current schema fingerprint still matches.

Client information

Qwen Code (v0.19.4)

Login information

No response

Anything else we need to know?

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

category/coreCore engine and logiccategory/performancePerformance and optimizationpriority/P2Medium - Moderately impactful, noticeable problemscope/cachingCaching mechanismsscope/token-managementToken handling and limitsstatus/needs-triageIssue needs to be triaged and labeledtype/bugSomething isn't working as expected

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions