You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Introduce a feature-gated stream-driven tool execution path for Qwen Code.
Today Qwen Code can surface ToolCallRequest / functionCall information while the model response is streaming, but the main execution paths still buffer tool calls and dispatch them after the model stream exits. This RFC proposes starting eligible tools as soon as their tool-call payload is complete, while still buffering tool results until the matching model functionCall is safely present in history.
This is similar in shape to Claude Code's in-band StreamingToolExecutor path, gated upstream behind tengu_streaming_tool_execution2.
This is an RFC rather than a PR because the change crosses several boundaries: stream parsing, Turn.run, CoreToolScheduler, history invariants, retry/fallback handling, and the React UI scheduler.
Status note: #4176 is in active review at the time of filing. This RFC builds on the invariant analysis in #4176, but is intentionally separate from that defensive fix.
Motivation
For tool-heavy turns, Qwen Code currently pays avoidable latency:
current:
model streams tool calls -> stream ends -> scheduler starts tools -> submit tool results
proposed:
model streams complete tool call A -> start safe tool A
model continues streaming -> start safe tool B when complete
stream ends -> wait only for remaining tools -> submit buffered tool results
The expected benefit is lower end-to-end latency for read/search/fetch-heavy turns and agent workflows, without changing the model-visible request shape.
This does not mean submitting functionResponse parts early. The API/history invariant still requires each functionResponse to be adjacent to its matching model functionCall.
Current State
Qwen Code already has several related pieces, but not in-band streaming tool execution:
Area
Current behavior
Stream parsing
Some providers can emit completed functionCall / ToolCallRequest during the stream.
Interactive UI
useGeminiStream collects tool requests during the stream and calls scheduleToolCalls(...) after the stream loop exits.
Headless CLI
Collects ToolCallRequestInfo[], finalizes the assistant output, then runs processToolCallBatch(...).
Tool scheduler
#2864 added safe parallel batching, but scheduling still starts after the tool-call batch is known.
Failure repair
#4176 repairs tool_use ↔ tool_result invariant failures caused by mid-stream drops and out-of-band scheduler timing.
So the missing piece is not tool-call streaming itself. The missing piece is an in-band owner that can start tools during the response stream, coordinate cancellation/discard on stream failure, and preserve history correctness.
Proposed Design
1. Add a feature-gated streaming executor
Introduce a StreamingToolExecutor-like coordinator owned by the core turn loop, not only by the React layer.
Responsibilities:
accept completed ToolCallRequestInfo as they arrive
This is non-negotiable for OpenAI/Anthropic-compatible providers and for resume/compression/repair behavior.
3. Start conservatively
Initial scope should only early-execute clearly safe tools:
read/list/glob/search
web fetch/search if classified safe
read-only shell commands only if already classified as read-only
possibly agent calls, but only after separate review
Do not early-execute in the first version:
edit/write tools
unknown tools
tools requiring confirmation unless the UI path is explicitly handled
batches containing structured_output until sibling suppression semantics are nailed down
4. Keep the old path as fallback
Add a flag such as:
experimental.streamingToolDispatch
QWEN_CODE_STREAMING_TOOL_DISPATCH=1
When disabled, the current post-stream scheduler path remains unchanged.
Important Constraints
This should not be implemented by simply calling scheduleToolCalls(...) immediately from the current ToolCallRequest handler.
Key risks:
History validity
functionResponse cannot be inserted before its matching model functionCall has been persisted.
Retry/fallback safety
If the stream retries after a tool started, the executor must cancel/discard or synthesize an error result without leaking orphan results into history.
Structured output
In --json-schema mode, structured_output can suppress sibling tools. Early execution must not run side-effecting sibling tools before the full batch semantics are known.
Some providers may only emit usable tool calls near the final chunk. The feature should degrade gracefully and remain gated until provider-specific behavior is verified.
Rollout Plan
Phase
Scope
Risk
Phase 1
Add a uniform completed-tool-call signal to provider stream conversion. Existing callers still buffer, so no behavior change.
Low
Phase 2
Add a streaming executor skeleton behind a disabled flag. It can accept calls and buffer results, but the old post-stream path remains default.
Medium
Phase 3
Enable early execution for read-only/concurrency-safe tools only. Preserve buffered result submission and current history shape.
Medium
Phase 4
Add retry/fallback/abort discard semantics and targeted tests for orphan prevention.
Medium
Phase 5
Evaluate broader tool classes and sibling-error abort once telemetry shows the safe path is stable.
Higher; defer
The feature should default off until telemetry and integration tests show no regression in tool_use ↔ tool_result invariant violations.
Alternatives Considered
A. Stay non-streaming and keep patching failure symptoms
#4176 is the correct defensive fix for the current architecture, but it does not reduce latency and it keeps the stream loop and tool scheduler split across layers.
B. Speculative dispatch on partial JSON deltas
Start a tool before its argument JSON is fully assembled. Rejected: the risk/value tradeoff is poor. The safer boundary is a complete tool-call payload, equivalent to Anthropic content_block_stop.
C. Stream-driven only for the Anthropic-compatible path
Tempting because that path already has a clear content_block_stop boundary. Rejected as a long-term default because maintaining divergent execution models by provider would keep the current cross-layer complexity.
Open Questions
Executor ownership
Should the new executor live entirely in core, or should the React scheduler be folded into the same ownership model?
Cancellation semantics
If the stream errors after a tool starts but before turn end, should the executor cancel in-flight tools, let safe tools finish and synthesize results, or choose per tool kind?
Result order
Should buffered functionResponse parts preserve dispatch order, model order, or completion order? The safest default seems model/dispatch order.
Provider support
Which providers actually emit complete tool calls before final stream termination? The feature should be measured per provider before default-on rollout.
Structured output
Should structured_output disable streaming dispatch for the entire batch, or only for sibling tools that are not provably read-only?
Acceptance Criteria
Feature flag/config controls the new path.
Existing post-stream tool scheduling remains the fallback.
Safe tools can begin execution before the model stream fully completes.
Tool results are not submitted to the model until matching function calls are present in history.
Retry/fallback/abort paths do not create orphan functionResponse entries.
Tests cover:
early execution of safe tools
unsafe tools preserving order
stream retry after early tool start
abort while early tool is running
structured output sibling suppression
headless and interactive paths
history invariant: every functionCall has exactly one matching functionResponse
Summary
Introduce a feature-gated stream-driven tool execution path for Qwen Code.
Today Qwen Code can surface
ToolCallRequest/functionCallinformation while the model response is streaming, but the main execution paths still buffer tool calls and dispatch them after the model stream exits. This RFC proposes starting eligible tools as soon as their tool-call payload is complete, while still buffering tool results until the matching modelfunctionCallis safely present in history.This is similar in shape to Claude Code's in-band
StreamingToolExecutorpath, gated upstream behindtengu_streaming_tool_execution2.This is an RFC rather than a PR because the change crosses several boundaries: stream parsing,
Turn.run,CoreToolScheduler, history invariants, retry/fallback handling, and the React UI scheduler.Motivation
For tool-heavy turns, Qwen Code currently pays avoidable latency:
The expected benefit is lower end-to-end latency for read/search/fetch-heavy turns and agent workflows, without changing the model-visible request shape.
This does not mean submitting
functionResponseparts early. The API/history invariant still requires eachfunctionResponseto be adjacent to its matching modelfunctionCall.Current State
Qwen Code already has several related pieces, but not in-band streaming tool execution:
functionCall/ToolCallRequestduring the stream.useGeminiStreamcollects tool requests during the stream and callsscheduleToolCalls(...)after the stream loop exits.ToolCallRequestInfo[], finalizes the assistant output, then runsprocessToolCallBatch(...).tool_use ↔ tool_resultinvariant failures caused by mid-stream drops and out-of-band scheduler timing.So the missing piece is not tool-call streaming itself. The missing piece is an in-band owner that can start tools during the response stream, coordinate cancellation/discard on stream failure, and preserve history correctness.
Proposed Design
1. Add a feature-gated streaming executor
Introduce a
StreamingToolExecutor-like coordinator owned by the core turn loop, not only by the React layer.Responsibilities:
ToolCallRequestInfoas they arrivegetCompletedResults()/getRemainingResults()discard()/ cancellation on retry, fallback, abort, or stream error2. Keep result submission buffered
Tool execution may start early, but tool results should not be submitted to the model until the matching model
functionCallis present in history.The invariant should remain:
This is non-negotiable for OpenAI/Anthropic-compatible providers and for resume/compression/repair behavior.
3. Start conservatively
Initial scope should only early-execute clearly safe tools:
Do not early-execute in the first version:
structured_outputuntil sibling suppression semantics are nailed down4. Keep the old path as fallback
Add a flag such as:
experimental.streamingToolDispatchQWEN_CODE_STREAMING_TOOL_DISPATCH=1When disabled, the current post-stream scheduler path remains unchanged.
Important Constraints
This should not be implemented by simply calling
scheduleToolCalls(...)immediately from the currentToolCallRequesthandler.Key risks:
History validity
functionResponsecannot be inserted before its matching modelfunctionCallhas been persisted.Retry/fallback safety
If the stream retries after a tool started, the executor must cancel/discard or synthesize an error result without leaking orphan results into history.
Structured output
In
--json-schemamode,structured_outputcan suppress sibling tools. Early execution must not run side-effecting sibling tools before the full batch semantics are known.React scheduler boundary
fix(core,cli): close tool_use↔tool_result invariant across all failure paths #4176 shows that the out-of-band React scheduler makes invariant repair harder. The RFC should decide whether the new executor lives fully in core, or whether the React scheduler is folded into the same ownership model.
Provider behavior
Some providers may only emit usable tool calls near the final chunk. The feature should degrade gracefully and remain gated until provider-specific behavior is verified.
Rollout Plan
The feature should default off until telemetry and integration tests show no regression in
tool_use ↔ tool_resultinvariant violations.Alternatives Considered
A. Stay non-streaming and keep patching failure symptoms
#4176 is the correct defensive fix for the current architecture, but it does not reduce latency and it keeps the stream loop and tool scheduler split across layers.
B. Speculative dispatch on partial JSON deltas
Start a tool before its argument JSON is fully assembled. Rejected: the risk/value tradeoff is poor. The safer boundary is a complete tool-call payload, equivalent to Anthropic
content_block_stop.C. Stream-driven only for the Anthropic-compatible path
Tempting because that path already has a clear
content_block_stopboundary. Rejected as a long-term default because maintaining divergent execution models by provider would keep the current cross-layer complexity.Open Questions
Executor ownership
Should the new executor live entirely in core, or should the React scheduler be folded into the same ownership model?
Cancellation semantics
If the stream errors after a tool starts but before turn end, should the executor cancel in-flight tools, let safe tools finish and synthesize results, or choose per tool kind?
Result order
Should buffered
functionResponseparts preserve dispatch order, model order, or completion order? The safest default seems model/dispatch order.Provider support
Which providers actually emit complete tool calls before final stream termination? The feature should be measured per provider before default-on rollout.
Structured output
Should
structured_outputdisable streaming dispatch for the entire batch, or only for sibling tools that are not provably read-only?Acceptance Criteria
functionResponseentries.functionCallhas exactly one matchingfunctionResponseReferences
fix(core,cli): close tool_use↔tool_result invariant across all failure pathstengu_streaming_tool_execution2, in-bandStreamingToolExecutor, and per-block dispatch atcontent_block_stopAsking For