Skip to content

[RFC] Stream-driven tool dispatch — align tool execution timing with upstream's StreamingToolExecutor #4387

Description

@BZ-D

Summary

Introduce a feature-gated stream-driven tool execution path for Qwen Code.

Today Qwen Code can surface ToolCallRequest / functionCall information while the model response is streaming, but the main execution paths still buffer tool calls and dispatch them after the model stream exits. This RFC proposes starting eligible tools as soon as their tool-call payload is complete, while still buffering tool results until the matching model functionCall is safely present in history.

This is similar in shape to Claude Code's in-band StreamingToolExecutor path, gated upstream behind tengu_streaming_tool_execution2.

This is an RFC rather than a PR because the change crosses several boundaries: stream parsing, Turn.run, CoreToolScheduler, history invariants, retry/fallback handling, and the React UI scheduler.

Status note: #4176 is in active review at the time of filing. This RFC builds on the invariant analysis in #4176, but is intentionally separate from that defensive fix.

Motivation

For tool-heavy turns, Qwen Code currently pays avoidable latency:

current:
model streams tool calls -> stream ends -> scheduler starts tools -> submit tool results

proposed:
model streams complete tool call A -> start safe tool A
model continues streaming -> start safe tool B when complete
stream ends -> wait only for remaining tools -> submit buffered tool results

The expected benefit is lower end-to-end latency for read/search/fetch-heavy turns and agent workflows, without changing the model-visible request shape.

This does not mean submitting functionResponse parts early. The API/history invariant still requires each functionResponse to be adjacent to its matching model functionCall.

Current State

Qwen Code already has several related pieces, but not in-band streaming tool execution:

Area Current behavior
Stream parsing Some providers can emit completed functionCall / ToolCallRequest during the stream.
Interactive UI useGeminiStream collects tool requests during the stream and calls scheduleToolCalls(...) after the stream loop exits.
Headless CLI Collects ToolCallRequestInfo[], finalizes the assistant output, then runs processToolCallBatch(...).
Tool scheduler #2864 added safe parallel batching, but scheduling still starts after the tool-call batch is known.
Failure repair #4176 repairs tool_use ↔ tool_result invariant failures caused by mid-stream drops and out-of-band scheduler timing.

So the missing piece is not tool-call streaming itself. The missing piece is an in-band owner that can start tools during the response stream, coordinate cancellation/discard on stream failure, and preserve history correctness.

Proposed Design

1. Add a feature-gated streaming executor

Introduce a StreamingToolExecutor-like coordinator owned by the core turn loop, not only by the React layer.

Responsibilities:

2. Keep result submission buffered

Tool execution may start early, but tool results should not be submitted to the model until the matching model functionCall is present in history.

The invariant should remain:

model(functionCall...) -> user(functionResponse...)

This is non-negotiable for OpenAI/Anthropic-compatible providers and for resume/compression/repair behavior.

3. Start conservatively

Initial scope should only early-execute clearly safe tools:

  • read/list/glob/search
  • web fetch/search if classified safe
  • read-only shell commands only if already classified as read-only
  • possibly agent calls, but only after separate review

Do not early-execute in the first version:

  • edit/write tools
  • unknown tools
  • tools requiring confirmation unless the UI path is explicitly handled
  • batches containing structured_output until sibling suppression semantics are nailed down

4. Keep the old path as fallback

Add a flag such as:

  • experimental.streamingToolDispatch
  • QWEN_CODE_STREAMING_TOOL_DISPATCH=1

When disabled, the current post-stream scheduler path remains unchanged.

Important Constraints

This should not be implemented by simply calling scheduleToolCalls(...) immediately from the current ToolCallRequest handler.

Key risks:

  1. History validity

    functionResponse cannot be inserted before its matching model functionCall has been persisted.

  2. Retry/fallback safety

    If the stream retries after a tool started, the executor must cancel/discard or synthesize an error result without leaking orphan results into history.

  3. Structured output

    In --json-schema mode, structured_output can suppress sibling tools. Early execution must not run side-effecting sibling tools before the full batch semantics are known.

  4. React scheduler boundary

    fix(core,cli): close tool_use↔tool_result invariant across all failure paths #4176 shows that the out-of-band React scheduler makes invariant repair harder. The RFC should decide whether the new executor lives fully in core, or whether the React scheduler is folded into the same ownership model.

  5. Provider behavior

    Some providers may only emit usable tool calls near the final chunk. The feature should degrade gracefully and remain gated until provider-specific behavior is verified.

Rollout Plan

Phase Scope Risk
Phase 1 Add a uniform completed-tool-call signal to provider stream conversion. Existing callers still buffer, so no behavior change. Low
Phase 2 Add a streaming executor skeleton behind a disabled flag. It can accept calls and buffer results, but the old post-stream path remains default. Medium
Phase 3 Enable early execution for read-only/concurrency-safe tools only. Preserve buffered result submission and current history shape. Medium
Phase 4 Add retry/fallback/abort discard semantics and targeted tests for orphan prevention. Medium
Phase 5 Evaluate broader tool classes and sibling-error abort once telemetry shows the safe path is stable. Higher; defer

The feature should default off until telemetry and integration tests show no regression in tool_use ↔ tool_result invariant violations.

Alternatives Considered

A. Stay non-streaming and keep patching failure symptoms

#4176 is the correct defensive fix for the current architecture, but it does not reduce latency and it keeps the stream loop and tool scheduler split across layers.

B. Speculative dispatch on partial JSON deltas

Start a tool before its argument JSON is fully assembled. Rejected: the risk/value tradeoff is poor. The safer boundary is a complete tool-call payload, equivalent to Anthropic content_block_stop.

C. Stream-driven only for the Anthropic-compatible path

Tempting because that path already has a clear content_block_stop boundary. Rejected as a long-term default because maintaining divergent execution models by provider would keep the current cross-layer complexity.

Open Questions

  1. Executor ownership

    Should the new executor live entirely in core, or should the React scheduler be folded into the same ownership model?

  2. Cancellation semantics

    If the stream errors after a tool starts but before turn end, should the executor cancel in-flight tools, let safe tools finish and synthesize results, or choose per tool kind?

  3. Result order

    Should buffered functionResponse parts preserve dispatch order, model order, or completion order? The safest default seems model/dispatch order.

  4. Provider support

    Which providers actually emit complete tool calls before final stream termination? The feature should be measured per provider before default-on rollout.

  5. Structured output

    Should structured_output disable streaming dispatch for the entire batch, or only for sibling tools that are not provably read-only?

Acceptance Criteria

  • Feature flag/config controls the new path.
  • Existing post-stream tool scheduling remains the fallback.
  • Safe tools can begin execution before the model stream fully completes.
  • Tool results are not submitted to the model until matching function calls are present in history.
  • Retry/fallback/abort paths do not create orphan functionResponse entries.
  • Tests cover:
    • early execution of safe tools
    • unsafe tools preserving order
    • stream retry after early tool start
    • abort while early tool is running
    • structured output sibling suppression
    • headless and interactive paths
    • history invariant: every functionCall has exactly one matching functionResponse

References

Asking For

  • Maintainer feedback on whether this direction is desirable.
  • A decision on executor ownership, especially the React scheduler boundary.
  • Provider-specific observations: which backends emit complete tool calls early enough for this to matter?
  • Agreement on an initial safe-tools-only scope before any implementation PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions