Skip to content

Background agent coordination gap: duplicate work, premature completion, and non-interactive send_message #8097

Description

@Aleks-0

What happened?

When running multiple background Explore subagents simultaneously and using send_message to communicate with them mid-flight, three related coordination failures occur:

1. Parent agent duplicates subagent work

After launching background Explore agents with agent({run_in_background: true}), the parent agent immediately starts doing its own grep_search / read_file / glob calls on the same topics — instead of waiting for results. This happens despite the instruction in the tool result: "Do not duplicate this agent's work... or briefly tell the user what you launched and end your response."

2. Premature completion: first subagent closing = session closing

When multiple background agents are running, the parent agent treats the first completion_notification as the signal that all agents are done. It then proceeds to compose a final summary, ignoring subsequent completions from the remaining agents. Data from later-completing agents is lost.

3. send_message is non-interactive for Explore agents

send_message is documented as a way to "continue a running, paused, or completed background task", but for Explore-type agents it behaves like a "message in a bottle" — the agent does not process incoming messages until it finishes its task. The parent agent sends send_message, receives "Message queued... the task will receive it at the next tool-round boundary", but no intermediate response comes back. All answers arrive bundled together inside the final completion_notification.

What did you expect to happen?

  1. After launching background agents, the parent agent should wait for completions instead of duplicating the same work with local tool calls.
  2. The parent agent should distinguish "1 of N completed" from "all N completed" and wait for all launched agents before composing a final answer.
  3. send_message should be genuinely interactive for any background agent, not just Monitor-type agents — an agent that receives a message during execution should respond to it before finishing its task.

Root cause analysis

All three failures share the same root: the parent agent has no "pause and wait for background agents" mechanism. Every signal (tool result, first completion, send_message acknowledgment) triggers immediate action rather than patience.

Duplicate work (1)

The "Do not duplicate" instruction is part of the tool's llmContent result string, mixed with technical metadata (task_id, output_file, read_file/tail hints). For the LLM, this is data, not a behavioral rule — and its weight competes with dozens of other tokens. The System Prompt has one line about duplication (prompts.ts:389), but it loses weight in long-context scenarios.

Relevant source locations:

  • packages/core/src/tools/agent/agent.ts:3647-3657 — tool result text with "Do not duplicate"
  • packages/core/src/core/prompts.ts:379-385 — System Prompt "Subagent Delegation" section
  • packages/core/src/tools/agent/agent.ts:903-928 — tool description usage notes

Premature completion (2)

The completion_notification XML from background-tasks.ts:1541-1610 arrives with <status>completed</status> — it contains no information about how many agents are still running. The parent agent has no state tracking "I launched N, only K have completed."

Relevant source locations:

  • packages/core/src/agents/background-tasks.ts:1541-1610 — completion notification format

send_message non-interactive (3)

The shouldWaitForExternalMessages predicate in agent.ts:3199 is wired to MonitorRegistry.hasRunningForOwner(). For Explore agents (the most commonly used background agent type), MonitorRegistry is empty, so the predicate always returns false. The reasoning loop in agent-core.ts:1075-1140 then skips the waitForExternalInputs block entirely and proceeds directly to finalization.

Relevant source locations:

  • packages/core/src/tools/agent/agent.ts:3191-3200 — setExternalMessageWaitPredicate wire-up
  • packages/core/src/agents/runtime/agent-core.ts:1075-1140 — reasoning loop with shouldWaitForExternalMessages check
  • packages/core/src/agents/background-tasks.ts:1363-1390 — queueMessage / message queue

Reproduction (items 2 and 3)

Three test runs on 2026-07-30:

Test Task duration send_message sent Result
Trivial task 7s 0 Clean completion
Medium task (4 parallel searches) 40s 1 mid-flight Answer bundled with completion
Long task (5 sequential searches) 72s 2 mid-flight Both answers bundled with completion

In all cases, send_message answers were delivered only inside the final completion_notification — no intermediate callback was generated.

Additional context

This is a design-level coordination gap, not a narrow code bug. Fixing it likely requires changes in three areas:

  1. Prompt engineering — separate "duplication prevention" into a dedicated prominent rule (not mixed with tool result data)
  2. Completion notification — include "K of N completed" metadata so the parent agent can reason about pending agents
  3. send_message interactivity — expand the shouldWaitForExternalMessages predicate scope, or introduce an explicit "pause for external input" mode for all agent types
中文

发生了什么?

当同时运行多个后台 Explore 子智能体并使用 send_message 与其进行中途通信时,会出现三个相关的协调失败:

1. 父智能体重复子智能体的工作

在使用 agent({run_in_background: true}) 启动后台 Explore 智能体后,父智能体立即开始在相同主题上执行自己的 grep_search / read_file / glob 调用 — 而不是等待结果。尽管工具结果中有提示:"不要重复此智能体的工作...或简要告诉用户你启动了什么并结束响应"。

2. 过早完成:第一个子智能体结束 = 整个会话结束

当多个后台智能体正在运行时,父智能体将第一个 completion_notification 视为所有智能体都已完成的信号。然后开始撰写最终摘要,忽略后续完成的智能体。后续完成智能体的数据会丢失。

3. send_message 对 Explore 智能体是非交互式的

send_message 被文档描述为一种 "继续运行中、已暂停或已完成的后台任务" 的方式,但对于 Explore 类型的智能体,它的行为类似于"瓶中信" — 智能体在完成任务之前不会处理收到的消息。父智能体发送 send_message,收到 "消息已排队...任务将在下一个工具轮次边界收到它",但没有中间响应返回。所有答案都捆绑在最终的 completion_notification 中一起到达。

期望的行为是什么?

  1. 启动后台智能体后,父智能体应该等待完成,而不是用本地工具调用重复相同的工作。
  2. 父智能体应该区分"N 个中已完成 1 个"和"所有 N 个已完成",并在撰写最终答案之前等待所有已启动的智能体。
  3. send_message 应该对任何后台智能体都真正是交互式的,而不仅仅是 Monitor 类型的智能体 — 在执行过程中收到消息的智能体应该在完成任务之前响应它。

根本原因分析

所有三个失败都有相同的根源:父智能体没有"暂停并等待后台智能体"的机制。 每个信号(工具结果、第一次完成、send_message 确认)都会触发立即行动,而不是耐心等待。

重复工作 (1)

"不要重复"指令是工具 llmContent 结果字符串的一部分,混杂着技术元数据(task_id、output_file、read_file/tail 提示)。对于 LLM 来说,这是数据,而不是行为规则 — 其权重与数十个其他 token 竞争。System Prompt 中有一行关于重复的内容(prompts.ts:389),但在长上下文场景中会失去权重。

相关源码位置:

  • packages/core/src/tools/agent/agent.ts:3647-3657 — 包含"不要重复"的工具结果文本
  • packages/core/src/core/prompts.ts:379-385 — System Prompt "子智能体委托" 部分
  • packages/core/src/tools/agent/agent.ts:903-928 — 工具描述使用说明

过早完成 (2)

来自 background-tasks.ts:1541-1610 的 completion_notification XML 以 <status>completed</status> 到达 — 它不包含关于还有多少智能体仍在运行的信息。父智能体没有状态跟踪"我启动了 N 个,只有 K 个已完成"。

相关源码位置:

  • packages/core/src/agents/background-tasks.ts:1541-1610 — 完成通知格式

send_message 非交互式 (3)

agent.ts:3199 中的 shouldWaitForExternalMessages 谓词绑定到 MonitorRegistry.hasRunningForOwner()。对于 Explore 智能体(最常用的后台智能体类型),MonitorRegistry 为空,因此谓词始终返回 false。agent-core.ts:1075-1140 中的推理循环随后完全跳过 waitForExternalInputs 块,直接进入最终化。

相关源码位置:

  • packages/core/src/tools/agent/agent.ts:3191-3200 — setExternalMessageWaitPredicate 连接
  • packages/core/src/agents/runtime/agent-core.ts:1075-1140 — 带 shouldWaitForExternalMessages 检查的推理循环
  • packages/core/src/agents/background-tasks.ts:1363-1390 — queueMessage / 消息队列

复现(第 2 和第 3 项)

2026-07-30 的三次测试运行:

测试 任务时长 发送的 send_message 结果
简单任务 7s 0 干净完成
中等任务(4 个并行搜索) 40s 1 个中途消息 答案与完成捆绑
长任务(5 个顺序搜索) 72s 2 个中途消息 两个答案都与完成捆绑

在所有情况下,send_message 的答案仅在最终的 completion_notification 中传递 — 没有生成中间回调。

附加上下文

这是一个设计层面的协调缺陷,而不是狭窄的代码错误。修复可能需要三个方面的更改:

  1. 提示工程 — 将"防止重复"分离为专门的突出规则(不与工具结果数据混合)
  2. 完成通知 — 包含"N 个中已完成 K 个"的元数据,使父智能体能够推理待处理的智能体
  3. send_message 交互性 — 扩展 shouldWaitForExternalMessages 谓词范围,或为所有智能体类型引入显式的"暂停等待外部输入"模式

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

category/coreCore engine and logicpriority/P2Medium - Moderately impactful, noticeable problemroadmap/multi-agentRoadmap: Multi-agent collaborationscope/coretype/bugSomething isn't working as expected

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions