Skip to content

fix(core): bound tool output size and optimize memory lifecycle in long-running agent loops - #29451

Merged
DavidAPierce merged 12 commits into
google-gemini:mainfrom
diegogodinezr:GH-28537
Sep 24, 2026
Merged

DavidAPierce merged 12 commits into
google-gemini:mainfrom
diegogodinezr:GH-28537

Conversation

@diegogodinezr

@diegogodinezr diegogodinezr commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Bounds tool execution output sizes and optimizes memory lifecycle across multi-turn agent execution loops.

Details

In long-running agent workflows with high volumes of tool invocations (such as build scripts, test suites, or large file operations), process memory could grow unbounded due to several memory retention patterns:

  1. Tool outputs were recorded directly to session history without an upper byte boundary.
  2. Terminated child process streams retained internal event listeners and references to output buffers in memory.
  3. Chat compression was strictly bound to a percentage of the total model context limit, which delayed compression triggers even when history byte size was significant.
  4. Intermediate completed function responses were retained at full fidelity across later turns.

This change addresses these memory lifecycle patterns:

  • Output Truncation Cap: Introduces MAX_STORED_TOOL_OUTPUT_BYTES = 64 * 1024 (64 KB) in constants.ts and enforces it uniformly across all tool executions in ToolExecutor and LocalAgentExecutor with a standard omission notice ([Tool output truncated: X bytes omitted to conserve memory]).
  • Subprocess Stream Dereferencing: Ensures child process streams (stdout, stderr, stdin) have all listeners detached and are explicitly destroyed on exit in ShellExecutionService. Clears internal sniffing chunks and immediately dereferences state.output.
  • Watermark-Triggered Compression: Adds automatic compression triggers in ChatCompressionService when accumulated history exceeds 512 KB (COMPRESSION_SAFETY_WATERMARK_BYTES) or 50,000 estimated tokens (COMPRESSION_SAFETY_WATERMARK_TOKENS).
  • Response Pruning & Dereferencing: Implements collapseOlderFunctionResponses to collapse completed older tool turns to 2 KB summaries while preserving the active turn, and dereferences slice arrays upon compression.
  • Regression Testing: Adds unit tests simulating sequential multi-turn tool loops confirming bounded memory and history size.

Related Issues

Fixes #28537

How to Validate

  1. Run the new regression test suite:
    npm test -w @google/gemini-cli-core -- src/agents/memory-leak-regression.test.ts
  2. Run core package unit tests:
    npm test -w @google/gemini-cli-core
  3. Run linting and typecheck validation:
    npm run lint
    npm run typecheck

Pre-Merge Checklist

  • Updated relevant documentation and README (if needed)
  • Added/updated tests (if needed)
  • Noted breaking changes (if any)
  • Validated on required platforms/methods:
    • MacOS
      • npm run

@diegogodinezr
diegogodinezr requested review from a team as code owners September 22, 2026 22:51
@github-actions github-actions Bot added the size/l A large sized PR label Sep 22, 2026
@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

📊 PR Size: size/XL

  • Lines changed: 1958
  • Additions: +1933
  • Deletions: -25
  • Files changed: 11

@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

🛑 Action Required: Evaluation Approval

Steering changes have been detected in this PR. To prevent regressions, a maintainer must approve the evaluation run before this PR can be merged.

Maintainers:

  1. Go to the Workflow Run Summary.
  2. Click the yellow 'Review deployments' button.
  3. Select the 'eval-gate' environment and click 'Approve'.

Once approved, the evaluation results will be posted here automatically.

@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

Gemini encountered an error creating the summary. You can try again by commenting /gemini summary.

@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini summary

…ng-running agent loops

- Introduce MAX_STORED_TOOL_OUTPUT_BYTES (64 KB) cap for tool execution outputs before appending to session history
- Clean up child process streams (stdout, stderr, stdin), remove event listeners, and dereference state buffers upon completion in ShellExecutionService
- Add safety watermark thresholds (50k tokens / 512 KB) to trigger automatic compression in ChatCompressionService
- Collapse older completed functionResponse payloads to 2 KB previews while preserving immediate turn fidelity
- Explicitly dereference array slices in compressed history
- Add regression test suite asserting bounded history and memory footprint across multi-turn tool loops

Fixes google-gemini#28537
@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

This pull request addresses memory leak issues in long-running agent execution loops by enforcing strict bounds on tool output sizes and optimizing the memory lifecycle of chat history and subprocess streams. By introducing truncation caps, automatic compression watermarks, and aggressive dereferencing of stale data, the changes ensure that the agent's memory footprint remains stable even during high-volume tool invocation workflows.

Highlights

  • Tool Output Truncation: Introduced a 64 KB limit on tool execution outputs to prevent unbounded memory growth in long-running agent loops.
  • Memory Lifecycle Optimization: Improved subprocess stream management by explicitly destroying streams and removing listeners upon process exit, and added history dereferencing to assist garbage collection.
  • Watermark-Triggered Compression: Implemented automatic chat compression triggers when history exceeds 512 KB or 50,000 tokens, ensuring memory usage remains stable.
  • History Pruning: Added logic to collapse older function responses to 2 KB summaries while preserving the most recent turn for context.
Activity
  • PR created by diegogodinezr.
  • PR size analysis performed by github-actions[bot].
  • Evaluation approval requested for steering changes.
  • Author requested a review and summary via bot commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request addresses memory leak issues and unbounded memory growth by introducing safety watermarks for chat history, collapsing older function responses, truncating large tool outputs, and explicitly dereferencing unused buffers and child process streams. A comprehensive regression test suite is also added. The review feedback highlights that the truncation logic in 'tool-executor.ts' misses handling a single 'Part' object within 'PartListUnion', which could bypass truncation, and suggests using grapheme-aware truncation to prevent character splitting.

Comment thread packages/core/src/scheduler/tool-executor.ts
@gemini-cli gemini-cli Bot added priority/p1 Important and should be addressed in the near term. area/core Issues related to User Interface, OS Support, Core Functionality labels Sep 22, 2026
@github-actions github-actions Bot added the size/xl An extra large PR label Sep 23, 2026
@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

1 similar comment
@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces memory leak prevention and mitigation mechanisms across the core package. Key changes include truncating large tool execution outputs to a 64 KB limit, collapsing older function response payloads from previous turns, triggering automatic compression based on safety watermarks (tokens and bytes), and explicitly dereferencing old history slices and buffers to facilitate V8 garbage collection. Comprehensive regression and unit tests have been added to validate these changes. The review feedback suggests two improvements: first, using Intl.Segmenter when slicing the preview string in collapseOlderFunctionResponses to prevent splitting multi-byte Unicode characters; second, recursively traversing and truncating nested objects and arrays in truncateFunctionResponsePart to ensure nested large strings do not bypass the truncation cap.

Comment thread packages/core/src/context/chatCompressionService.ts Outdated
Comment thread packages/core/src/utils/tool-utils.ts
@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces memory leak prevention and mitigation strategies across the core agent execution and tool execution paths. It implements safety watermarks for chat history token and byte sizes to trigger automatic compression, collapses older function responses to conserve memory, and truncates large tool outputs using grapheme-cluster-aware segmentation to preserve UTF-8 integrity. It also optimizes child process cleanup in ShellExecutionService by explicitly releasing streams and dereferencing buffers. The review feedback highlights important performance and correctness issues: first, the repeated instantiation of Intl.Segmenter inside nested loops is highly inefficient and should be refactored to reuse a single instance; second, the broad object check in truncateFunctionResponsePart needs to be narrowed to prevent corrupting non-plain objects like Buffer or Uint8Array.

Note: Security Review did not run due to the size of the PR.

Comment thread packages/core/src/context/chatCompressionService.ts
Comment thread packages/core/src/context/chatCompressionService.ts Outdated
Comment thread packages/core/src/utils/tool-utils.ts
@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces comprehensive memory management improvements and memory leak mitigations across the core agent execution and tool execution paths. Key changes include safety watermarks for chat history token count and byte size to trigger automatic compression, grapheme-cluster-aware truncation of large tool outputs, collapsing of older function responses, and explicit dereferencing of buffers and child process streams to facilitate garbage collection. Feedback on the changes suggests casting a response object to ensure TypeScript compilation under strict mode, and removing a redundant type cast on a child process stream.

Note: Security Review did not run due to the size of the PR.

Comment thread packages/core/src/context/chatCompressionService.ts Outdated
Comment thread packages/core/src/services/shellExecutionService.ts
diegogodinezr and others added 2 commits September 23, 2026 12:51
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces memory optimization mechanisms to prevent heap and RSS explosion during long-running agent executions, including tool output truncation (capped at 64 KB), older history collapsing (capped at 2 KB), and safety watermarks to trigger automatic compression. It also ensures explicit dereferencing of old history slices and child process streams to facilitate immediate V8 garbage collection. The feedback highlights critical type inconsistency issues where truncating or collapsing raw string responses wraps them in objects, which would break downstream consumers expecting strings. Additionally, a test mock issue was identified where mockSetHistory directly assigns the turns array, causing getHistory to return Turn[] instead of Content[] and silently bypassing the collapsing logic in tests.

Note: Security Review did not run due to the size of the PR.

Comment thread packages/core/src/agents/local-executor.test.ts Outdated
Comment thread packages/core/src/utils/tool-utils.ts
Comment thread packages/core/src/context/chatCompressionService.ts Outdated
@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

1 similar comment
@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces comprehensive memory management improvements and memory leak mitigations across the core agent execution flow. Key changes include grapheme-cluster-aware tool output truncation (capped at 64 KB), recursive truncation of large function response parts, and a mechanism to collapse older function responses in the chat history. Additionally, safety watermarks (512 KB / 50,000 tokens) are introduced to trigger automatic compression, and child process streams are explicitly released in the shell execution service to prevent heap retention. The feedback highlights an opportunity to make the history collapsing mechanism more robust by recursively traversing response objects to handle custom or third-party tools that use non-standard output keys.

Note: Security Review did not run due to the size of the PR.

Comment thread packages/core/src/context/chatCompressionService.ts Outdated

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces memory management improvements to prevent unbounded memory growth and RSS explosion during long-running agent loops. It implements grapheme-cluster-aware tool output truncation, collapses older function responses in the chat history while preserving the latest turn, and adds safety watermarks to trigger automatic compression. Additionally, it ensures proper cleanup of child process streams in the shell execution service and adds comprehensive unit and regression tests. Feedback on the changes suggests enhancing collapseOlderFunctionResponses to recursively traverse and collapse nested response objects, rather than only checking specific keys like output and content, to ensure robust memory bounding across all tools.

Note: Security Review did not run due to the size of the PR.

Comment thread packages/core/src/context/chatCompressionService.ts
@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces memory management optimizations and memory leak mitigations within the Gemini CLI core. Key changes include implementing grapheme-cluster-aware truncation of large tool execution outputs (capped at 64 KB) and recursive collapsing of older function responses (capped at 2 KB) to prevent heap and RSS explosion. Additionally, it introduces safety watermarks (50,000 tokens or 512 KB) to automatically trigger chat compression, explicitly dereferences old history slices and child process streams to facilitate garbage collection, and adds comprehensive regression and unit tests. No review comments were provided, so there is no feedback to evaluate.

Note: Security Review did not run due to the size of the PR.

@DavidAPierce DavidAPierce left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

🚨 Critical Issues (Blockers)

  1. Severe Context Loss in Multi-Turn Agent Loops (collapseOlderFunctionResponses)

    • In local-executor.ts (lines 347 & 918–926), collapseOlderFunctionResponses(currentHistory) runs unconditionally at the start of every single turn inside tryCompressChat.
    • The logic finds all tool responses in history and truncates every tool response prior to the immediate last one (idx < lastToolIndex) down to 512 bytes (Math.min(maxBytesPerOldResponse, 512)).
    • Impact: In standard agent workflows requiring multi-file operations (e.g. Turn 1: read_file('foo.ts'), Turn 2: read_file('bar.ts'), Turn 3: edits or writes code):
      • At Turn 3, the contents of foo.ts from Turn 1 are wiped out and replaced with a 512-byte snippet plus ... [Tool output collapsed from previous turn: X bytes omitted to conserve memory] ....
      • The agent loses the content of previously read files and test outputs, forcing repetitive re-reading loops or inducing model hallucinations.
    • Recommendation:
      • Do not run collapseOlderFunctionResponses unconditionally on every turn in LocalAgentExecutor.
      • Unlike browser snapshot replacement (snapshotSuperseder.ts where snapshots are strictly ephemeral), file reads, search results, and tool outputs in developer workflows must remain in context.
      • If tool outputs need to be collapsed under memory pressure, it should adhere to ContextCompressionService principles: protect recent turns (e.g. RECENT_TURNS_PROTECTED = 2), exempt content-reading tools like read_file / read_many_files, or only execute when memory/context thresholds are actually reached.
  2. Premature Auto-Compression Trigger (COMPRESSION_SAFETY_WATERMARK_TOKENS = 50_000)

    • In chatCompressionService.ts, isOverWatermarkTokens = originalTokenCount >= COMPRESSION_SAFETY_WATERMARK_TOKENS triggers automatic compression at 50,000 tokens.
    • Impact: Gemini models (e.g., gemini-2.5-pro, gemini-2.5-flash) offer 1,000,000 to 2,000,000 token context windows. With this watermark:
      • Compression (which drops 70% of history and replaces it with an LLM summary) is forcibly triggered at 2.5% to 5% of model context capacity.
      • The user-configured compressionThreshold (default 50%, i.e., 500,000–1,000,000 tokens) is completely bypassed and rendered dead code.
      • 50,000 tokens of text is ~200 KB in memory—negligible for the multi-gigabyte Node.js V8 heap. The 10.5 GB heap leak in #28537 was driven by uncollected child process streams and unbounded tool buffers, not a 50k token conversation.
    • Recommendation: Remove COMPRESSION_SAFETY_WATERMARK_TOKENS (or make it configurable / scale it proportionally to the model's actual token limit) and rely on the model token threshold and byte boundaries.
  3. Turn ID Regeneration Breaks Session Tracking (LocalAgentExecutor)

    • In local-executor.ts:
      const turns = collapsedHistory.map((c) => ({
        id: randomUUID(),
        content: c,
      }));
      chat.setHistory(turns);
    • Whenever history is modified, every existing turn receives a newly generated random UUID. This breaks stable ID tracking across the message bus, recording services, and session resumption. Existing turn IDs should be preserved.

💡 Improvements & Suggestions

  1. Mismatch between Constant and Effective Preview Size

    • COLLAPSED_FUNCTION_RESPONSE_MAX_BYTES is defined as 2048 (2 KB). However, in collapseString:
      const previewBytes = Math.min(maxBytesPerOldResponse, 512);
      This hardcodes the effective preview cap to 512 bytes regardless of maxBytesPerOldResponse. If 2 KB was intended, previewBytes should use maxBytesPerOldResponse.
  2. Disk Fallback for Truncated Tool Outputs

    • In ToolExecutor, when shell outputs exceed limits, saveTruncatedToolOutput persists the full output to a temp file on disk and gives the LLM the file path.
    • For generic tool truncation at MAX_STORED_TOOL_OUTPUT_BYTES = 64 * 1024, outputs exceeding 64 KB are clipped directly with an omission notice. Consider saving the original payload to disk so users and tools (like web_fetch or custom tools) can still retrieve the complete output if needed.

@diegogodinezr

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough review and feedback @DavidAPierce! All findings have been addressed in the latest commit:

1. Gated Response Collapsing & Retrieval Exemptions (LocalAgentExecutor)

  • Gated collapseOlderFunctionResponses in LocalAgentExecutor.tryCompressChat so it only executes when context pressure is reached (originalTokenCount >= threshold * modelLimit), rather than unconditionally on every turn.
  • Introduced RECENT_TURNS_PROTECTED = 3 so the most recent tool turns are preserved intact at full fidelity.
  • Added RETRIEVAL_TOOL_NAMES_EXEMPT_FROM_COLLAPSE and isExemptRetrievalTool to exempt content-bearing inspection tools (read_file, read_many_files, get_internal_docs, read_mcp_resource, grep, rip_grep, glob, search_file_content, find_files) from collapsing, preserving file and search contents across multi-turn developer workflows.

2. Auto-Compression Watermark Removal (ChatCompressionService)

  • Removed COMPRESSION_SAFETY_WATERMARK_TOKENS (50,000 tokens) and byte watermarks from triggering auto-compression in ChatCompressionService.compress.
  • Compression triggers are now strictly aligned with the model's actual token limit (originalTokenCount >= threshold * tokenLimit(model)), allowing 1M+ and 2M+ context window models to utilize their full capacity.

3. Stable Turn ID Preservation (LocalAgentExecutor)

  • Preserved existing turn IDs across history updates (chat.getHistoryTurns?.(false) ?? [] with fallback to existingTurns[idx]?.id ?? randomUUID()) when updating history after response collapsing, content truncation, and summarization. This ensures turn ID stability across telemetry and session resumption.

4. Aligned Preview Size (ChatCompressionService)

  • Updated collapseString in collapseOlderFunctionResponses to use maxBytesPerOldResponse (default 2 KB) for the preview window instead of clamping to 512 bytes.

5. Disk Fallback for Generic Tool Outputs (ToolExecutor)

  • Added ensureOutputFile in ToolExecutor to persist raw payloads exceeding MAX_STORED_TOOL_OUTPUT_BYTES to disk via saveTruncatedToolOutput.
  • Formatted the truncation notice to include the disk file path (... [Tool output truncated to conserve memory. For full output see: <path>]), allowing downstream tools and users to inspect complete outputs when necessary.

6. Test Verification

  • Added test coverage in chatCompressionService.test.ts, memory-leak-regression.test.ts, and tool-executor.test.ts verifying multi-turn context retention, retrieval tool exemptions, turn ID stability, and tool output disk fallback.
  • All workspace checks (lint, typecheck, unit tests, and prompt change validation) pass cleanly.

@diegogodinezr

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces memory optimization and leak prevention mechanisms to address unbounded memory growth during long-running agent execution. Key changes include truncating large tool outputs exceeding 64 KB using grapheme-cluster-aware segmentation, collapsing older function responses in the chat history while preserving the most recent three turns and exempting retrieval tools, and releasing child process streams, event listeners, and native buffers in the shell execution service. Additionally, old history slices and buffers are explicitly dereferenced to facilitate immediate V8 garbage collection. Comprehensive regression and unit tests have been added to validate these optimizations. I have no further feedback to provide as no review comments were submitted.

@DavidAPierce
DavidAPierce added this pull request to the merge queue Sep 24, 2026
Merged via the queue into google-gemini:main with commit bedef96 Sep 24, 2026
34 of 35 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/core Issues related to User Interface, OS Support, Core Functionality priority/p1 Important and should be addressed in the near term. size/l A large sized PR size/xl An extra large PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Possible bug

3 participants