Skip to content

feat(evals): add tool call formatter and integrate failure summaries - #28305

Merged
gundermanc merged 4 commits into
google-gemini:mainfrom
Gsoc26:feat/eval-failure-triage
Aug 12, 2026
Merged

gundermanc merged 4 commits into
google-gemini:mainfrom
Gsoc26:feat/eval-failure-triage

Conversation

@ved015

@ved015 ved015 commented Jul 7, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds tool-call timeline formatting and failure summary diagnostics to behavioral evaluations. When an eval fails, the test runner now automatically prints a compact, numbered timeline of the agent's tool calls (with arguments, status, and error details) directly inside the console failure message, eliminating the need to manually parse telemetry logs. Also introduces a comprehensive eval review checklist to serve as a contributor guide.

Details

  • scripts/utils/tool-log-formatter.ts — pure formatToolLogChain() utility that formats telemetry logs (TestRig.readToolLogs()) into a human-readable chain, truncating long arguments and nested objects for readability
  • scripts/tests/tool-log-formatter.test.ts — 12 unit tests verifying formatting, argument truncation, invalid JSON parsing, and numbering padding
  • evals/test-helper.ts — integrated the formatter into the internalEvalTest catch block to append the tool call chain to failing assertion errors
  • evals/test-helper.test.ts — added unit test verifying that assertion failures successfully output the tool log chain
  • evals/README.md — appended an Eval Review Checklist detailing acceptance criteria, local run expectations, assertion quality, and anti-patterns
  • 13 unit + integration tests for the formatter and test-helper harness

Pre-Merge Checklist

  • Updated relevant documentation and README (if needed)
  • Added/updated tests (if needed)
  • Noted breaking changes (if any)
  • Validated on required platforms/methods:
    • MacOS
      • npm run
      • npx
      • Docker
      • Podman
      • Seatbelt
    • Windows
      • npm run
      • npx
      • Docker
    • Linux
      • npm run
      • npx
      • Docker

Fixes: #28696

@ved015
ved015 requested review from a team as code owners July 7, 2026 22:15
@github-actions github-actions Bot added the size/l A large sized PR label Jul 7, 2026
@github-actions

github-actions Bot commented Jul 7, 2026 •

Copy link
Copy Markdown

📊 PR Size: size/L

  • Lines changed: 403
  • Additions: +402
  • Deletions: -1
  • Files changed: 5

@github-actions

github-actions Bot commented Jul 7, 2026 •

Copy link
Copy Markdown

🛑 Action Required: Evaluation Approval

Steering changes have been detected in this PR. To prevent regressions, a maintainer must approve the evaluation run before this PR can be merged.

Maintainers:

  1. Go to the Workflow Run Summary.
  2. Click the yellow 'Review deployments' button.
  3. Select the 'eval-gate' environment and click 'Approve'.

Once approved, the evaluation results will be posted here automatically.

Comment thread evals/README.md Outdated
Comment thread evals/README.md Outdated
Comment thread evals/README.md Outdated
Comment thread evals/README.md Outdated
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enhances the developer experience for behavioral evaluations by providing immediate, actionable diagnostics when tests fail. By automatically surfacing a formatted timeline of tool calls and their associated errors directly in the console, developers can debug agent behavior without manually parsing logs. Additionally, the inclusion of a structured review checklist standardizes eval quality and helps maintain consistency across the test suite.

Highlights

  • Tool Call Diagnostics: Implemented a new utility to format tool call telemetry into a human-readable chain, which is now automatically appended to console failure messages during behavioral evaluations.
  • Eval Review Checklist: Added a comprehensive contributor guide to the documentation, outlining acceptance criteria, local run expectations, and common anti-patterns for behavioral evals.
  • Testing: Added 13 new unit and integration tests to ensure the robustness of the tool log formatter and the error reporting integration.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-cli gemini-cli Bot added the status/need-issue Pull requests that need to have an associated issue. label Jul 7, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces an Eval Review Checklist to the documentation and enhances eval test failure reporting by appending the tool call chain to assertion error messages. A new utility, tool-log-formatter.ts, was added along with corresponding tests to format these tool logs. The review feedback highlights two critical issues: first, a potential crash in the log formatter if the tool arguments JSON parses to null or a non-object primitive; second, a potential TypeError when directly mutating error.message if the error object is frozen or read-only. Both issues should be addressed using the provided code suggestions to ensure robustness.

Comment thread scripts/utils/tool-log-formatter.ts
Comment thread evals/test-helper.ts
@gemini-cli

gemini-cli Bot commented Jul 15, 2026

Copy link
Copy Markdown
Contributor

Hi there! Thank you for your interest in contributing to Gemini CLI.

To ensure we maintain high code quality and focus on our prioritized roadmap, we only guarantee review and consideration of pull requests for issues that are explicitly labeled as 'help wanted'.

This PR will be closed in 7 days if it remains without that designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding.

Comment thread scripts/utils/tool-log-formatter.ts
@gemini-cli

gemini-cli Bot commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

This pull request is being closed as it has been open for 14 days without a 'help wanted' designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding.

@gemini-cli gemini-cli Bot closed this Jul 22, 2026
@gundermanc gundermanc reopened this Aug 5, 2026
@gemini-cli gemini-cli Bot added help wanted We will accept PRs from all issues marked as "help wanted". Thanks for your support! priority/p3 Backlog - a good idea but not currently a priority. area/core Issues related to User Interface, OS Support, Core Functionality and removed status/need-issue Pull requests that need to have an associated issue. labels Aug 5, 2026
auto-merge was automatically disabled August 12, 2026 05:28

Head branch was pushed to by a user without write access

@gundermanc
gundermanc enabled auto-merge August 12, 2026 05:38
@gundermanc
gundermanc added this pull request to the merge queue Aug 12, 2026
Merged via the queue into google-gemini:main with commit 4238b0b Aug 12, 2026
34 of 35 checks passed

This branch had an error being deployed

1 failed deployment
eval-gate — db4bd094 Deployed Aug 12, 2026 by ved015 via Evaluate Steering & Regressions #1767
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/core Issues related to User Interface, OS Support, Core Functionality help wanted We will accept PRs from all issues marked as "help wanted". Thanks for your support! priority/p3 Backlog - a good idea but not currently a priority. size/l A large sized PR status/pr-nudge-sent

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GSoC project issue

2 participants