Repository navigation
feat(evals): add tool call formatter and integrate failure summaries - #28305
Conversation
|
📊 PR Size: size/L
|
🛑 Action Required: Evaluation ApprovalSteering changes have been detected in this PR. To prevent regressions, a maintainer must approve the evaluation run before this PR can be merged. Maintainers:
Once approved, the evaluation results will be posted here automatically. |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request enhances the developer experience for behavioral evaluations by providing immediate, actionable diagnostics when tests fail. By automatically surfacing a formatted timeline of tool calls and their associated errors directly in the console, developers can debug agent behavior without manually parsing logs. Additionally, the inclusion of a structured review checklist standardizes eval quality and helps maintain consistency across the test suite. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces an Eval Review Checklist to the documentation and enhances eval test failure reporting by appending the tool call chain to assertion error messages. A new utility, tool-log-formatter.ts, was added along with corresponding tests to format these tool logs. The review feedback highlights two critical issues: first, a potential crash in the log formatter if the tool arguments JSON parses to null or a non-object primitive; second, a potential TypeError when directly mutating error.message if the error object is frozen or read-only. Both issues should be addressed using the provided code suggestions to ensure robustness.
|
Hi there! Thank you for your interest in contributing to Gemini CLI. To ensure we maintain high code quality and focus on our prioritized roadmap, we only guarantee review and consideration of pull requests for issues that are explicitly labeled as 'help wanted'. This PR will be closed in 7 days if it remains without that designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding. |
|
This pull request is being closed as it has been open for 14 days without a 'help wanted' designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding. |
Head branch was pushed to by a user without write access
4238b0b
Summary
Adds tool-call timeline formatting and failure summary diagnostics to behavioral evaluations. When an eval fails, the test runner now automatically prints a compact, numbered timeline of the agent's tool calls (with arguments, status, and error details) directly inside the console failure message, eliminating the need to manually parse telemetry logs. Also introduces a comprehensive eval review checklist to serve as a contributor guide.
Details
scripts/utils/tool-log-formatter.ts— pureformatToolLogChain()utility that formats telemetry logs (TestRig.readToolLogs()) into a human-readable chain, truncating long arguments and nested objects for readabilityscripts/tests/tool-log-formatter.test.ts— 12 unit tests verifying formatting, argument truncation, invalid JSON parsing, and numbering paddingevals/test-helper.ts— integrated the formatter into theinternalEvalTestcatch block to append the tool call chain to failing assertion errorsevals/test-helper.test.ts— added unit test verifying that assertion failures successfully output the tool log chainevals/README.md— appended anEval Review Checklistdetailing acceptance criteria, local run expectations, assertion quality, and anti-patternsPre-Merge Checklist
Fixes: #28696