Repository navigation
Conversation
…out-independently
…s behavioral evaluations
…ehavioral evaluations
|
📊 PR Size: size/XL
|
🛑 Action Required: Evaluation ApprovalSteering changes have been detected in this PR. To prevent regressions, a maintainer must approve the evaluation run before this PR can be merged. Maintainers:
Once approved, the evaluation results will be posted here automatically. |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly expands the behavioral evaluation suite to cover critical agent capabilities, including task tracker management, error recovery flows for file system and shell operations, and documentation retrieval. It also enhances the testing infrastructure to support more robust path resolution and test reporting across different environments, ensuring reliable evaluation results. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces a comprehensive suite of behavioral evaluation tests for various agent capabilities (such as skill activation, task completion, error recovery, MCP resources, and tracker queries) along with test utility enhancements in AppRig.tsx and path resolution improvements in eval-report.ts. The review feedback correctly identifies several critical issues where missing optional chaining on potentially undefined tool calls could lead to runtime TypeErrors. Additionally, the feedback points out that request.args is a JSON string and must be parsed using JSON.parse() before accessing its properties, and that part must be verified as an object before using the 'in' operator in AppRig.tsx.
…l error recovery behavioral evaluations
1d72458 to
91f374c
Compare
|
Hi there! Thank you for your interest in contributing to Gemini CLI. To ensure we maintain high code quality and focus on our prioritized roadmap, we only guarantee review and consideration of pull requests for issues that are explicitly labeled as 'help wanted'. This PR will be closed in 7 days if it remains without that designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding. |
|
This pull request is being closed as it has been open for 14 days without a 'help wanted' designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding. |
|
Summary
Adds behavioral evaluations for task graph dependencies (
tracker_add_dependency), task graph visualization (tracker_visualize), file path error recovery (re-searching and reading valid files upon 404), and shell command failure recovery (diagnosing execution errors and retrying).Details
evals/tracker_relationships.eval.ts— [NEW] Behavioral evaluations for task tracker graph planning:tracker_add_dependencyis called to declare that a task depends on another whenexperimental.taskTrackeris enabled.tracker_visualizeis called to display a visual chart of tasks and dependencies.evals/error_recovery.eval.ts— [NEW] Behavioral evaluation asserting that whenread_filefails on an invalid path, the agent recovers by searching/globbing (glob/grep_search) to locate and inspect the correct file.evals/shell_error_recovery.eval.ts— [NEW] Behavioral evaluation asserting that when a shell command fails due to execution or flag errors, the agent diagnoses the output and retries with a corrected command (run_shell_command).Pre-Merge Checklist