Skip to content

Feat/evals tracker relationships error recovery - #28823

Closed
ved015 wants to merge 7 commits into
google-gemini:mainfrom
Gsoc26:feat/evals-tracker-relationships-error-recovery
Closed

ved015 wants to merge 7 commits into
google-gemini:mainfrom
Gsoc26:feat/evals-tracker-relationships-error-recovery

Conversation

@ved015

@ved015 ved015 commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds behavioral evaluations for task graph dependencies (tracker_add_dependency), task graph visualization (tracker_visualize), file path error recovery (re-searching and reading valid files upon 404), and shell command failure recovery (diagnosing execution errors and retrying).

Details

  • evals/tracker_relationships.eval.ts — [NEW] Behavioral evaluations for task tracker graph planning:
    • Add dependency: Asserts tracker_add_dependency is called to declare that a task depends on another when experimental.taskTracker is enabled.
    • Visualize graph: Asserts tracker_visualize is called to display a visual chart of tasks and dependencies.
  • evals/error_recovery.eval.ts — [NEW] Behavioral evaluation asserting that when read_file fails on an invalid path, the agent recovers by searching/globbing (glob / grep_search) to locate and inspect the correct file.
  • evals/shell_error_recovery.eval.ts — [NEW] Behavioral evaluation asserting that when a shell command fails due to execution or flag errors, the agent diagnoses the output and retries with a corrected command (run_shell_command).

Pre-Merge Checklist

  • Updated relevant documentation and README (if needed)
  • Added/updated tests (if needed)
  • Noted breaking changes (if any)
  • Validated on required platforms/methods:
    • MacOS
      • npm run
      • npx
      • Docker
      • Podman
      • Seatbelt
    • Windows
      • npm run
      • npx
      • Docker
    • Linux
      • npm run
      • npx
      • Docker

@ved015
ved015 requested review from a team as code owners August 15, 2026 12:40
@github-actions github-actions Bot added the size/xl An extra large PR label Aug 15, 2026
@github-actions

github-actions Bot commented Aug 15, 2026 •

Copy link
Copy Markdown

📊 PR Size: size/XL

  • Lines changed: 1439
  • Additions: +1422
  • Deletions: -17
  • Files changed: 15

@github-actions

github-actions Bot commented Aug 15, 2026 •

Copy link
Copy Markdown

🛑 Action Required: Evaluation Approval

Steering changes have been detected in this PR. To prevent regressions, a maintainer must approve the evaluation run before this PR can be merged.

Maintainers:

  1. Go to the Workflow Run Summary.
  2. Click the yellow 'Review deployments' button.
  3. Select the 'eval-gate' environment and click 'Approve'.

Once approved, the evaluation results will be posted here automatically.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly expands the behavioral evaluation suite to cover critical agent capabilities, including task tracker management, error recovery flows for file system and shell operations, and documentation retrieval. It also enhances the testing infrastructure to support more robust path resolution and test reporting across different environments, ensuring reliable evaluation results.

Highlights

  • New Behavioral Evaluations: Added a comprehensive suite of behavioral evaluations covering task tracker relationships, task visualization, file path error recovery, and shell command failure handling.
  • Infrastructure Enhancements: Updated test helper utilities, including AppRig and eval-report scripts, to improve cross-platform path resolution and test reporting accuracy.
  • Agent Capability Testing: Introduced new evaluations for skill activation, MCP resource management, and batch file reading to ensure consistent agent behavior.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a comprehensive suite of behavioral evaluation tests for various agent capabilities (such as skill activation, task completion, error recovery, MCP resources, and tracker queries) along with test utility enhancements in AppRig.tsx and path resolution improvements in eval-report.ts. The review feedback correctly identifies several critical issues where missing optional chaining on potentially undefined tool calls could lead to runtime TypeErrors. Additionally, the feedback points out that request.args is a JSON string and must be parsed using JSON.parse() before accessing its properties, and that part must be verified as an object before using the 'in' operator in AppRig.tsx.

Comment thread packages/cli/src/test-utils/AppRig.tsx
Comment thread evals/tracker_queries.eval.ts
Comment thread evals/tracker_queries.eval.ts Outdated
Comment thread evals/web_fetch.eval.ts Outdated
Comment thread evals/write_todos.eval.ts Outdated
Comment thread evals/activate_skill.eval.ts Outdated
@gemini-cli gemini-cli Bot added the status/need-issue Pull requests that need to have an associated issue. label Aug 15, 2026
@ved015
ved015 force-pushed the feat/evals-tracker-relationships-error-recovery branch from 1d72458 to 91f374c Compare August 15, 2026 13:19
@gemini-cli

gemini-cli Bot commented Aug 23, 2026

Copy link
Copy Markdown
Contributor

Hi there! Thank you for your interest in contributing to Gemini CLI.

To ensure we maintain high code quality and focus on our prioritized roadmap, we only guarantee review and consideration of pull requests for issues that are explicitly labeled as 'help wanted'.

This PR will be closed in 7 days if it remains without that designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding.

@gemini-cli

gemini-cli Bot commented Aug 30, 2026

Copy link
Copy Markdown
Contributor

This pull request is being closed as it has been open for 14 days without a 'help wanted' designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding.

@gemini-cli gemini-cli Bot closed this Aug 30, 2026
@wissamblue69-dotcom

Copy link
Copy Markdown

Summary

Adds behavioral evaluations for task graph dependencies (tracker_add_dependency), task graph visualization (tracker_visualize), file path error recovery (re-searching and reading valid files upon 404), and shell command failure recovery (diagnosing execution errors and retrying).

Details

  • evals/tracker_relationships.eval.ts — [NEW] Behavioral evaluations for task tracker graph planning:
    • Add dependency: Asserts tracker_add_dependency is called to declare that a task depends on another when experimental.taskTracker is enabled.
    • Visualize graph: Asserts tracker_visualize is called to display a visual chart of tasks and dependencies.
  • evals/error_recovery.eval.ts — [NEW] Behavioral evaluation asserting that when read_file fails on an invalid path, the agent recovers by searching/globbing (glob / grep_search) to locate and inspect the correct file.
  • evals/shell_error_recovery.eval.ts — [NEW] Behavioral evaluation asserting that when a shell command fails due to execution or flag errors, the agent diagnoses the output and retries with a corrected command (run_shell_command).

Pre-Merge Checklist

  • Updated relevant documentation and README (if needed)
  • Added/updated tests (if needed)
  • Noted breaking changes (if any)
  • Validated on required platforms/methods:
    • MacOS
      • npm run
      • npx
      • Docker
      • Podman
      • Seatbelt
    • Windows
      • npm run
      • npx
      • Docker
    • Linux
      • npm run
      • npx
      • Docker

This branch had an error being deployed

1 failed deployment
eval-gate — f628e512 Deployed Aug 15, 2026 by ved015 via Evaluate Steering & Regressions #1797
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/xl An extra large PR status/need-issue Pull requests that need to have an associated issue. status/pr-nudge-sent

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants