Skip to content

Feat/evals todos tasks tracker - #28822

Closed
ved015 wants to merge 5 commits into
google-gemini:mainfrom
Gsoc26:feat/evals-todos-tasks-tracker
Closed

ved015 wants to merge 5 commits into
google-gemini:mainfrom
Gsoc26:feat/evals-todos-tasks-tracker

Conversation

@ved015

@ved015 ved015 commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds behavioral evaluations for task planning (write_todos), task completion signaling (complete_task), and task tracker status querying (tracker_list_tasks and tracker_get_task).

Details

  • evals/write_todos.eval.ts — [NEW] Behavioral evaluation asserting that write_todos is invoked to structure TODO items when given a multi-step refactoring prompt.
  • evals/complete_task.eval.ts — [NEW] Behavioral evaluation asserting that complete_task is invoked to submit final findings upon completing a task flow.
  • evals/tracker_queries.eval.ts — [NEW] Behavioral evaluations for task tracker read queries:
    • List tasks: Asserts tracker_list_tasks is called to list active tasks when experimental.taskTracker is enabled.
    • Get task details: Asserts tracker_get_task is called with the target task ID to fetch specific task details.

Pre-Merge Checklist

  • Updated relevant documentation and README (if needed)
  • Added/updated tests (if needed)
  • Noted breaking changes (if any)
  • Validated on required platforms/methods:
    • MacOS
      • npm run
      • npx
      • Docker
      • Podman
      • Seatbelt
    • Windows
      • npm run
      • npx
      • Docker
    • Linux
      • npm run
      • npx
      • Docker

@ved015
ved015 requested review from a team as code owners August 15, 2026 12:31
@github-actions github-actions Bot added the size/xl An extra large PR label Aug 15, 2026
@github-actions

Copy link
Copy Markdown

📊 PR Size: size/XL

  • Lines changed: 1066
  • Additions: +1049
  • Deletions: -17
  • Files changed: 12

@github-actions

Copy link
Copy Markdown

🛑 Action Required: Evaluation Approval

Steering changes have been detected in this PR. To prevent regressions, a maintainer must approve the evaluation run before this PR can be merged.

Maintainers:

  1. Go to the Workflow Run Summary.
  2. Click the yellow 'Review deployments' button.
  3. Select the 'eval-gate' environment and click 'Approve'.

Once approved, the evaluation results will be posted here automatically.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces a robust suite of behavioral evaluations designed to verify agent tool usage and decision-making logic. By adding these tests, the codebase gains better coverage for critical workflows. Additionally, the PR upgrades the testing infrastructure with improved inspection helpers and a new path normalization utility, which resolves cross-platform path issues in evaluation reporting and ensures more reliable test results.

Highlights

  • New Behavioral Evaluations: Added comprehensive behavioral evaluation suites for core agent tools, including activate_skill, complete_task, get_internal_docs, mcp_resources, read_many_files, tracker_queries, web_fetch, and write_todos.
  • Test Infrastructure Upgrades: Enhanced AppRig with new helper methods like getToolCalls and getLastModelTextResponse to facilitate more robust verification of agent behavior during tests.
  • Path Normalization Utility: Implemented a new getRelativePath utility to handle cross-platform and CI-specific path inconsistencies, ensuring evaluation reports are consistent across different environments.
  • Report Summarization Improvements: Updated report summarization logic to filter out non-passed/failed assertion statuses and correctly match test policies across different absolute checkouts.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds several behavioral evaluation tests to verify agent capabilities and tool usage, enhances the AppRig test utility, and refactors the evaluation reporting script to use checkout-independent relative paths. The review feedback suggests improving Windows path resolution in getRelativePath by normalizing drive letter casing, adding a test case for casing mismatches, and using the CoreToolCallStatus.Success enum instead of a hardcoded string in AppRig.

Comment on lines +88 to +111
export function getRelativePath(filePath: string, rootDir: string): string {
const normalizedPath = filePath.replace(/\\/g, '/');
const root = path.resolve(rootDir).replace(/\\/g, '/');

// Handle Windows absolute paths with drive letters on POSIX systems
if (/^[a-zA-Z]:\//.test(normalizedPath)) {
if (normalizedPath.startsWith(root + '/')) {
return normalizedPath.slice(root.length + 1);
}
const match = normalizedPath.match(
/(?:^|\/)(evals|packages|scripts|integration-tests|memory-tests)\/(.+)$/,
);
if (match) {
return `${match[1]}/${match[2]}`;
}
return normalizedPath.split('/').pop() || '';
}

const absolute = path.resolve(rootDir, filePath).replace(/\\/g, '/');
if (absolute.startsWith(root + '/')) {
return absolute.slice(root.length + 1);
} else if (absolute === root) {
return '';
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Windows drive letters can have inconsistent casing (e.g., C:/ vs c:/) depending on how the path was resolved or retrieved. This casing mismatch can cause normalizedPath.startsWith(root + '/') to evaluate to false, failing to correctly resolve relative paths on Windows developer machines. Normalizing the drive letter casing to lowercase for both paths ensures robust cross-platform path resolution.

export function getRelativePath(filePath: string, rootDir: string): string {
  let normalizedPath = filePath.replace(/\\/g, '/');
  let root = path.resolve(rootDir).replace(/\\/g, '/');

  // Normalize Windows drive letter casing to avoid mismatch (e.g. C:/ vs c:/)
  if (/^[a-zA-Z]:/.test(normalizedPath)) {
    normalizedPath = normalizedPath[0].toLowerCase() + normalizedPath.slice(1);
  }
  if (/^[a-zA-Z]:/.test(root)) {
    root = root[0].toLowerCase() + root.slice(1);
  }

  // Handle Windows absolute paths with drive letters on POSIX systems
  if (/^[a-zA-Z]:\//.test(normalizedPath)) {
    if (normalizedPath.startsWith(root + '/')) {
      return normalizedPath.slice(root.length + 1);
    }
    const match = normalizedPath.match(
      /(?:^|\/)(evals|packages|scripts|integration-tests|memory-tests)\/(.+)$/,
    );
    if (match) {
      return `${match[1]}/${match[2]}`;
    }
    return normalizedPath.split('/').pop() || '';
  }

  let absolute = path.resolve(rootDir, filePath).replace(/\\/g, '/');
  if (/^[a-zA-Z]:/.test(absolute)) {
    absolute = absolute[0].toLowerCase() + absolute.slice(1);
  }
  if (absolute.startsWith(root + '/')) {
    return absolute.slice(root.length + 1);
  } else if (absolute === root) {
    return '';
  }

Comment on lines +101 to +115
describe('getRelativePath', () => {
it('correctly handles Windows drive letter paths on any platform', () => {
const winPath = 'C:\\coding\\gemini-cli\\evals\\my_test.eval.ts';
expect(getRelativePath(winPath, '/home/runner/work/repo')).toBe(
'evals/my_test.eval.ts',
);
});

it('extracts relative path when starts with rootDir', () => {
const fullPath = '/home/runner/work/repo/evals/my_test.eval.ts';
expect(getRelativePath(fullPath, '/home/runner/work/repo')).toBe(
'evals/my_test.eval.ts',
);
});
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Add a unit test to verify that getRelativePath correctly handles Windows drive letter casing mismatches (e.g., c:\... vs C:\...).

  describe('getRelativePath', () => {
    it('correctly handles Windows drive letter paths on any platform', () => {
      const winPath = 'C:\\coding\\gemini-cli\\evals\\my_test.eval.ts';
      expect(getRelativePath(winPath, '/home/runner/work/repo')).toBe(
        'evals/my_test.eval.ts',
      );
    });

    it('handles Windows drive letter casing mismatches gracefully', () => {
      const winPath = 'c:\\coding\\gemini-cli\\evals\\my_test.eval.ts';
      expect(getRelativePath(winPath, 'C:\\coding\\gemini-cli')).toBe(
        'evals/my_test.eval.ts',
      );
    });

    it('extracts relative path when starts with rootDir', () => {
      const fullPath = '/home/runner/work/repo/evals/my_test.eval.ts';
      expect(getRelativePath(fullPath, '/home/runner/work/repo')).toBe(
        'evals/my_test.eval.ts',
      );
    });
  });

Comment on lines +488 to +498
async waitForCompletedToolCall(toolName: string, timeout = 10000) {
await this.waitUntil(
() => this.toolCalls.some(
(call) => call.request.name === toolName && call.status === 'success',
),
{
timeout,
message: `Timed out waiting for tool call "${toolName}" to complete with success status`,
},
);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Avoid using the hardcoded string 'success' when checking the tool call status. Instead, use the imported CoreToolCallStatus.Success enum to maintain type safety and consistency with the rest of the file.

Suggested change
async waitForCompletedToolCall(toolName: string, timeout = 10000) {
await this.waitUntil(
() => this.toolCalls.some(
(call) => call.request.name === toolName && call.status === 'success',
),
{
timeout,
message: `Timed out waiting for tool call "${toolName}" to complete with success status`,
},
);
}
async waitForCompletedToolCall(toolName: string, timeout = 10000) {
await this.waitUntil(
() => this.toolCalls.some(
(call) => call.request.name === toolName && call.status === CoreToolCallStatus.Success,
),
{
timeout,
message: `Timed out waiting for tool call "${toolName}" to complete with success status`,
},
);
}

@gemini-cli

gemini-cli Bot commented Aug 23, 2026

Copy link
Copy Markdown
Contributor

Hi there! Thank you for your interest in contributing to Gemini CLI.

To ensure we maintain high code quality and focus on our prioritized roadmap, we only guarantee review and consideration of pull requests for issues that are explicitly labeled as 'help wanted'.

This PR will be closed in 7 days if it remains without that designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding.

@gemini-cli

gemini-cli Bot commented Aug 30, 2026

Copy link
Copy Markdown
Contributor

This pull request is being closed as it has been open for 14 days without a 'help wanted' designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding.

@gemini-cli gemini-cli Bot closed this Aug 30, 2026

This branch had an error being deployed

1 failed deployment
eval-gate — dfb0b2a5 Deployed Aug 15, 2026 by ved015 via Evaluate Steering & Regressions #1792
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/xl An extra large PR status/need-issue Pull requests that need to have an associated issue. status/pr-nudge-sent

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant