Repository navigation
feat(caretaker-evals): add triage evaluation framework and judge runner - #28530
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request establishes a comprehensive evaluation framework for the Caretaker Agent's issue triage pipeline. By leveraging parallel execution and automated LLM-based grading, the system enables rigorous benchmarking of triage performance against curated golden datasets, facilitating faster iteration and improved quality assurance for the agent's decision-making process. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
📊 PR Size: size/L
|
There was a problem hiding this comment.
Code Review
This pull request introduces an evaluation benchmark runner and judge suite for the triage agent under tools/caretaker-agent/evals/triage/. It includes modules for dataset streaming from Firestore, GitHub API utilities, summary report generation, Git worktree management, and an LLM-as-a-judge evaluation module. The review feedback highlights critical issues regarding concurrency and runtime safety that must be addressed: a TypeError when calling process_issue_triage with an unsupported keyword argument, race conditions from sharing environment variables across threads (requiring a switch to ProcessPoolExecutor), potential Git lockfile collisions during concurrent worktree operations, a potential AttributeError when handling None values in Firestore documents, and thread-safety concerns with a global API client.
|
ready for review! @adamfweidman @gundermanc |
Summary
Adds the core evaluation framework, LLM-as-a-Judge rubric, and parallel Git Worktree benchmark runner for the Caretaker Agent issue triage pipeline under
tools/caretaker-agent/evals/triage/.Details
runner.py): Executes evaluation runs across golden issue test cases usingProcessPoolExecutor(for process-level environment variable isolation) and isolated Git Worktrees (worktrees.py) for thread-safe parallel checkouts.judge.py,judge.md): Scores predicted Workable Specs against ground-truth Golden Specs using a 4-criterion 0–2 Rubric scale via Gemini API.dataset.py,summary.py,github_api.py): Streams test cases from Firestore with safenullbody handling, resolves target commit SHAs via GitHub REST API, and formats structured JSON / Markdown scorecards underresults/.Related
Requires PR #28524 to be merged first.
Pre-Merge Checklist