Repository navigation
Robust component level evalutions #24353
Copy link
Copy link
Closed as not planned
Labels
aiq/eval_infraarea/agentIssues related to Core Agent, Tools, Memory, Sub-Agents, Hooks, Agent QualityIssues related to Core Agent, Tools, Memory, Sub-Agents, Hooks, Agent Qualitykind/customer-issueIssues that were reported by customersIssues that were reported by customerspriority/p1Important and should be addressed in the near term.Important and should be addressed in the near term.status/bot-triagedworkstream-rollupLabel used to tag epics and features that are associated with one of the three primary workstreamsLabel used to tag epics and features that are associated with one of the three primary workstreams🔒 maintainer only⛔ Do not contribute. Internal roadmap item.⛔ Do not contribute. Internal roadmap item.
Description
Activity
Metadata
Metadata
Assignees
Labels
aiq/eval_infraarea/agentIssues related to Core Agent, Tools, Memory, Sub-Agents, Hooks, Agent QualityIssues related to Core Agent, Tools, Memory, Sub-Agents, Hooks, Agent Qualitykind/customer-issueIssues that were reported by customersIssues that were reported by customerspriority/p1Important and should be addressed in the near term.Important and should be addressed in the near term.status/bot-triagedworkstream-rollupLabel used to tag epics and features that are associated with one of the three primary workstreamsLabel used to tag epics and features that are associated with one of the three primary workstreams🔒 maintainer only⛔ Do not contribute. Internal roadmap item.⛔ Do not contribute. Internal roadmap item.
Component Level Evaluations
This EPIC is a follow up for #15300 which introduced the concept of "behavioral evals" tests to the repository.
Since that issue's creation we have generate 76 behavioral eval tests, we run them for 6 supported Gemini models, and have created some basic infrastructure for naming, running, and tracking the runs over time.
This EPIC tracks:
Behavioral Evals vs. Component Level Evals
As part of this work the existing behavioral evals tests suite will be expanded and refactored and tagged into the following suites. Each suite will be tagged with a specific feature area for targeted validation on applicable PRs.
Behavioral -- a small number of medium sized evaluations of user facing bugs, misbehaviors, and feature-level prompted behaviors. Typically judged deterministically. e.g.: "was tool X called with the expected parameters".
Component evals -- A large number of relatively small evaluations that test a specific components performance at a specific tasks. e.g.: "compression prompt recall".
Hero tests -- A small number of relatively large evaluations that test the end to end GCLI's performance at specific high value task. e.g.: "does Gemini CLI successfully complete a long horizon task, like building a compiler from scratch over 100+ turns".