You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Gemini CLI should make use of behavioral eval tests #15300
Gemini CLI tests are currently split into unit and integration tests, with unit tests covering individual modules, and integration tests validating the E2E CLI functionality.
There is also a de facto 3rd form of tests present, such as save_memory.test.ts, which are effectively behavioral evals.
What are behavioral evals
Behavioral evals are tests that accept a prompt and run the agent and then check for some desirable outcome. Typically the validations are simple, like ensuring that a tool or function call happened. These tests are valuable for use as:
A feedback loop for validating changes to prompts, tool descriptions, and other quality impacting factors.
A stake in the ground that prevents regressions to behaviors implemented through 'steering' the model as the prompt, model, and toolset change.
A tool for use in refining model steering behaviors, such as the automatic invocation of subagents in Support subagents in extensions #14846, whether or not to commit user changes, and which memories the agent should choose to store with the save_memory tool.
How are behavioral evals different than industry benchmarks
Behavioral evals tend to encompass simpler scenarios and validations, acting more akin to unit tests of the behaviors implied by the prompt rather than 'challenges' for the agent to solve.
Behavioral evals perform simpler validations and so should pass more consistently than typical benchmarks.
Behavioral evals should generally only be checked in a passing state, though being non-deterministic agent tests, they may not pass 100% of the time. Notably, a behavioral eval might pass for one specific model + prompt combination, but not another.
Benchmarks measure an agent's performance against a standardized set of challenges. Behavioral evals, however, validate specific behaviors that we wish to preserve and/or refine from build to build and across models, prompts, and configurations.
How are behavioral evals different than integration tests
Integration tests are expected to succeed consistently. A single integration test failure often blocks the completion of a PR or build.
Behavioral evals should be run multiple times each run to ensure that the agent exhibits the same behavior with reasonable consistency.
Behavioral evals will probably not run during the CI build but rather on a schedule or perhaps in response to changes in certain key files.
What are some items that may benefit from behavioral evals
[Agents] Post V1.0 Work #3132 -- behavioral evals can be used to ensure that system prompt steering as to when to use subagents works in several key scenarios.
What
Gemini CLI tests are currently split into unit and integration tests, with unit tests covering individual modules, and integration tests validating the E2E CLI functionality.
There is also a de facto 3rd form of tests present, such as save_memory.test.ts, which are effectively behavioral evals.
What are behavioral evals
Behavioral evals are tests that accept a prompt and run the agent and then check for some desirable outcome. Typically the validations are simple, like ensuring that a tool or function call happened. These tests are valuable for use as:
How are behavioral evals different than industry benchmarks
Behavioral evals tend to encompass simpler scenarios and validations, acting more akin to unit tests of the behaviors implied by the prompt rather than 'challenges' for the agent to solve.
Behavioral evals perform simpler validations and so should pass more consistently than typical benchmarks.
Behavioral evals should generally only be checked in a passing state, though being non-deterministic agent tests, they may not pass 100% of the time. Notably, a behavioral eval might pass for one specific model + prompt combination, but not another.
Benchmarks measure an agent's performance against a standardized set of challenges. Behavioral evals, however, validate specific behaviors that we wish to preserve and/or refine from build to build and across models, prompts, and configurations.
How are behavioral evals different than integration tests
What are some items that may benefit from behavioral evals
eslint --fixwhen available rather than attempting large fixes itself.Details