Skip to content

Gemini CLI should make use of behavioral eval tests #15300

Description

@gundermanc

What

Gemini CLI tests are currently split into unit and integration tests, with unit tests covering individual modules, and integration tests validating the E2E CLI functionality.

There is also a de facto 3rd form of tests present, such as save_memory.test.ts, which are effectively behavioral evals.

What are behavioral evals

Behavioral evals are tests that accept a prompt and run the agent and then check for some desirable outcome. Typically the validations are simple, like ensuring that a tool or function call happened. These tests are valuable for use as:

  • A feedback loop for validating changes to prompts, tool descriptions, and other quality impacting factors.
  • A stake in the ground that prevents regressions to behaviors implemented through 'steering' the model as the prompt, model, and toolset change.
  • A tool for use in refining model steering behaviors, such as the automatic invocation of subagents in Support subagents in extensions #14846, whether or not to commit user changes, and which memories the agent should choose to store with the save_memory tool.

How are behavioral evals different than industry benchmarks

  • Behavioral evals tend to encompass simpler scenarios and validations, acting more akin to unit tests of the behaviors implied by the prompt rather than 'challenges' for the agent to solve.

  • Behavioral evals perform simpler validations and so should pass more consistently than typical benchmarks.

  • Behavioral evals should generally only be checked in a passing state, though being non-deterministic agent tests, they may not pass 100% of the time. Notably, a behavioral eval might pass for one specific model + prompt combination, but not another.

  • Benchmarks measure an agent's performance against a standardized set of challenges. Behavioral evals, however, validate specific behaviors that we wish to preserve and/or refine from build to build and across models, prompts, and configurations.

How are behavioral evals different than integration tests

  • Integration tests are expected to succeed consistently. A single integration test failure often blocks the completion of a PR or build.
  • Behavioral evals should be run multiple times each run to ensure that the agent exhibits the same behavior with reasonable consistency.
  • Behavioral evals will probably not run during the CI build but rather on a schedule or perhaps in response to changes in certain key files.

What are some items that may benefit from behavioral evals

Details

  • Should use vitest
  • Should repeat at least a few times.
  • Should not be a required part of the build. Perhaps runs asynchronously, as needed (based on directory changed), or manually and user invoked.
  • Should emit logs in a machine readable format that is used to monitor the trend of how many tests pass all iterations over time.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area/agentIssues related to Core Agent, Tools, Memory, Sub-Agents, Hooks, Agent Qualityworkstream-rollupLabel used to tag epics and features that are associated with one of the three primary workstreams🔒 maintainer only⛔ Do not contribute. Internal roadmap item.

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions