You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
v0.55.1: --model gemini-3.1-pro-preview silently served by a 2.5-series model under oauth-personal auth #28825
Requesting gemini-3.1-pro-preview with "Login with Google" credentials that lack 3.1 entitlement returns a successful response from a different model, with no error, no warning, and no indication in the CLI output that substitution occurred.
This is adjacent to #26938 (closed, v0.41.2, aux-role leakage) but distinct: this is entitlement-driven substitution of the primary model on the current release, and the substitute target is not stable.
Requested gemini-3.1-pro-preview, served gemini-2.5-flash — two generations down.
The substitute is not consistent
Same credentials, same model flag, two environments:
Environment
Requested
Actually served
Docker container (harbor agent runner)
gemini-3.1-pro-preview
gemini-2.5-pro
Host shell, GOOGLE_CLOUD_PROJECT set
gemini-3.1-pro-preview
gemini-2.5-flash
So it isn't a fixed documented fallback that a caller could account for — what you receive varies with context.
Things ruled out
Wrong model ID — gemini-3.1-pro-preview is current per the API docs, and the bundle contains 57 references to it, so the CLI knows the model.
Stale CLI — reproduced on 0.55.1 (@latest).
Missing project — reproduced with GOOGLE_CLOUD_PROJECT set.
Preview-features setting — {"previewFeatures": true} and {"general":{"previewFeatures":true}} both still yield 2.5-flash.
Why this matters beyond cost control
We hit this while running a benchmark comparison across three model families. The harness requested 3.1 Pro; every trial silently executed on a 2.5-series model and produced normal-looking scores. Had we not recorded the observed model per trial, we would have published results labelled "Gemini 3.1 Pro" that were measured on a different, older model.
Any published benchmark number produced through the CLI on personal credentials is exposed to this, and nothing in the CLI's own output would reveal it.
Requests
Error rather than substitute. If the requested model is unavailable to the credential, fail with a clear message.
Failing that, warn on stderr whenever the served model differs from the requested one.
Requesting
gemini-3.1-pro-previewwith "Login with Google" credentials that lack 3.1 entitlement returns a successful response from a different model, with no error, no warning, and no indication in the CLI output that substitution occurred.This is adjacent to #26938 (closed, v0.41.2, aux-role leakage) but distinct: this is entitlement-driven substitution of the primary model on the current release, and the substitute target is not stable.
Repro
Response arrives normally (
PONG). The session log records what actually served it:Requested
gemini-3.1-pro-preview, servedgemini-2.5-flash— two generations down.The substitute is not consistent
Same credentials, same model flag, two environments:
gemini-3.1-pro-previewgemini-2.5-proGOOGLE_CLOUD_PROJECTsetgemini-3.1-pro-previewgemini-2.5-flashSo it isn't a fixed documented fallback that a caller could account for — what you receive varies with context.
Things ruled out
gemini-3.1-pro-previewis current per the API docs, and the bundle contains 57 references to it, so the CLI knows the model.@latest).GOOGLE_CLOUD_PROJECTset.{"previewFeatures": true}and{"general":{"previewFeatures":true}}both still yield 2.5-flash.Why this matters beyond cost control
We hit this while running a benchmark comparison across three model families. The harness requested 3.1 Pro; every trial silently executed on a 2.5-series model and produced normal-looking scores. Had we not recorded the observed model per trial, we would have published results labelled "Gemini 3.1 Pro" that were measured on a different, older model.
Any published benchmark number produced through the CLI on personal credentials is exposed to this, and nothing in the CLI's own output would reveal it.
Requests
--strict-modelflag (proposed in [bug] --model pin silently leaks: hardcoded aux-role models + single-element last-resort fallback override user selection #26938) that exits non-zero rather than falling back. For automated and benchmarking use this is the difference between a usable tool and an unusable one.At minimum, the served model should be surfaced somewhere the caller can see without parsing
~/.gemini/tmpsession logs after the fact.Environment
@google/gemini-cli0.55.1, Node 22.11.0, Ubuntu 22.04 x86_64oauth-personal(Login with Google), account without 3.1 entitlement