Repository navigation
fix(quota): surface the limit and reset window the server reports - #29429
sabhishek13-py wants to merge 2 commits into
Conversation
When a request is rejected with RESOURCE_EXHAUSTED, the Cloud Code API reports which limit was hit and when it clears, in ErrorInfo.metadata: quotaResetTimeStamp, quotaResetDelay and uiMessage. None of those were read. The reset window only ever came from RetryInfo, so a rejection that arrives without RetryInfo told the user nothing about when access returns, and the server's own explanation was replaced by a generic "Usage limit reached" line. That is the gap behind google-gemini#29425: retrieveUserQuota reported 95.8-100% remaining for every model while every request was rejected, because the limit doing the rejecting is not one of the buckets that endpoint returns. The error was the only place that limit was visible, and the CLI discarded it. - Parse quotaResetTimeStamp, quotaResetDelay and uiMessage, and carry them on TerminalQuotaError and RetryableQuotaError. - Treat the reset instant as a fact to report, and a delay as something to wait on: a delay is only derived from the timestamp when the window is short enough to sit through, so a reset hours away is shown, never slept on. RetryInfo still wins when present. - Show the reset instant rather than now + delay, with the date when it is not today, and keep the server's explanation when it marked the message user-facing. - Report the most constrained bucket per model instead of whichever one came last, since a model's buckets can disagree and the server enforces the smallest remainder. Keep a spent bucket in the display rather than dropping the model when its limit cannot be derived.
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request improves the accuracy and transparency of quota limit reporting in the Gemini CLI. By parsing and surfacing server-provided metadata from Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
📊 PR Size: size/L
|
There was a problem hiding this comment.
Code Review
This pull request improves quota error handling and reporting in the Gemini CLI by extracting and propagating absolute reset timestamps, reset delays, and user-facing flags from Google API errors. The CLI UI now prefers absolute reset times, formats them clearly (including dates for multi-day resets), strips redundant server countdowns, and displays server-provided explanations when marked user-facing. Additionally, the quota tracking logic in Config is updated to track and report the most constrained quota bucket (e.g., tokens vs. requests) per model, preventing the CLI from showing remaining quota when a model is actually exhausted. Relevant unit tests and troubleshooting documentation have been added. I have no feedback to provide as there are no review comments.
Note: Security Review did not run due to the size of the PR.
A real rejection reported on google-gemini#29425 phrases its countdown as "Resets in 9h22m20s." rather than "Your quota will reset after ...", so surfacing the server's explanation left that countdown next to the reset instant the CLI renders from the metadata. Two countdowns, and the one in the message goes stale while the dialog is open. Strip both phrasings, and add the reported rejection as a fixture so the exact payload is covered: ErrorInfo metadata with quotaResetTimeStamp, quotaResetDelay and uiMessage, alongside a RetryInfo that still supplies the delay.
|
Hi there! Thank you for your interest in contributing to Gemini CLI. To ensure we maintain high code quality and focus on our prioritized roadmap, we only guarantee review and consideration of pull requests for issues that are explicitly labeled as 'help wanted'. This PR will be closed in 7 days if it remains without that designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding. |
|
This pull request is being closed as it has been open for 14 days without a 'help wanted' designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding. |
Fixes #29425
Problem
When the Cloud Code API rejects a request with
RESOURCE_EXHAUSTED, it says which limit was hit and when it clears, inErrorInfo.metadata:quotaResetTimeStamp,quotaResetDelayanduiMessage. None of those were read anywhere outside test fixtures. The reset window came only fromRetryInfo, so a rejection that arrives withoutRetryInfotold the user nothing about when access returns, and the server's own explanation ("Individual quota reached. Please upgrade your subscription to increase your limits.") was replaced by a genericUsage limit reached for all Pro models.That is the gap in #29425.
retrieveUserQuotareported 95.8–100% remaining for Flash, Pro and Flash Lite while every request was rejected with the samequotaResetTimeStamp, because the limit doing the rejecting is not one of the buckets that endpoint returns. The rejection was the only place that limit was visible, and the CLI discarded it.This does not diagnose the server-side entitlement question in the issue — it makes the limit and its reset window visible to the user instead of leaving them with a quota display that contradicts what the server enforces.
Changes
googleQuotaErrors.ts: parsequotaResetTimeStamp,quotaResetDelayanduiMessage, and carry them onTerminalQuotaError/RetryableQuotaErrorasresetTimeandisUserFacingMessage.RetryInfostill wins when present, so existing classification is unchanged.useQuotaAndFallback.ts: render the absolute reset instant rather thannow + delay, including the date when the reset is not today, and keep the server's explanation when it marked the message user-facing (its trailing countdown is dropped so two reset claims cannot disagree).config.ts: report the most constrained bucket per model rather than whichever one came last. A model can have several buckets (one pertokenType) and the server enforces the smallest remainder. A spent bucket whose limit cannot be derived now stays in the display as exhausted instead of dropping the model from it entirely.docs/resources/troubleshooting.md: entry for "Usage limit reached" while the usage bars still show remaining quota.Testing
packages/core:googleQuotaErrors.test.ts(53 passing, 5 new covering the metadata paths: reset window withoutRetryInfo, a past timestamp deriving no delay,RetryInfopreferred over metadata,quotaResetDelayfallback, and a short timestamp that does become a delay);config.test.ts(240 passing, 2 new for the binding bucket and for keeping an exhausted model visible).packages/cli:useQuotaAndFallback.test.ts(37 passing, 2 new for the reset instant with a date and for not echoing a message the server did not mark user-facing);ModelQuotaDisplay.test.tsxpassing.retry.test.ts,handler.test.ts,flashFallback.test.ts,sharedProjectThrottling.test.ts,googleErrors.test.ts,errors.test.ts— all passing. Lint and Prettier clean on every changed file.