Skip to content

fix(quota): surface the limit and reset window the server reports - #29429

Closed
sabhishek13-py wants to merge 2 commits into
google-gemini:mainfrom
sabhishek13-py:fix/quota-reset-visibility
Closed

sabhishek13-py wants to merge 2 commits into
google-gemini:mainfrom
sabhishek13-py:fix/quota-reset-visibility

Conversation

@sabhishek13-py

Copy link
Copy Markdown

Fixes #29425

Problem

When the Cloud Code API rejects a request with RESOURCE_EXHAUSTED, it says which limit was hit and when it clears, in ErrorInfo.metadata: quotaResetTimeStamp, quotaResetDelay and uiMessage. None of those were read anywhere outside test fixtures. The reset window came only from RetryInfo, so a rejection that arrives without RetryInfo told the user nothing about when access returns, and the server's own explanation ("Individual quota reached. Please upgrade your subscription to increase your limits.") was replaced by a generic Usage limit reached for all Pro models.

That is the gap in #29425. retrieveUserQuota reported 95.8–100% remaining for Flash, Pro and Flash Lite while every request was rejected with the same quotaResetTimeStamp, because the limit doing the rejecting is not one of the buckets that endpoint returns. The rejection was the only place that limit was visible, and the CLI discarded it.

This does not diagnose the server-side entitlement question in the issue — it makes the limit and its reset window visible to the user instead of leaving them with a quota display that contradicts what the server enforces.

Changes

  • googleQuotaErrors.ts: parse quotaResetTimeStamp, quotaResetDelay and uiMessage, and carry them on TerminalQuotaError / RetryableQuotaError as resetTime and isUserFacingMessage.
  • The reset instant is treated as a fact to report; a delay is treated as something to wait on. A delay is only derived from the timestamp when the window is within the existing 5-minute retryable bound, so a reset hours away is shown to the user and never slept on. RetryInfo still wins when present, so existing classification is unchanged.
  • useQuotaAndFallback.ts: render the absolute reset instant rather than now + delay, including the date when the reset is not today, and keep the server's explanation when it marked the message user-facing (its trailing countdown is dropped so two reset claims cannot disagree).
  • config.ts: report the most constrained bucket per model rather than whichever one came last. A model can have several buckets (one per tokenType) and the server enforces the smallest remainder. A spent bucket whose limit cannot be derived now stays in the display as exhausted instead of dropping the model from it entirely.
  • docs/resources/troubleshooting.md: entry for "Usage limit reached" while the usage bars still show remaining quota.

Testing

  • packages/core: googleQuotaErrors.test.ts (53 passing, 5 new covering the metadata paths: reset window without RetryInfo, a past timestamp deriving no delay, RetryInfo preferred over metadata, quotaResetDelay fallback, and a short timestamp that does become a delay); config.test.ts (240 passing, 2 new for the binding bucket and for keeping an exhausted model visible).
  • packages/cli: useQuotaAndFallback.test.ts (37 passing, 2 new for the reset instant with a date and for not echoing a message the server did not mark user-facing); ModelQuotaDisplay.test.tsx passing.
  • Also re-ran retry.test.ts, handler.test.ts, flashFallback.test.ts, sharedProjectThrottling.test.ts, googleErrors.test.ts, errors.test.ts — all passing. Lint and Prettier clean on every changed file.

When a request is rejected with RESOURCE_EXHAUSTED, the Cloud Code API
reports which limit was hit and when it clears, in ErrorInfo.metadata:
quotaResetTimeStamp, quotaResetDelay and uiMessage. None of those were
read. The reset window only ever came from RetryInfo, so a rejection
that arrives without RetryInfo told the user nothing about when access
returns, and the server's own explanation was replaced by a generic
"Usage limit reached" line.

That is the gap behind google-gemini#29425: retrieveUserQuota reported 95.8-100%
remaining for every model while every request was rejected, because the
limit doing the rejecting is not one of the buckets that endpoint
returns. The error was the only place that limit was visible, and the
CLI discarded it.

- Parse quotaResetTimeStamp, quotaResetDelay and uiMessage, and carry
  them on TerminalQuotaError and RetryableQuotaError.
- Treat the reset instant as a fact to report, and a delay as something
  to wait on: a delay is only derived from the timestamp when the window
  is short enough to sit through, so a reset hours away is shown, never
  slept on. RetryInfo still wins when present.
- Show the reset instant rather than now + delay, with the date when it
  is not today, and keep the server's explanation when it marked the
  message user-facing.
- Report the most constrained bucket per model instead of whichever one
  came last, since a model's buckets can disagree and the server
  enforces the smallest remainder. Keep a spent bucket in the display
  rather than dropping the model when its limit cannot be derived.
@sabhishek13-py
sabhishek13-py requested review from a team as code owners September 20, 2026 17:20
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request improves the accuracy and transparency of quota limit reporting in the Gemini CLI. By parsing and surfacing server-provided metadata from RESOURCE_EXHAUSTED errors, the CLI can now correctly inform users about specific subscription-wide caps or individual limits that were previously hidden. It also ensures that the most restrictive quota bucket is prioritized in the UI, aligning the displayed status with the actual server-side enforcement.

Highlights

  • Quota Error Metadata Parsing: Updated googleQuotaErrors.ts to parse quotaResetTimeStamp, quotaResetDelay, and uiMessage from API error metadata, surfacing these details in TerminalQuotaError and RetryableQuotaError.
  • Improved User Feedback: Updated useQuotaAndFallback.ts to display the absolute reset time reported by the server, ensuring users receive accurate information about when their access will be restored.
  • Constrained Bucket Reporting: Refined config.ts to track and report the most constrained quota bucket per model, preventing the CLI from incorrectly indicating available quota when a specific limit has been reached.
  • Documentation Update: Added a troubleshooting entry to docs/resources/troubleshooting.md explaining why users might see "Usage limit reached" errors even when usage bars show remaining quota.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions github-actions Bot added the size/l A large sized PR label Sep 20, 2026
@github-actions

github-actions Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

📊 PR Size: size/L

  • Lines changed: 607
  • Additions: +583
  • Deletions: -24
  • Files changed: 7

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request improves quota error handling and reporting in the Gemini CLI by extracting and propagating absolute reset timestamps, reset delays, and user-facing flags from Google API errors. The CLI UI now prefers absolute reset times, formats them clearly (including dates for multi-day resets), strips redundant server countdowns, and displays server-provided explanations when marked user-facing. Additionally, the quota tracking logic in Config is updated to track and report the most constrained quota bucket (e.g., tokens vs. requests) per model, preventing the CLI from showing remaining quota when a model is actually exhausted. Relevant unit tests and troubleshooting documentation have been added. I have no feedback to provide as there are no review comments.

Note: Security Review did not run due to the size of the PR.

@gemini-cli gemini-cli Bot added priority/p1 Important and should be addressed in the near term. area/enterprise Issues related to Telemetry, Policy, Quota / Licensing labels Sep 20, 2026
A real rejection reported on google-gemini#29425 phrases its countdown as "Resets in
9h22m20s." rather than "Your quota will reset after ...", so surfacing
the server's explanation left that countdown next to the reset instant
the CLI renders from the metadata. Two countdowns, and the one in the
message goes stale while the dialog is open.

Strip both phrasings, and add the reported rejection as a fixture so the
exact payload is covered: ErrorInfo metadata with quotaResetTimeStamp,
quotaResetDelay and uiMessage, alongside a RetryInfo that still supplies
the delay.
@gemini-cli

gemini-cli Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Hi there! Thank you for your interest in contributing to Gemini CLI.

To ensure we maintain high code quality and focus on our prioritized roadmap, we only guarantee review and consideration of pull requests for issues that are explicitly labeled as 'help wanted'.

This PR will be closed in 7 days if it remains without that designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding.

@gemini-cli

gemini-cli Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

This pull request is being closed as it has been open for 14 days without a 'help wanted' designation. We encourage you to find and contribute to existing 'help wanted' issues in our backlog! Thank you for your understanding.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/enterprise Issues related to Telemetry, Policy, Quota / Licensing priority/p1 Important and should be addressed in the near term. size/l A large sized PR status/pr-nudge-sent

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Code Assist Standard: QUOTA_EXHAUSTED despite 95.8% Flash / 100% Pro quota remaining

1 participant