Repository navigation
feat(diagnostics): Native V8 Memory & Profiling Suite - #24303
Mustafa0216 wants to merge 1 commit into
Conversation
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces a comprehensive, native diagnostic suite for the Gemini CLI, designed to autonomously investigate memory leaks, execution bloat, and performance regressions. By utilizing V8's internal APIs and the Chrome DevTools Protocol, the suite provides a robust, zero-compilation alternative to traditional debugging tools, enabling the CLI to self-diagnose and generate actionable reports for LLM-driven remediation. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces a native diagnostic suite for the Gemini CLI, providing tools for V8 heap snapshot analysis, core dumps, and performance tracing. Feedback focuses on critical security and stability improvements: sanitizing class and branch names in reports to prevent prompt injection, refactoring snapshot logic to avoid memory leaks and Promise anti-patterns, implementing streaming for large heap files to prevent OOM crashes, and handling symbolic links to avoid infinite loops during directory traversal.
| topRealLeakers.forEach((leaker, index) => { | ||
| markdownTable += `| ${index + 1} | \`${leaker.className}\` | **${leaker.mb} MB** | ${leaker.bytes} | \n`; | ||
| }); |
There was a problem hiding this comment.
This script extracts class names directly from a V8 heap snapshot and embeds them into a Markdown report without any sanitization. Since this report is intended to be read and summarized by an AI agent, an attacker who can influence the application's memory can perform a prompt injection attack. Malicious strings in the heap could trick the AI into performing unintended actions or providing misleading summaries.
| topRealLeakers.forEach((leaker, index) => { | |
| markdownTable += `| ${index + 1} | \`${leaker.className}\` | **${leaker.mb} MB** | ${leaker.bytes} | \n`; | |
| }); | |
| topRealLeakers.forEach((leaker, index) => { | |
| const safeClassName = leaker.className.replace(/[\x60|]/g, "\\$&"); | |
| markdownTable += "| " + (index + 1) + " | " + safeClassName + " | **" + leaker.mb + " MB** | " + leaker.bytes + " | \n"; | |
| }); |
| const report = ` | ||
| ### Differential Bloat Analysis (Native File Traversal) | ||
| **Branch Target:** \`${gitOutput}\` | ||
|
|
||
| - **Compiled Bundles Output (./bundle/):** ${distMB} MB | ||
| - **Dependency Weights (./node_modules/):** ${depsMB} MB | ||
|
|
||
| *This replaces the need for native C++ \`bloaty\` footprint mapping by directly crawling the V8-compiled production payload footprints on your disk using OS bindings.* | ||
| `; |
There was a problem hiding this comment.
The script retrieves the current git branch name and embeds it directly into a Markdown report. Branch names are untrusted input that can be controlled by an attacker. Since this report is consumed by an AI agent, a branch name containing LLM instructions could lead to a prompt injection attack.
const safeGitOutput = gitOutput.replace(/[\x60|]/g, "\\$&");
const report = "\n### Differential Bloat Analysis (Native File Traversal)\n**Branch Target:** " + safeGitOutput + "\n\n- **Compiled Bundles Output (./bundle/):** " + distMB + " MB\n- **Dependency Weights (./node_modules/):** " + depsMB + " MB\n\n*This replaces the need for native C++ bloaty footprint mapping by directly crawling the V8-compiled production payload footprints on your disk using OS bindings.*\n";| async function takeSnapshot(client, filename) { | ||
| return new Promise(async (resolve, reject) => { | ||
| const stream = fs.createWriteStream(filename); | ||
|
|
||
| // Listen for the chunks of the snapshot | ||
| const chunkListener = client.HeapProfiler.addHeapSnapshotChunk(({ chunk }) => { | ||
| stream.write(chunk); | ||
| }); | ||
|
|
||
| try { | ||
| // takeHeapSnapshot Promise formally resolves ONLY after all chunks are sent | ||
| await client.HeapProfiler.takeHeapSnapshot({ reportProgress: false }); | ||
|
|
||
| // Explicitly flush and close the write stream | ||
| stream.end(() => { | ||
| // Cleanup the listener so we don't leak memory on consecutive snapshots | ||
| chunkListener(); // chrome-remote-interface returns an unsubscriber function | ||
| resolve(filename); | ||
| }); | ||
| } catch (err) { | ||
| stream.end(); | ||
| reject(err); | ||
| } | ||
| }); | ||
| } |
There was a problem hiding this comment.
The takeSnapshot function uses an async executor inside a new Promise, which is an anti-pattern. Errors thrown during the asynchronous execution of takeHeapSnapshot may not be properly captured by the promise's rejection logic. Furthermore, the chunkListener is not unsubscribed if an error occurs, which could lead to memory leaks or duplicate listeners in subsequent calls.
function takeSnapshot(client, filename) {
return new Promise((resolve, reject) => {
const stream = fs.createWriteStream(filename);
const chunkListener = client.HeapProfiler.addHeapSnapshotChunk(({ chunk }) => {
stream.write(chunk);
});
client.HeapProfiler.takeHeapSnapshot({ reportProgress: false })
.then(() => {
stream.end(() => {
chunkListener();
resolve(filename);
});
})
.catch((err) => {
chunkListener();
stream.end();
reject(err);
});
});
}| const rawData = fs.readFileSync(snapshotPath, 'utf8'); | ||
| snapshot = JSON.parse(rawData); |
There was a problem hiding this comment.
Reading a heap snapshot file entirely into a string using fs.readFileSync and then parsing it with JSON.parse is not scalable for massive snapshots. V8 has a hard limit on string length, and JSON.parse is extremely memory-intensive. For a tool designed to diagnose memory leaks, this implementation will likely crash with ERR_STRING_TOO_LONG or an OOM error when analyzing large heaps.
| const stats = fs.statSync(filePath); | ||
| if (stats.isDirectory()) { |
There was a problem hiding this comment.
Using fs.statSync follows symbolic links. In environments with circular symlinks or complex node_modules structures, this recursive function will enter an infinite loop and crash with a stack overflow. It is safer to use fs.lstatSync and explicitly skip symbolic links to ensure the tool remains robust.
| const stats = fs.statSync(filePath); | |
| if (stats.isDirectory()) { | |
| const stats = fs.lstatSync(filePath); | |
| if (stats.isSymbolicLink()) continue; | |
| if (stats.isDirectory()) { |
|
Thanks for the update |
Proposed Pull Request
Repository:
google/gemini-cliBranch:
feat/diagnostic-suitePR Title
feat(diagnostics): Native V8 Memory & Profiling SuitePR Body
GSoC 2026: Proposed PR for Terminal-Integrated Performance & Memory Investigation Companion #23365
This PR introduces a zero-dependency, pure-JavaScript diagnostic architecture designed to autonomously root-cause memory pipeline leaks and execution bloat within the Gemini CLI without relying on outdated node compatible extensions and very un optimal DAP protocols.
Key Additions:
analyze_memory.jswhich natively parses and computes massive V8 integer graph arrays to mathematically isolate Top 5 object memory leaks.9229to manually trigger physical garbage closures and snapshot bounds seamlessly.Usage & Agent Integration
Because the
SKILL.mddefinition is included directly in the/diagnosticsdirectory, this suite operates entirely natively. To extract memory limits, you must first expose the V8 WebSocket on the target CLI:You can then trigger the diagnostics in two distinct ways:
1. Traditional Execution:
Open a second terminal and trigger the full CI/CD orchestration script directly:
2. As a Native Gemini CLI Skill:
While chatting with the Gemini CLI prompt inside the initial
--inspectterminal, you can seamlessly command the AI to background the profiler directly on itself:Files to be Committed
Only the strict architecture files and dependencies will be tracked. All generated
.heapsnapshot,json, and.mdartifacts dumped during runtime will be intentionally ignored.package.json(adds chrome-remote-interface & source-map-explorer)package-lock.jsondiagnostics/3-snapshot.jsdiagnostics/analyze_memory.jsdiagnostics/gdb_batch.jsdiagnostics/orchestrate.jsdiagnostics/profile_bloat.jsdiagnostics/root_cause.jsdiagnostics/to_perfetto.jsdiagnostics/README.mddiagnostics/SKILL.mdUser Review Required
Please read the Title and Body above. If you approve of this exact PR format, click Approve and I will instantly run the
gitandghsequences to push the code and open the real Pull Request on GitHub!