What would you like to be added?
Use linear array reconstruction in truncateHistoryToBudget in packages/core/src/context/chatCompressionService.ts.
The helper correctly visits messages and their parts from newest to oldest to prioritize recent tool responses. It currently restores chronological order by calling unshift for each processed part and each reconstructed message. For large histories, repeated front insertion adds avoidable array movement.
Append processed parts and messages with push, reverse each completed parts array once, and reverse the completed history once. Preserve the existing backwards traversal so token-budget decisions and truncation-file creation order remain unchanged.
Why is this needed?
For H messages with P_i parts per message, repeated front insertion can require O(H² + sum(P_i²)) array-reconstruction work. Appending and reversing once requires O(H + sum(P_i)) array work. Token estimation, serialization, and file I/O retain their existing costs.
This helper runs before history is split for compression, so the overhead applies to the whole curated history, including messages that are subsequently summarized.
Additional context
Observed in commit 2fe7c2d3f065dc40ad573d50b2091116f8a4aa18:
|
|
|
if (content.parts) { |
|
// Process parts of the message backwards as well. |
|
for (let j = content.parts.length - 1; j >= 0; j--) { |
|
const part = content.parts[j]; |
|
|
|
if (part.functionResponse) { |
|
const responseObj = part.functionResponse.response; |
|
// Ensure we have a string representation to truncate. |
|
// If the response is an object, we try to extract a primary string field (output or content). |
|
let contentStr: string; |
|
if (typeof responseObj === 'string') { |
|
contentStr = responseObj; |
|
} else if (responseObj && typeof responseObj === 'object') { |
|
if ( |
|
'output' in responseObj && |
|
// eslint-disable-next-line no-restricted-syntax |
|
typeof responseObj['output'] === 'string' |
|
) { |
|
contentStr = responseObj['output']; |
|
} else if ( |
|
'content' in responseObj && |
|
// eslint-disable-next-line no-restricted-syntax |
|
typeof responseObj['content'] === 'string' |
|
) { |
|
contentStr = responseObj['content']; |
|
} else { |
|
contentStr = JSON.stringify(responseObj, null, 2); |
|
} |
|
} else { |
|
contentStr = JSON.stringify(responseObj, null, 2); |
|
} |
|
|
|
const tokens = estimateTokenCountSync([{ text: contentStr }]); |
|
|
|
if ( |
|
functionResponseTokenCounter + tokens > |
|
COMPRESSION_FUNCTION_RESPONSE_TOKEN_BUDGET |
|
) { |
|
try { |
|
// Budget exceeded: Truncate this response. |
|
const { outputFile } = await saveTruncatedToolOutput( |
|
contentStr, |
|
part.functionResponse.name ?? 'unknown_tool', |
|
config.getNextCompressionTruncationId(), |
|
config.storage.getProjectTempDir(), |
|
); |
|
|
|
const truncatedMessage = formatTruncatedToolOutput( |
|
contentStr, |
|
outputFile, |
|
config.getTruncateToolOutputThreshold(), |
|
); |
|
|
|
newParts.unshift({ |
|
functionResponse: { |
|
// eslint-disable-next-line @typescript-eslint/no-misused-spread |
|
...part.functionResponse, |
|
response: { output: truncatedMessage }, |
|
}, |
|
}); |
|
|
|
// Count the small truncated placeholder towards the budget. |
|
functionResponseTokenCounter += estimateTokenCountSync([ |
|
{ text: truncatedMessage }, |
|
]); |
|
} catch (error) { |
|
// Fallback: if truncation fails, keep the original part to avoid data loss in the chat. |
|
debugLogger.debug('Failed to truncate history to budget:', error); |
|
newParts.unshift(part); |
|
functionResponseTokenCounter += tokens; |
|
} |
|
} else { |
|
// Within budget: keep the full response. |
|
functionResponseTokenCounter += tokens; |
|
newParts.unshift(part); |
|
} |
|
} else { |
|
// Non-tool response part: always keep. |
|
newParts.unshift(part); |
|
} |
|
} |
|
} |
|
|
|
// Reconstruct the message with processed (potentially truncated) parts. |
|
truncatedHistory.unshift({ ...content, parts: newParts }); |
|
} |
|
|
|
return truncatedHistory; |
The proposed patch is limited to array reconstruction. Regression coverage should check message and part ordering, newest-first budget priority within a multi-part message, and preservation of the original part when saving truncated output fails.
The expected benefit is lower local reconstruction overhead for long histories; this does not imply a whole-CLI speedup or reduced model/network latency.
What would you like to be added?
Use linear array reconstruction in
truncateHistoryToBudgetinpackages/core/src/context/chatCompressionService.ts.The helper correctly visits messages and their parts from newest to oldest to prioritize recent tool responses. It currently restores chronological order by calling
unshiftfor each processed part and each reconstructed message. For large histories, repeated front insertion adds avoidable array movement.Append processed parts and messages with
push, reverse each completed parts array once, and reverse the completed history once. Preserve the existing backwards traversal so token-budget decisions and truncation-file creation order remain unchanged.Why is this needed?
For H messages with P_i parts per message, repeated front insertion can require
O(H² + sum(P_i²))array-reconstruction work. Appending and reversing once requiresO(H + sum(P_i))array work. Token estimation, serialization, and file I/O retain their existing costs.This helper runs before history is split for compression, so the overhead applies to the whole curated history, including messages that are subsequently summarized.
Additional context
Observed in commit
2fe7c2d3f065dc40ad573d50b2091116f8a4aa18:gemini-cli/packages/core/src/context/chatCompressionService.ts
Lines 375 to 463 in 2fe7c2d
The proposed patch is limited to array reconstruction. Regression coverage should check message and part ordering, newest-first budget priority within a multi-part message, and preservation of the original part when saving truncated output fails.
The expected benefit is lower local reconstruction overhead for long histories; this does not imply a whole-CLI speedup or reduced model/network latency.