Skip to content

linearize-chat-compression-history-reconstruction #29511

Description

@harshitgupta31415

What would you like to be added?

Use linear array reconstruction in truncateHistoryToBudget in packages/core/src/context/chatCompressionService.ts.

The helper correctly visits messages and their parts from newest to oldest to prioritize recent tool responses. It currently restores chronological order by calling unshift for each processed part and each reconstructed message. For large histories, repeated front insertion adds avoidable array movement.

Append processed parts and messages with push, reverse each completed parts array once, and reverse the completed history once. Preserve the existing backwards traversal so token-budget decisions and truncation-file creation order remain unchanged.

Why is this needed?

For H messages with P_i parts per message, repeated front insertion can require O(H² + sum(P_i²)) array-reconstruction work. Appending and reversing once requires O(H + sum(P_i)) array work. Token estimation, serialization, and file I/O retain their existing costs.

This helper runs before history is split for compression, so the overhead applies to the whole curated history, including messages that are subsequently summarized.

Additional context

Observed in commit 2fe7c2d3f065dc40ad573d50b2091116f8a4aa18:

if (content.parts) {
// Process parts of the message backwards as well.
for (let j = content.parts.length - 1; j >= 0; j--) {
const part = content.parts[j];
if (part.functionResponse) {
const responseObj = part.functionResponse.response;
// Ensure we have a string representation to truncate.
// If the response is an object, we try to extract a primary string field (output or content).
let contentStr: string;
if (typeof responseObj === 'string') {
contentStr = responseObj;
} else if (responseObj && typeof responseObj === 'object') {
if (
'output' in responseObj &&
// eslint-disable-next-line no-restricted-syntax
typeof responseObj['output'] === 'string'
) {
contentStr = responseObj['output'];
} else if (
'content' in responseObj &&
// eslint-disable-next-line no-restricted-syntax
typeof responseObj['content'] === 'string'
) {
contentStr = responseObj['content'];
} else {
contentStr = JSON.stringify(responseObj, null, 2);
}
} else {
contentStr = JSON.stringify(responseObj, null, 2);
}
const tokens = estimateTokenCountSync([{ text: contentStr }]);
if (
functionResponseTokenCounter + tokens >
COMPRESSION_FUNCTION_RESPONSE_TOKEN_BUDGET
) {
try {
// Budget exceeded: Truncate this response.
const { outputFile } = await saveTruncatedToolOutput(
contentStr,
part.functionResponse.name ?? 'unknown_tool',
config.getNextCompressionTruncationId(),
config.storage.getProjectTempDir(),
);
const truncatedMessage = formatTruncatedToolOutput(
contentStr,
outputFile,
config.getTruncateToolOutputThreshold(),
);
newParts.unshift({
functionResponse: {
// eslint-disable-next-line @typescript-eslint/no-misused-spread
...part.functionResponse,
response: { output: truncatedMessage },
},
});
// Count the small truncated placeholder towards the budget.
functionResponseTokenCounter += estimateTokenCountSync([
{ text: truncatedMessage },
]);
} catch (error) {
// Fallback: if truncation fails, keep the original part to avoid data loss in the chat.
debugLogger.debug('Failed to truncate history to budget:', error);
newParts.unshift(part);
functionResponseTokenCounter += tokens;
}
} else {
// Within budget: keep the full response.
functionResponseTokenCounter += tokens;
newParts.unshift(part);
}
} else {
// Non-tool response part: always keep.
newParts.unshift(part);
}
}
}
// Reconstruct the message with processed (potentially truncated) parts.
truncatedHistory.unshift({ ...content, parts: newParts });
}
return truncatedHistory;

The proposed patch is limited to array reconstruction. Regression coverage should check message and part ordering, newest-first budget priority within a multi-part message, and preservation of the original part when saving truncated output fails.

The expected benefit is lower local reconstruction overhead for long histories; this does not imply a whole-CLI speedup or reduced model/network latency.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/agentIssues related to Core Agent, Tools, Memory, Sub-Agents, Hooks, Agent Qualitystatus/need-triageIssues that need to be triaged by the triage automation.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions