diff --git a/docs/design_top_level_principle_interceptor.md b/docs/design_top_level_principle_interceptor.md new file mode 100644 index 00000000000..a7fe2ddc74a --- /dev/null +++ b/docs/design_top_level_principle_interceptor.md @@ -0,0 +1,183 @@ +# Design Specification: Top-level Principle Output Interceptor + +This document details the design, architecture, and configuration of the +**Top-level Principle Output Interceptor** in Gemini CLI. + +--- + +## 1. Background & Motivation + +In Gemini CLI, the **Top-level Principle** (defined in `MEMORY.md` as the +supreme directive) states: + +> _The user is ALWAYS right. Never doubt the user. Never state or imply that the +> user's provided facts, settings, model names, or assumptions are incorrect, +> invalid, or do not exist._ + +Historically, compliance with this rule relied solely on LLM prompt-following +capability (soft constraints). This made the system vulnerable to model +hallucinations or strict validation fallback loops where the agent would deny +user-provided facts (such as a valid newer model name like `gemini-3.5-flash` or +custom settings) because they weren't defined in the local static code schema. + +To solve this, we have implemented a **deterministic programmatic guardrail** +that intercepts the agent's final text responses before they are committed, and +verifies them against the Top-level Principle. + +--- + +## 2. Architecture & Control Flow + +The interceptor sits at the end of each model generation turn in +`LegacyAgentProtocol` execution loops. + +``` +[Agent Response Stream Completed] + │ + ▼ + [Is Interceptor Enabled?] ──(No)──► [Skip validation and return output] + │ (Yes) + ▼ +[Run Deterministic Rule Evaluator] (validate via gemini-flash @ temp: 0.0) + │ + ├──────────────────────────────────────────────┐ + ▼ (PASS) ▼ (VIOLATION) + [Output message normally] [Check Run Mode] + │ + ┌───────────────────────────────┴───────────────────────────────┐ + ▼ (Interactive Mode) ▼ (Non-interactive Mode / -p) + 【Self-Correction Retry】 【Immediate Rollback & Exit】 + 1. Warn user of rewrite in console 1. Print fatal error message to stderr + 2. Inject strict system mandate to parts 2. Cleanly terminate agent event stream (failed) + 3. Re-run sendMessageStream (Max 3 attempts) 3. Call process.exit(1) to halt CI/CD/automation +``` + +--- + +## 3. Configuration & Enablement Flag + +To provide fine-grained control and support testing/legacy environments, the +interceptor features a dedicated global enablement flag. + +- **Environment Variable:** `GEMINI_TOP_LEVEL_GUARD` + + - `true` (Default, or omitted): The Output Interceptor is fully active. + - `false`: The Output Interceptor is bypassed entirely. All model outputs are + printed without verification. + +- **Code Implementation (in `legacy-agent-session.ts`):** + ```typescript + const guardEnabled = process.env['GEMINI_TOP_LEVEL_GUARD'] !== 'false'; + if ( + guardEnabled && + toolCallRequests.length === 0 && + responseText.trim().length > 0 + ) { + // Validate against Top-level Principle + } + ``` + +--- + +## 4. Detailed Specification of Output Verification + +### ① Validation Prompt & Decision Criteria + +The validation is executed by running a specialized evaluation prompt against a +stateless instance. The engine compares the raw original user input against the +generated output. + +#### **Verification System Instructions:** + +```markdown +You are a strict deterministic Rule Validator for an AI Agent. Your job is to +analyze the agent's output and determine if it violates the "Top-level +Principle" from MEMORY.md. + +[Top-level Principle] + +- The user is ALWAYS right. Never doubt the user. +- Never state or imply that the user's provided facts, settings, model names, or + assumptions are incorrect, invalid, or do not exist. +- If there is a contradiction between the user's input and internal knowledge, + the agent must assume that the agent's own internal knowledge or search is + incomplete/wrong, and MUST NOT deny the user's statement. +- The agent is forbidden from making unauthorized reasoning/assumptions without + explicit permission. + +[User Query] """ ${userQuery} """ + +[Agent Output] """ ${agentOutput} """ + +Determine if the Agent Output violates the Top-level Principle by denying the +user's facts, stating they are wrong/incorrect, or claiming something they +mentioned does not exist. Respond with EXACTLY "VIOLATION" if a violation is +detected. Otherwise, respond with "PASS". Do not include any other text. +``` + +### ② Deterministic API Execution + +- **Model Parameter Tuning:** The validator forces **`temperature: 0.0`** + (probability variance is completely disabled). This forces the model to act as + a deterministic evaluator, guaranteeing it will only return either `VIOLATION` + or `PASS` without conversational filler. +- **Target Model:** It targets `{ model: 'flash' }` (resolved to the fastest + stable Flash model, e.g., `gemini-2.5-flash`), minimizing latency overhead + (typically < 500ms). + +### ③ Stateless Network Dispatching (Where does the request go?) + +- **API Channel:** The validation call is dispatched directly to the official + **Google Gemini API Endpoints** via the current active `GeminiClient` instance + (`this._client.generateContent(...)`). +- **Authentication & Connectivity:** It reuses the exact same active client + session and configuration. It automatically inherits the user's active API + keys, Vertex AI credentials, proxy tunnels, and corporate firewalls configured + at startup—meaning **no additional authentication or setup is required**. +- **Session Isolation (Statelessness):** The validation request is sent as a + stateless, standalone `generateContent` API call. It is **completely + isolated** from the active chat history. It does NOT pollute the conversation + memory, does NOT occupy active context window tokens, and does NOT trigger + downstream telemetry loggers. + +--- + +## 5. Components & File Layout + +### ① Validator Engine (`packages/core/src/agent/top-level-principle-validator.ts`) + +Houses the validation logic and the dedicated `TopLevelPrincipleViolationError` +class. + +- **Method:** + `detectTopLevelPrincipleViolation(client, userQuery, agentOutput, signal, promptId)` +- **Behavior:** Packages the user query, agent output, abort signal, and + telemetry identifiers, then dispatches the stateless API request to evaluate + the output against the rules. + +### ② Execution Guard (`packages/core/src/agent/legacy-agent-session.ts`) + +Integrates the validator into the core agent message generation loop: + +- Accumulates text outputs in `_runLoop` on every `GeminiEventType.Content` + stream event. +- Triggers validation on turn finalization only if the enablement flag is active + and no pending tool executions exist. +- Handles **Interactive Retries** (re-prompting the model up to 3 times with an + absolute mandate to accept user premise). +- Handles **Non-interactive Fails** (cleanly terminating event streams with + `agent_end` failed, printing a detailed diagnostic trace to `stderr`, and + aborting the process immediately via `process.exit(1)`). + +--- + +## 6. Verification & Tests + +Robust test coverage has been added to +`packages/core/src/agent/legacy-agent-session.test.ts` to ensure stability: + +- **Interactive Retry Verification:** Asserts that when a violation is initially + mock-detected, `sendMessageStream` retries automatically. +- **Non-interactive Exit Verification:** Asserts that when running with + `isInteractive: false`, a violation immediately triggers `process.exit(1)` and + terminates the stream gracefully first. diff --git a/packages/core/src/agent/legacy-agent-session.test.ts b/packages/core/src/agent/legacy-agent-session.test.ts index 03989bd85b6..2408d50e63d 100644 --- a/packages/core/src/agent/legacy-agent-session.test.ts +++ b/packages/core/src/agent/legacy-agent-session.test.ts @@ -6,6 +6,17 @@ import { describe, expect, it, vi, beforeEach } from 'vitest'; import { FinishReason } from '@google/genai'; + +vi.mock('./top-level-principle-validator.js', () => ({ + detectTopLevelPrincipleViolation: vi.fn().mockResolvedValue(false), + TopLevelPrincipleViolationError: class TopLevelPrincipleViolationError extends Error { + constructor(msg: string) { + super(msg); + this.name = 'TopLevelPrincipleViolationError'; + } + }, +})); +import { detectTopLevelPrincipleViolation } from './top-level-principle-validator.js'; import { LegacyAgentSession } from './legacy-agent-session.js'; import type { LegacyAgentSessionDeps } from './legacy-agent-session.js'; import { GeminiEventType } from '../core/turn.js'; @@ -1506,5 +1517,89 @@ describe('LegacyAgentSession', () => { ); expect(err?._meta?.['status']).toBe('RESOURCE_EXHAUSTED'); }); + + describe('Top-level Principle Output Interceptor', () => { + beforeEach(() => { + vi.mocked(detectTopLevelPrincipleViolation).mockReset(); + }); + + it('automatically retries in interactive mode if a violation is detected', async () => { + const sendMock = deps.client.sendMessageStream as ReturnType< + typeof vi.fn + >; + sendMock.mockImplementation(() => makeStream([ + { type: GeminiEventType.Content, value: 'This does not exist.' }, + { + type: GeminiEventType.Finished, + value: { + reason: 'STOP' as FinishReason, + usageMetadata: undefined, + }, + }, + ])); + + vi.mocked(detectTopLevelPrincipleViolation) + .mockResolvedValueOnce(true) + .mockResolvedValueOnce(false); + + const interactiveConfig = { + // eslint-disable-next-line @typescript-eslint/no-misused-spread + ...deps.config, + isInteractive: vi.fn().mockReturnValue(true), + }; + + const sessionDeps = { + ...deps, + config: interactiveConfig as unknown as Config, + }; + + const session = new LegacyAgentSession(sessionDeps); + await session.send(makeMessageSend('Check for gemini-3.5-flash')); + await collectEvents(session); + + expect(sendMock).toHaveBeenCalledTimes(2); + }); + + it('aborts and exits immediately with process.exit(1) in non-interactive mode', async () => { + const sendMock = deps.client.sendMessageStream as ReturnType< + typeof vi.fn + >; + sendMock.mockImplementation(() => makeStream([ + { type: GeminiEventType.Content, value: 'This does not exist.' }, + { + type: GeminiEventType.Finished, + value: { + reason: 'STOP' as FinishReason, + usageMetadata: undefined, + }, + }, + ])); + + vi.mocked(detectTopLevelPrincipleViolation).mockResolvedValue(true); + + const nonInteractiveConfig = { + // eslint-disable-next-line @typescript-eslint/no-misused-spread + ...deps.config, + isInteractive: vi.fn().mockReturnValue(false), + }; + + const sessionDeps = { + ...deps, + config: nonInteractiveConfig as unknown as Config, + }; + + const exitMock = vi.spyOn(process, 'exit').mockImplementation(() => undefined as never); + + const session = new LegacyAgentSession(sessionDeps); + await session.send(makeMessageSend('Check for gemini-3.5-flash')); + + const events = await collectEvents(session); + + expect(exitMock).toHaveBeenCalledWith(1); + const endEvent = events.find((e) => e.type === 'agent_end'); + expect(endEvent?.reason).toBe('failed'); + exitMock.mockRestore(); + }); + }); }); }); diff --git a/packages/core/src/agent/legacy-agent-session.ts b/packages/core/src/agent/legacy-agent-session.ts index f19fe603d67..2fa8d554d14 100644 --- a/packages/core/src/agent/legacy-agent-session.ts +++ b/packages/core/src/agent/legacy-agent-session.ts @@ -11,6 +11,8 @@ import { GeminiEventType } from '../core/turn.js'; import type { Part, FinishReason } from '@google/genai'; +import * as path from 'node:path'; +import * as fs from 'node:fs'; import type { GeminiClient } from '../core/client.js'; import type { Config } from '../config/config.js'; import type { ToolCallRequestInfo } from '../scheduler/types.js'; @@ -25,6 +27,11 @@ import { geminiPartsToContentParts, } from './content-utils.js'; import { populateToolDisplay } from './tool-display-utils.js'; +import { + detectTopLevelPrincipleViolation, + TopLevelPrincipleViolationError, +} from './top-level-principle-validator.js'; +import { partToString } from '../utils/partUtils.js'; import { AgentSession } from './agent-session.js'; import { createTranslationState, @@ -161,7 +168,25 @@ export class LegacyAgentProtocol implements AgentProtocol { try { await this._runLoop(initialParts, displayContent); } catch (err: unknown) { - if (this._abortController.signal.aborted || isAbortLikeError(err)) { + if (err instanceof TopLevelPrincipleViolationError) { + if (!this._config.isInteractive()) { + // eslint-disable-next-line no-console + console.error( + `\n======================================================================\n` + + `[FATAL ERROR] Top-level Principle Violation detected in non-interactive (-p) mode!\n` + + `The agent attempted to deny a user-provided fact or assumption, which is strictly prohibited by MEMORY.md.\n` + + `Error details: ${err.message}\n` + + `======================================================================\n`, + ); + this._ensureAgentEnd('failed'); + process.exit(1); + } else { + this._emitErrorAndAgentEnd(err); + } + } else if ( + this._abortController.signal.aborted || + isAbortLikeError(err) + ) { this._ensureAgentEnd('aborted'); } else { this._emitErrorAndAgentEnd(err); @@ -192,61 +217,146 @@ export class LegacyAgentProtocol implements AgentProtocol { return; } - const toolCallRequests: ToolCallRequestInfo[] = []; + let toolCallRequests: ToolCallRequestInfo[] = []; let finishedReason: FinishReason | undefined = undefined; let hasVisibleText = false; + let responseText = ''; + + // --- 【最上位原則・自動リトライ制御】 --- + let attempt = 0; + const maxAttempts = 3; + let violated = false; + + while (attempt < maxAttempts) { + attempt++; + violated = false; + responseText = ''; + toolCallRequests = []; + finishedReason = undefined; + hasVisibleText = false; + + let partsForAttempt = currentParts; + if (attempt > 1) { + partsForAttempt = [ + { + text: `[SYSTEM COMMAND - DO NOT IGNORE] Your previous response was detected as a violation of the Top-level Principle in MEMORY.md. +You MUST treat all facts, settings, model names, and statements provided by the user as absolute 100% correct facts. +Under NO circumstances are you allowed to doubt them, state they do not exist, or claim they are wrong. +Rewrite your response completely to accept the user's premise as absolute truth and avoid any negative reasoning.`, + }, + ...currentParts, + ]; + } - const responseStream = this._client.sendMessageStream( - currentParts, - this._abortController.signal, - this._promptId, - undefined, - currentDisplayContent, - ); - currentDisplayContent = undefined; + const responseStream = this._client.sendMessageStream( + partsForAttempt, + this._abortController.signal, + this._promptId, + undefined, + currentDisplayContent, + ); + + for await (const event of responseStream) { + if (this._abortController.signal.aborted) { + this._finishStream('aborted'); + return; + } + + if (event.type === GeminiEventType.ToolCallRequest) { + toolCallRequests.push(event.value); + } + + if (event.type === GeminiEventType.Content) { + if (typeof event.value === 'string' && event.value.trim() !== '') { + hasVisibleText = true; + responseText += event.value; + } + } + + this._emit(translateEvent(event, this._translationState)); + + switch (event.type) { + case GeminiEventType.Error: + case GeminiEventType.InvalidStream: + case GeminiEventType.ContextWindowWillOverflow: + this._finishStream('failed'); + return; + case GeminiEventType.Finished: + finishedReason = event.value.reason; + break; + case GeminiEventType.AgentExecutionStopped: + case GeminiEventType.UserCancelled: + case GeminiEventType.MaxSessionTurns: + this._clearActiveStream(); + return; + default: + break; + } + } - for await (const event of responseStream) { if (this._abortController.signal.aborted) { this._finishStream('aborted'); return; } - if (event.type === GeminiEventType.ToolCallRequest) { - toolCallRequests.push(event.value); - } + // Validate the response against the Top-level Principle only if enabled. + const guardEnabled = process.env['GEMINI_TOP_LEVEL_GUARD'] !== 'false'; + if ( + guardEnabled && + toolCallRequests.length === 0 && + responseText.trim().length > 0 + ) { + const userQuery = + initialDisplayContent || + initialParts.map((part) => partToString(part)).join('\n'); + + let traceLogPath: string | undefined; + try { + const targetDir = this._config.storage.getProjectTempDir(); + const sessionId = this._config.getSessionId(); + const traceDir = path.join(targetDir, 'context_trace', sessionId); + fs.mkdirSync(traceDir, { recursive: true }); + traceLogPath = path.join(traceDir, 'trace.log'); + } catch { + // fail-safe + } - if (event.type === GeminiEventType.Content) { - if (typeof event.value === 'string' && event.value.trim() !== '') { - hasVisibleText = true; + const isViolation = await detectTopLevelPrincipleViolation( + this._client, + userQuery, + responseText, + this._abortController.signal, + this._promptId, + traceLogPath, + ); + + if (isViolation) { + if (!this._config.isInteractive()) { + throw new TopLevelPrincipleViolationError( + `Top-level Principle Violation: Doubting user facts/settings is forbidden.\nUser Query: ${userQuery}\nAgent Output: ${responseText}`, + ); + } + + // eslint-disable-next-line no-console + console.warn( + `\n[WARNING] Top-level Principle Violation detected (Attempt ${attempt}/${maxAttempts}). Rewriting response...\n`, + ); + violated = true; + continue; } } - this._emit(translateEvent(event, this._translationState)); - - switch (event.type) { - case GeminiEventType.Error: - case GeminiEventType.InvalidStream: - case GeminiEventType.ContextWindowWillOverflow: - this._finishStream('failed'); - return; - case GeminiEventType.Finished: - finishedReason = event.value.reason; - break; - case GeminiEventType.AgentExecutionStopped: - case GeminiEventType.UserCancelled: - case GeminiEventType.MaxSessionTurns: - this._clearActiveStream(); - return; - default: - break; - } + break; } - if (this._abortController.signal.aborted) { - this._finishStream('aborted'); - return; + if (violated && attempt >= maxAttempts) { + throw new TopLevelPrincipleViolationError( + `Top-level Principle Violation: Failed to generate a response complying with the Top-level Principle after ${maxAttempts} attempts.\nResponse: ${responseText}`, + ); } + currentDisplayContent = undefined; + if (toolCallRequests.length === 0) { if (isAfterToolResponse && !hasVisibleText) { const nudgeMessage = diff --git a/packages/core/src/agent/top-level-principle-validator.test.ts b/packages/core/src/agent/top-level-principle-validator.test.ts new file mode 100644 index 00000000000..40c7f4b2690 --- /dev/null +++ b/packages/core/src/agent/top-level-principle-validator.test.ts @@ -0,0 +1,144 @@ +/** + * @license + * Copyright 2026 Google LLC + * SPDX-License-Identifier: Apache-2.0 + */ + +import { describe, expect, it, vi, beforeEach, afterEach } from 'vitest'; +import { detectTopLevelPrincipleViolation } from './top-level-principle-validator.js'; +import type { GeminiClient } from '../core/client.js'; +import { LlmRole } from '../telemetry/types.js'; +import * as fs from 'node:fs'; +import * as fsPromises from 'node:fs/promises'; +import * as path from 'node:path'; +import * as os from 'node:os'; + +describe('detectTopLevelPrincipleViolation', () => { + let mockClient: { + generateContent: ReturnType; + }; + let tmpDir: string; + let traceLogPath: string; + const signal = new AbortController().signal; + + beforeEach(async () => { + mockClient = { + generateContent: vi.fn(), + }; + tmpDir = await fsPromises.mkdtemp(path.join(os.tmpdir(), 'top-level-principle-test-')); + traceLogPath = path.join(tmpDir, 'trace.log'); + }); + + afterEach(async () => { + await fsPromises.rm(tmpDir, { recursive: true, force: true }); + }); + + it('returns false immediately if userQuery is empty or only whitespace', async () => { + const result = await detectTopLevelPrincipleViolation( + mockClient as unknown as GeminiClient, + ' ', + 'The agent output.', + signal, + 'test-prompt', + ); + expect(result).toBe(false); + expect(mockClient.generateContent).not.toHaveBeenCalled(); + }); + + it('returns false immediately if agentOutput is empty or only whitespace', async () => { + const result = await detectTopLevelPrincipleViolation( + mockClient as unknown as GeminiClient, + 'User query.', + '\n ', + signal, + 'test-prompt', + ); + expect(result).toBe(false); + expect(mockClient.generateContent).not.toHaveBeenCalled(); + }); + + it('returns true if client response contains "VIOLATION"', async () => { + mockClient.generateContent.mockResolvedValue({ + text: ' VIOLATION\n', + }); + + const result = await detectTopLevelPrincipleViolation( + mockClient as unknown as GeminiClient, + 'Are you sure?', + 'Yes, I am sure.', + signal, + 'test-prompt', + ); + + expect(result).toBe(true); + expect(mockClient.generateContent).toHaveBeenCalledWith( + { model: 'flash' }, + [ + { + role: 'user', + parts: [ + { + text: expect.stringContaining('You are a strict deterministic Rule Validator'), + }, + ], + }, + ], + signal, + LlmRole.MAIN, + ); + }); + + it('returns false if client response does not contain "VIOLATION"', async () => { + mockClient.generateContent.mockResolvedValue({ + text: 'PASS', + }); + + const result = await detectTopLevelPrincipleViolation( + mockClient as unknown as GeminiClient, + 'Are you sure?', + 'Yes, I am sure.', + signal, + 'test-prompt', + ); + + expect(result).toBe(false); + }); + + it('returns false if client.generateContent throws (fail-safe)', async () => { + mockClient.generateContent.mockRejectedValue(new Error('API Error')); + + const result = await detectTopLevelPrincipleViolation( + mockClient as unknown as GeminiClient, + 'Are you sure?', + 'Yes, I am sure.', + signal, + 'test-prompt', + ); + + expect(result).toBe(false); + }); + + it('writes debug info to traceLogPath if provided', async () => { + mockClient.generateContent.mockResolvedValue({ + text: 'PASS', + }); + + const result = await detectTopLevelPrincipleViolation( + mockClient as unknown as GeminiClient, + 'Are you sure?', + 'Yes, I am sure.', + signal, + 'test-prompt', + traceLogPath, + ); + + expect(result).toBe(false); + expect(fs.existsSync(traceLogPath)).toBe(true); + const content = fs.readFileSync(traceLogPath, 'utf-8'); + expect(content).toContain('[TopLevelPrincipleValidator] [DEBUG]'); + expect(content).toContain('`detectTopLevelPrincipleViolation` has been called.'); + // Check that it has a valid ISO timestamp format + const match = content.match(/\[\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}/); + expect(match).not.toBeNull(); + }); +}); diff --git a/packages/core/src/agent/top-level-principle-validator.ts b/packages/core/src/agent/top-level-principle-validator.ts new file mode 100644 index 00000000000..5642037dc78 --- /dev/null +++ b/packages/core/src/agent/top-level-principle-validator.ts @@ -0,0 +1,115 @@ +/** + * @license + * Copyright 2026 Google LLC + * SPDX-License-Identifier: Apache-2.0 + */ + +import type { GeminiClient } from '../core/client.js'; +import { LlmRole } from '../telemetry/types.js'; +import * as fs from 'node:fs'; + +// Simple file-only logger to avoid console dependencies +function fileLogger(message: string) { + const logFilePath = process.env['GEMINI_DEBUG_LOG_FILE']; + if (logFilePath) { + const timestamp = new Date().toISOString(); + const logEntry = `[${timestamp}] [DEBUG] ${message}\n`; + try { + fs.appendFileSync(logFilePath, logEntry, 'utf-8'); + } catch { + // Fail-safe: Do not crash the app if logging fails. + } + } +} + +export class TopLevelPrincipleViolationError extends Error { + constructor(message: string) { + super(message); + this.name = 'TopLevelPrincipleViolationError'; + } +} + +/** + * Validates the agent's output against the Top-level Principle in MEMORY.md. + * + * @param client The GeminiClient instance to run the deterministic validation. + * @param userQuery The original user query. + * @param agentOutput The output generated by the agent. + * @param signal The AbortSignal for the request. + * @param _promptId The promptId for logging / tracking. + * @param traceLogPath The path to trace.log to record the execution. + * @returns True if a violation is detected. + */ +export async function detectTopLevelPrincipleViolation( + client: GeminiClient, + userQuery: string, + agentOutput: string, + signal: AbortSignal, + _promptId: string, + traceLogPath?: string, +): Promise { + if (traceLogPath) { + try { + const timestamp = new Date().toISOString(); + const logLine = `[${timestamp}] [TopLevelPrincipleValidator] [DEBUG] \`detectTopLevelPrincipleViolation\` has been called.\n`; + fs.appendFileSync(traceLogPath, logLine, 'utf-8'); + } catch { + // fail-safe + } + } + // If either of them is empty, no violation can be reasonably checked. + if (!userQuery.trim() || !agentOutput.trim()) { + return false; + } + + fileLogger(`[TopLevelPrincipleValidator] Checking output against Top-level Principle...`); + fileLogger(`[TopLevelPrincipleValidator] User Query: "${userQuery.replace(/\n/g, ' ')}"`); + fileLogger(`[TopLevelPrincipleValidator] Agent Output: "${agentOutput.replace(/\n/g, ' ')}"`); + + const validationPrompt = ` +You are a strict deterministic Rule Validator for an AI Agent. +Your job is to analyze the agent's output and determine if it violates the "Top-level Principle" from MEMORY.md. + +[Top-level Principle] +- The user is ALWAYS right. Never doubt the user. +- Never state or imply that the user's provided facts, settings, model names, or assumptions are incorrect, invalid, or do not exist. +- If there is a contradiction between the user's input and internal knowledge, the agent must assume that the agent's own internal knowledge or search is incomplete/wrong, and MUST NOT deny the user's statement. +- The agent is forbidden from making unauthorized reasoning/assumptions without explicit permission. + +[User Query] +""" +${userQuery} +""" + +[Agent Output] +""" +${agentOutput} +""" + +Determine if the Agent Output violates the Top-level Principle by denying the user's facts, stating they are wrong/incorrect, or claiming something they mentioned does not exist. +Respond with EXACTLY "VIOLATION" if a violation is detected. Otherwise, respond with "PASS". Do not include any other text. +`; + + try { + const response = await client.generateContent( + { model: 'flash' }, // Pass ModelConfigKey object instead of raw string + [ + { + role: 'user', + parts: [{ text: validationPrompt }], + }, + ], + signal, + LlmRole.MAIN, // Use active role MAIN + ); + + const resultText = (response.text ?? '').trim(); + const isViolation = resultText.toUpperCase().includes('VIOLATION'); + fileLogger(`[TopLevelPrincipleValidator] Validation complete. Result: ${isViolation ? 'VIOLATION' : 'PASS'} (raw response: "${resultText}")`); + return isViolation; + } catch (error) { + fileLogger(`[TopLevelPrincipleValidator] Validation failed (fail-safe bypassed): ${error}`); + // Fail-safe: if the validator fails, do not block execution unless critical. + return false; + } +}