Skip to content

Commit 9ba7ecd

Browse files
committed
Add cookbook: check citation faithfulness in RAG with a zero-token gate
Adds examples/citation_faithfulness_check.ipynb: a deterministic verbatim gate that catches fabricated, franken-, and misattributed citations before any judge call, plus an optional burden-of-proof judge for the remaining ambiguous cases. Review feedback addressed: the registry entry uses the existing uppercase RAG tag, and 'found' is framed as safe to pass to support checking rather than safe to surface. Rebased onto main to resolve the registry.yaml conflict; the Whisper-to-GPT-Transcribe entry added upstream is preserved.
1 parent 0a796c4 commit 9ba7ecd

3 files changed

Lines changed: 237 additions & 0 deletions

File tree

‎authors.yaml‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -627,3 +627,8 @@ saptob:
627627
name: "Saptarshi Banerjee"
628628
website: "https://www.linkedin.com/in/saptarshi-banerjee-83472679"
629629
avatar: "https://avatars.githubusercontent.com/saptob"
630+
631+
palo-alto-ai-research-lab:
632+
name: "Palo Alto AI Research Lab"
633+
website: "https://github.com/Palo-Alto-AI-Research-Lab"
634+
avatar: "https://avatars.githubusercontent.com/u/194927794?v=4"
Lines changed: 220 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,220 @@
1+
{
2+
"cells": [
3+
{
4+
"cell_type": "markdown",
5+
"id": "b40669c5",
6+
"metadata": {},
7+
"source": [
8+
"# Check citation faithfulness in RAG with a zero-token gate\n",
9+
"\n",
10+
"When a RAG answer cites its sources, the citations can fail in ways that read as\n",
11+
"completely authoritative — and an LLM asked \"does this quote support the claim?\"\n",
12+
"waves them through, because they *look* fluent and supportive:\n",
13+
"\n",
14+
"- **fabricated** — a quoted span that appears in no retrieved document;\n",
15+
"- **frankenquote** — every word is real, but the exact span was never written\n",
16+
" contiguously in the source;\n",
17+
"- **misattributed** — a real span, but attributed to the wrong document.\n",
18+
"\n",
19+
"This cookbook shows a *cheap deterministic detector → expensive judge* pattern for\n",
20+
"catching these before the answer reaches a user:\n",
21+
"\n",
22+
"1. Ask the model, via **Structured Outputs**, to return each claim with the\n",
23+
" `document_id` it relies on and a short **verbatim quote** from that document.\n",
24+
"2. Run a **0-token verbatim gate** (pure Python — no model, no API key) that checks\n",
25+
" each quote really appears in the cited document. This alone rejects the three\n",
26+
" failure modes above and runs offline.\n",
27+
"3. Only for quotes that pass the gate, optionally call a **burden-of-proof judge**\n",
28+
" to decide whether the quote actually *supports* the claim (a right quote can\n",
29+
" still be the wrong evidence). Fabrications never reach the judge, so they cost\n",
30+
" zero tokens.\n",
31+
"\n",
32+
"The gate is inlined here; its standalone, framework-agnostic version (gate +\n",
33+
"burden-of-proof judge) lives at\n",
34+
"[`verbatim-citation-gate`](https://github.com/Palo-Alto-AI-Research-Lab/verbatim-citation-gate)."
35+
]
36+
},
37+
{
38+
"cell_type": "markdown",
39+
"id": "34a868bf",
40+
"metadata": {},
41+
"source": [
42+
"## The verbatim gate (deterministic, runs offline)"
43+
]
44+
},
45+
{
46+
"cell_type": "code",
47+
"execution_count": null,
48+
"id": "0ca719b2",
49+
"metadata": {},
50+
"outputs": [],
51+
"source": [
52+
"import re\n",
53+
"\n",
54+
"\n",
55+
"def normalize(text: str) -> str:\n",
56+
" \"\"\"Case/typography/whitespace-insensitive form for verbatim matching.\"\"\"\n",
57+
" text = text.lower()\n",
58+
" text = re.sub(r\"[‘’]\", \"'\", text)\n",
59+
" text = re.sub(r\"[“”]\", '\"', text)\n",
60+
" text = re.sub(r\"[–—]\", \"-\", text)\n",
61+
" text = re.sub(r\"[^a-z0-9%.]+\", \" \", text)\n",
62+
" return \" \".join(text.split())\n",
63+
"\n",
64+
"\n",
65+
"def gate(quote: str, cited_doc_id: str, docs: dict) -> str:\n",
66+
" \"\"\"Return 'found' | 'misattributed' | 'not_found'. Fails closed on empty quotes.\"\"\"\n",
67+
" q = normalize(quote)\n",
68+
" if not q:\n",
69+
" return \"not_found\"\n",
70+
" cited = docs.get(cited_doc_id)\n",
71+
" if cited is not None and q in normalize(cited):\n",
72+
" return \"found\"\n",
73+
" if any(q in normalize(t) for d, t in docs.items() if d != cited_doc_id):\n",
74+
" return \"misattributed\"\n",
75+
" return \"not_found\""
76+
]
77+
},
78+
{
79+
"cell_type": "markdown",
80+
"id": "55121e4a",
81+
"metadata": {},
82+
"source": [
83+
"### Verify it offline on structured citations\n",
84+
"\n",
85+
"No API key needed here — this is the shape of citations Structured Outputs returns,\n",
86+
"with one faithful citation and the three planted failure modes."
87+
]
88+
},
89+
{
90+
"cell_type": "code",
91+
"execution_count": null,
92+
"id": "659b55f5",
93+
"metadata": {},
94+
"outputs": [],
95+
"source": [
96+
"DOCS = {\n",
97+
" \"doc_0\": \"GPT-4o has a context window of 128,000 tokens.\",\n",
98+
" \"doc_1\": \"text-embedding-3-large produces embeddings with 3,072 dimensions.\",\n",
99+
"}\n",
100+
"\n",
101+
"CITATIONS = [\n",
102+
" {\"claim\": \"GPT-4o supports a 128k context.\", \"document_id\": \"doc_0\",\n",
103+
" \"quote\": \"context window of 128,000 tokens\"}, # faithful\n",
104+
" {\"claim\": \"The large embedding model outputs 3072 dims.\", \"document_id\": \"doc_0\",\n",
105+
" \"quote\": \"embeddings with 3,072 dimensions\"}, # wrong doc\n",
106+
" {\"claim\": \"GPT-4o outputs 3072-dim embeddings.\", \"document_id\": \"doc_1\",\n",
107+
" \"quote\": \"GPT-4o produces 3,072 dimensions\"}, # frankenquote\n",
108+
" {\"claim\": \"GPT-4o has a 1M token context.\", \"document_id\": \"doc_0\",\n",
109+
" \"quote\": \"context window of 1,000,000 tokens\"}, # fabricated\n",
110+
"]\n",
111+
"\n",
112+
"for c in CITATIONS:\n",
113+
" status = gate(c[\"quote\"], c[\"document_id\"], DOCS)\n",
114+
" flag = \"PASS\" if status == \"found\" else \"FLAG\"\n",
115+
" print(f\"{flag} [{status:>13}] {c['claim']}\")"
116+
]
117+
},
118+
{
119+
"cell_type": "markdown",
120+
"id": "a846f6eb",
121+
"metadata": {},
122+
"source": [
123+
"`found` citations have cleared existence and attribution, so they are safe to *pass\n",
124+
"on to support checking* — not automatically safe to surface: a real, correctly\n",
125+
"attributed quote can still fail to support the claim it is attached to (see the\n",
126+
"burden-of-proof judge below). `misattributed` and `not_found` should be flagged or\n",
127+
"dropped outright — all decided deterministically, for zero tokens.\n"
128+
]
129+
},
130+
{
131+
"cell_type": "markdown",
132+
"id": "abb6ca3c",
133+
"metadata": {},
134+
"source": [
135+
"## Generate the citations with Structured Outputs\n",
136+
"\n",
137+
"With an API key, ask the model to answer **and** return structured citations, then\n",
138+
"run the same gate over them. Uses the Responses API with a Pydantic schema; needs\n",
139+
"`OPENAI_API_KEY` (not run in CI)."
140+
]
141+
},
142+
{
143+
"cell_type": "code",
144+
"execution_count": null,
145+
"id": "604ff28e",
146+
"metadata": {},
147+
"outputs": [],
148+
"source": [
149+
"# pip install openai pydantic\n",
150+
"import os\n",
151+
"\n",
152+
"if not os.getenv(\"OPENAI_API_KEY\"):\n",
153+
" print(\"Set OPENAI_API_KEY to run the live example.\")\n",
154+
"else:\n",
155+
" from openai import OpenAI\n",
156+
" from pydantic import BaseModel\n",
157+
"\n",
158+
" class Citation(BaseModel):\n",
159+
" claim: str\n",
160+
" document_id: str\n",
161+
" quote: str\n",
162+
"\n",
163+
" class CitedAnswer(BaseModel):\n",
164+
" answer: str\n",
165+
" citations: list[Citation]\n",
166+
"\n",
167+
" client = OpenAI()\n",
168+
" library = \"\\n\".join(f\"[{doc_id}] {text}\" for doc_id, text in DOCS.items())\n",
169+
" resp = client.responses.parse(\n",
170+
" model=\"gpt-4.1\",\n",
171+
" input=[\n",
172+
" {\"role\": \"system\", \"content\": \"Answer using only the library. For each claim, cite the document_id \"\n",
173+
" \"and a short quote copied verbatim from that document.\"},\n",
174+
" {\"role\": \"user\", \"content\": f\"Library:\\n{library}\\n\\nQuestion: What context window does GPT-4o have, \"\n",
175+
" \"and how many dimensions does text-embedding-3-large output?\"},\n",
176+
" ],\n",
177+
" text_format=CitedAnswer,\n",
178+
" )\n",
179+
" cited = resp.output_parsed\n",
180+
" print(cited.answer, \"\\n\")\n",
181+
" for c in cited.citations:\n",
182+
" status = gate(c.quote, c.document_id, DOCS)\n",
183+
" flag = \"PASS\" if status == \"found\" else \"FLAG\"\n",
184+
" print(f\"{flag} [{status:>13}] {c.quote!r} -> {c.document_id}\")"
185+
]
186+
},
187+
{
188+
"cell_type": "markdown",
189+
"id": "026179f4",
190+
"metadata": {},
191+
"source": [
192+
"## Optional: a burden-of-proof judge for the ambiguous case\n",
193+
"\n",
194+
"The gate settles whether a quote *exists*. Whether a real, correctly-attributed\n",
195+
"quote actually **supports** its claim is a judgment call best given to a model — but\n",
196+
"with the burden of proof on the citation: the verdict defaults to *unsupported*, the\n",
197+
"model's outside knowledge is inadmissible, and an unparseable verdict fails closed.\n",
198+
"Because only `found` quotes reach it, fabricated and misattributed citations cost\n",
199+
"zero judge calls.\n",
200+
"\n",
201+
"A ready-made, model-agnostic version of this judge is in\n",
202+
"[`verbatim-citation-gate`](https://github.com/Palo-Alto-AI-Research-Lab/verbatim-citation-gate);\n",
203+
"plug your `client.responses.create` call into its `llm_call` hook."
204+
]
205+
}
206+
],
207+
"metadata": {
208+
"kernelspec": {
209+
"display_name": "Python 3",
210+
"language": "python",
211+
"name": "python3"
212+
},
213+
"language_info": {
214+
"name": "python",
215+
"version": "3.11"
216+
}
217+
},
218+
"nbformat": 4,
219+
"nbformat_minor": 5
220+
}

‎registry.yaml‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -18,6 +18,18 @@
1818
- migration
1919
- whisper
2020

21+
- title: Check citation faithfulness in RAG with a zero-token gate
22+
path: examples/citation_faithfulness_check.ipynb
23+
slug: citation-faithfulness-check
24+
description: Catch fabricated, franken-, and misattributed RAG citations with a deterministic verbatim gate plus an optional burden-of-proof judge, using Structured Outputs.
25+
date: 2026-07-24
26+
authors:
27+
- palo-alto-ai-research-lab
28+
tags:
29+
- RAG
30+
- structured-outputs
31+
- evals
32+
2133
- title: Enabling Long-Term Agent Memory with Oracle AI Agent Memory
2234
path: examples/vector_databases/oracle_db/deep_research_openai_agents.ipynb
2335
slug: deep-research-openai-agents

0 commit comments

Comments
 (0)