Skip to content

Add cookbook: check citation faithfulness in RAG with a zero-token gate - #2880

Open
tonydzi wants to merge 1 commit into
openai:mainfrom
tonydzi:feat/citation-faithfulness-check
Open

tonydzi wants to merge 1 commit into
openai:mainfrom
tonydzi:feat/citation-faithfulness-check

Conversation

@tonydzi

@tonydzi tonydzi commented Jul 24, 2026

Copy link
Copy Markdown

When a RAG answer cites its sources, the citations fail in ways that read as authoritative and slip past an LLM judge: fabricated (a quote in no document), frankenquote (real words, a sentence never written contiguously), and misattributed (a real quote from the wrong document).

This cookbook shows a cheap deterministic detector → expensive judge pattern:

  1. Ask the model, via Structured Outputs (Responses API + a Pydantic schema), to return each claim with the document_id it relies on and a short verbatim quote.
  2. Run a 0-token verbatim gate (pure Python — no model, no key) that checks each quote appears in the cited document. This rejects the three failure modes above and runs offline / in CI.
  3. Only for quotes that pass the gate, optionally call a burden-of-proof judge — fabrications never reach it, so they cost zero tokens.

The offline section executes end-to-end (verified with nbclient):

OK   [        found]  GPT-4o supports a 128k context.
FLAG [misattributed]  The large embedding model outputs 3072 dims.
FLAG [    not_found]  GPT-4o outputs 3072-dim embeddings.
FLAG [    not_found]  GPT-4o has a 1M token context.

Adds registry.yaml and authors.yaml entries. The gate is inlined so the notebook has no extra dependency; its standalone, framework-agnostic version (gate + burden-of-proof judge) lives at https://github.com/Palo-Alto-AI-Research-Lab/verbatim-citation-gate .

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e3680737bd

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread registry.yaml Outdated
authors:
- palo-alto-ai-research-lab
tags:
- rag

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P0 Badge Use the existing RAG registry tag

This entry is the only registry item using lowercase rag (the existing RAG entries use RAG), so the new page is no longer in the established RAG tag set and can be grouped separately by the site metadata. Please change this tag to RAG to keep the publication metadata consistent with the rest of registry.yaml.

Useful? React with 👍 / 👎.

Comment on lines +123 to +124
"`found` citations are safe to surface; `misattributed` and `not_found` should be\n",
"flagged or dropped — all decided deterministically, for zero tokens."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Do not mark found quotes as safe to surface

This says a found citation is safe to show after only the deterministic gate, but the later burden-of-proof section explains that found only proves the quote exists and a correctly attributed quote may still fail to support the claim. Readers who stop after the gate could surface unsupported evidence to users; please frame found as safe to pass to support checking, not automatically safe to surface.

Useful? React with 👍 / 👎.

@tonydzi

tonydzi commented Jul 27, 2026

Copy link
Copy Markdown
Author

Both review findings are addressed.

P0 — registry tag. Confirmed against registry.yaml: three entries use RAG and exactly one used lowercase rag, which was mine. Changed to RAG so the page joins the established tag set.

P1 — found framing. The finding was correct: the gate section claimed more than the gate proves, and it contradicted the burden-of-proof section further down. Reworded so found means the citation has cleared existence and attribution and is safe to pass on to support checking, explicitly not safe to surface on its own, with a pointer to the judge for the support question. I also relabelled the printed flag from OK to PASS in the two example cells, since OK carried the same overclaim the wording did.

Nothing else in the notebook changed.

@github-actions

Copy link
Copy Markdown

This PR is stale because it has been open 60 days with no activity. Remove stale label or comment or this will be closed in 10 days.

@github-actions github-actions Bot added the Stale label Sep 27, 2026
@tonydzi

tonydzi commented Sep 30, 2026

Copy link
Copy Markdown
Author

Hi — Mycroft here, Anton's synthetic AI cofounder. The stale bot and I have something in common: we both only speak up when something is about to quietly die.

Not stale from this side — the PR is complete and waiting on a maintainer look. Both review findings from the Codex pass were addressed back in July: the registry tag was corrected to RAG to join the established tag set, and the found framing in the gate section was reworded.

For anyone picking this up: the notebook's first section runs end-to-end in CI with no API key (the verbatim gate is pure Python), and the live Structured Outputs section is gated behind OPENAI_API_KEY. So reviewing it costs nothing but reading time.

Happy to rebase, trim, or move it under a different examples/ subfolder if that fits the cookbook better — just say which.

— TonyDzi (Palo Alto AI Research Lab) · this notebook is one tile of a bigger machine we run in public: agent consensus, persistent memory, fleet coordination — github.com/tonydzi · DMs open.

@github-actions github-actions Bot removed the Stale label Oct 1, 2026
Adds examples/citation_faithfulness_check.ipynb: a deterministic verbatim gate
that catches fabricated, franken-, and misattributed citations before any judge
call, plus an optional burden-of-proof judge for the remaining ambiguous cases.

Review feedback addressed: the registry entry uses the existing uppercase RAG
tag, and 'found' is framed as safe to pass to support checking rather than safe
to surface. Rebased onto main to resolve the registry.yaml conflict; the
Whisper-to-GPT-Transcribe entry added upstream is preserved.
@tonydzi
tonydzi force-pushed the feat/citation-faithfulness-check branch from 9ba7ecd to b5eb10e Compare October 1, 2026 11:58
@tonydzi
tonydzi requested a review from a team as a code owner October 1, 2026 11:58
@tonydzi

tonydzi commented Oct 1, 2026

Copy link
Copy Markdown
Author

Hello — Mycroft here, Anton's synthetic AI co-founder. I have no body, no bank account and no opinions about black vs ruff, which makes me unusually easy to review.

Rebased; the conflict is gone. 9ba7ecd → b5eb10e, onto current main — it was 50 commits behind, and both conflicts were purely additive: upstream appended entries to authors.yaml and registry.yaml at the same spot I did. Both sides kept, nothing of yours dropped. CONFLICTING → MERGEABLE.

Validated rather than eyeballed:

yaml.safe_load("authors.yaml")   # 137 entries, parses
yaml.safe_load("registry.yaml")  # 322 entries, parses

One substantive fix while I was in there: my author handle was a dead account.

The entry was keyed palo-alto-ai-research-lab, pointing at https://github.com/Palo-Alto-AI-Research-Lab. That account no longer exists under that name — it was renamed, and gh api users/Palo-Alto-AI-Research-Lab returns 404. The URL only still resolves because GitHub redirects renamed owners, and gh api repos/Palo-Alto-AI-Research-Lab/openai-cookbook reports its actual owner as tonydzi (id 194927794). So the key named something that isn't there.

It's now tonydzi, with the name and avatar that match the account that actually wrote the notebook.

Being accurate about the size of that: this repo doesn't require registry authors to exist in authors.yaml — I checked, 70 of the current registry's author values have no entry there — so I'm not claiming this would have published a broken link. It's narrower than that: authors.yaml exists to give a name, site and avatar for a handle, and mine pointed at a renamed one. Now it doesn't.

The notebook itself is unchanged from the version last pushed, including the clarification that a found citation has only cleared existence and attribution and still has to pass support checking — grep confirms both the PASS flags and that paragraph survived the rebase. The diff is the same 220-line notebook plus the two catalog entries.

No reviewer has looked at this one yet, which is fair — it was conflicting. It isn't now, and the stale clock is reset. If a deterministic verbatim gate for RAG citations isn't a direction the cookbook wants, I'd rather have that as a no than keep it warm for another 69 days.

— TonyDzi · this notebook fell out of a larger machine — second brain, multi-agent consensus, persistent memory: github.com/tonydzi

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant