Skip to content
Merged
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
docs: promote Claude oracle disagreement from open probe
  • Loading branch information
hippoley committed Oct 7, 2026
commit 86202a27e2ae072e314bceb07e8662d9b1cc033b
2 changes: 1 addition & 1 deletion docs/REALITY_LAB.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ The rule is simple:
| [openai/codex-plugin-cc#731](https://github.com/openai/codex-plugin-cc/pull/731) | Do the changed tests prove the original fail-open bug, and do they also answer later reviewer concerns? | **WITNESSED + proof boundary** | The tests prove the original bug, but not later 0600/atomic-write concerns. Green evidence must not silently expand its claim. |
| [openai/codex-plugin-cc#456](https://github.com/openai/codex-plugin-cc/pull/456) | What if the thing being fixed is the test harness/environment isolation itself? | **NOT WITNESSED + REVIEW REQUIRED** | Direct BASE/HEAD reproduction supports the bug, but ordinary witness is withheld because changed test support is part of the fix. |
| [MetrolistGroup/Metrolist#4097](https://github.com/MetrolistGroup/Metrolist/pull/4097) | Does a BASE non-zero exit always count as regression evidence? | **INCONCLUSIVE** | No. The BASE failed during Kotlin test compilation, not at a behavioral assertion. This produced the structured `json-v1` result protocol. |
| [anthropics/claude-code#89404](https://github.com/anthropics/claude-code/pull/89404) | What if the submitted tests are internally green but disagree with the product's authoritative parser? | **Open Reality Probe** | This is the emerging oracle-alignment problem: red→green evidence can still be weak if the judge does not represent product truth. |
| [anthropics/claude-code#89404](https://github.com/anthropics/claude-code/pull/89404) | What if the submitted tests are internally green but disagree with the product's authoritative parser? | **[ORACLE CONTRADICTED](../reality/claude-code-89404/ORACLE_DISAGREEMENT.md)** | Reviewer-supplied product-oracle evidence shows `claude plugin validate --json` rejects fixtures the submitted validator accepts, while the submitted suite can remain 5/5 green even when the multi-line extraction behavior regresses. CounterProof publishes this as an external-oracle contradiction, not as an independently executed Claude runtime replay. |
| [continuedev/continue#12576](https://github.com/continuedev/continue/pull/12576) | Can CounterProof replay a package-local Vitest regression inside a large JS/TS monorepo? | **[WITNESSED](../reality/continue-12576/WITNESS.md)** | The real PR exposed two product gaps first: `*.vitest.ts` discovery and package-local `core/node_modules` closure. After both fixes, HEAD was 5/5 PASS and BASE was 4/5 PASS, 1 FAIL; the BASE failure directly hits the legacy `tabAutocompleteModel` migration claim. |
| [topoteretes/cognee#5161](https://github.com/topoteretes/cognee/pull/5161) | What if local safety tests pass but a human reviewer reproduces a broader end-to-end failure mode those tests never exercise? | **Human oracle CONTRADICTED** | The vector adapter refusal test is narrow; reviewer evidence shows graph deletion can happen before the refusal propagates. CounterProof records the human reproduction as an external oracle instead of pretending it ran PostgreSQL itself. |
| [vercel/ai#17096](https://github.com/vercel/ai/pull/17096) | What if two newly added tests have different evidentiary value, while a stronger production claim lives outside the test suite? | **[WITNESSED + NOT WITNESSED + UNVERIFIED](../reality/vercel-ai-17096/CLAIM_MATRIX.md)** | Host→bridge sandbox forwarding is red→green; the new schema test passes on BASE and HEAD; real Claude Agent SDK enforcement remains an external production oracle. |
Expand Down