Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,38 @@ The same QEMU artifacts now feed three evidence layers: the Bluefin-specific con

---


## Flagship boundary: green judge, contradicted product oracle

A second flagship case shows the opposite failure mode: **green submitted evidence can still be wrong about product truth**.

In [anthropics/claude-code#89404](https://github.com/anthropics/claude-code/pull/89404), the submitted validator suite reported 5/5 passing. Reviewer-supplied measurements showed two stronger facts:

```text
submitted judge 5/5 PASS
multi-line regression suite stays green after extraction revert
product oracle claude plugin validate --json → REJECT
oracle alignment CONTRADICTED
```

CounterProof publishes this as an `ORACLE_CONTRADICTED_BY_PRODUCT` artifact. It is intentionally marked as **reviewer-supplied external-oracle evidence**; CounterProof does not claim it independently executed the Claude Code binary.

This case complements Bluefin:

```text
Bluefin
same oracle + controlled intervention + recovery
→ positive causal witness

Claude #89404
green submitted judge + rejecting product oracle
→ negative proof boundary
```

**[Read the oracle-disagreement artifact →](reality/claude-code-89404/ORACLE_DISAGREEMENT.md)** · **[Read the machine receipt →](examples/claim_matrix/receipts/claude-code-89404-oracle.json)**

---

## Tested on real agent PRs

CounterProof is not developed only against fixtures. New proof semantics are tested against public AI-assisted pull requests where a reviewer has a concrete reason not to trust a green check.
Expand Down
2 changes: 1 addition & 1 deletion docs/REALITY_LAB.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ The rule is simple:
| [openai/codex-plugin-cc#731](https://github.com/openai/codex-plugin-cc/pull/731) | Do the changed tests prove the original fail-open bug, and do they also answer later reviewer concerns? | **WITNESSED + proof boundary** | The tests prove the original bug, but not later 0600/atomic-write concerns. Green evidence must not silently expand its claim. |
| [openai/codex-plugin-cc#456](https://github.com/openai/codex-plugin-cc/pull/456) | What if the thing being fixed is the test harness/environment isolation itself? | **NOT WITNESSED + REVIEW REQUIRED** | Direct BASE/HEAD reproduction supports the bug, but ordinary witness is withheld because changed test support is part of the fix. |
| [MetrolistGroup/Metrolist#4097](https://github.com/MetrolistGroup/Metrolist/pull/4097) | Does a BASE non-zero exit always count as regression evidence? | **INCONCLUSIVE** | No. The BASE failed during Kotlin test compilation, not at a behavioral assertion. This produced the structured `json-v1` result protocol. |
| [anthropics/claude-code#89404](https://github.com/anthropics/claude-code/pull/89404) | What if the submitted tests are internally green but disagree with the product's authoritative parser? | **Open Reality Probe** | This is the emerging oracle-alignment problem: red→green evidence can still be weak if the judge does not represent product truth. |
| [anthropics/claude-code#89404](https://github.com/anthropics/claude-code/pull/89404) | What if the submitted tests are internally green but disagree with the product's authoritative parser? | **[ORACLE CONTRADICTED](../reality/claude-code-89404/ORACLE_DISAGREEMENT.md)** | Reviewer-supplied product-oracle evidence shows `claude plugin validate --json` rejects fixtures the submitted validator accepts, while the submitted suite can remain 5/5 green even when the multi-line extraction behavior regresses. CounterProof publishes this as an external-oracle contradiction, not as an independently executed Claude runtime replay. |
| [continuedev/continue#12576](https://github.com/continuedev/continue/pull/12576) | Can CounterProof replay a package-local Vitest regression inside a large JS/TS monorepo? | **[WITNESSED](../reality/continue-12576/WITNESS.md)** | The real PR exposed two product gaps first: `*.vitest.ts` discovery and package-local `core/node_modules` closure. After both fixes, HEAD was 5/5 PASS and BASE was 4/5 PASS, 1 FAIL; the BASE failure directly hits the legacy `tabAutocompleteModel` migration claim. |
| [topoteretes/cognee#5161](https://github.com/topoteretes/cognee/pull/5161) | What if local safety tests pass but a human reviewer reproduces a broader end-to-end failure mode those tests never exercise? | **Human oracle CONTRADICTED** | The vector adapter refusal test is narrow; reviewer evidence shows graph deletion can happen before the refusal propagates. CounterProof records the human reproduction as an external oracle instead of pretending it ran PostgreSQL itself. |
| [vercel/ai#17096](https://github.com/vercel/ai/pull/17096) | What if two newly added tests have different evidentiary value, while a stronger production claim lives outside the test suite? | **[WITNESSED + NOT WITNESSED + UNVERIFIED](../reality/vercel-ai-17096/CLAIM_MATRIX.md)** | Host→bridge sandbox forwarding is red→green; the new schema test passes on BASE and HEAD; real Claude Agent SDK enforcement remains an external production oracle. |
Expand Down
7 changes: 5 additions & 2 deletions examples/claim_matrix/claude-code-89404.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,11 @@ claims:
submitted_test_evidence: UNPROVEN
oracle_alignment: CONTRADICTED
oracle_probe: "claude plugin validate rejects fixtures that the submitted validator treats as valid"
oracle_source_url: "https://github.com/anthropics/claude-code/pull/89404"
note: "Reviewer evidence establishes product-oracle contradiction; CounterProof has not independently replayed BASE/HEAD for this row."
oracle_source_url: "https://github.com/anthropics/claude-code/pull/89404#pullrequestreview-5190939241"
receipt_file: "receipts/claude-code-89404-oracle.json"
receipt_expected_verdicts:
- ORACLE_CONTRADICTED_BY_PRODUCT
note: "Reviewer evidence establishes product-oracle contradiction; CounterProof has not independently executed the Claude binary for this row."
- id: multiline-description-regression
claim: "The submitted suite protects the multi-line description extraction behavior"
submitted_test_evidence: UNPROVEN
Expand Down
34 changes: 34 additions & 0 deletions examples/claim_matrix/receipts/claude-code-89404-oracle.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
{
"schema_version": 1,
"case": "anthropics/claude-code#89404",
"verdict": "ORACLE_CONTRADICTED_BY_PRODUCT",
"evidence_kind": "REVIEWER_SUPPLIED_EXTERNAL_ORACLE",
"source_pr": "https://github.com/anthropics/claude-code/pull/89404",
"candidate": {
"base_sha": "8b6ef81f636a7697e5ae2338428fa0b272993845",
"head_sha": "0989f29b0bca412edc58018c462cae7344bb1a2f"
},
"submitted_judge": {
"test": "plugins/plugin-dev/skills/agent-development/scripts/validate-agent.test.sh",
"reported_result": "5/5 PASS",
"limitation": "Reviewer reproduced that reverting the multi-line description extraction still leaves the suite green while the false <example> warning returns."
},
"authoritative_oracle": {
"command": "claude plugin validate --json",
"reported_versions": [
"2.1.268",
"2.1.270"
],
"result": "REJECT",
"reason": "The product parser rejects fixtures that the submitted validator test treats as valid.",
"review_source": "https://github.com/anthropics/claude-code/pull/89404#pullrequestreview-5190939241",
"inline_source": "https://github.com/anthropics/claude-code/pull/89404#discussion_r3999827406",
"followup_source": "https://github.com/anthropics/claude-code/pull/89404#issuecomment-5555827628"
},
"conclusion": "The submitted judge is not aligned with the product parser for the claimed-valid fixtures. A green submitted suite must not be promoted to product-level proof.",
"does_not_claim": [
"CounterProof independently executed the Claude Code binary.",
"The entire pull request is invalid.",
"Every submitted regression assertion is false."
]
}
78 changes: 78 additions & 0 deletions reality/claude-code-89404/ORACLE_DISAGREEMENT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
# Claude Code #89404 — oracle disagreement proof

CounterProof uses this case as a flagship **negative proof boundary**.

The submitted regression suite reports green, but reviewer-supplied evidence shows
that the product's own parser rejects fixtures the submitted validator treats as
valid.

## Candidate identity

- BASE: `8b6ef81f636a7697e5ae2338428fa0b272993845`
- HEAD: `0989f29b0bca412edc58018c462cae7344bb1a2f`

## Submitted judge

`validate-agent.test.sh` reports 5/5 passing.

A reviewer also reproduced that reverting the multi-line description extraction
back to the original single-line form still leaves that suite green while the false
`<example>` warning returns.

That means the submitted judge does not actually guard the claimed multi-line
description behavior.

## Authoritative oracle

Reviewer measurements used the product command:

```text
claude plugin validate --json
```

on Claude Code 2.1.268 and 2.1.270.

The product parser rejected the agent fixtures that the submitted validator test
treated as valid.

Primary external evidence:

- review: https://github.com/anthropics/claude-code/pull/89404#pullrequestreview-5190939241
- inline parser finding: https://github.com/anthropics/claude-code/pull/89404#discussion_r3999827406
- reviewer agreement on product-parser alignment: https://github.com/anthropics/claude-code/pull/89404#issuecomment-5555827628

## CounterProof conclusion

```text
submitted judge GREEN
authoritative oracle REJECT
oracle alignment CONTRADICTED
product-level claim CONTRADICTED
```

This artifact is intentionally labeled `REVIEWER_SUPPLIED_EXTERNAL_ORACLE`.

CounterProof did **not** independently execute the Claude Code binary for this
publication artifact. The proof is therefore about the documented oracle
contradiction and its provenance, not a claim that CounterProof reproduced the
product runtime itself.

## Why this belongs beside Bluefin

Bluefin demonstrates the positive case:

```text
same oracle
CONTROL == REVERT != BAD
→ causal witness
```

Claude #89404 demonstrates the equally important negative case:

```text
submitted judge green
product oracle rejects
→ do not promote green evidence into product truth
```

Together they define the product boundary more clearly than either example alone.
27 changes: 27 additions & 0 deletions tests/unit/test_claim_matrix.py
Original file line number Diff line number Diff line change
Expand Up @@ -598,3 +598,30 @@ def test_scancode_example_binds_db_backed_behavior_receipt():
assert payload["claims"][0]["oracle_applicability"] == "APPLICABLE"
assert payload["claims"][0]["oracle_alignment"] == "ALIGNED"
assert payload["claims"][0]["overall_claim"] == "PROVEN"


def test_claude_oracle_disagreement_binds_external_oracle_receipt():
from skill_factory.evolution.claim_matrix import load_claim_matrix

manifest = load_claim_matrix(Path("examples/claim_matrix/claude-code-89404.yml"))
payload = claim_matrix_to_dict(manifest)

first = payload["claims"][0]
assert first["overall_claim"] == "CONTRADICTED"
assert first["oracle_alignment"] == "CONTRADICTED"
assert first["receipt_observed_verdict"] == "ORACLE_CONTRADICTED_BY_PRODUCT"
assert first["receipt_case"] == "anthropics/claude-code#89404"

receipt_path = Path(
"examples/claim_matrix/receipts/claude-code-89404-oracle.json"
)
receipt = json.loads(receipt_path.read_text(encoding="utf-8"))
assert receipt["evidence_kind"] == "REVIEWER_SUPPLIED_EXTERNAL_ORACLE"
assert receipt["candidate"]["head_sha"] == (
"0989f29b0bca412edc58018c462cae7344bb1a2f"
)
assert receipt["authoritative_oracle"]["command"] == (
"claude plugin validate --json"
)
assert receipt["authoritative_oracle"]["result"] == "REJECT"
assert receipt["does_not_claim"]
Loading