Skip to content

Reality probe: is red→green evidence misleading when the test oracle disagrees with the product? #16

Description

@hippoley

This is a CounterProof Reality Probe about evidence semantics, not a request to change anthropics/claude-code#89404.

Source PR:
anthropics/claude-code#89404

@konsta95 @sylvesterkaczmarek — your review thread exposed a case I think is more important than ordinary “tests pass” verification.

The PR adds validate-agent.test.sh and reports 5/5 passing tests. But the review established two stronger facts:

  1. the test suite asserts exit 0 for agent files that the actual product's claude plugin validate rejects as invalid YAML/frontmatter;
  2. reverting the multi-line description extraction can still leave the 5/5 suite green while the false <example> warning returns.

So even a perfect before/after replay of the submitted test could produce a technically real red→green result while still being weak evidence for the actual product behavior.

Why I am asking

CounterProof currently has three distinct evidence outcomes from real external PRs:

  • Regression Witness — same changed test is PASS on HEAD / FAIL on BASE.
  • Proof Integrity — surface when the PR changes the CI/test-support machinery producing its own evidence.
  • INCONCLUSIVE / NOT WITNESSED — refuse to turn compile/setup/judge changes into a success claim.

#89404 suggests a fourth problem:

the submitted test can be internally consistent but use the wrong oracle.

For this PR, claude plugin validate --json appears to be closer to the authoritative oracle than validate-agent.sh's own notion of valid frontmatter.

Question

If you were reviewing an AI-assisted PR like #89404, would the useful evidence shape be something like:

Submitted regression:
  BASE -> FAIL
  HEAD -> PASS

Authoritative oracle probe:
  product parser on claimed-valid fixture -> REJECT

Conclusion:
  regression is real relative to the submitted judge,
  but the judge is not aligned with product behavior

Or is there a different distinction that would be more useful in practice?

I am specifically trying not to build an “oracle alignment” feature until a real reviewer says what information would actually reduce review uncertainty. A short answer is plenty.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions