This is a CounterProof Reality Probe about evidence semantics, not a request to change anthropics/claude-code#89404.
Source PR:
anthropics/claude-code#89404
@konsta95 @sylvesterkaczmarek — your review thread exposed a case I think is more important than ordinary “tests pass” verification.
The PR adds validate-agent.test.sh and reports 5/5 passing tests. But the review established two stronger facts:
- the test suite asserts exit 0 for agent files that the actual product's
claude plugin validate rejects as invalid YAML/frontmatter;
- reverting the multi-line description extraction can still leave the 5/5 suite green while the false
<example> warning returns.
So even a perfect before/after replay of the submitted test could produce a technically real red→green result while still being weak evidence for the actual product behavior.
Why I am asking
CounterProof currently has three distinct evidence outcomes from real external PRs:
- Regression Witness — same changed test is PASS on HEAD / FAIL on BASE.
- Proof Integrity — surface when the PR changes the CI/test-support machinery producing its own evidence.
- INCONCLUSIVE / NOT WITNESSED — refuse to turn compile/setup/judge changes into a success claim.
#89404 suggests a fourth problem:
the submitted test can be internally consistent but use the wrong oracle.
For this PR, claude plugin validate --json appears to be closer to the authoritative oracle than validate-agent.sh's own notion of valid frontmatter.
Question
If you were reviewing an AI-assisted PR like #89404, would the useful evidence shape be something like:
Submitted regression:
BASE -> FAIL
HEAD -> PASS
Authoritative oracle probe:
product parser on claimed-valid fixture -> REJECT
Conclusion:
regression is real relative to the submitted judge,
but the judge is not aligned with product behavior
Or is there a different distinction that would be more useful in practice?
I am specifically trying not to build an “oracle alignment” feature until a real reviewer says what information would actually reduce review uncertainty. A short answer is plenty.
This is a CounterProof Reality Probe about evidence semantics, not a request to change
anthropics/claude-code#89404.Source PR:
anthropics/claude-code#89404
@konsta95 @sylvesterkaczmarek — your review thread exposed a case I think is more important than ordinary “tests pass” verification.
The PR adds
validate-agent.test.shand reports 5/5 passing tests. But the review established two stronger facts:claude plugin validaterejects as invalid YAML/frontmatter;<example>warning returns.So even a perfect before/after replay of the submitted test could produce a technically real red→green result while still being weak evidence for the actual product behavior.
Why I am asking
CounterProof currently has three distinct evidence outcomes from real external PRs:
#89404 suggests a fourth problem:
For this PR,
claude plugin validate --jsonappears to be closer to the authoritative oracle thanvalidate-agent.sh's own notion of valid frontmatter.Question
If you were reviewing an AI-assisted PR like #89404, would the useful evidence shape be something like:
Or is there a different distinction that would be more useful in practice?
I am specifically trying not to build an “oracle alignment” feature until a real reviewer says what information would actually reduce review uncertainty. A short answer is plenty.