This is a CounterProof Reality Probe, not a claim that openai/codex-plugin-cc#456 is wrong.
Source PR: openai/codex-plugin-cc#456
@Caleb0796 @BlueHarmel — I used #456 because it has a useful property that ordinary regression-witness tooling tends to mishandle: the bug is real, but the PR fixes the test harness / environment isolation itself.
Direct external reproduction
I reproduced the relevant environment condition on a clean GitHub runner by setting the same kind of ambient variables a live Claude Code session exposes.
Exact commits:
- BASE:
db52e28f4d9ded852ab3942cea316258ae4ef346
- HEAD:
0615c5db852dc2471092f73fb41291a7a32d3223
Running pristine tests/state.test.mjs directly:
- BASE: 2 pass / 1 fail
- HEAD: 3 pass / 0 fail
So the environmental regression is independently reproducible.
Reality run:
https://github.com/hippoley/CounterProof/actions/runs/35969738266
Why CounterProof intentionally does NOT mint a Regression Witness
The PR changes:
tests/state.test.mjs
tests/helpers.mjs
CounterProof's normal before/after replay overlays changed test/support files onto BASE so both sides are judged by the same submitted regression evidence.
For #456 that means the changed tests/helpers.mjs — which is itself part of the fix — also lands on BASE.
Result:
- HEAD + PR test/support: PASS
- BASE + PR test/support: PASS
- NOT WITNESSED
That refusal is intentional. A green BASE replay after the changed test harness has been copied back is not independent evidence about the old harness.
Proof-integrity signal
This real PR also exposed a CounterProof bug: the first run called this situation CLEAN.
That is now fixed in CounterProof main. The same #456 replay reports:
- Proof Integrity: REVIEW-REQUIRED
- finding:
test-support-changed
- path:
tests/helpers.mjs
- risk: MEDIUM
Fix:
#14
The distinction I am testing
For a reviewer, the evidence now says two things at once:
- Direct environment reproduction supports the PR's reported test-harness bug.
- A standard Regression Witness is withheld because the judge/test support itself changed.
Would that distinction be useful when reviewing a test-infrastructure PR like #456, or would you prefer a different evidence model for “the judge is the thing being fixed”?
A one-line answer is enough; I am using real maintainer reactions to decide whether CounterProof needs a first-class “judge-fix proof” mode at all.
This is a CounterProof Reality Probe, not a claim that
openai/codex-plugin-cc#456is wrong.Source PR: openai/codex-plugin-cc#456
@Caleb0796 @BlueHarmel — I used #456 because it has a useful property that ordinary regression-witness tooling tends to mishandle: the bug is real, but the PR fixes the test harness / environment isolation itself.
Direct external reproduction
I reproduced the relevant environment condition on a clean GitHub runner by setting the same kind of ambient variables a live Claude Code session exposes.
Exact commits:
db52e28f4d9ded852ab3942cea316258ae4ef3460615c5db852dc2471092f73fb41291a7a32d3223Running pristine
tests/state.test.mjsdirectly:So the environmental regression is independently reproducible.
Reality run:
https://github.com/hippoley/CounterProof/actions/runs/35969738266
Why CounterProof intentionally does NOT mint a Regression Witness
The PR changes:
tests/state.test.mjstests/helpers.mjsCounterProof's normal before/after replay overlays changed test/support files onto BASE so both sides are judged by the same submitted regression evidence.
For #456 that means the changed
tests/helpers.mjs— which is itself part of the fix — also lands on BASE.Result:
That refusal is intentional. A green BASE replay after the changed test harness has been copied back is not independent evidence about the old harness.
Proof-integrity signal
This real PR also exposed a CounterProof bug: the first run called this situation
CLEAN.That is now fixed in CounterProof main. The same #456 replay reports:
test-support-changedtests/helpers.mjsFix:
#14
The distinction I am testing
For a reviewer, the evidence now says two things at once:
Would that distinction be useful when reviewing a test-infrastructure PR like #456, or would you prefer a different evidence model for “the judge is the thing being fixed”?
A one-line answer is enough; I am using real maintainer reactions to decide whether CounterProof needs a first-class “judge-fix proof” mode at all.