Behavioral proof for agent-generated pull requests.
Click the panel. Break the proof. Change the judge. See what survives.
Works beside Claude Code · Codex · Copilot · Cursor · PR-Agent · human-written PRs — Counterproof verifies the evidence, not the author.
A green CI run proves that your code passes now.
It does not prove that the regression test added by the same coding agent would have caught the bug before the fix.
Counterproof asks that missing question.
PR code + PR test → PASS
old code + the same test → FAIL
test / CI judge unchanged → CLEAN
↓
REGRESSION WITNESSED
That turns:
“the agent says it fixed the bug”
into:
this exact test behaves differently before and after the fix.
The browser experience lets you play with different evidence situations instead of reading another architecture diagram.
REAL REGRESSION
HEAD passes / BASE fails
→ strong before-vs-after evidence
WEAK TEST
HEAD passes / BASE also passes
→ the test does not witness the claimed fix
JUDGE CHANGED
the regression evidence exists
but CI / test machinery changed too
→ reviewer attention required
FULL SUITE
the full suite differs
but the changed test was not isolated
→ weaker evidence than an exact witness
The browser scenarios are fixtures. Real evidence comes from the CLI / GitHub Action.
Click the walkthrough to open the live Proof Lab.
Take tests changed in a pull request.
Run them on the PR.
Then replay the same tests against the pre-change code.
PR HEAD BASE
same changed test PASS FAIL
\ /
\ /
WITNESSED
If the same test already passes on BASE, Counterproof does not manufacture a success story.
It says the proof is weak.
The simplest production entry point on main is the Regression Witness Action:
name: Counterproof
on:
pull_request:
permissions:
contents: read
pull-requests: write
jobs:
proof:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
ref: ${{ github.event.pull_request.head.sha }}
fetch-depth: 0
- uses: actions/setup-python@v5
with:
python-version: "3.11"
# Install your project dependencies first.
- uses: hippoley/SkillFactory/actions/witness@main
with:
test-command: "python -m pytest -q {tests}"
require-witness: "true"
require-clean-integrity: "true"Counterproof will:
1. find tests added or modified by the PR
2. run them on PR HEAD
3. create a detached worktree at BASE
4. overlay the PR test/support files
5. run the same evidence on old code
6. inspect whether the PR changed the judge
7. write one sticky proof comment
No hosted service.
No API key.
No LLM is required for this proof path.
A passing test is weaker evidence if the same PR also weakens the system that evaluates it.
Counterproof's Proof Integrity Guard surfaces changes such as:
deleted test → review
new skip / xfail → review
continue-on-error: true → review
pytest ... || true → review
pull_request trigger removed → review
test / coverage / CI config changed → surface explicitly
A finding does not mean “malicious PR.”
It means:
the evidence surface changed, so the reviewer should not treat the green result as independent evidence.
Coding agents are getting very good at producing:
code
+ tests
+ green CI
+ a confident explanation
That is useful.
It also creates a new review problem:
the system proposing the fix can now help produce the evidence for its own fix.
Counterproof does not solve that by adding another model.
It adds a deterministic before/after experiment.
Git history
+
your existing test runner
+
the same evidence on both sides
=
something a reviewer can inspect
AI reviewers and Counterproof answer different questions.
| AI reviewer | Counterproof | |
|---|---|---|
| Main question | “Does this diff look suspicious?” | “Does this evidence distinguish before from after?” |
| Core input | code + model context | Git history + tests |
| Main output | suggestions / comments | replayable behavioral evidence |
| LLM required | usually | no for PR proof |
| Changed test/CI judge | not the core primitive | explicitly surfaced |
| Can refuse a story | model-dependent | yes — weak/inconclusive evidence stays weak |
Use Counterproof next to Claude Code, Codex, Copilot, Cursor, PR-Agent, or a human engineer.
It does not care who wrote the patch.
python -m pip install "git+https://github.com/hippoley/SkillFactory.git"
counterproof demoOpen:
http://127.0.0.1:8765
Or just use the public version:
Regression Witness is the smallest useful entry point.
Counterproof also contains an experimental runtime for falsifiable agent self-improvement.
Instead of asking only:
“what lesson should the agent remember?”
it asks:
which explanation survives an experiment, and what evidence earns the right to change behavior?
RAW TRACE
↓
Decision Capsule + Outcome Receipt
↓
typed evidence
↓
competing hypotheses
↓
Probe Contracts
↓
same cases × multiple interventions
↓
survived / falsified / inconclusive
↓
unique survivor?
↙ ↘
yes no
↓ ↓
select remain ambiguous
↓
Behavior Proof
↓
promotion gate
counterproof prove examples/traces/tenant_failure.json \
--replay-manifest examples/replay_suite.json \
--surface policy \
--out BEHAVIOR_PROOF.mdcounterproof discriminate examples/traces/tenant_failure.json \
--experiment-manifest examples/discrimination_suite.json \
--surface policy \
--surface skill \
--surface prompt \
--out DISCRIMINATION.md \
--json-out DISCRIMINATION.jsoncounterproof evolve examples/traces/tenant_failure.json \
--experiment-manifest examples/discrimination_suite.json \
--surface policy \
--surface skill \
--surface prompt \
--out EVOLUTION_REVIEW.md \
--packet-out EVOLVED_PACKET.jsonCounterproof is allowed to return ambiguity.
No winner is better than a fake winner.
For the deeper architecture, see docs/COUNTERPROOF.md.
Counterproof keeps a runtime truth table instead of pretending roadmap items are finished.
counterproof auditThe current project includes tested paths for:
Regression Witness
Proof Integrity Guard
raw JSON / JSONL trace ingestion
real subprocess replay
Probe Contracts
multi-intervention discrimination
pre-registered PASS / FAIL predictions
fitness vs diagnostic case semantics
reviewed adapter binding
structured probe results
Proof Receipt + source fingerprints
clean wheel installation
browser interaction smoke tests
And it still has clear research gaps:
live agent-framework trace adapters
automatic trustworthy domain-test synthesis
arbitrary world snapshot / restore
generic live mutation executors
shadow / canary rollout
automatic mutation rollback
That distinction is intentional.
If the evidence cannot be reproduced, it should not become a stronger claim.
CI and test configuration are evidence-producing machinery.
Ambiguity is a valid result.
That is the point.
skill_factory/evolution/
├── trace.py
├── replay.py
├── discriminate.py
├── probe_planner.py
├── adapter_binding.py
├── receipt.py
├── capabilities.py
├── report.py
└── cli.py
actions/
└── witness/
└── action.yml
site/
├── index.html
├── standalone.html
├── app.js
├── styles.css
└── data/
examples/
├── traces/
├── replay/
└── *_suite.json
A 1280×640 social card is included at:
assets/social-preview.svg
Use it for the repository social preview, launch posts, HN/X screenshots, or release notes.
The most valuable contributions are not “more AI.”
They are things that make proof harder to fake:
- a runner adapter that preserves before/after semantics;
- a real PR fixture that breaks an assumption;
- an integrity rule with low false-positive cost;
- a reproduction where Counterproof overclaims;
- a better discriminating probe;
- an adapter for a real agent runtime.
If Counterproof labels weak evidence as strong evidence, that is a bug.
Apache-2.0.
Counterproof
Claim nothing you can't replay.