Skip to content
Public template

About

Prove AI-assisted PR claims with replayable BASE→HEAD evidence, scoped claim matrices, oracle checks, and auditable receipts.

Topics

Resources

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

Counterproof

Your coding agent says it fixed the bug. Prove it.

Behavioral proof for agent-generated pull requests.

CI Python License No LLM No API key

Open Counterproof Proof Lab

Click the panel. Break the proof. Change the judge. See what survives.

Works beside Claude Code · Codex · Copilot · Cursor · PR-Agent · human-written PRs — Counterproof verifies the evidence, not the author.


A green CI run proves that your code passes now.

It does not prove that the regression test added by the same coding agent would have caught the bug before the fix.

Counterproof asks that missing question.

PR code + PR test         → PASS
old code + the same test  → FAIL
test / CI judge unchanged → CLEAN
                             ↓
                     REGRESSION WITNESSED

That turns:

“the agent says it fixed the bug”

into:

this exact test behaves differently before and after the fix.


See it before you install it

The browser experience lets you play with different evidence situations instead of reading another architecture diagram.

REAL REGRESSION
HEAD passes / BASE fails
→ strong before-vs-after evidence

WEAK TEST
HEAD passes / BASE also passes
→ the test does not witness the claimed fix

JUDGE CHANGED
the regression evidence exists
but CI / test machinery changed too
→ reviewer attention required

FULL SUITE
the full suite differs
but the changed test was not isolated
→ weaker evidence than an exact witness

The browser scenarios are fixtures. Real evidence comes from the CLI / GitHub Action.


Counterproof proof walkthrough

Click the walkthrough to open the live Proof Lab.


The fastest useful thing Counterproof does

Take tests changed in a pull request.

Run them on the PR.

Then replay the same tests against the pre-change code.

                         PR HEAD        BASE
same changed test          PASS          FAIL
                              \          /
                               \        /
                              WITNESSED

If the same test already passes on BASE, Counterproof does not manufacture a success story.

It says the proof is weak.


Drop it into a PR

The simplest production entry point on main is the Regression Witness Action:

name: Counterproof

on:
  pull_request:

permissions:
  contents: read
  pull-requests: write

jobs:
  proof:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          ref: ${{ github.event.pull_request.head.sha }}
          fetch-depth: 0

      - uses: actions/setup-python@v5
        with:
          python-version: "3.11"

      # Install your project dependencies first.
      - uses: hippoley/SkillFactory/actions/witness@main
        with:
          test-command: "python -m pytest -q {tests}"
          require-witness: "true"
          require-clean-integrity: "true"

Counterproof will:

1. find tests added or modified by the PR
2. run them on PR HEAD
3. create a detached worktree at BASE
4. overlay the PR test/support files
5. run the same evidence on old code
6. inspect whether the PR changed the judge
7. write one sticky proof comment

No hosted service.

No API key.

No LLM is required for this proof path.


It also checks whether the PR changed the judge

A passing test is weaker evidence if the same PR also weakens the system that evaluates it.

Counterproof's Proof Integrity Guard surfaces changes such as:

deleted test                         → review
new skip / xfail                     → review
continue-on-error: true              → review
pytest ... || true                   → review
pull_request trigger removed         → review
test / coverage / CI config changed  → surface explicitly

A finding does not mean “malicious PR.”

It means:

the evidence surface changed, so the reviewer should not treat the green result as independent evidence.


Why this matters now

Coding agents are getting very good at producing:

code
+ tests
+ green CI
+ a confident explanation

That is useful.

It also creates a new review problem:

the system proposing the fix can now help produce the evidence for its own fix.

Counterproof does not solve that by adding another model.

It adds a deterministic before/after experiment.

Git history
+
your existing test runner
+
the same evidence on both sides
=
something a reviewer can inspect

Not another AI reviewer

AI reviewers and Counterproof answer different questions.

AI reviewer Counterproof
Main question “Does this diff look suspicious?” “Does this evidence distinguish before from after?”
Core input code + model context Git history + tests
Main output suggestions / comments replayable behavioral evidence
LLM required usually no for PR proof
Changed test/CI judge not the core primitive explicitly surfaced
Can refuse a story model-dependent yes — weak/inconclusive evidence stays weak

Use Counterproof next to Claude Code, Codex, Copilot, Cursor, PR-Agent, or a human engineer.

It does not care who wrote the patch.


Run the playground locally

python -m pip install "git+https://github.com/hippoley/SkillFactory.git"

counterproof demo

Open:

http://127.0.0.1:8765

Or just use the public version:


The deeper runtime

Regression Witness is the smallest useful entry point.

Counterproof also contains an experimental runtime for falsifiable agent self-improvement.

Instead of asking only:

“what lesson should the agent remember?”

it asks:

which explanation survives an experiment, and what evidence earns the right to change behavior?

RAW TRACE
   ↓
Decision Capsule + Outcome Receipt
   ↓
typed evidence
   ↓
competing hypotheses
   ↓
Probe Contracts
   ↓
same cases × multiple interventions
   ↓
survived / falsified / inconclusive
   ↓
unique survivor?
  ↙           ↘
yes            no
 ↓              ↓
select          remain ambiguous
 ↓
Behavior Proof
 ↓
promotion gate

One-command proof

counterproof prove examples/traces/tenant_failure.json \
  --replay-manifest examples/replay_suite.json \
  --surface policy \
  --out BEHAVIOR_PROOF.md

Active discrimination

counterproof discriminate examples/traces/tenant_failure.json \
  --experiment-manifest examples/discrimination_suite.json \
  --surface policy \
  --surface skill \
  --surface prompt \
  --out DISCRIMINATION.md \
  --json-out DISCRIMINATION.json

Guarded evolution

counterproof evolve examples/traces/tenant_failure.json \
  --experiment-manifest examples/discrimination_suite.json \
  --surface policy \
  --surface skill \
  --surface prompt \
  --out EVOLUTION_REVIEW.md \
  --packet-out EVOLVED_PACKET.json

Counterproof is allowed to return ambiguity.

No winner is better than a fake winner.

For the deeper architecture, see docs/COUNTERPROOF.md.


What is real today?

Counterproof keeps a runtime truth table instead of pretending roadmap items are finished.

counterproof audit

The current project includes tested paths for:

Regression Witness
Proof Integrity Guard
raw JSON / JSONL trace ingestion
real subprocess replay
Probe Contracts
multi-intervention discrimination
pre-registered PASS / FAIL predictions
fitness vs diagnostic case semantics
reviewed adapter binding
structured probe results
Proof Receipt + source fingerprints
clean wheel installation
browser interaction smoke tests

And it still has clear research gaps:

live agent-framework trace adapters
automatic trustworthy domain-test synthesis
arbitrary world snapshot / restore
generic live mutation executors
shadow / canary rollout
automatic mutation rollback

That distinction is intentional.


The design rules

Claim nothing you can't replay.

If the evidence cannot be reproduced, it should not become a stronger claim.

A changed judge is part of the change.

CI and test configuration are evidence-producing machinery.

No winner is better than a fake winner.

Ambiguity is a valid result.

Evidence should survive outside the model that produced the patch.

That is the point.


Repository map

skill_factory/evolution/
├── trace.py
├── replay.py
├── discriminate.py
├── probe_planner.py
├── adapter_binding.py
├── receipt.py
├── capabilities.py
├── report.py
└── cli.py

actions/
└── witness/
    └── action.yml

site/
├── index.html
├── standalone.html
├── app.js
├── styles.css
└── data/

examples/
├── traces/
├── replay/
└── *_suite.json

Share it

A 1280×640 social card is included at:

assets/social-preview.svg

Use it for the repository social preview, launch posts, HN/X screenshots, or release notes.


Contributing

The most valuable contributions are not “more AI.”

They are things that make proof harder to fake:

  • a runner adapter that preserves before/after semantics;
  • a real PR fixture that breaks an assumption;
  • an integrity rule with low false-positive cost;
  • a reproduction where Counterproof overclaims;
  • a better discriminating probe;
  • an adapter for a real agent runtime.

If Counterproof labels weak evidence as strong evidence, that is a bug.


License

Apache-2.0.


Green is a state. Proof is a relationship between before and after.

Counterproof

Claim nothing you can't replay.

About

Prove AI-assisted PR claims with replayable BASE→HEAD evidence, scoped claim matrices, oracle checks, and auditable receipts.

Topics

Resources

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages