Skip to content

Standards boundary note: in-toto AI provenance vs claim proof #131

Description

@hippoley

Why this note exists

I tried to contribute two narrow reviewer-evidence boundaries directly to:

The connected GitHub integration cannot write to that repository (403 Resource not accessible by integration), so I am recording the concrete interoperability notes here and mentioning the proposal authors rather than opening parallel duplicate issues.

@csheargm @carmithersh — these are intended as small boundary observations from real reviewer-side replay cases, not requests to expand either predicate.


1. For in-toto/attestation#604

Current terminology already distinguishes:

self-reported acceptance
vs
re-executed acceptance

One additional boundary showed up in a real AI-assisted PR:

submitted / producer judge
  -> green

same submitted judge re-executed
  -> still green

authoritative product parser
  -> rejects the claimed-valid fixture

Case:
anthropics/claude-code#89404

Reviewer-facing artifact:
https://github.com/hippoley/CounterProof/blob/main/reality/claude-code-89404/ORACLE_DISAGREEMENT.md

The useful invariant is:

Re-executing an acceptance contract proves the result of that contract under the verifier; it does not prove that the contract is the authoritative oracle for every downstream product claim.

I do not think the Generation predicate should absorb claim-review semantics. A non-goal/composition note is probably enough:

Generation attestation
  -> what was generated + embedded acceptance contract

verifier re-execution
  -> did that exact contract pass?

claim-level review
  -> is that contract the right oracle for the claim?

2. For in-toto/attestation#600

The proposed Agentic Process Evidence is strong for answering:

which agent/model/tools ran
which context artifacts were present
which humans were accountable
which logs belong to the process
did the process complete?

A separate reviewer question is:

What does that execution actually establish about a specific software claim?

A real case:
aboutcode-org/scancode.io#2207

The exact submitted Django test:

HEAD -> PASS
BASE + exact same submitted test -> FAIL
because BASE still requires input_location

Reviewer handoff:
https://github.com/hippoley/CounterProof/blob/main/reality/scancode-2207/MAINTAINER_HANDOFF.md

That is stronger than result: COMPLETED, but still deliberately narrower than “the PR should merge.”

The boundary I would keep is:

process evidence
  -> what happened during the agentic process

candidate evidence
  -> which exact candidate produced the result

claim evidence
  -> what property the evidence establishes

oracle evidence
  -> whether the judge is authoritative for that claim

Again, I would not add claim-verification fields to the process predicate.

A small non-goal may be enough:

result: COMPLETED attests process completion under this evidence model. It does not by itself attest that a specific software claim, requirement, or fix has been proven.


Question

Do either of these boundaries conflict with the intended scope of #604 / #600?

If not, I will treat them as downstream composition rules in CounterProof rather than proposing new in-toto fields.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions