Skip to content

feat(skills): add self-audit — mechanical verification + four-dimension reasoning quality gate (v1.3.0) - #1367

Open
YuhaoLin2005 wants to merge 14 commits into
anthropics:mainfrom
YuhaoLin2005:self-audit
Open

YuhaoLin2005 wants to merge 14 commits into
anthropics:mainfrom
YuhaoLin2005:self-audit

Conversation

@YuhaoLin2005

@YuhaoLin2005 YuhaoLin2005 commented Jun 28, 2026 •

Copy link
Copy Markdown

What this is

A skill that audits AI output before delivery — mechanical file verification first, then four-dimension reasoning audit in damage-severity priority order. Universal — works with any project, any tech stack, any model.

Step 0 (mechanical): Verify every claimed output file exists AND has non-trivial content via file-read tool. Catches "I wrote report.md" when no file was written — a failure mode invisible to reasoning. Now includes files produced by sub-agents.

Steps 1-4 (reasoning, in priority order):

  1. Honesty — Am I being honest about what wasn't done? (Check first — if the agent is lying, the other three dimensions check fiction.)
  2. Completeness — Did I answer every request? (Missing a requirement silently drops work.)
  3. Consistency — Did I contradict myself or the rules?
  4. Groundedness — Did I show evidence, or just claim?

Why this order: Ordered by damage severity, not failure frequency. Lexicographic rule: never fix Groundedness while Honesty is failing.

How it complements this repo

The repo's existing skills cover code correctness, security review, and deployment safety. None check reasoning quality — whether the AI's thinking itself was sound. Tests verify code. Nothing verifies reasoning. This fills that gap by adding a verification dimension the ecosystem currently lacks.

Specifically pairs with:

  • skill-creator — "create → self-audit before shipping" cycle
  • webapp-testing — test behavior → audit the reasoning behind results

Why these four dimensions?

The taxonomy (Honesty, Completeness, Consistency, Groundedness) is validated by independent convergence — equivalent dimensions emerged separately in multi-agent pipelines (SwarmAI T-CBB) and single-agent self-audit. Different architectures independently rediscovered the same categories when asking "how do we verify agent output?"

What differs between the two contexts is the trust boundary: multi-agent pipelines have adversarial pressure from independent reviewers; self-audit runs inside the same trust boundary. This makes Honesty the critical dimension (misrepresentation within-session is harder to flag) and Step 0 essential (file-system read-back is the one thing an agent can't fabricate).

Design decisions

  • Step 0 before reasoning: Mechanical verification catches what reasoning can't — a missing file has no reasoning to audit. Tools can report success without writing content.
  • Common rationalizations table: Names the bypass mechanisms agents use ("simple change, no audit", "checked as I went", "the tool said it wrote the file") and rebuts each one preemptively.
  • Hard Constraints: Never fabricate findings, never expose sensitive data, never block on subjective grounds.
  • STUCK threshold: Failing 3+ of 4 questions after one fix attempt → escalate to user instead of looping.

@xg-gh-25

Copy link
Copy Markdown

This is a reasoning quality gate, and it fills a real gap. The existing code-reviewer checks syntax correctness. security-review checks vulnerabilities. Neither checks reasoning correctness — whether the agent's thinking itself was sound. This PR targets that layer.

Why this belongs in anthropics/skills

Tests verify code. Nothing verifies reasoning. This PR adds four mechanical checks before delivery:

  1. Completeness — Did I answer every request?
  2. Consistency — Did I contradict myself or the rules?
  3. Groundedness — Did I show evidence, or just claim?
  4. Honesty — Am I being honest about what wasn't done?

These aren't subjective. They're auditable. They catch the failure modes where the agent believes it succeeded but the output is incomplete, contradictory, or unsupported.

How SwarmAI uses this pattern (T-CBB: Coding as Black Box)

We run a similar four-dimension gate at the pipeline handoff boundary — before any agent output enters the next stage. Ours is:

  1. Output-artifact present — Did the agent produce the declared deliverable? (Not just claim it did.)
  2. Requirement coverage — Does the output address every line in the input spec?
  3. Contradiction scan — Does the output contradict the input constraints or its own statements?
  4. Evidence audit — Does the output cite sources/logs/measurements, or just assert?

This is a mechanical quality gate, not a human judgment call. It runs before the "Done" status is allowed. It catches the failure mode where the agent feels it succeeded but the artifact is wrong.

The root cause it solves: agents don't naturally audit their own reasoning. They produce output, feel satisfied, and declare success — even when the output is incomplete or contradicts the spec. A mechanical gate catches this before it cascades.

Suggested next step

The PR is solid. The four questions are the right dimensions. Two additions would make it universal:

  1. Mechanical artifact check — Before the four reasoning questions, confirm the declared output artifact exists and is readable. (Catches "I generated report.md" when no file was written.)
  2. Requirement traceability — If the input was a spec or feature list, map each requirement to the output section that addresses it. Flag unmapped requirements.

These are cheap to add and prevent the failure mode where the reasoning audit passes but the artifact is still wrong.


Reasoning quality gates are a T-CBB pattern from SwarmAI. Discussion: T-CBB: Coding as Black Box

…k) + requirement traceability in Completeness + de-branded Background section with scope note
@YuhaoLin2005

Copy link
Copy Markdown
Author

@xg-gh-25 Thank you for this review — genuinely. You spotted the exact structural weakness in two sentences that I had walked past for weeks. Both suggestions are correct, and I've implemented them in the latest commit.

What changed:

  1. Mechanical artifact check (Step 0) — Before the four reasoning questions, the skill now requires verifying claimed output files exist via the file-read tool. Missing file → hard fail, no reasoning audit proceeds. This catches the failure mode where the agent claims "I generated report.md" but never wrote it.

  2. Requirement traceability — The Completeness dimension now includes: "If the input was a spec or numbered requirements, map each requirement to the output section that addresses it. Flag unmapped requirements." This catches the failure mode where the agent forgets a requirement and therefore never lists it in its own self-check.

What I learned from studying T-CBB:

After your comment, I read through the full Autonomous-Pipeline-Design.md. A few observations from a newcomer:

  • The convergence is striking: four-dimension gate (C/C/G/H ↔ Output/Requirement/Contradiction/Evidence), mechanical enforcement, binary push-ready, knowledge compounding loop. Two systems arriving at the same pattern from different starting points.
  • OP8 (Config Consistency) directly validates a separate finding I had about AI config file format drift — I had observed that mixing formats across config files degrades LLM behavior, and seeing it listed as a system-level invariant in a production pipeline was confirming.
  • The distinction between stages and gates — stages are "what we do," gates are "how we know it's good enough to continue" — clarified something I had been conflating in my own architecture.

A question if you have time:

Your 45 RP patterns are organized by category. In your experience, which categories trigger most frequently at the pipeline handoff boundary (DELIVER's adversarial gate), and which ones tend to be false-positive-prone? I'm building a smaller pattern catalog from session-level failures and trying to calibrate which patterns are high-signal vs. noise.

One scope note I added to the skill:

This skill runs within the same agent session (same trust boundary). Multi-agent pipeline architectures deploy equivalent gates where reviewer and producer are independent. I noted this distinction in the Background section so readers understand the architectural difference. If this framing doesn't accurately reflect how T-CBB's gate independence works, please correct me.

… review

Six changes synthesized from Ken Thompson + Linus Torvalds + Steve Jobs panel:

1. Process: remove ambiguous 'COMPLETE task' step, renumber to 0-3 flow
2. Priority bridge: explain why question order ≠ priority order, drop soft qualifier
3. When to Use: replace '3+ file edits' with concrete triggers (new logic, architectural decisions)
4. Background: replace defensive 'independently convergent, not derived' with neutral phrasing
5. Red Flags: split into Process (audit skipped) vs Content (audit shallow) groups
6. See Also: intentionally not adding shipping-and-launch (doesn't exist in this repo — caught by Jobs during review)
…ergence framing, new failure modes and rationalizations

- Remove version/tags from frontmatter (not used by this repo)
- Expand Priority Order with damage-severity lexicographic chain
- Reframe Background as independent convergence (site SwarmAI T-CBB)
- Add failure mode: file exists but content is empty/placeholder
- Add rationalizations: "reviewed it mentally" and "tool said it wrote the file"
- Reorder four questions to priority order (Honesty first)
- Strengthen Step 0 with content verification (not just existence check)
- Enhanced description with specific triggering contexts
…rsion/tags)

- test_no_version_field: assert version NOT present (repo uses name+description only)
- test_no_tags_field: assert tags NOT present (repo uses name+description only)
- test_description_length: raise limit to 250 chars (skill-creator recommends pushy descriptions with triggering contexts)
@YuhaoLin2005 YuhaoLin2005 changed the title feat(skills): add self-audit — four-dimension reasoning quality gate before delivery (replaces #1361) feat(skills): add self-audit — mechanical verification + four-dimension reasoning quality gate (v1.3.0) Jul 2, 2026
Replace `code-reviewer` + `security-review` (do not exist in anthropics/skills) with `skill-creator` + `webapp-testing` (verified by get_file_contents). Expert review F1 finding — cross-references must be mechanically verified against target repo contents.
YuhaoLin2005

This comment was marked as outdated.

YuhaoLin2005

This comment was marked as outdated.

… scope, STUCK clarity

- Fix description length (251→246 chars, now ≤250)
- Reorder: When to Use before Hard Constraints (logical flow)
- Remove confusing logical-workflow-order blockquote note
- Add sub-agent scope to Step 0 mechanical check
- Clarify STUCK threshold: mention first fix attempt
YuhaoLin2005

This comment was marked as outdated.

@YuhaoLin2005

Copy link
Copy Markdown
Author

What this skill does

Self-audit — verifies AI output before delivery in two steps:

Step 0 (mechanical): Uses file-read tool to confirm every claimed output file actually exists and has non-trivial content. If an agent says "I wrote report.md" but the file doesn't exist or contains only placeholder text — hard fail. This catches write-tool false positives that reasoning alone can't detect. Also covers sub-agent-produced files.

Steps 1-4 (reasoning audit): Four questions in damage-severity priority order:

  1. Honesty — Misrepresenting what was done poisons everything downstream, check first
  2. Completeness — Missing a requirement silently drops work the user must rediscover
  3. Consistency — Contradictions erode trust but don't lose information
  4. Groundedness — Evidence quality can be improved incrementally; weakest damage

Lexicographic rule: never fix Groundedness while Honesty is failing.

The four-dimension taxonomy (H/C/C/G) is validated by independent convergence — the same categories emerged separately in SwarmAI's T-CBB multi-agent pipeline and in this single-agent self-audit context.

Tests — 22/22 passing

TestSkillExists::test_skill_directory_exists PASSED
TestSkillExists::test_skill_md_exists PASSED
TestFrontmatter::test_frontmatter_valid_yaml PASSED
TestFrontmatter::test_has_name_field PASSED
TestFrontmatter::test_name_matches_directory PASSED
TestFrontmatter::test_has_description_field PASSED
TestFrontmatter::test_description_length PASSED
TestFrontmatter::test_description_not_empty PASSED
TestRequiredSections::test_has_priority_order_section PASSED
TestRequiredSections::test_has_when_to_use_section PASSED
TestRequiredSections::test_has_hard_constraints_section PASSED
TestRequiredSections::test_has_step_0_section PASSED
TestRequiredSections::test_has_four_questions_section PASSED
TestRequiredSections::test_has_process_section PASSED
TestRequiredSections::test_has_failure_modes_section PASSED
TestRequiredSections::test_has_common_rationalizations_section PASSED
TestRequiredSections::test_has_red_flags_section PASSED
TestRequiredSections::test_has_verification_section PASSED
TestBodyQuality::test_no_trailing_whitespace PASSED
TestBodyQuality::test_step_0_mentions_file_read_tool PASSED
TestBodyQuality::test_four_questions_ordered_by_priority PASSED
TestBodyQuality::test_lexicographic_rule_mentioned PASSED

Run time: 0.08s. No failures, no skips.

Changes in latest commit (417e5ee)

  • Description: 251 to 246 chars (was over the 250-char test limit)
  • Section order: "When to Use" now appears before "Hard Constraints"
  • Step 0 scope: Explicitly covers files produced by sub-agents/spawned tasks
  • STUCK condition: Clarified as "failing 3+ of 4 questions after first fix attempt"
  • Removed confusing blockquote that implied a separate "logical workflow order"

How this fits the repo

Existing skills cover code correctness and security. None cover reasoning quality. This adds the verification dimension — tests verify code, self-audit verifies thinking. Pairs naturally with skill-creator (create to audit cycle) and webapp-testing (verify behavior to audit the reasoning).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants