Repository navigation
feat(skills): add self-audit — mechanical verification + four-dimension reasoning quality gate (v1.3.0) - #1367
YuhaoLin2005 wants to merge 14 commits into
Conversation
…rences/patterns.md + evals/evals.json
…ead See-Also, clarify Hard Constraints, expand Verification
…for GPG compliance
|
This is a reasoning quality gate, and it fills a real gap. The existing Why this belongs in
|
…k) + requirement traceability in Completeness + de-branded Background section with scope note
|
@xg-gh-25 Thank you for this review — genuinely. You spotted the exact structural weakness in two sentences that I had walked past for weeks. Both suggestions are correct, and I've implemented them in the latest commit. What changed:
What I learned from studying T-CBB: After your comment, I read through the full Autonomous-Pipeline-Design.md. A few observations from a newcomer:
A question if you have time: Your 45 RP patterns are organized by category. In your experience, which categories trigger most frequently at the pipeline handoff boundary (DELIVER's adversarial gate), and which ones tend to be false-positive-prone? I'm building a smaller pattern catalog from session-level failures and trying to calibrate which patterns are high-signal vs. noise. One scope note I added to the skill: This skill runs within the same agent session (same trust boundary). Multi-agent pipeline architectures deploy equivalent gates where reviewer and producer are independent. I noted this distinction in the Background section so readers understand the architectural difference. If this framing doesn't accurately reflect how T-CBB's gate independence works, please correct me. |
… review Six changes synthesized from Ken Thompson + Linus Torvalds + Steve Jobs panel: 1. Process: remove ambiguous 'COMPLETE task' step, renumber to 0-3 flow 2. Priority bridge: explain why question order ≠ priority order, drop soft qualifier 3. When to Use: replace '3+ file edits' with concrete triggers (new logic, architectural decisions) 4. Background: replace defensive 'independently convergent, not derived' with neutral phrasing 5. Red Flags: split into Process (audit skipped) vs Content (audit shallow) groups 6. See Also: intentionally not adding shipping-and-launch (doesn't exist in this repo — caught by Jobs during review)
…ergence framing, new failure modes and rationalizations - Remove version/tags from frontmatter (not used by this repo) - Expand Priority Order with damage-severity lexicographic chain - Reframe Background as independent convergence (site SwarmAI T-CBB) - Add failure mode: file exists but content is empty/placeholder - Add rationalizations: "reviewed it mentally" and "tool said it wrote the file" - Reorder four questions to priority order (Honesty first) - Strengthen Step 0 with content verification (not just existence check) - Enhanced description with specific triggering contexts
…rsion/tags) - test_no_version_field: assert version NOT present (repo uses name+description only) - test_no_tags_field: assert tags NOT present (repo uses name+description only) - test_description_length: raise limit to 250 chars (skill-creator recommends pushy descriptions with triggering contexts)
Replace `code-reviewer` + `security-review` (do not exist in anthropics/skills) with `skill-creator` + `webapp-testing` (verified by get_file_contents). Expert review F1 finding — cross-references must be mechanically verified against target repo contents.
… scope, STUCK clarity - Fix description length (251→246 chars, now ≤250) - Reorder: When to Use before Hard Constraints (logical flow) - Remove confusing logical-workflow-order blockquote note - Add sub-agent scope to Step 0 mechanical check - Clarify STUCK threshold: mention first fix attempt
What this skill doesSelf-audit — verifies AI output before delivery in two steps: Step 0 (mechanical): Uses file-read tool to confirm every claimed output file actually exists and has non-trivial content. If an agent says "I wrote report.md" but the file doesn't exist or contains only placeholder text — hard fail. This catches write-tool false positives that reasoning alone can't detect. Also covers sub-agent-produced files. Steps 1-4 (reasoning audit): Four questions in damage-severity priority order:
Lexicographic rule: never fix Groundedness while Honesty is failing. The four-dimension taxonomy (H/C/C/G) is validated by independent convergence — the same categories emerged separately in SwarmAI's T-CBB multi-agent pipeline and in this single-agent self-audit context. Tests — 22/22 passingRun time: 0.08s. No failures, no skips. Changes in latest commit (417e5ee)
How this fits the repoExisting skills cover code correctness and security. None cover reasoning quality. This adds the verification dimension — tests verify code, self-audit verifies thinking. Pairs naturally with skill-creator (create to audit cycle) and webapp-testing (verify behavior to audit the reasoning). |
What this is
A skill that audits AI output before delivery — mechanical file verification first, then four-dimension reasoning audit in damage-severity priority order. Universal — works with any project, any tech stack, any model.
Step 0 (mechanical): Verify every claimed output file exists AND has non-trivial content via file-read tool. Catches "I wrote report.md" when no file was written — a failure mode invisible to reasoning. Now includes files produced by sub-agents.
Steps 1-4 (reasoning, in priority order):
Why this order: Ordered by damage severity, not failure frequency. Lexicographic rule: never fix Groundedness while Honesty is failing.
How it complements this repo
The repo's existing skills cover code correctness, security review, and deployment safety. None check reasoning quality — whether the AI's thinking itself was sound. Tests verify code. Nothing verifies reasoning. This fills that gap by adding a verification dimension the ecosystem currently lacks.
Specifically pairs with:
skill-creator— "create → self-audit before shipping" cyclewebapp-testing— test behavior → audit the reasoning behind resultsWhy these four dimensions?
The taxonomy (Honesty, Completeness, Consistency, Groundedness) is validated by independent convergence — equivalent dimensions emerged separately in multi-agent pipelines (SwarmAI T-CBB) and single-agent self-audit. Different architectures independently rediscovered the same categories when asking "how do we verify agent output?"
What differs between the two contexts is the trust boundary: multi-agent pipelines have adversarial pressure from independent reviewers; self-audit runs inside the same trust boundary. This makes Honesty the critical dimension (misrepresentation within-session is harder to flag) and Step 0 essential (file-system read-back is the one thing an agent can't fabricate).
Design decisions