I build small practical tools and AI-agent evaluation projects.
A compact open-source benchmark for testing whether AI coding agents can find business-logic bugs from written product rules.
- Product-rule source of truth in
SPEC.md - Intentionally flawed payment, order, wallet, webhook, user, and pricing code
- Baseline tests documenting seeded defects
- Maintainer answer key and 100-point scoring rubric
- Landing page: https://dmatut7.github.io/shoppay-audit-benchmark/
If you are testing Codex-style agents, try running an audit against it and compare the result with the scoring guide.