Skip to content
Dmatut7Public

About

Dmatut7 GitHub profile README

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Hi, I'm Dmatut7

I build small practical tools and AI-agent evaluation projects.

Featured project

A compact open-source benchmark for testing whether AI coding agents can find business-logic bugs from written product rules.

  • Product-rule source of truth in SPEC.md
  • Intentionally flawed payment, order, wallet, webhook, user, and pricing code
  • Baseline tests documenting seeded defects
  • Maintainer answer key and 100-point scoring rubric
  • Landing page: https://dmatut7.github.io/shoppay-audit-benchmark/

If you are testing Codex-style agents, try running an audit against it and compare the result with the scoring guide.

About

Dmatut7 GitHub profile README

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors