Skip to content
#

skill-evaluation

Here are 43 public repositories matching this topic...

oh-my-knowledge

OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.

  • Updated Oct 8, 2026
  • TypeScript
skill-harness

How do you KNOW if a 'skill' is any good - deterministically? This runs controlled skill-vs-no-skill experiments. It doesn't just ask, “Did the skill work?” It asks “Did we actually give the skill to the model, did the comparison remain fair, what changed, what did it cost, and do we have enough evidence to say that confidently?”

  • Updated Oct 7, 2026
  • Python

Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, trigger quality, functional quality, regression protection, baseline value, model variance, rollout safety. Never gradients.

  • Updated Oct 8, 2026
  • TypeScript

Description Evidence-driven evaluation, benchmarking, and verified repair for Agent Skills across Codex, Claude Code, Gemini CLI, and Antigravity.

  • Updated Sep 6, 2026
  • Python
trigger-doctor

🩺 Behavioral testing for skill triggers — diagnose why a skill never fires (or fires too often), prescribe a fixed description, save a regression suite. One SKILL.md, 79 agents (Claude Code, Hermes, Codex, Gemini...)

  • Updated Oct 6, 2026
  • Python

SQS (Skill Quality Suite): linter, validator, security scanner and eval harness for AI Agent Skills (SKILL.md) on Claude Code, Codex, Cursor and Gemini CLI - and it tells you which skill to fix or write next from your own sessions

  • Updated Sep 26, 2026
  • Python

Add this topic to your repo

To associate your repository with the skill-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more