Repository navigation
fix(skill-creator): isolate trigger-eval command files from the live project registry - #1261
alvingarcia wants to merge 3 commits into
Conversation
…project registry
run_single_query wrote {skill}-skill-{8hex}.md into the nearest real
project's .claude/commands/ (find_project_root + cwd=project_root), so
during the parallel eval window every concurrent Claude Code session
rooted at that project saw the synthetic variants in its skill registry
and typeahead (up to num_workers at once, default 10).
Write each variant under a per-query tempfile.mkdtemp() root instead and
run claude -p with cwd=eval_root — commands are discovered from cwd's
.claude/, so the variant is visible to the eval subprocess only. The
finally now rmtrees the whole temp root (replacing the per-file unlink).
Fixes anthropics#1260
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
…ensics A worker killed mid-query (SIGKILL) never reaches the rmtree, so a stranded skill-eval-* tempdir should identify which skill/run left it. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
Friendly ping on this one — it's been open ~3.5 weeks with no maintainer review. It's a small (+15/−8, single file), self-contained fix for the registry-pollution bug in #1260, with a green smoke test (variant discovered/triggered from the isolated root, live project Happy to fold in two optional refinements from that #1260 discussion if you'd like them in this PR:
I can also remove the now-unused |
Emit an info-level line when each trigger-eval variant is created under its throwaway tempdir root, making the hermetic-by-design intent explicit in eval logs. Uses a module logger (no handler configured), so it stays silent by default and only surfaces when the caller enables logging. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
Update: I've folded in both refinements from the #1260 discussion so the diff is final and ready to review.
Still a small, single-file change ( I deliberately left the dead-code cleanup of |
Summary
Fixes #1260 — the trigger eval writes its synthetic
{skill}-skill-{8hex}.mdcommand filesinto the user's live project
.claude/commands/(viafind_project_root()+cwd=project_root),so during the parallel eval window (default 10 workers) every concurrent Claude Code session
rooted at that project sees up to ~10 phantom hash-suffixed skill variants in its registry,
typeahead, and system-reminder skill listings.
This change isolates each eval variant in a per-query throwaway project root:
tempfile.mkdtemp()/.claude/commands/claude -psubprocess runs withcwd=eval_root(it discovers commands from its cwd's.claude/; global auth/config in~/.claudeare unaffected by cwd)finallyremoves the whole temp root (shutil.rmtree, replacing the per-fileunlink)The eval becomes hermetic: no live-registry pollution at any worker count, and project-level
commands/skills in the user's repo can no longer confound the trigger measurement. (This also
resolves the project-level case of anthropics/claude-plugins-official#632; skills installed
globally in
~/.claude/skills/are still loaded regardless of cwd and would need a HOMEoverride — out of scope here.)
run_loop.pyinherits the fix — it delegates torun_eval().find_project_root()and theproject_rootparameter are now unused for command placement; left intact to keep this diffminimal, happy to remove them here or in a follow-up if preferred.
Test plan
python3 -m py_compile scripts/run_eval.pycode reproducibly wrote variants into the project's
.claude/commands/):1-query eval set,
--num-workers 1 --runs-per-query 1→ 1/1 PASS (variant discovered andtriggered from the isolated root), the live project's
.claude/commands/was never created,and no
skill-eval-*temp dirs were left behind.🤖 Generated with Claude Code