Repository navigation
mcp-builder: update evaluation.py default model to claude-sonnet-5 - #1724
ExpertVagabond wants to merge 1 commit into
Conversation
The evaluation harness defaulted to claude-3-7-sonnet-20250219 in both the run_evaluation() signature and the -m/--model argparse option, so anyone running evaluation.py without -m scored their MCP server against a model several generations old. The troubleshooting section in reference/evaluation.md recommended that same snapshot as the "more capable" alternative, so following the advice changed nothing. Changes: - Point both defaults at claude-sonnet-5, the current Sonnet-class id that skills/claude-api/shared/model-migration.md lists as the replacement for claude-3-7-sonnet-20250219. Model ids in this repo are bare aliases with no date suffix. - Update the usage text in reference/evaluation.md to match. - Reword the timeout troubleshooting bullet to suggest a step up from the default (-m claude-opus-5) instead of the default itself. - Print the model id in the run header so every score is attributable to a baseline. - Update the --help example that passed claude-3-5-sonnet-20241022 to use claude-opus-5, so the example does not point at a retired snapshot. Fixes anthropics#1716
|
Bumping this since #1716 is still open and the diff is small. What it does: evaluation.py and reference/evaluation.md both defaulted to claude-3-7-sonnet-20250219, so the harness scored MCP tool descriptions against a three-generation-old model, and the low-score troubleshooting advice pointed back at that same snapshot. This moves both defaults to claude-sonnet-5, repoints the troubleshooting bullet at claude-opus-5 so it suggests a step up rather than the default, prints the model id in the run header so a score is attributable to a baseline, and drops the retired claude-3-5-sonnet-20241022 from the --help example. The ids are bare aliases with no date suffix, matching what this repo's own skills/claude-api/shared/model-migration.md lists as the replacement, so the default won't break on the next snapshot retirement. Still merges clean, no conflicts. py_compile passes and --help shows the new default. Happy to rebase or split it up if that makes review easier. |
What
skills/mcp-builder/scripts/evaluation.pydefaulted toclaude-3-7-sonnet-20250219in both therun_evaluation()signature and the-m/--modelargparse option, andreference/evaluation.mdrecommended that same snapshot as the "more capable" alternative under troubleshooting. This PR points both defaults atclaude-sonnet-5, updates the usage text inreference/evaluation.mdto match, and rewords the troubleshooting bullet to suggest a step up from the default (-m claude-opus-5) rather than the default itself. It also prints the model id in the run header so a score is always attributable to a baseline, and updates the--helpexample that passedclaude-3-5-sonnet-20241022so it no longer points at a retired snapshot.Why
The harness exists to tell an MCP author whether their tool descriptions are good enough for a current agent to use, and scoring against a three-generation-old model measures the wrong thing.
claude-sonnet-5is the id this repo's ownskills/claude-api/shared/model-migration.mdlists as the replacement forclaude-3-7-sonnet-20250219, and model ids in this repo are bare aliases with no date suffix, so the default will not break when an old snapshot is retired.Fixes #1716
Verification
python3 -m py_compile skills/mcp-builder/scripts/evaluation.pyexits 0.evaluation.py --help(argparse, withconnectionsstubbed) prints-m, --model MODEL Claude model to use (default: claude-sonnet-5)and exits 0.grep -rn "claude-3-7\|claude-3-5" skills/mcp-builderreturns no matches.--helpdirectly withmcp2.1.1 installed fails inconnections.py(streamablehttp_clientimport). That reproduces identically onmainand is unrelated to this change.