Skip to content

mcp-builder: update evaluation.py default model to claude-sonnet-5 - #1724

Open
ExpertVagabond wants to merge 1 commit into
anthropics:mainfrom
ExpertVagabond:fix/mcp-builder-eval-model-default
Open

ExpertVagabond wants to merge 1 commit into
anthropics:mainfrom
ExpertVagabond:fix/mcp-builder-eval-model-default

Conversation

@ExpertVagabond

Copy link
Copy Markdown

What

skills/mcp-builder/scripts/evaluation.py defaulted to claude-3-7-sonnet-20250219 in both the run_evaluation() signature and the -m/--model argparse option, and reference/evaluation.md recommended that same snapshot as the "more capable" alternative under troubleshooting. This PR points both defaults at claude-sonnet-5, updates the usage text in reference/evaluation.md to match, and rewords the troubleshooting bullet to suggest a step up from the default (-m claude-opus-5) rather than the default itself. It also prints the model id in the run header so a score is always attributable to a baseline, and updates the --help example that passed claude-3-5-sonnet-20241022 so it no longer points at a retired snapshot.

Why

The harness exists to tell an MCP author whether their tool descriptions are good enough for a current agent to use, and scoring against a three-generation-old model measures the wrong thing. claude-sonnet-5 is the id this repo's own skills/claude-api/shared/model-migration.md lists as the replacement for claude-3-7-sonnet-20250219, and model ids in this repo are bare aliases with no date suffix, so the default will not break when an old snapshot is retired.

Fixes #1716

Verification

  • python3 -m py_compile skills/mcp-builder/scripts/evaluation.py exits 0.
  • evaluation.py --help (argparse, with connections stubbed) prints -m, --model MODEL Claude model to use (default: claude-sonnet-5) and exits 0.
  • grep -rn "claude-3-7\|claude-3-5" skills/mcp-builder returns no matches.
  • Note: running --help directly with mcp 2.1.1 installed fails in connections.py (streamablehttp_client import). That reproduces identically on main and is unrelated to this change.

The evaluation harness defaulted to claude-3-7-sonnet-20250219 in both
the run_evaluation() signature and the -m/--model argparse option, so
anyone running evaluation.py without -m scored their MCP server against
a model several generations old. The troubleshooting section in
reference/evaluation.md recommended that same snapshot as the "more
capable" alternative, so following the advice changed nothing.

Changes:
- Point both defaults at claude-sonnet-5, the current Sonnet-class id
  that skills/claude-api/shared/model-migration.md lists as the
  replacement for claude-3-7-sonnet-20250219. Model ids in this repo are
  bare aliases with no date suffix.
- Update the usage text in reference/evaluation.md to match.
- Reword the timeout troubleshooting bullet to suggest a step up from
  the default (-m claude-opus-5) instead of the default itself.
- Print the model id in the run header so every score is attributable
  to a baseline.
- Update the --help example that passed claude-3-5-sonnet-20241022 to
  use claude-opus-5, so the example does not point at a retired snapshot.

Fixes anthropics#1716
@ExpertVagabond

Copy link
Copy Markdown
Author

Bumping this since #1716 is still open and the diff is small.

What it does: evaluation.py and reference/evaluation.md both defaulted to claude-3-7-sonnet-20250219, so the harness scored MCP tool descriptions against a three-generation-old model, and the low-score troubleshooting advice pointed back at that same snapshot. This moves both defaults to claude-sonnet-5, repoints the troubleshooting bullet at claude-opus-5 so it suggests a step up rather than the default, prints the model id in the run header so a score is attributable to a baseline, and drops the retired claude-3-5-sonnet-20241022 from the --help example.

The ids are bare aliases with no date suffix, matching what this repo's own skills/claude-api/shared/model-migration.md lists as the replacement, so the default won't break on the next snapshot retirement.

Still merges clean, no conflicts. py_compile passes and --help shows the new default.

Happy to rebase or split it up if that makes review easier.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

mcp-builder: evaluation.py defaults to claude-3-7-sonnet-20250219, and the low-score advice points back at that same model

1 participant