Repository navigation
Conversation
…lt model evaluation.py line 117 passed list[TextContent] to json.dumps, which raises TypeError because TextContent is a pydantic object. The surrounding try/except silently converted every tool call into a fabricated error, so the agent saw nothing but errors and evaluation scored 0/N regardless of server quality. Fix: extract .text from TextContent blocks for list results; fall through to str() when no .text attributes are present. dict results still use json.dumps; everything else uses str(). Also updates the hardcoded default model from claude-3-7-sonnet-20250219 (retired, returns 404) to claude-haiku-4-5 in both run_evaluation() and the argparse default. Stock runs no longer die before the eval starts. Fixes anthropics#1390
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Two bugs in
evaluation.pythat cause every evaluation run to score 0/N regardless of server quality.Bug 1 — TextContent not JSON serializable (line 117)
connection.call_tool()returnslist[mcp.types.TextContent](pydantic objects). The original code passed this tojson.dumps(), which raisesTypeError: Object of type TextContent is not JSON serializableon every call. Thetry/exceptaround it silently converted theTypeErrorinto a fabricated"Error executing tool ..."message fed back to the model, which then answered NOT_FOUND on every task.Fix: for list results, join
.textfrom any TextContent blocks; fall back tostr()if none have.text. dict results keepjson.dumps; everything else usesstr().Bug 2 — Retired default model (lines 223, 324)
claude-3-7-sonnet-20250219returns404 not_found_error, so a stock run dies before the serialization bug is even reached.Fix: default to
claude-haiku-4-5in bothrun_evaluation()and the argparse default.Files changed
skills/mcp-builder/scripts/evaluation.py— 1 file, 11 insertions, 3 deletionsTest plan
python evaluation.py eval.xml(no-mflag) no longer 404sFixes #1390