Skip to content

fix(mcp-builder): serialize TextContent results, update retired default model - #1588

Open
Xsidz wants to merge 1 commit into
anthropics:mainfrom
Xsidz:fix/mcp-builder-eval-textcontent-serialization
Open

Xsidz wants to merge 1 commit into
anthropics:mainfrom
Xsidz:fix/mcp-builder-eval-textcontent-serialization

Conversation

@Xsidz

@Xsidz Xsidz commented Aug 16, 2026

Copy link
Copy Markdown

What

Two bugs in evaluation.py that cause every evaluation run to score 0/N regardless of server quality.

Bug 1 — TextContent not JSON serializable (line 117)

connection.call_tool() returns list[mcp.types.TextContent] (pydantic objects). The original code passed this to json.dumps(), which raises TypeError: Object of type TextContent is not JSON serializable on every call. The try/except around it silently converted the TypeError into a fabricated "Error executing tool ..." message fed back to the model, which then answered NOT_FOUND on every task.

Fix: for list results, join .text from any TextContent blocks; fall back to str() if none have .text. dict results keep json.dumps; everything else uses str().

Bug 2 — Retired default model (lines 223, 324)

claude-3-7-sonnet-20250219 returns 404 not_found_error, so a stock run dies before the serialization bug is even reached.

Fix: default to claude-haiku-4-5 in both run_evaluation() and the argparse default.

Files changed

  • skills/mcp-builder/scripts/evaluation.py — 1 file, 11 insertions, 3 deletions

Test plan

  • Tool call against a live MCP server returns text content, not a fabricated error
  • python evaluation.py eval.xml (no -m flag) no longer 404s
  • dict-returning tools still serialize with json.dumps

Fixes #1390

…lt model

evaluation.py line 117 passed list[TextContent] to json.dumps, which raises
TypeError because TextContent is a pydantic object. The surrounding
try/except silently converted every tool call into a fabricated error, so
the agent saw nothing but errors and evaluation scored 0/N regardless of
server quality.

Fix: extract .text from TextContent blocks for list results; fall through
to str() when no .text attributes are present. dict results still use
json.dumps; everything else uses str().

Also updates the hardcoded default model from claude-3-7-sonnet-20250219
(retired, returns 404) to claude-haiku-4-5 in both run_evaluation() and
the argparse default. Stock runs no longer die before the eval starts.

Fixes anthropics#1390
Copilot AI lite review requested due to automatic review settings August 16, 2026 12:23

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

mcp-builder: evaluation.py scores 0/N against any real MCP server (TextContent not JSON serializable, swallowed into fabricated tool errors)

2 participants