Repository navigation
Images in tool results bypass model modality check (text-only models receive base64 images) #39824
Description
Activity
github-actions commented on Jul 31, 2026
This issue might be a duplicate of existing issues. Please check:
- GLM-5.2 session broken when model foolishly tries to view a screenshot #34113: Same root cause — GLM-5.2 (text-only model) session permanently broken after tool result contains an image; identical symptom of every subsequent message replaying the same error
- Non-vision models blocked from passing images to vision-capable MCP tools #29216: Related — non-vision models and image handling in tool results / MCP tool dispatch
If these don't fully cover your report (your issue adds useful root-cause analysis of the unsupportedParts + supportsMediaInToolResult interaction), it may still be worth linking them for context.
Fix: gate supportsMediaInToolResult on model capabilities
I ran into the same breakage (text-only model receiving base64 images from tool results, 400'ing the session). The root cause described above is spot-on. Here's a minimal fix that makes supportsMediaInToolResult honor model.capabilities.input so unsupported media gets extracted into a synthetic user message where unsupportedParts can catch it — exactly what the function's own comment already says it should do:
// Only apply this workaround if the model actually supports that media input - otherwise unsupportedParts() will turn it into a user-visible error.
Why it's npm-specific
supportsMediaInToolResult decides whether media in a tool result stays in the tool-role message or is extracted into a synthetic user message. The decision is based only on model.api.npm, never on model.capabilities.input. For @ai-sdk/openai it returns true unconditionally, so the image stays in the tool message — and unsupportedParts only inspects user-role messages (if (msg.role !== "user") return msg), so it never gets filtered.
With the default @ai-sdk/openai-compatible, the function returns false, the image is extracted to a user message, and unsupportedParts correctly replaces it with an error. That's why the bug only surfaces when npm: "@ai-sdk/openai" is set explicitly.
npm |
supportsMediaInToolResult(image) |
image stays in | unsupportedParts catches it? |
sent to text-only model? |
|---|---|---|---|---|
@ai-sdk/openai |
true (unconditional) |
tool msg |
no | yes (bug) |
@ai-sdk/openai-compatible (default) |
false |
user msg |
yes | no (error text) |
Change 1 — packages/opencode/src/provider/transform.ts
Export the existing mimeToModality helper so it can be reused:
- function mimeToModality(mime: string): Modality | undefined {
+ export function mimeToModality(mime: string): Modality | undefined {
if (mime.startsWith("image/")) return "image"
if (mime.startsWith("audio/")) return "audio"
if (mime.startsWith("video/")) return "video"
if (mime === "application/pdf") return "pdf"
return undefined
}Change 2 — packages/opencode/src/session/message-v2.ts
Add the capability check at the top of supportsMediaInToolResult:
import { isMedia } from "@/util/media"
+ import { mimeToModality } from "@/provider/transform"
...
const supportsMediaInToolResult = (attachment: { mime: string }) => {
+ // If the model lacks this modality, extract it so unsupportedParts() turns it
+ // into a user-visible error instead of sending unsupported media to the model.
+ const modality = mimeToModality(attachment.mime)
+ if (modality && !model.capabilities.input[modality]) return false
if (model.api.npm === "@ai-sdk/anthropic") return true
if (model.api.npm === "@ai-sdk/openai") return true
...
}Result
Decision order becomes: model capabilities first, then SDK package name.
| scenario | modality | capabilities.input[modality] |
returns | outcome |
|---|---|---|---|---|
@ai-sdk/openai + text-only model + image |
image |
false |
false |
extracted to user msg → unsupportedParts turns into error ✓ |
@ai-sdk/openai + image-capable model + image |
image |
true |
true |
stays in tool msg, sent normally ✓ |
@ai-sdk/openai-compatible (default) + text-only + image |
image |
false |
false |
unchanged ✓ |
| non-media attachment | undefined |
— | falls through to npm logic | unchanged ✓ |
This also covers the other unconditional-true branches (@ai-sdk/anthropic, @ai-sdk/amazon-bedrock/mantle, @ai-sdk/google-vertex/anthropic) which had the same latent issue.
Type safety
mimeToModality returns Modality | undefined where Modality = "text"|"audio"|"image"|"video"|"pdf" (from models-dev.ts). model.capabilities.input is ProviderModalities (provider.ts) with exactly those keys, so model.capabilities.input[modality] typechecks as boolean. The @/provider/transform import adds no circular dependency (transform.ts only depends on ai, remeda, provider and models-dev types).
Happy to open a PR with this if it's useful.
Summary
When a model is configured as text-only (
modalities.input: ["text"], i.e.capabilities.input.image === false), image attachments inside tool results are still sent to the API as base64 data URLs. The API then rejects the entire request with a 400 error, which permanently breaks the session — every subsequent message replays the same image-bearing tool result.Root cause
There are two code paths that should strip unsupported media, but only one of them checks the model's input capabilities, and the other is where tool-result images actually live:
unsupportedParts(packages/opencode/src/provider/transform.ts) — correctly checksmodel.capabilities.input[modality]and replaces unsupportedimage/fileparts with a text error. But it only inspects top-level parts ofusermessages. It does not look insidetool-resultparts.toModelMessagesEffect(packages/opencode/src/session/message-v2.ts) — decides whether to keep media embedded in a tool result or extract it into a synthetic user message, viasupportsMediaInToolResult. That function returnstruefor@ai-sdk/openai(and others), so the image stays embedded in the tool-result part — exactly whereunsupportedPartsnever looks.The result: for providers where
supportsMediaInToolResultistrue, images in tool results bypass the modality gate entirely.Reproduction
Configure a text-only model (e.g. via a custom OpenAI-compatible provider whose endpoint rejects image input):
Provider npm:
@ai-sdk/openai.In a session, have the agent use the
readtool on a PNG file (e.g. a screenshot). The tool result carries the image aspart.state.attachments[].url = "data:image/png;base64,...".The next LLM call fails with an API 400, e.g.:
Every subsequent message in the session replays the same tool-result history and fails identically — the session is bricked.
Expected behaviour
capabilities.input.image === falseshould be honored regardless of where the image sits in the message history. Two possible fixes:unsupportedPartsto also walk the content oftool-resultparts and strip/replace unsupported media there.supportsMediaInToolResultcapability-aware: returnfalsewhenmodel.capabilities.input.imageisfalse, so the image gets extracted into a synthetic user message whereunsupportedPartscan catch it.Option B is the smaller change and reuses the existing extraction + filtering pipeline.
Environment
npm: "@ai-sdk/openai", baseURL pointing at a Volcano Engine (Ark) Responses-API endpointglm-5.2,modalities.input: ["text"]Workaround
Manually edit the session database (
~/.local/share/opencode/opencode.db,parttable) to strip thedata:image/...;base64,...URL from the offendingstate.attachmentsentry, replacing it with adata:text/plain;stripped,...placeholder.