Skip to content

Images in tool results bypass model modality check (text-only models receive base64 images) #39824

Description

@Yiipu

Summary

When a model is configured as text-only (modalities.input: ["text"], i.e. capabilities.input.image === false), image attachments inside tool results are still sent to the API as base64 data URLs. The API then rejects the entire request with a 400 error, which permanently breaks the session — every subsequent message replays the same image-bearing tool result.

Root cause

There are two code paths that should strip unsupported media, but only one of them checks the model's input capabilities, and the other is where tool-result images actually live:

  1. unsupportedParts (packages/opencode/src/provider/transform.ts) — correctly checks model.capabilities.input[modality] and replaces unsupported image/file parts with a text error. But it only inspects top-level parts of user messages. It does not look inside tool-result parts.

  2. toModelMessagesEffect (packages/opencode/src/session/message-v2.ts) — decides whether to keep media embedded in a tool result or extract it into a synthetic user message, via supportsMediaInToolResult. That function returns true for @ai-sdk/openai (and others), so the image stays embedded in the tool-result part — exactly where unsupportedParts never looks.

The result: for providers where supportsMediaInToolResult is true, images in tool results bypass the modality gate entirely.

Reproduction

  1. Configure a text-only model (e.g. via a custom OpenAI-compatible provider whose endpoint rejects image input):

    "glm-5.2": {
      "modalities": { "input": ["text"], "output": ["text"] }
    }

    Provider npm: @ai-sdk/openai.

  2. In a session, have the agent use the read tool on a PNG file (e.g. a screenshot). The tool result carries the image as part.state.attachments[].url = "data:image/png;base64,...".

  3. The next LLM call fails with an API 400, e.g.:

    Model only support text input
    statusCode: 400
    isRetryable: false
    
  4. Every subsequent message in the session replays the same tool-result history and fails identically — the session is bricked.

Expected behaviour

capabilities.input.image === false should be honored regardless of where the image sits in the message history. Two possible fixes:

  • Option A — extend unsupportedParts to also walk the content of tool-result parts and strip/replace unsupported media there.
  • Option B — make supportsMediaInToolResult capability-aware: return false when model.capabilities.input.image is false, so the image gets extracted into a synthetic user message where unsupportedParts can catch it.

Option B is the smaller change and reuses the existing extraction + filtering pipeline.

Environment

  • opencode 1.18.5 (Homebrew)
  • Provider: custom, npm: "@ai-sdk/openai", baseURL pointing at a Volcano Engine (Ark) Responses-API endpoint
  • Model: glm-5.2, modalities.input: ["text"]

Workaround

Manually edit the session database (~/.local/share/opencode/opencode.db, part table) to strip the data:image/...;base64,... URL from the offending state.attachments entry, replacing it with a data:text/plain;stripped,... placeholder.

Activity

github-actions commented on Jul 31, 2026

@github-actions
Contributor

This issue might be a duplicate of existing issues. Please check:

If these don't fully cover your report (your issue adds useful root-cause analysis of the unsupportedParts + supportsMediaInToolResult interaction), it may still be worth linking them for context.

ivaNov-0042 commented on Aug 10, 2026

@ivaNov-0042

Fix: gate supportsMediaInToolResult on model capabilities

I ran into the same breakage (text-only model receiving base64 images from tool results, 400'ing the session). The root cause described above is spot-on. Here's a minimal fix that makes supportsMediaInToolResult honor model.capabilities.input so unsupported media gets extracted into a synthetic user message where unsupportedParts can catch it — exactly what the function's own comment already says it should do:

// Only apply this workaround if the model actually supports that media input - otherwise unsupportedParts() will turn it into a user-visible error.

Why it's npm-specific

supportsMediaInToolResult decides whether media in a tool result stays in the tool-role message or is extracted into a synthetic user message. The decision is based only on model.api.npm, never on model.capabilities.input. For @ai-sdk/openai it returns true unconditionally, so the image stays in the tool message — and unsupportedParts only inspects user-role messages (if (msg.role !== "user") return msg), so it never gets filtered.

With the default @ai-sdk/openai-compatible, the function returns false, the image is extracted to a user message, and unsupportedParts correctly replaces it with an error. That's why the bug only surfaces when npm: "@ai-sdk/openai" is set explicitly.

npm supportsMediaInToolResult(image) image stays in unsupportedParts catches it? sent to text-only model?
@ai-sdk/openai true (unconditional) tool msg no yes (bug)
@ai-sdk/openai-compatible (default) false user msg yes no (error text)

Change 1 — packages/opencode/src/provider/transform.ts

Export the existing mimeToModality helper so it can be reused:

- function mimeToModality(mime: string): Modality | undefined {
+ export function mimeToModality(mime: string): Modality | undefined {
    if (mime.startsWith("image/")) return "image"
    if (mime.startsWith("audio/")) return "audio"
    if (mime.startsWith("video/")) return "video"
    if (mime === "application/pdf") return "pdf"
    return undefined
  }

Change 2 — packages/opencode/src/session/message-v2.ts

Add the capability check at the top of supportsMediaInToolResult:

  import { isMedia } from "@/util/media"
+ import { mimeToModality } from "@/provider/transform"
  ...
  const supportsMediaInToolResult = (attachment: { mime: string }) => {
+   // If the model lacks this modality, extract it so unsupportedParts() turns it
+   // into a user-visible error instead of sending unsupported media to the model.
+   const modality = mimeToModality(attachment.mime)
+   if (modality && !model.capabilities.input[modality]) return false
    if (model.api.npm === "@ai-sdk/anthropic") return true
    if (model.api.npm === "@ai-sdk/openai") return true
    ...
  }

Result

Decision order becomes: model capabilities first, then SDK package name.

scenario modality capabilities.input[modality] returns outcome
@ai-sdk/openai + text-only model + image image false false extracted to user msg → unsupportedParts turns into error ✓
@ai-sdk/openai + image-capable model + image image true true stays in tool msg, sent normally ✓
@ai-sdk/openai-compatible (default) + text-only + image image false false unchanged ✓
non-media attachment undefined — falls through to npm logic unchanged ✓

This also covers the other unconditional-true branches (@ai-sdk/anthropic, @ai-sdk/amazon-bedrock/mantle, @ai-sdk/google-vertex/anthropic) which had the same latent issue.

Type safety

mimeToModality returns Modality | undefined where Modality = "text"|"audio"|"image"|"video"|"pdf" (from models-dev.ts). model.capabilities.input is ProviderModalities (provider.ts) with exactly those keys, so model.capabilities.input[modality] typechecks as boolean. The @/provider/transform import adds no circular dependency (transform.ts only depends on ai, remeda, provider and models-dev types).

Happy to open a PR with this if it's useful.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions