Skip to content

dots3-note-prev (Image-Text-to-Text) incorrectly detected as non-vision, images rejected by unsupportedParts() #43392

Description

@Ezra-Yu

Description

When uploading an image while using dots3-note-prev, the model receives this error instead of the image:

ERROR: Cannot read clipboard (this model does not support image input). Inform the user.

This is wrong — dots3-note-prev is a multimodal omni model (280B-A16B omni MoE) that natively supports image, audio, and video input. It should accept images.

Root Cause

unsupportedParts() in packages/opencode/src/provider/transform.ts:438-440 strips image parts when model.capabilities.input.image is false. Capability detection in packages/opencode/src/provider/provider.ts:1267 derives this from the external model config:

image: model.modalities?.input?.includes("image") ?? false,

If dots3-note-prev is not listed in the model config with modalities.input including "image", it defaults to false and images get rejected — even though the model actually supports them.

Steps to Reproduce

  1. Configure dots3-note-prev as the model
  2. Upload/paste an image
  3. Observe: error text instead of image processing

Expected Behavior

Images should pass through to the model since dots3-note-prev natively supports image input (as well as audio and video).

Evidence

Environment

  • Model: dots3-note-prev (280B-A16B omni MoE, Image-Text-to-Text + audio + video)

Related

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions