Description
When uploading an image while using dots3-note-prev, the model receives this error instead of the image:
ERROR: Cannot read clipboard (this model does not support image input). Inform the user.
This is wrong — dots3-note-prev is a multimodal omni model (280B-A16B omni MoE) that natively supports image, audio, and video input. It should accept images.
Root Cause
unsupportedParts() in packages/opencode/src/provider/transform.ts:438-440 strips image parts when model.capabilities.input.image is false. Capability detection in packages/opencode/src/provider/provider.ts:1267 derives this from the external model config:
image: model.modalities?.input?.includes("image") ?? false,
If dots3-note-prev is not listed in the model config with modalities.input including "image", it defaults to false and images get rejected — even though the model actually supports them.
Steps to Reproduce
- Configure
dots3-note-prev as the model
- Upload/paste an image
- Observe: error text instead of image processing
Expected Behavior
Images should pass through to the model since dots3-note-prev natively supports image input (as well as audio and video).
Evidence
Environment
- Model:
dots3-note-prev (280B-A16B omni MoE, Image-Text-to-Text + audio + video)
Related
Description
When uploading an image while using
dots3-note-prev, the model receives this error instead of the image:This is wrong —
dots3-note-previs a multimodal omni model (280B-A16B omni MoE) that natively supports image, audio, and video input. It should accept images.Root Cause
unsupportedParts()inpackages/opencode/src/provider/transform.ts:438-440strips image parts whenmodel.capabilities.input.imageisfalse. Capability detection inpackages/opencode/src/provider/provider.ts:1267derives this from the external model config:If
dots3-note-previs not listed in the model config withmodalities.inputincluding"image", it defaults tofalseand images get rejected — even though the model actually supports them.Steps to Reproduce
dots3-note-prevas the modelExpected Behavior
Images should pass through to the model since
dots3-note-prevnatively supports image input (as well as audio and video).Evidence
image-text-to-textEnvironment
dots3-note-prev(280B-A16B omni MoE, Image-Text-to-Text + audio + video)Related