Repository navigation
feature request: read tool should support audio and video attachments #22260
Description
Activity
github-actions commented on Apr 13, 2026
This issue doesn't fully meet our contributing guidelines.
What needs to be fixed:
- No issue template was used. This appears to be a feature request, so please use the Feature Request template, which requires: a title starting with
[FEATURE]:, a verification checkbox confirming the feature hasn't been suggested before, and a description of the enhancement.
Please edit this issue to address the above within 2 hours, or it will be automatically closed.
This issue might be a duplicate of existing issues. Please check:
- [FEATURE]: Native Multimodal Context Support (Video/Audio) #10531: [FEATURE]: Native Multimodal Context Support (Video/Audio) — requests audio and video support in OpenCode, which overlaps significantly with this request.
If you believe this was flagged incorrectly, please let a maintainer know.
github-actions commented on Apr 13, 2026
This issue has been automatically closed because it was not updated to meet our contributing guidelines within the 2-hour window.
Feel free to open a new issue that follows our issue templates.
yeah this could be cool
Reminder: some LLM provider doesn’t support base64 message body exceeding few megabytes, such as 10M. Not sure it’s true for all but at least for dashscope (alibaba).
We’ll still need a hard cap around media file to prevent complaints from LLM provider.
Whichever way this goes, we need something to bring OpenCode up to snuff to use the new multimodal models. Google's got the new Gemma 4 12B, Xiaomi, Minimax. Everybody is starting to be full multimodal, and we have no way to feed those audio and video files into models using OpenCode. I looked at doing it, and the only way I can see to do it requires forking OpenCode, which I don't wish to do. I would love for OpenCode simply to be able to support the rich range of media that modern LLMs are supporting now as standard features .
@rekram1-node Could I reopen the PR for review? However, might need to check first to see how many conflicts there are with the main branch. Additionally, we may need to discuss the details regarding the handling of large files further, as this is an issue I actually encountered in my personal project.
+1 — concrete local-model use case to add weight to this request.
Setup: I run google/gemma-4-e4b locally on macOS via LM Studio, connected to OpenCode Desktop as a custom @ai-sdk/openai-compatible provider. Gemma 4 E2B/E4B natively understand audio and video, and the local backends are catching up: llama.cpp already merged Gemma 4 audio encoder support (ggml-org/llama.cpp#21421) and routes input_audio content blocks on its OpenAI-compatible /v1/chat/completions endpoint; vLLM supports audio for E2B/E4B plus video via frame extraction.
Current blocker is on the OpenCode side: V2 prompt attachments and the read tool only pass UTF-8 text and PNG/JPEG/GIF/WebP images to the model — audio/video files never reach the provider, so even a fully audio-capable local backend cannot be exercised from the TUI/Desktop. I verified this end-to-end: image input works (configured modalities: { input: [text, image] } in opencode.json), while an input_audio test succeeds at the backend level but there is no way to attach the file in OpenCode.
Suggestion: when extending attachment media types, gate them behind the existing per-model modalities config (e.g. input: [text, image, audio, video]) so text-only/vision-only models stay protected, and surface a clear error like the current image path does. This would unlock fully-local multimodal workflows (meeting transcription, video analysis) with zero cloud dependency.
github-actions commented on Sep 21, 2026
To stay organized issues are automatically closed after 60 days of no activity. If the issue is still relevant please open a new one.
The built-in read tool can return images and PDFs as file attachments, but audio and video files are rejected as binary files.
This prevents agents from inspecting local media files through active retrieval. The read tool should attach supported audio/video files as model-native media attachments and fail clearly for oversized files.