OpenCode Manager
Mobile-first web interface for OpenCode AI agents. Manage, control, and code with multiple agents from any device with Git integration and real-time chat.
I design and deploy private LLM systems—including local inference, RAG, and agent workflows—without sending sensitive data outside your environment.
Run open-source and custom models inside your environment, tuned for latency, throughput, and cost.
Ground models in internal knowledge without moving sensitive documents into public AI services.
Connect private models to your applications, data, and workflows with controlled tool access.
Open-source tools and plugins I've shipped.
Mobile-first web interface for OpenCode AI agents. Manage, control, and code with multiple agents from any device with Git integration and real-time chat.
Neovim plugin that captures console outputs as virtual text inline with your code. Automatic framework detection and comprehensive debugging.
OpenCode plugin that lets text-only models work with images by routing each image through a vision-capable model and returning a text description.
Plugin for autonomous dev loops, planning, and worktree sandbox environments in OpenCode.
Comprehensive text-to-speech plugin for Neovim with macOS native speech synthesis and OpenAI-compatible TTS endpoints.
Neovim plugin that lets coding agents point your editor at a line and explain why you are looking at it.
I build production LLM infrastructure, full-stack applications, APIs, and performance-critical systems. I'm also behind the open-source work featured here, including OpenCode Manager and consolelog.nvim. From first conversation to final delivery, you work directly with the person responsible for the outcome. No account manager. No handoff. No theater.
How I distilled Qwen-Image-2.1 into an 8-step LoRA covering text-to-image, editing and transparent images: a first DMD2 run that drifted, the trajectory-anchored rebuild that fixed it, and the evaluation that told them apart.
Measured throughput, KV capacity, and long-context reasoning performance of Qwen3.8-Flash-Next FP8 — a 180B-parameter hybrid-attention MoE with 6B active per token — on four RTX PRO 6000 Blackwell GPUs served by vLLM.
Measured throughput, capacity limits, and two failed reasoning checks from serving Ornith-1.5-35B-A3B-FP8 on two RTX 6000 Ada GPUs with SGLang.
GLM 5.3 Flash's answer to a single-shot WebGL build prompt — a GPU raymarched Mandelbulb fractal with orbit controls in one HTML file.
WebGL GPU raymarched Mandelbulb fractal
Qwen3.8-27B FP8's answer to a single-shot Three.js build prompt — a full interactive orrery in one HTML file.
Create a 3D solar system with Three.js in a single HTML JavaScript file.
Qwen3.8-27B FP8's answer to a single-shot Three.js build prompt — a draggable 3×3 cube with scramble and solve in one HTML file.
Create a single self-contained HTML Three.js Rubik's cube emulation with solve and scramble functionality that moves and works like the real thing.