Java Unified Neural Orchestration
Distributed LLM inference and fine-tuning. Pure Java. No Python, no GIL, no Spring.
Juno runs distributed LLM inference and LoRA fine-tuning on the JVM with CUDA and ROCm GPU acceleration via Panama FFI. No Python runtime, no sidecar processes.
Features:
- Distributed Inference: pipeline and tensor parallelism over gRPC
- GPU Acceleration: CUDA 12.x + cuBLAS, ROCm 6+ + rocBLAS
- LoRA Fine-Tuning: GPU training, DoRA, GGUF merge
- OpenAI-Compatible API and Juno-native REST API
- Programmatic API: via JVM facade:
JunoPlayer,LoraTrainer,JunoHttpClient - Observability: JFR events, per-node health dash-board
- Supported Models and quantizations
- Performance reports
POST /v1/vision/chat(blocking + SSE), registered automatically on./juno local --api-port Nwhen the loaded model is a LLaVA-family model- Requires
--mmproj-path PATHpointing at the model's separate mmproj GGUF (real GGUF releases never merge the CLIP encoder into the base LLM file) "model"in the request body can be omitted —--localmode only ever loads one model, so it resolves unambiguously without it- See juno-documentation, Part 12: Vision (Image-to-Text)
Add the BOM from Maven Central at version 0.1.1:
<dependencyManagement>
<dependencies>
<dependency>
<groupId>cab.ml</groupId>
<artifactId>juno-bom</artifactId>
<version>0.1.1</version>
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>Single-JVM quickstart with juno-cookbook LocalChat:
lc = LocalChat.builder(Path.of(MODEL_PATH)).nodeCount(1).useGpu(false)
.samplingParams(SamplingParams.defaults().withMaxTokens(64).withTemperature(0.7f)).build();
String reply = lc.chat("Hello, how are you?");Full API and streaming examples: 1.3 Quickstart: JVM Embedding and 4.6 Programmatic API. Cookbook: juno-cookbook.
Build from source and run a local interactive console:
git clone https://github.com/ml-cab/juno.git && cd juno
mvn clean package -DskipTestsDownload a .gguf or .llamafile then run the interactive console:
Linux / macOS:
./juno local --model-path models/tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf
or Windows:
juno.bat local --model-path models\tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf
REST alongside the REPL via setting api-port
./juno local --model-path models/... --api-port 8080
LoRA training is a separate mode with various artificial options :
./juno lora --model-path models/...
Merge a trained adapter into a stand-alone GGUF:
./juno merge --model-path models/...
See 1.2 Quickstart: Local, Part 3. CLI Reference, 4.3 Training Guide, and 4.5 Merging Adapters.
Run juno-master as the coordinator and juno-node on each worker over gRPC. AWS automation
scripts are provided under scripts/aws/.
See 6.1 On-Prem Cluster and 6.2 AWS Deployment.
| Module | Role |
|---|---|
juno-bom |
Maven BOM. aligned versions for all cab.ml artifacts |
api |
OpenAPI spec, protobuf/gRPC contracts |
registry |
Shard planning, model registry |
coordinator |
Scheduler, generation loop, REST |
node |
Transformer handlers, GGUF, GPU matmul (CUDA + ROCm via Panama FFI) |
lora |
Adapter tensors, optimizer |
tokenizer, sampler, kvcache, health, metrics |
Shared infrastructure |
juno-player |
CLI REPL and cluster harness |
juno-node, juno-master |
Shaded deploy jars |
Architecture and design decisions: Part 2. Architecture. Full module map: 2.6 Module Map.
Full documentation reference: ml.cab/juno-documentation
| Topic | Page |
|---|---|
| Requirements | 1.1 Requirements |
| Quickstart: Local | 1.2 Quickstart: Local |
| Quickstart: JVM | 1.3 Quickstart: JVM Embedding |
| Supported models | 1.4 Supported Models |
| CLI flags | 3.2 Flags |
| LoRA training | 4.3 Training Guide |
| REST API | 5.2 OpenAI-Compatible API |
| On-prem cluster | 6.1 On-Prem Cluster |
| AWS deployment | 6.2 AWS Deployment |
| Windows | 6.3 Windows |
| Performance | 7.2 Performance Methodology |
| Contributing | 10.1 Contributing |
| Legal | 9.1 License and Patents |
| Security | 10.3 Security Policy |
| Release notes | 11.1 Release Notes |
| Changelog | 11.2 Changelog |
JDK 25+, Maven 3.9+. GPU nodes: CUDA 12.x + NVIDIA driver or ROCm 6+ + AMD driver. CPU-only inference requires neither.
Windows: juno.bat at the project root requires JDK 25+ on PATH or JAVA_HOME set. See
6.3 Windows.
Full requirements: 1.1 Requirements.
Apache 2.0. See LICENSE.
