Skip to content
ml-cabPublic

About

JUNO: Java Unified Neural Orchestration

Topics

Resources

Security policy

Stars

11 stars

Watchers

0 watching

Forks

Latest commit

 

History

155 Commits

Folders and files

Repository files navigation

Juno

Java Unified Neural Orchestration

Distributed LLM inference and fine-tuning. Pure Java. No Python, no GIL, no Spring.

Java 25+ Maven CUDA ROCm License

1. What is Juno

Juno runs distributed LLM inference and LoRA fine-tuning on the JVM with CUDA and ROCm GPU acceleration via Panama FFI. No Python runtime, no sidecar processes.

Features:

1.1 What's new?

Vision (image-to-text)

  • POST /v1/vision/chat (blocking + SSE), registered automatically on ./juno local --api-port N when the loaded model is a LLaVA-family model
  • Requires --mmproj-path PATH pointing at the model's separate mmproj GGUF (real GGUF releases never merge the CLIP encoder into the base LLM file)
  • "model" in the request body can be omitted — --local mode only ever loads one model, so it resolves unambiguously without it
  • See juno-documentation, Part 12: Vision (Image-to-Text)

2. How to use

2.1 JVM Integration

Add the BOM from Maven Central at version 0.1.1:

<dependencyManagement>
  <dependencies>
    <dependency>
      <groupId>cab.ml</groupId>
      <artifactId>juno-bom</artifactId>
      <version>0.1.1</version>
      <type>pom</type>
      <scope>import</scope>
    </dependency>
  </dependencies>
</dependencyManagement>

Single-JVM quickstart with juno-cookbook LocalChat:

lc = LocalChat.builder(Path.of(MODEL_PATH)).nodeCount(1).useGpu(false)
        .samplingParams(SamplingParams.defaults().withMaxTokens(64).withTemperature(0.7f)).build();

String reply = lc.chat("Hello, how are you?");

Full API and streaming examples: 1.3 Quickstart: JVM Embedding and 4.6 Programmatic API. Cookbook: juno-cookbook.

2.2 Local player and LoRA

Build from source and run a local interactive console:

git clone https://github.com/ml-cab/juno.git && cd juno
mvn clean package -DskipTests

Download a .gguf or .llamafile then run the interactive console:

Linux / macOS:

./juno local --model-path models/tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf

or Windows:

juno.bat local --model-path models\tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf

REST alongside the REPL via setting api-port

./juno local --model-path models/... --api-port 8080

LoRA training is a separate mode with various artificial options :

./juno lora --model-path models/...

Merge a trained adapter into a stand-alone GGUF:

./juno merge --model-path models/...

Juno local console running TinyLlama-1.1B, with CPU and memory usage shown alongside

See 1.2 Quickstart: Local, Part 3. CLI Reference, 4.3 Training Guide, and 4.5 Merging Adapters.

2.3 On-prem and cloud orchestration

Run juno-master as the coordinator and juno-node on each worker over gRPC. AWS automation scripts are provided under scripts/aws/.

See 6.1 On-Prem Cluster and 6.2 AWS Deployment.

3. Modules

Module Role
juno-bom Maven BOM. aligned versions for all cab.ml artifacts
api OpenAPI spec, protobuf/gRPC contracts
registry Shard planning, model registry
coordinator Scheduler, generation loop, REST
node Transformer handlers, GGUF, GPU matmul (CUDA + ROCm via Panama FFI)
lora Adapter tensors, optimizer
tokenizer, sampler, kvcache, health, metrics Shared infrastructure
juno-player CLI REPL and cluster harness
juno-node, juno-master Shaded deploy jars

Architecture and design decisions: Part 2. Architecture. Full module map: 2.6 Module Map.

4. Documentation

Full documentation reference: ml.cab/juno-documentation

Topic Page
Requirements 1.1 Requirements
Quickstart: Local 1.2 Quickstart: Local
Quickstart: JVM 1.3 Quickstart: JVM Embedding
Supported models 1.4 Supported Models
CLI flags 3.2 Flags
LoRA training 4.3 Training Guide
REST API 5.2 OpenAI-Compatible API
On-prem cluster 6.1 On-Prem Cluster
AWS deployment 6.2 AWS Deployment
Windows 6.3 Windows
Performance 7.2 Performance Methodology
Contributing 10.1 Contributing
Legal 9.1 License and Patents
Security 10.3 Security Policy
Release notes 11.1 Release Notes
Changelog 11.2 Changelog

Requirements

JDK 25+, Maven 3.9+. GPU nodes: CUDA 12.x + NVIDIA driver or ROCm 6+ + AMD driver. CPU-only inference requires neither.

Windows: juno.bat at the project root requires JDK 25+ on PATH or JAVA_HOME set. See 6.3 Windows.

Full requirements: 1.1 Requirements.

License

Apache 2.0. See LICENSE.

About

JUNO: Java Unified Neural Orchestration

Topics

Resources

Security policy

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages