Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ Learn all about Gemini CLI in our [documentation](https://geminicli.com/docs/).
- **🔌 Extensible**: MCP (Model Context Protocol) support for custom
integrations.
- **💻 Terminal-first**: Designed for developers who live in the command line.
- **⚡ SGLang Inference Server**: Direct connection to local or remote SGLang servers (Kimi-K3, DeepSeek, Qwen). See [SGLang Setup Guide](docs/sglang.md).
- **🛡️ Open source**: Apache 2.0 licensed.

## 📦 Installation
Expand Down Expand Up @@ -71,6 +72,29 @@ conda activate gemini_env
npm install -g @google/gemini-cli
```

#### Build from Source with SGLang Support (Linux / macOS)

To connect Gemini CLI to local or remote SGLang / OpenAI inference servers (e.g., Moonshot Kimi-K3):

```bash
# 1. Clone the branch
git clone -b feat/sglang-support https://github.com/shivajid/gemini-cli.git
cd gemini-cli

# 2. Install dependencies & build binary bundle
npm install
npm run build
npm run bundle
npm link

# 3. Configure and run with SGLang
export SGLANG_BASE_URL="http://127.0.0.1:30100/v1"
export GEMINI_MODEL="moonshotai/Kimi-K3"
gemini
```

See the [SGLang Setup Guide](docs/sglang.md) for full Linux prerequisites and troubleshooting.

## Release Channels

See [Releases](https://www.geminicli.com/docs/changelogs) for more details.
Expand Down
189 changes: 189 additions & 0 deletions docs/sglang.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,189 @@
# Connecting Gemini CLI to SGLang Server

Gemini CLI includes native support for connecting directly to local or remote **SGLang inference servers** (such as [Moonshot Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3), DeepSeek-V3/R1, or Qwen models).

This integration leverages OpenAI-compatible `/v1/chat/completions` endpoints with full support for:
- ⚡ **Streaming responses** (`stream: true`)
- 🧠 **Reasoning thought traces** (e.g. `delta.reasoning_content`) rendered cleanly in the CLI thinking box
- 🛠️ **Built-in tools & MCP function calling** with recursive schema conversion (handling Gemini uppercase types to standard JSON Schema)
- 🛑 **Interactive stream cancellation** (`ESC` key support)
- 🔁 **Multi-turn conversation history** with persistent tool call identifiers

---

## 1. Prerequisites on Linux (Debian / Ubuntu / COS)

Before building, ensure you have **Node.js 20+**, **npm**, and build essentials installed on your Linux machine:

### Install Node.js 20 and Build Tools

```bash
# Update package lists
sudo apt-get update

# Install git, curl, and native compilation tools
sudo apt-get install -y git curl build-essential python3

# Install Node.js 20.x LTS via NodeSource
curl -fsSL https://deb.nodesource.com/setup_20.x | sudo -E bash -
sudo apt-get install -y nodejs

# Verify versions
node -v # Should be v20.x or higher
npm -v # Should be 10.x or higher
```

### Install GKE / Kubernetes Client Tools (If running on GKE)

```bash
# Install kubectl and GKE auth plugin
sudo apt-get install -y kubectl google-cloud-cli-gke-gcloud-auth-plugin

# Configure cluster credentials
gcloud container clusters get-credentials <CLUSTER_NAME> \
--region <REGION> \
--project <PROJECT_ID>
```

---

## 2. Building `feat/sglang-support` from Source

Clone the repository, switch to the `feat/sglang-support` branch, install dependencies, and compile:

```bash
# 1. Clone the repository and checkout the feat/sglang-support branch
git clone -b feat/sglang-support https://github.com/shivajid/gemini-cli.git
cd gemini-cli

# 2. Install workspace dependencies
npm install

# 3. Compile all packages (including @google/gemini-cli-core)
npm run build

# 4. Generate the standalone CLI binary bundle
npm run bundle

# 5. Link globally so the `gemini` command is available system-wide
npm link
```

> **Tip for updating**: To pull future updates, simply run:
> ```bash
> git pull origin feat/sglang-support
> npm run build && npm run bundle && npm link
> ```

---

## 3. Quick Start & Connecting to SGLang

### Step 1: Start SGLang Port-Forwarding

If your SGLang server is running in Kubernetes / GKE, forward the API port to your local machine:

```bash
# Forward port 30100 from your SGLang leader pod
kubectl port-forward -n <NAMESPACE> pod/<SGLANG_LEADER_POD> 30100:30100 &
```

Verify that the server is reachable:

```bash
curl http://127.0.0.1:30100/v1/models
```

---

### Step 2: Set Environment Variables

```bash
export SGLANG_BASE_URL="http://127.0.0.1:30100/v1"
export GEMINI_MODEL="moonshotai/Kimi-K3"
export GEMINI_DEFAULT_AUTH_TYPE="sglang"
```

> **Note**: Always use `http://127.0.0.1:30100/v1` instead of `localhost` on Linux containers to avoid DNS resolution issues.

---

### Step 3: Configure Settings (Optional)

Create or update `~/.gemini/settings.json` to persist the SGLang authentication:

```bash
mkdir -p ~/.gemini
cat << 'EOF' > ~/.gemini/settings.json
{
"general": {
"enableAutoUpdateNotification": false
},
"security": {
"auth": {
"selectedType": "sglang"
}
}
}
EOF
```

---

### Step 4: Run Gemini CLI

Start an interactive chat session:

```bash
gemini
```

Or pass an immediate prompt:

```bash
gemini "Hello Kimi-K3! List the files in this directory."
```

---

## 4. Interactive Authentication Menu

If you run `gemini` without predefined settings, or type `/auth` inside an active session:

```
? How would you like to authenticate for this project?
● 1. SGLang Server (Local / Remote Kimi-K3)
2. Sign in with Google
3. Use Gemini API Key
4. Vertex AI
```

Select **`1. SGLang Server (Local / Remote Kimi-K3)`** to bypass Google credentials and route traffic directly to your SGLang endpoint.

---

## 5. Configuration Reference

| Variable / Setting | Description | Default |
|---|---|---|
| `SGLANG_BASE_URL` | Base URL of the OpenAI-compatible SGLang server | `http://127.0.0.1:30100/v1` |
| `OPENAI_BASE_URL` | Secondary fallback base URL | `http://127.0.0.1:30100/v1` |
| `GEMINI_MODEL` / `SGLANG_MODEL` | Served model name on SGLang | `moonshotai/Kimi-K3` |
| `GEMINI_DEFAULT_AUTH_TYPE` | Default auth method (`sglang`, `oauth-personal`, `gemini-api-key`) | `oauth-personal` |
| `enableAutoUpdateNotification` | Set to `false` in `settings.json` to hide git update banners | `true` |

---

## 6. Troubleshooting

### 1. `socket.gaierror: [Errno -2] Name or service not known`
- **Cause**: Linux environment does not resolve `localhost` in `/etc/hosts`.
- **Fix**: Use numeric IP `http://127.0.0.1:30100/v1`.

### 2. `Connection refused`
- **Cause**: SGLang server is initializing or `kubectl port-forward` terminated.
- **Fix**: Check `kubectl get pods -n <namespace>` and restart port-forwarding.

### 3. `Model "moonshotai/Kimi-K3" was not found`
- **Cause**: Saved setting in `~/.gemini/settings.json` is still set to Google API (`gemini-api-key` or `oauth-personal`).
- **Fix**: Run `/auth` and select **SGLang Server**, or set `"selectedType": "sglang"` in `~/.gemini/settings.json`.
3 changes: 2 additions & 1 deletion packages/cli/src/config/auth.ts
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,8 @@ export async function validateAuthMethod(
loadEnvironment(loadSettings().merged, process.cwd());
if (
authMethod === AuthType.LOGIN_WITH_GOOGLE ||
authMethod === AuthType.COMPUTE_ADC
authMethod === AuthType.COMPUTE_ADC ||
authMethod === AuthType.SGLANG
) {
return null;
}
Expand Down
5 changes: 5 additions & 0 deletions packages/cli/src/ui/auth/AuthDialog.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,11 @@ export function AuthDialog({
}: AuthDialogProps): React.JSX.Element {
const [exiting, setExiting] = useState(false);
let items = [
{
label: 'SGLang Server (Local / Remote Kimi-K3)',
value: AuthType.SGLANG,
key: AuthType.SGLANG,
},
{
label: 'Sign in with Google',
value: AuthType.LOGIN_WITH_GOOGLE,
Expand Down
8 changes: 6 additions & 2 deletions packages/cli/src/ui/auth/useAuth.ts
Original file line number Diff line number Diff line change
Expand Up @@ -29,8 +29,12 @@ export async function validateAuthMethodWithSettings(
if (settings.merged.security.auth.useExternal) {
return null;
}
// If using Gemini API key, we don't validate it here as we might need to prompt for it.
if (authType === AuthType.USE_GEMINI) {
// If using Gemini API key or SGLang, we don't validate it here as we might need to prompt for it.
if (
authType === AuthType.USE_GEMINI ||
authType === AuthType.SGLANG ||
(authType as string) === 'sglang'
) {
return null;
}
return validateAuthMethod(authType);
Expand Down
5 changes: 5 additions & 0 deletions packages/core/src/config/models.ts
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,7 @@ export const PREVIEW_GEMINI_FLASH_LITE_MODEL = 'none';

export const GEMMA_4_31B_IT_MODEL = 'gemma-4-31b-it';
export const GEMMA_4_26B_A4B_IT_MODEL = 'gemma-4-26b-a4b-it';
export const KIMI_K3_MODEL = 'moonshotai/Kimi-K3';

export const VALID_GEMINI_MODELS = new Set([
PREVIEW_GEMINI_MODEL,
Expand All @@ -99,6 +100,7 @@ export const VALID_GEMINI_MODELS = new Set([

GEMMA_4_31B_IT_MODEL,
GEMMA_4_26B_A4B_IT_MODEL,
KIMI_K3_MODEL,
]);

/** @deprecated Use GEMINI_MODEL_ALIAS_AUTO instead. */
Expand Down Expand Up @@ -549,6 +551,9 @@ export function isActiveModel(
useCustomToolModel: boolean = false,
experimentalGemma: boolean = true,
): boolean {
if (model === KIMI_K3_MODEL || model.startsWith('moonshotai/') || model.includes('kimi')) {
return true;
}
if (!VALID_GEMINI_MODELS.has(model) || model === 'none') {
return false;
}
Expand Down
45 changes: 41 additions & 4 deletions packages/core/src/core/contentGenerator.ts
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@ import { getVersion, resolveModel } from '../../index.js';
import type { LlmRole } from '../telemetry/llmRole.js';
import { ModelMappingContentGenerator } from './modelMappingContentGenerator.js';
import { CCPA_AI_MODEL_MAPPINGS } from '../config/models.js';
import { SglangContentGenerator } from './sglangContentGenerator.js';

/**
* Interface abstracting the core functionalities for generating content and counting tokens.
Expand Down Expand Up @@ -67,17 +68,22 @@ export enum AuthType {
LEGACY_CLOUD_SHELL = 'cloud-shell',
COMPUTE_ADC = 'compute-default-credentials',
GATEWAY = 'gateway',
SGLANG = 'sglang',
}

/**
* Detects the best authentication type based on environment variables.
*
* Checks in order:
* 1. GOOGLE_GENAI_USE_GCA=true -> LOGIN_WITH_GOOGLE
* 2. GOOGLE_GENAI_USE_VERTEXAI=true -> USE_VERTEX_AI
* 3. GEMINI_API_KEY -> USE_GEMINI
* 1. SGLANG_BASE_URL or OPENAI_BASE_URL -> SGLANG
* 2. GOOGLE_GENAI_USE_GCA=true -> LOGIN_WITH_GOOGLE
* 3. GOOGLE_GENAI_USE_VERTEXAI=true -> USE_VERTEX_AI
* 4. GEMINI_API_KEY -> USE_GEMINI
*/
export function getAuthTypeFromEnv(): AuthType | undefined {
if (process.env['SGLANG_BASE_URL'] || process.env['OPENAI_BASE_URL']) {
return AuthType.SGLANG;
}
if (process.env['GOOGLE_GENAI_USE_GCA'] === 'true') {
return AuthType.LOGIN_WITH_GOOGLE;
}
Expand Down Expand Up @@ -166,7 +172,8 @@ export async function createContentGeneratorConfig(
// (WSL/SSH/Docker/CI) keytar can block indefinitely on its functional probe.
if (
authType === AuthType.LOGIN_WITH_GOOGLE ||
authType === AuthType.COMPUTE_ADC
authType === AuthType.COMPUTE_ADC ||
authType === AuthType.SGLANG
) {
return contentGeneratorConfig;
}
Expand Down Expand Up @@ -409,6 +416,36 @@ export async function createContentGenerator(
});
return new LoggingContentGenerator(googleGenAI.models, gcConfig);
}
if (
config.authType === AuthType.SGLANG ||
String(config.authType).toLowerCase() === 'sglang'
) {
const baseUrl =
config.baseUrl ||
process.env['SGLANG_BASE_URL'] ||
process.env['OPENAI_BASE_URL'] ||
'http://127.0.0.1:30100/v1';
validateBaseUrl(baseUrl);
// Prefer an explicitly configured served-model name; never forward
// internal gemini-* aliases to the SGLang server.
const configuredModel = (
gcConfig as { getModel?: () => string }
).getModel?.();
const isServedModelName = (m?: string): m is string =>
!!m &&
m !== 'auto' &&
m !== 'none' &&
!m.startsWith('gemini') &&
!m.startsWith('gemma');
const modelName =
process.env['SGLANG_MODEL'] ||
(isServedModelName(configuredModel) ? configuredModel : undefined) ||
'moonshotai/Kimi-K3';
return new LoggingContentGenerator(
new SglangContentGenerator(baseUrl, modelName),
gcConfig,
);
Comment on lines +444 to +447

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

With SglangContentGenerator properly implementing the ContentGenerator interface, we can remove the unsafe as never type assertion.

Suggested change
return new LoggingContentGenerator(
new SglangContentGenerator(baseUrl, modelName) as never,
gcConfig,
);
return new LoggingContentGenerator(
new SglangContentGenerator(baseUrl, modelName),
gcConfig,
);

}
throw new Error(
`Error creating contentGenerator: Unsupported authType: ${config.authType}`,
);
Expand Down
Loading