Skip to content
This repository was archived by the owner on Sep 23, 2026. It is now read-only.
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,8 @@ Only write entries that are worth mentioning to users.

## Unreleased

- Kosong: Stop Kimi from automatically sending the legacy `reasoning_effort` parameter when configuring thinking — requests now use `thinking.type` exclusively while preserving explicit legacy passthrough

## 1.47.0 (2026-06-05)

- Shell: Guide users to the new standalone Kimi Code — adds a `/upgrade` command that installs it (migrating your config & sessions automatically), a welcome-screen nudge, and a once-per-day tip shown on exit
Expand Down
7 changes: 7 additions & 0 deletions docs/en/release-notes/breaking-changes.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,13 @@ This page documents breaking changes in Kimi Code CLI releases and provides migr

## Unreleased

### Kimi no longer sends legacy `reasoning_effort` automatically

Kimi thinking configuration now uses `thinking.type` exclusively. `Kimi.with_thinking(...)` no longer adds the legacy `reasoning_effort` parameter to requests.

- **Affected**: Applications using Kosong's Kimi provider with older Kimi-compatible endpoints that require `reasoning_effort`
- **Migration**: Update the endpoint to support `thinking.type`. If an older endpoint still requires the legacy parameter, pass it explicitly with `with_generation_kwargs(reasoning_effort="...")`

## 1.43.0

### MCP OAuth token cache moved to `~/.kimi/mcp-oauth/`
Expand Down
2 changes: 2 additions & 0 deletions docs/en/release-notes/changelog.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,8 @@ This page documents the changes in each Kimi Code CLI release.

## Unreleased

- Kosong: Stop Kimi from automatically sending the legacy `reasoning_effort` parameter when configuring thinking — requests now use `thinking.type` exclusively while preserving explicit legacy passthrough

## 1.47.0 (2026-06-05)

- Shell: Guide users to the new standalone Kimi Code — adds a `/upgrade` command that installs it (migrating your config & sessions automatically), a welcome-screen nudge, and a once-per-day tip shown on exit
Expand Down
7 changes: 7 additions & 0 deletions docs/zh/release-notes/breaking-changes.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,13 @@

## 未发布

### Kimi 不再自动发送旧版 `reasoning_effort`

Kimi 的 Thinking 配置现在仅使用 `thinking.type`。`Kimi.with_thinking(...)` 不再自动把旧版 `reasoning_effort` 参数加入请求。

- **受影响**:通过 Kosong 的 Kimi 供应商连接仍要求 `reasoning_effort` 的旧版 Kimi 兼容端点的应用
- **迁移**:将端点升级为支持 `thinking.type`。如果旧版端点仍要求该参数,请通过 `with_generation_kwargs(reasoning_effort="...")` 显式传入

## 1.43.0

### MCP OAuth token 缓存迁移到 `~/.kimi/mcp-oauth/`
Expand Down
2 changes: 2 additions & 0 deletions docs/zh/release-notes/changelog.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,8 @@

## 未发布

- Kosong:配置 Thinking 模式时不再自动向 Kimi 请求发送旧版 `reasoning_effort` 参数——请求现在仅使用 `thinking.type`,同时保留显式传递旧版参数的兼容能力

## 1.47.0 (2026-06-05)

- Shell:引导用户升级到新版独立 Kimi Code——新增 `/upgrade` 命令一键安装(自动迁移现有配置与会话),并新增欢迎界面提示与每天一次的退出提示
Expand Down
1 change: 1 addition & 0 deletions packages/kosong/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@

## Unreleased

- Kimi: Stop automatically sending the legacy `reasoning_effort` parameter when configuring thinking — requests now use `thinking.type` exclusively while preserving explicit legacy passthrough
- Kimi: Preserve empty-string `reasoning_content` as `ThinkPart(think="")` in both streaming and non-streaming responses — previously the truthy check dropped empty deltas, conflating "reasoned but empty" with "no reasoning at all"; the stored (empty) ThinkPart is what makes `_convert_message` emit `reasoning_content` on the next request, so preserved-thinking backends that require the field on every assistant message no longer 400 after a reason-free turn

## 0.53.0 (2026-04-28)
Expand Down
5 changes: 3 additions & 2 deletions packages/kosong/src/kosong/chat_provider/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -130,8 +130,9 @@ def input(self) -> int:
- **OpenAI**: ``xhigh`` is accepted natively for reasoning-capable models
after ``gpt-5.1-codex-max`` and passes through unchanged. ``max`` is
Anthropic-specific and clamps to ``xhigh`` (OpenAI's ceiling).
- **Kimi / Gemini**: ``xhigh`` and ``max`` clamp to ``high`` (no native
support).
- **Kimi**: requests only serialize thinking as enabled or disabled; the
caller-provided effort remains unchanged as provider state.
- **Gemini**: ``xhigh`` and ``max`` clamp to ``high`` (no native support).
Comment on lines +133 to +135
"""


Expand Down
31 changes: 7 additions & 24 deletions packages/kosong/src/kosong/chat_provider/kimi.py
Original file line number Diff line number Diff line change
Expand Up @@ -96,7 +96,7 @@ class GenerationKwargs(TypedDict, total=False):
stop: str | list[str] | None
prompt_cache_key: str | None
reasoning_effort: str | None
"""Legacy thinking parameter. Use `extra_body.thinking` instead."""
"""Legacy explicit passthrough. `with_thinking` uses `extra_body.thinking` instead."""
extra_body: ExtraBody | None

def __init__(
Expand Down Expand Up @@ -131,25 +131,16 @@ def __init__(
)
"""The underlying `AsyncOpenAI` client."""
self._generation_kwargs: Kimi.GenerationKwargs = {}
self._thinking_effort: ThinkingEffort | None = None
"""Thinking state kept separately from parameters serialized onto the wire."""

@property
def model_name(self) -> str:
return self.model

@property
def thinking_effort(self) -> ThinkingEffort | None:
reasoning_effort = self._generation_kwargs.get("reasoning_effort")
if reasoning_effort is None:
return None
match reasoning_effort:
case "low":
return "low"
case "medium":
return "medium"
case "high":
return "high"
case _:
return "off"
return self._thinking_effort

async def generate(
self,
Expand Down Expand Up @@ -202,23 +193,15 @@ def on_retryable_error(self, error: BaseException) -> bool:
return True

def with_thinking(self, effort: ThinkingEffort) -> Self:
match effort:
case "off":
reasoning_effort = None
case "low":
reasoning_effort = "low"
case "medium":
reasoning_effort = "medium"
case "high" | "xhigh" | "max":
# Kimi's API caps at "high"; xhigh/max are Anthropic-specific.
reasoning_effort = "high"
return self.with_generation_kwargs(reasoning_effort=reasoning_effort).with_extra_body(
new_self = self.with_extra_body(
{
"thinking": {
"type": "enabled" if effort != "off" else "disabled",
}
}
)
new_self._thinking_effort = effort
return new_self

def with_generation_kwargs(self, **kwargs: Unpack[GenerationKwargs]) -> Self:
"""
Expand Down
57 changes: 54 additions & 3 deletions packages/kosong/tests/api_snapshot_tests/test_kimi.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,12 +4,14 @@
from collections.abc import AsyncIterator
from typing import Any

import pytest
import respx
from common import COMMON_CASES, Case, make_chat_completion_response, run_test_cases
from httpx import Response
from inline_snapshot import snapshot
from openai.types.chat import ChatCompletion, ChatCompletionChunk

from kosong.chat_provider import ThinkingEffort
from kosong.chat_provider.kimi import Kimi, KimiStreamedMessage
from kosong.message import Message, TextPart, ThinkPart, ToolCall
from kosong.tooling import Tool
Expand Down Expand Up @@ -507,19 +509,68 @@ async def test_kimi_generation_overrides_take_precedence_over_provider_kwargs():
assert provider.model_parameters["max_completion_tokens"] == 8000


async def test_kimi_with_thinking():
@pytest.mark.parametrize(
("effort", "expected_type"),
[
("off", "disabled"),
("low", "enabled"),
("medium", "enabled"),
("high", "enabled"),
("xhigh", "enabled"),
("max", "enabled"),
],
)
async def test_kimi_with_thinking_omits_legacy_reasoning_effort(
effort: ThinkingEffort,
expected_type: str,
):
with respx.mock(base_url="https://api.moonshot.ai") as mock:
mock.post("/v1/chat/completions").mock(
return_value=Response(200, json=make_chat_completion_response())
)
provider = Kimi(
model="kimi-k2-turbo-preview", api_key="test-key", stream=False
).with_thinking("high")
).with_thinking(effort)
stream = await provider.generate("", [], [Message(role="user", content="Think")])
async for _ in stream:
pass
body = json.loads(mock.calls.last.request.content.decode())
assert body["reasoning_effort"] == snapshot("high")
assert "reasoning_effort" not in body
assert body["thinking"] == {"type": expected_type}
assert provider.thinking_effort == effort


def test_kimi_thinking_effort_preserves_caller_value_without_mapping():
provider = Kimi(model="kimi-k2-turbo-preview", api_key="test-key", stream=False)

assert provider.thinking_effort is None
for configured, expected in (
(provider.with_thinking("off"), "off"),
(provider.with_thinking("low"), "low"),
(provider.with_thinking("medium"), "medium"),
(provider.with_thinking("high"), "high"),
(provider.with_thinking("xhigh"), "xhigh"),
(provider.with_thinking("max"), "max"),
):
assert configured.thinking_effort == expected
assert "reasoning_effort" not in configured.model_parameters


def test_kimi_explicit_legacy_reasoning_effort_is_independent_from_thinking_state():
provider = Kimi(
model="kimi-k2-turbo-preview", api_key="test-key", stream=False
).with_generation_kwargs(reasoning_effort="medium")

assert provider.thinking_effort is None
assert provider.model_parameters["reasoning_effort"] == "medium"

configured = provider.with_thinking("high")
assert configured.thinking_effort == "high"
assert configured.model_parameters["reasoning_effort"] == "medium"

updated_legacy = configured.with_generation_kwargs(reasoning_effort="low")
assert updated_legacy.thinking_effort == "high"
assert updated_legacy.model_parameters["reasoning_effort"] == "low"


async def test_kimi_with_extra_body_thinking_deep_merge():
Expand Down
44 changes: 42 additions & 2 deletions tests/core/test_create_llm.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,13 @@
from kosong.contrib.chat_provider.openai_responses import OpenAIResponses
from pydantic import SecretStr

from kimi_cli.config import LLMModel, LLMProvider
from kimi_cli.llm import augment_provider_with_env_vars, compute_max_completion_tokens, create_llm
from kimi_cli.config import Config, LLMModel, LLMProvider
from kimi_cli.llm import (
augment_provider_with_env_vars,
clone_llm_with_model_alias,
compute_max_completion_tokens,
create_llm,
)


def test_augment_provider_with_env_vars_kimi(monkeypatch):
Expand Down Expand Up @@ -523,6 +528,8 @@ def test_create_llm_kimi_thinking_keep_not_set_omits_field(monkeypatch):
thinking = extra_body.get("thinking") or {}
assert "keep" not in thinking
assert thinking.get("type") == "enabled"
assert llm.chat_provider.thinking_effort == "high"
assert "reasoning_effort" not in llm.chat_provider.model_parameters


def test_create_llm_kimi_thinking_keep_empty_string_omits_field(monkeypatch):
Expand Down Expand Up @@ -586,6 +593,9 @@ def test_create_llm_kimi_thinking_keep_skipped_when_thinking_off(monkeypatch):
extra_body = llm.chat_provider.model_parameters.get("extra_body") or {}
thinking = extra_body.get("thinking") or {}
assert "keep" not in thinking
assert thinking.get("type") == "disabled"
assert llm.chat_provider.thinking_effort == "off"
assert "reasoning_effort" not in llm.chat_provider.model_parameters


def test_create_llm_kimi_thinking_keep_skipped_when_no_thinking_branch(monkeypatch):
Expand All @@ -603,6 +613,8 @@ def test_create_llm_kimi_thinking_keep_skipped_when_no_thinking_branch(monkeypat
# with no thinking key. Both are acceptable; what must hold is "no keep".
thinking = extra_body.get("thinking") or {}
assert "keep" not in thinking
assert llm.chat_provider.thinking_effort is None
assert "reasoning_effort" not in llm.chat_provider.model_parameters


def test_create_llm_kimi_thinking_keep_injected_on_explicit_thinking_true(monkeypatch):
Expand All @@ -626,3 +638,31 @@ def test_create_llm_kimi_thinking_keep_injected_on_explicit_thinking_true(monkey
assert llm.chat_provider.model_parameters.get("extra_body") == snapshot(
{"thinking": {"type": "enabled", "keep": "all"}}
)
assert llm.chat_provider.thinking_effort == "high"
assert "reasoning_effort" not in llm.chat_provider.model_parameters


def test_clone_llm_with_model_alias_preserves_kimi_thinking_off():
provider, model = _make_kimi_plain_model()
model.capabilities = {"thinking"}
llm = create_llm(provider, model, thinking=False)
assert llm is not None

target_model = model.model_copy(update={"model": "kimi-code"})
config = Config(
models={"target": target_model},
providers={"kimi": provider},
)
cloned = clone_llm_with_model_alias(
llm,
config,
"target",
session_id="test-session",
oauth=None,
)

assert cloned is not None
assert isinstance(cloned.chat_provider, Kimi)
assert cloned.chat_provider.thinking_effort == "off"
assert cloned.chat_provider.model_parameters["extra_body"] == {"thinking": {"type": "disabled"}}
assert "reasoning_effort" not in cloned.chat_provider.model_parameters
58 changes: 57 additions & 1 deletion tests/e2e/test_kimi_empty_tool_call_content_e2e.py
Original file line number Diff line number Diff line change
Expand Up @@ -223,14 +223,15 @@ async def _run_kimi_print_json(
share_dir: Path,
work_dir: Path,
prompt: str,
thinking: bool | None = None,
) -> tuple[int, str, str]:
env = os.environ.copy()
env["KIMI_SHARE_DIR"] = str(share_dir)
env["KIMI_DISABLE_TELEMETRY"] = "1"
env["COLUMNS"] = "120"
env["LINES"] = "40"

process = await asyncio.create_subprocess_exec(
args = [
sys.executable,
"-m",
"kimi_cli.cli",
Expand All @@ -244,6 +245,12 @@ async def _run_kimi_print_json(
str(config_path),
"--work-dir",
str(work_dir),
]
if thinking is not None:
args.append("--thinking" if thinking else "--no-thinking")

process = await asyncio.create_subprocess_exec(
*args,
cwd=str(_repo_root()),
env=env,
stdout=asyncio.subprocess.PIPE,
Expand All @@ -261,6 +268,20 @@ def _extract_final_message(stdout: str) -> dict[str, Any]:
return cast(dict[str, Any], json.loads(lines[-1]))


def _assert_thinking_request_body(body: dict[str, Any], *, expected_type: str) -> None:
assert "reasoning_effort" not in body
assert body["thinking"] == {"type": expected_type}


def _assert_session_artifacts(share_dir: Path) -> None:
context_files = list((share_dir / "sessions").rglob("context.jsonl"))
wire_files = list((share_dir / "sessions").rglob("wire.jsonl"))
assert len(context_files) == 1
assert len(wire_files) == 1
assert context_files[0].stat().st_size > 0
assert wire_files[0].stat().st_size > 0


async def test_kimi_compat_endpoint_accepts_tool_call_history_without_empty_content(
tmp_path: Path, mock_kimi_compat_server: MockKimiCompatServer
) -> None:
Expand Down Expand Up @@ -288,6 +309,41 @@ async def test_kimi_compat_endpoint_accepts_tool_call_history_without_empty_cont
}

assert len(mock_kimi_compat_server.requests) == 2
for body in mock_kimi_compat_server.requests:
_assert_thinking_request_body(body, expected_type="disabled")
assistant_message = _find_assistant_tool_call_message(mock_kimi_compat_server.requests[1])
assert assistant_message is not None
assert "content" not in assistant_message
_assert_session_artifacts(share_dir)


async def test_kimi_thinking_uses_type_without_legacy_reasoning_effort(
tmp_path: Path, mock_kimi_compat_server: MockKimiCompatServer
) -> None:
work_dir = tmp_path / "work"
work_dir.mkdir()
(work_dir / "sample.txt").write_text("hello from sample\n", encoding="utf-8")

share_dir = tmp_path / "share"
share_dir.mkdir()

config_path = tmp_path / "config.json"
_write_kimi_config(config_path, base_url=f"{mock_kimi_compat_server.base_url}/v1")

return_code, stdout, stderr = await _run_kimi_print_json(
config_path=config_path,
share_dir=share_dir,
work_dir=work_dir,
prompt="Read sample.txt with ReadFile and then confirm success.",
thinking=True,
)

assert return_code == 0, f"stdout:\n{stdout}\nstderr:\n{stderr}"
assert _extract_final_message(stdout) == {
"role": "assistant",
"content": "Read finished.",
}
assert len(mock_kimi_compat_server.requests) == 2
for body in mock_kimi_compat_server.requests:
_assert_thinking_request_body(body, expected_type="enabled")
_assert_session_artifacts(share_dir)
Loading