You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Sep 23, 2026. It is now read-only.
Repository navigation
This repository was archived by the owner on Sep 23, 2026. It is now read-only.
When using kimi-cli with an OpenAI-compatible server (e.g. sglang or vllm) that separates reasoning/thinking content into a dedicated response field, the openai_legacy provider drops all reasoning content because reasoning_key is never passed to the OpenAILegacy constructor.
The underlying kosong library's OpenAILegacy class already supports a reasoning_key constructor parameter that tells the stream handler which field to read thinking content from. But create_llm() in llm.py never passes it:
When the model produces a response consisting solely of reasoning/thinking (no text content, no tool calls), the stream yields zero parts and kosong/_generate.py raises APIEmptyResponseError("The API returned an empty response."). The agent crashes even though the server returned a valid response.
Proposed fix
Add an optional reasoning_key: str | None = None field to LLMProvider in config.py.
Pass reasoning_key=provider.reasoning_key when constructing OpenAILegacy in create_llm().
This allows users to configure e.g.:
[providers.vllm]
type = "openai_legacy"base_url = "http://localhost:8000/v1"api_key = "dummy"reasoning_key = "reasoning_content"
Description
When using kimi-cli with an OpenAI-compatible server (e.g. sglang or vllm) that separates reasoning/thinking content into a dedicated response field, the
openai_legacyprovider drops all reasoning content becausereasoning_keyis never passed to theOpenAILegacyconstructor.The underlying kosong library's
OpenAILegacyclass already supports areasoning_keyconstructor parameter that tells the stream handler which field to read thinking content from. Butcreate_llm()inllm.pynever passes it:Impact
When the model produces a response consisting solely of reasoning/thinking (no text content, no tool calls), the stream yields zero parts and
kosong/_generate.pyraisesAPIEmptyResponseError("The API returned an empty response."). The agent crashes even though the server returned a valid response.Proposed fix
reasoning_key: str | None = Nonefield toLLMProviderinconfig.py.reasoning_key=provider.reasoning_keywhen constructingOpenAILegacyincreate_llm().This allows users to configure e.g.:
PR: #1154