Skip to content

fix(core): turn off Workers AI thinking through the chat template - #52017

Merged
rekram1-node merged 1 commit into
v2from
workers-ai-thinking
Sep 29, 2026
Merged

rekram1-node merged 1 commit into
v2from
workers-ai-thinking

Conversation

@rekram1-node

Copy link
Copy Markdown
Collaborator

Issue for this PR

No linked issue.

Type of change

  • Bug fix
  • New feature
  • Refactor / code improvement
  • Documentation

What does this PR do?

Workers AI used the generic Chat variant builder, which sends every catalog effort as reasoning_effort. That caused two problems:

  • The none variant does not turn thinking off. The Chat endpoint accepts reasoning_effort: "none" but ignores it on deepseek-v4-flash, deepseek-v4-pro, gemma-4 and glm-5.2; only kimi-k2.6 honors it. (Cloudflare's model metadata lists none as supported, which is where the catalog value comes from.)
  • Models that only offer thinking on/off get no off switch. glm-4.7-flash had no variants at all; qwen3.8 and nemotron had no way to turn thinking off.

Workers AI now gets its own builder, like NVIDIA, Baseten and the other hosting providers already have:

  • none and toggle-off send chat_template_kwargs: { enable_thinking: false, thinking: false }, and toggle-on sends both as true. Cloudflare's per-model input schemas document chat_template_kwargs.enable_thinking as the reasoning switch. Kimi's template reads thinking instead, so both keys are sent.
  • Other effort levels still send reasoning_effort, which works.

How did you verify your code works?

  • New case in packages/core/test/variant.test.ts (fails without the change). bun typecheck passes in packages/core.
  • Live, opencode run --standalone -m cloudflare-workers-ai/<model>#<variant> through a logging proxy, counting reasoning in the response and in the stored message:
Model Before After
deepseek-v4-flash #none reasons (164 chars) no reasoning
deepseek-v4-pro #none reasons no reasoning
gemma-4 #none reasons (373 chars) no reasoning
glm-5.2 #none reasons (142 chars) no reasoning
kimi-k2.6 #none no reasoning no reasoning
glm-4.7-flash #none / #thinking variant unavailable off / reasons (750 chars)
qwen3.8 #none, nemotron #none variant unavailable no reasoning
  • Multi-step bash tool loops with deepseek-v4-flash #none, kimi-k2.6 #none, qwen3.8 #none and glm-4.7-flash #thinking all complete, and prompt cache reads still appear on later steps.

Screenshots / recordings

Not applicable.

Checklist

  • I have tested my changes locally
  • I have not included unrelated changes in this PR

@rekram1-node
rekram1-node merged commit 6712cc4 into v2 Sep 29, 2026
10 checks passed
@rekram1-node
rekram1-node deleted the workers-ai-thinking branch September 29, 2026 05:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant