Skip to content

feat(ai): support Azure Foundry Chat Completions deployments - #9714

Merged
davidbrai merged 5 commits into
earendil-works:mainfrom
jsanter27:feat/azure-chat-completions
Oct 5, 2026
Merged

davidbrai merged 5 commits into
earendil-works:mainfrom
jsanter27:feat/azure-chat-completions

Conversation

@jsanter27

@jsanter27 jsanter27 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Closes #9645. The Azure provider only implemented the Responses API, so Foundry deployments that use Chat Completions (DeepSeek V4 Pro) couldn't work. Picks up @davidbrai's suggestion in #6228 to expand the Azure provider to other APIs.

Only deepseek-v4-pro is in the built-in catalog, but the endpoint resolution is model-agnostic, so other Foundry deployments now route correctly when declared in models.json under the azure provider with api: "openai-completions": Llama 3.3 70B, Mistral Large 3, Grok 4.6, Kimi K3, Phi-4, etc. (full list)

Decisions

  • Kept one Azure provider dispatching both APIs by model.api, as fireworks and cloudflare-ai-gateway already do, rather than a second provider: one Foundry resource, one key, one base URL.
  • azureStreams resolves the Azure endpoint onto the model before dispatch, since openai-completions reads it off model.baseUrl and Azure models ship with an empty baseUrl.
  • AZURE_OPENAI_DEPLOYMENT_NAME_MAP now applies to Chat Completions too: the provider swaps the deployment name into the request via onPayload, so both APIs resolve deployments from the same map without changes to the shared openai-completions adapter.
  • Renamed the provider azure-openai-responses to azure (per @davidbrai), since it now serves more than the Responses API. The azure-openai-responses api id is unchanged.

Impact

Azure users gain Chat Completions deployments. Breaking: auth.json and models.json entries keyed on azure-openai-responses must move to azure, existing sessions bust their cache on the next turn, and old clients keep a frozen Azure catalog. Env vars are unchanged.

Verification

  • 15 tests added to packages/ai/test/azure-openai-completions.test.ts.
  • 28 live checks against a DeepSeek V4 Pro deployment on Foundry: all seven thinking levels, every cacheRetention value plus PI_CACHE_RETENTION=long, multi-turn reasoning replay, tool calling, etc.
  • Deployment name map tested against a real Foundry deployment: a mapped name routes to it, and a name with no deployment returns DeploymentNotFound.

@davidbrai

Copy link
Copy Markdown
Contributor

@jsanter27 I want to change the provider name to azure as part of this change.

things it will break:

  1. stored auth in auth.json will need to be fixed or do a relogin
  2. models.json any custom / overrides will need to be fixed
  3. existing sessions: pi will think it's a different model so cache will be busted and reasoning converted to text
  4. old clients will start getting 404 responses to pi.dev model catalog updates, but that just means their catalog will be frozen

I think these are the main things, and they are acceptable. adding provider aliases is lot of complexity for a narrow usecase.

wdyt? makes sense to you?

@jsanter27

Copy link
Copy Markdown
Contributor Author

@davidbrai That sounds good with me, I've updated per your suggestion!

The Azure provider only implemented the Responses API, so Foundry
deployments that speak Chat Completions could not be used. On Chat
Completions the deepseek/deepseek-v4-pro model sends `thinking` and
`prompt_cache_key`, both of which Azure rejects with 400.

The provider now dispatches openai-completions alongside
azure-openai-responses. A stream wrapper resolves the Azure endpoint
onto the model first, since the shared openai-completions
implementation reads it off the model.

The request-shape differences are catalog compat rather than api code:
reasoning_effort instead of DeepSeek's thinking field, no prompt cache
parameters, reasoning_content kept on assistant turns so the prefix
Foundry cached stays byte-identical, and effort clamped to the
low/medium/high the deployment accepts.

The provider is renamed from azure-openai-responses to azure, since it
now serves more than the Responses API. The azure-openai-responses api
id is unchanged. Breaking: auth.json and models.json entries keyed on
azure-openai-responses must move to azure, and existing sessions see a
different provider on their next turn.

closes earendil-works#9645
@jsanter27
jsanter27 force-pushed the feat/azure-chat-completions branch from a2ca476 to 605a537 Compare September 30, 2026 14:32
@jsanter27

Copy link
Copy Markdown
Contributor Author

Hey @davidbrai just wanted to follow up and get your thoughts on my latest implementation

@davidbrai davidbrai self-assigned this Oct 2, 2026
@davidbrai

Copy link
Copy Markdown
Contributor

@jsanter27 did some basic tests and seem to work.
should the completions path also support AZURE_OPENAI_DEPLOYMENT_NAME_MAP like the responses one does. wdyt?

@jsanter27

Copy link
Copy Markdown
Contributor Author

@davidbrai Yep makes sense, I'd considered it originally but wanted to keep the first pass as small as possible. My thinking is to have the Azure provider swap in the deployment name from the map via onPayload, just before the request goes out. Both the Responses and Chat Completions paths would resolve the deployment name the same way from the same map, so the behaviour stays consistent across the two.

I also weighed a dedicated Azure Chat Completions adapter, like the Responses one, but it'd mean a new api id and rewiring the existing completions compat, which felt like a lot for this. Sound good?

@davidbrai

davidbrai commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

@jsanter27 yes go ahead with the onPayload approach. please test it works on real deployments, my azure setup is lacking at the moment :)

The Responses path already sent the deployment name from
AZURE_OPENAI_DEPLOYMENT_NAME_MAP. The Chat Completions path sent the
catalog model id, so any deployment not named after the model 404'd.

The Azure provider now swaps the deployment name into the request via
onPayload before dispatch, then runs the caller's own onPayload on it.
model.id stays the catalog id. The map helpers move to
azure-openai-config.ts so the provider can share them without
importing the SDK.
@jsanter27

Copy link
Copy Markdown
Contributor Author

@davidbrai Updated the PR and ran my fork against a few real azure deployments I have access to and verified it maps correctly!

…ename

Breaking: the azure-openai-responses provider is now azure. Users rename
the key in auth.json, models.json, and settings.json; SDK users switch to
getModel("azure", ...) and providers/azure.

Also note Azure Foundry Chat Completions support (DeepSeek V4 Pro).
@davidbrai
davidbrai merged commit a37306d into earendil-works:main Oct 5, 2026
linull24 pushed a commit to linull24/pi that referenced this pull request Oct 5, 2026
…l-works#9714)

* feat(ai): support Azure Foundry Chat Completions deployments

The Azure provider only implemented the Responses API, so Foundry
deployments that speak Chat Completions could not be used. On Chat
Completions the deepseek/deepseek-v4-pro model sends `thinking` and
`prompt_cache_key`, both of which Azure rejects with 400.

The provider now dispatches openai-completions alongside
azure-openai-responses. A stream wrapper resolves the Azure endpoint
onto the model first, since the shared openai-completions
implementation reads it off the model.

The request-shape differences are catalog compat rather than api code:
reasoning_effort instead of DeepSeek's thinking field, no prompt cache
parameters, reasoning_content kept on assistant turns so the prefix
Foundry cached stays byte-identical, and effort clamped to the
low/medium/high the deployment accepts.

The provider is renamed from azure-openai-responses to azure, since it
now serves more than the Responses API. The azure-openai-responses api
id is unchanged. Breaking: auth.json and models.json entries keyed on
azure-openai-responses must move to azure, and existing sessions see a
different provider on their next turn.

closes earendil-works#9645

* feat(ai): apply the Azure deployment name map to Chat Completions

The Responses path already sent the deployment name from
AZURE_OPENAI_DEPLOYMENT_NAME_MAP. The Chat Completions path sent the
catalog model id, so any deployment not named after the model 404'd.

The Azure provider now swaps the deployment name into the request via
onPayload before dispatch, then runs the caller's own onPayload on it.
model.id stays the catalog id. The map helpers move to
azure-openai-config.ts so the provider can share them without
importing the SDK.

* docs(ai,coding-agent): add changelog entries for the Azure provider rename

Breaking: the azure-openai-responses provider is now azure. Users rename
the key in auth.json, models.json, and settings.json; SDK users switch to
getModel("azure", ...) and providers/azure.

Also note Azure Foundry Chat Completions support (DeepSeek V4 Pro).

---------

Co-authored-by: David Brailovsky <[email protected]>
@jsanter27
jsanter27 deleted the feat/azure-chat-completions branch October 5, 2026 14:35
zhouzhaonan pushed a commit to zhouzhaonan/pi that referenced this pull request Oct 7, 2026
…l-works#9714)

* feat(ai): support Azure Foundry Chat Completions deployments

The Azure provider only implemented the Responses API, so Foundry
deployments that speak Chat Completions could not be used. On Chat
Completions the deepseek/deepseek-v4-pro model sends `thinking` and
`prompt_cache_key`, both of which Azure rejects with 400.

The provider now dispatches openai-completions alongside
azure-openai-responses. A stream wrapper resolves the Azure endpoint
onto the model first, since the shared openai-completions
implementation reads it off the model.

The request-shape differences are catalog compat rather than api code:
reasoning_effort instead of DeepSeek's thinking field, no prompt cache
parameters, reasoning_content kept on assistant turns so the prefix
Foundry cached stays byte-identical, and effort clamped to the
low/medium/high the deployment accepts.

The provider is renamed from azure-openai-responses to azure, since it
now serves more than the Responses API. The azure-openai-responses api
id is unchanged. Breaking: auth.json and models.json entries keyed on
azure-openai-responses must move to azure, and existing sessions see a
different provider on their next turn.

closes earendil-works#9645

* feat(ai): apply the Azure deployment name map to Chat Completions

The Responses path already sent the deployment name from
AZURE_OPENAI_DEPLOYMENT_NAME_MAP. The Chat Completions path sent the
catalog model id, so any deployment not named after the model 404'd.

The Azure provider now swaps the deployment name into the request via
onPayload before dispatch, then runs the caller's own onPayload on it.
model.id stays the catalog id. The map helpers move to
azure-openai-config.ts so the provider can share them without
importing the SDK.

* docs(ai,coding-agent): add changelog entries for the Azure provider rename

Breaking: the azure-openai-responses provider is now azure. Users rename
the key in auth.json, models.json, and settings.json; SDK users switch to
getModel("azure", ...) and providers/azure.

Also note Azure Foundry Chat Completions support (DeepSeek V4 Pro).

---------

Co-authored-by: David Brailovsky <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Azure: support Chat Completions deployments (DeepSeek V4 Pro on Foundry)

2 participants