Repository navigation
feat(ai): support Azure Foundry Chat Completions deployments - #9714
Conversation
c5080a1 to
5f188f2
Compare
|
@jsanter27 I want to change the provider name to azure as part of this change. things it will break:
I think these are the main things, and they are acceptable. adding provider aliases is lot of complexity for a narrow usecase. wdyt? makes sense to you? |
|
@davidbrai That sounds good with me, I've updated per your suggestion! |
e15211c to
a2ca476
Compare
The Azure provider only implemented the Responses API, so Foundry deployments that speak Chat Completions could not be used. On Chat Completions the deepseek/deepseek-v4-pro model sends `thinking` and `prompt_cache_key`, both of which Azure rejects with 400. The provider now dispatches openai-completions alongside azure-openai-responses. A stream wrapper resolves the Azure endpoint onto the model first, since the shared openai-completions implementation reads it off the model. The request-shape differences are catalog compat rather than api code: reasoning_effort instead of DeepSeek's thinking field, no prompt cache parameters, reasoning_content kept on assistant turns so the prefix Foundry cached stays byte-identical, and effort clamped to the low/medium/high the deployment accepts. The provider is renamed from azure-openai-responses to azure, since it now serves more than the Responses API. The azure-openai-responses api id is unchanged. Breaking: auth.json and models.json entries keyed on azure-openai-responses must move to azure, and existing sessions see a different provider on their next turn. closes earendil-works#9645
a2ca476 to
605a537
Compare
|
Hey @davidbrai just wanted to follow up and get your thoughts on my latest implementation |
|
@jsanter27 did some basic tests and seem to work. |
|
@davidbrai Yep makes sense, I'd considered it originally but wanted to keep the first pass as small as possible. My thinking is to have the Azure provider swap in the deployment name from the map via onPayload, just before the request goes out. Both the Responses and Chat Completions paths would resolve the deployment name the same way from the same map, so the behaviour stays consistent across the two. I also weighed a dedicated Azure Chat Completions adapter, like the Responses one, but it'd mean a new api id and rewiring the existing completions compat, which felt like a lot for this. Sound good? |
|
@jsanter27 yes go ahead with the onPayload approach. please test it works on real deployments, my azure setup is lacking at the moment :) |
The Responses path already sent the deployment name from AZURE_OPENAI_DEPLOYMENT_NAME_MAP. The Chat Completions path sent the catalog model id, so any deployment not named after the model 404'd. The Azure provider now swaps the deployment name into the request via onPayload before dispatch, then runs the caller's own onPayload on it. model.id stays the catalog id. The map helpers move to azure-openai-config.ts so the provider can share them without importing the SDK.
|
@davidbrai Updated the PR and ran my fork against a few real azure deployments I have access to and verified it maps correctly! |
…ename
Breaking: the azure-openai-responses provider is now azure. Users rename
the key in auth.json, models.json, and settings.json; SDK users switch to
getModel("azure", ...) and providers/azure.
Also note Azure Foundry Chat Completions support (DeepSeek V4 Pro).
…l-works#9714) * feat(ai): support Azure Foundry Chat Completions deployments The Azure provider only implemented the Responses API, so Foundry deployments that speak Chat Completions could not be used. On Chat Completions the deepseek/deepseek-v4-pro model sends `thinking` and `prompt_cache_key`, both of which Azure rejects with 400. The provider now dispatches openai-completions alongside azure-openai-responses. A stream wrapper resolves the Azure endpoint onto the model first, since the shared openai-completions implementation reads it off the model. The request-shape differences are catalog compat rather than api code: reasoning_effort instead of DeepSeek's thinking field, no prompt cache parameters, reasoning_content kept on assistant turns so the prefix Foundry cached stays byte-identical, and effort clamped to the low/medium/high the deployment accepts. The provider is renamed from azure-openai-responses to azure, since it now serves more than the Responses API. The azure-openai-responses api id is unchanged. Breaking: auth.json and models.json entries keyed on azure-openai-responses must move to azure, and existing sessions see a different provider on their next turn. closes earendil-works#9645 * feat(ai): apply the Azure deployment name map to Chat Completions The Responses path already sent the deployment name from AZURE_OPENAI_DEPLOYMENT_NAME_MAP. The Chat Completions path sent the catalog model id, so any deployment not named after the model 404'd. The Azure provider now swaps the deployment name into the request via onPayload before dispatch, then runs the caller's own onPayload on it. model.id stays the catalog id. The map helpers move to azure-openai-config.ts so the provider can share them without importing the SDK. * docs(ai,coding-agent): add changelog entries for the Azure provider rename Breaking: the azure-openai-responses provider is now azure. Users rename the key in auth.json, models.json, and settings.json; SDK users switch to getModel("azure", ...) and providers/azure. Also note Azure Foundry Chat Completions support (DeepSeek V4 Pro). --------- Co-authored-by: David Brailovsky <[email protected]>
…l-works#9714) * feat(ai): support Azure Foundry Chat Completions deployments The Azure provider only implemented the Responses API, so Foundry deployments that speak Chat Completions could not be used. On Chat Completions the deepseek/deepseek-v4-pro model sends `thinking` and `prompt_cache_key`, both of which Azure rejects with 400. The provider now dispatches openai-completions alongside azure-openai-responses. A stream wrapper resolves the Azure endpoint onto the model first, since the shared openai-completions implementation reads it off the model. The request-shape differences are catalog compat rather than api code: reasoning_effort instead of DeepSeek's thinking field, no prompt cache parameters, reasoning_content kept on assistant turns so the prefix Foundry cached stays byte-identical, and effort clamped to the low/medium/high the deployment accepts. The provider is renamed from azure-openai-responses to azure, since it now serves more than the Responses API. The azure-openai-responses api id is unchanged. Breaking: auth.json and models.json entries keyed on azure-openai-responses must move to azure, and existing sessions see a different provider on their next turn. closes earendil-works#9645 * feat(ai): apply the Azure deployment name map to Chat Completions The Responses path already sent the deployment name from AZURE_OPENAI_DEPLOYMENT_NAME_MAP. The Chat Completions path sent the catalog model id, so any deployment not named after the model 404'd. The Azure provider now swaps the deployment name into the request via onPayload before dispatch, then runs the caller's own onPayload on it. model.id stays the catalog id. The map helpers move to azure-openai-config.ts so the provider can share them without importing the SDK. * docs(ai,coding-agent): add changelog entries for the Azure provider rename Breaking: the azure-openai-responses provider is now azure. Users rename the key in auth.json, models.json, and settings.json; SDK users switch to getModel("azure", ...) and providers/azure. Also note Azure Foundry Chat Completions support (DeepSeek V4 Pro). --------- Co-authored-by: David Brailovsky <[email protected]>
Summary
Closes #9645. The Azure provider only implemented the Responses API, so Foundry deployments that use Chat Completions (DeepSeek V4 Pro) couldn't work. Picks up @davidbrai's suggestion in #6228 to expand the Azure provider to other APIs.
Only
deepseek-v4-prois in the built-in catalog, but the endpoint resolution is model-agnostic, so other Foundry deployments now route correctly when declared inmodels.jsonunder theazureprovider withapi: "openai-completions": Llama 3.3 70B, Mistral Large 3, Grok 4.6, Kimi K3, Phi-4, etc. (full list)Decisions
model.api, asfireworksandcloudflare-ai-gatewayalready do, rather than a second provider: one Foundry resource, one key, one base URL.azureStreamsresolves the Azure endpoint onto the model before dispatch, sinceopenai-completionsreads it offmodel.baseUrland Azure models ship with an empty baseUrl.AZURE_OPENAI_DEPLOYMENT_NAME_MAPnow applies to Chat Completions too: the provider swaps the deployment name into the request viaonPayload, so both APIs resolve deployments from the same map without changes to the sharedopenai-completionsadapter.azure-openai-responsestoazure(per @davidbrai), since it now serves more than the Responses API. Theazure-openai-responsesapi id is unchanged.Impact
Azure users gain Chat Completions deployments. Breaking:
auth.jsonandmodels.jsonentries keyed onazure-openai-responsesmust move toazure, existing sessions bust their cache on the next turn, and old clients keep a frozen Azure catalog. Env vars are unchanged.Verification
packages/ai/test/azure-openai-completions.test.ts.cacheRetentionvalue plusPI_CACHE_RETENTION=long, multi-turn reasoning replay, tool calling, etc.DeploymentNotFound.