Repository navigation
Releases: openvinotoolkit/model_server
Release list
OpenVINO Model Server 2026.4.1
This is a hotfix release including bug fixes in the server component and OpenVINO Runtime.
Improvements included:
- Updated OpenVINO Runtime to 2026.4.1 version.
- DFlash algorithm can not accept requests without setting max_tokens parameter. Earlier, deployment with DFlash algorithm, it was mandatory to set
max_tokensin all request to LLM models. Check Speculative Decoding Demo for instructions how to run it. - Fixed duplicated reporting of model
/v1/modelsendpoint when Idle model management is enabled. #4604
Deprecation notice
Starting from next release the following capabilities are planned to be dropped:
- packages targeted on Ubuntu22 will not be published. Best effort support for building will be supported until 2027.0. There is planned adding Ubuntu26 packages.
- MediaPipe calculators using TensorFlow dependency will be removed from OVMS. In result demos iris_tracking and holistic will be removed. Other calculators with improved performance will be added instead.
You can use OpenVINO Model Server public Docker images based on Ubuntu via the following command:
docker pull openvino/model_server:2026.4.1- CPU device support with image based on Ubuntu 24.04docker pull openvino/model_server:2026.4.1-gpu- GPU, NPU and CPU device support with image based on Ubuntu 24.04
or use provided binary packages. Only packages with suffix _python_on have support for python.
There is also additional distribution channel via https://storage.openvinotoolkit.org/repositories/openvino_model_server/packages/2026.4.1/
OpenVINO Model Server 2026.4.0
New models highlights:
Added support for Qwen3.8-27B model
Improved accuracy for Gemma-4-26B-a4B on GPU accelerators
Added support for Muse-Glimmer-30B model including native function calling
Added support for Kokoro model in speech generation pipeline with execution on NPU accelerator
New capabilities:
Introduced preview support for idle model management, enabling models to be automatically unloaded to free RAM when not in use.
Added option to relax input count validation in KServe API --disable_input_count_validation
Performance:
Added preview support for MTP and DFlash
Added support for Eagle-3 speculative decoding, including tree drafting and chain drafting
Fixes:
Resolved issues in NPU execution for Qwen3-Embedding and Qwen3 when context length exceeds 8k tokens. Added configuration parameter --max_length for context length configuration to reduce memory usage in NPU
Fixed problem in structured output and handling "#" character
Security related improvements
Internal changes and improvements:
Updated Azure SDK and removed boost dependency from the Linux build
You can use OpenVINO Model Server public Docker images based on Ubuntu via the following command:
docker pull openvino/model_server:2026.4.0- CPU device support with image based on Ubuntu 24.04docker pull openvino/model_server:2026.4.0-gpu- GPU, NPU and CPU device support with image based on Ubuntu 24.04
or use provided binary packages. Only packages with suffix _python_on have support for python.
There is also additional distribution channel via https://storage.openvinotoolkit.org/repositories/openvino_model_server/packages/2026.4.0/
OpenVINO Model Server 2026.3.1
This is a hotfix release including minor changes in the server component and a few improvements from OpenVINO Runtime.
Improvements included:
- Corrections in input validation for KServer and audio endpoints
- Fixed structured output generation including # char in the content
- Addressed security vulnerabilities
- Updated OpenVINO Runtime to 2026.3.1 version
- Kokoro model can be used with NPU accelerator
- Gemma4 model has accuracy improvements
Known issues:
- Model Qwen3.6-27B is supported only when exported via optimum-intel with OpenVINO 2026.3 version. The default version in https://huggingface.co/OpenVINO/Qwen3.8-27B-int4-ov fails to load unless branch 2026.3.1 is used. The latest weekly package drop and docker image
openvino/model_server:weeklydon't have such limitation - Model Muse-Glimmer-30B is supported without native tool parsing. Tool parser has been added to 2026.4.0. The latest weekly package drop and docker image
openvino/model_server:weeklydon't have such limitation - OMNI models are supported as a preview feature with limitations - MOE models are not supported yet, concurrency is not possible - requests will be queued in processing.
- NPU execution on LLM has a limit on max prompt parameter 8k tokens. Model will fail to load with higher parameter. It is expected to be fixed in next release.
- NPU plugin for Qwen3-embeddings and Qwen4-reranking fails to load the model when allowed context exceeds 8k tokens. In order to user such models, the workaround is to edit config.json in the model folder and set "max_position_embeddings" to be below 8k. It is expected to be fixed in next release.
You can use OpenVINO Model Server public Docker images based on Ubuntu via the following command:
docker pull openvino/model_server:2026.3.1- CPU device support with image based on Ubuntu 24.04docker pull openvino/model_server:2026.3.1-gpu- GPU, NPU and CPU device support with image based on Ubuntu 24.04
or use provided binary packages. Only packages with suffix _python_on have support for python.
There is also additional distribution channel via https://storage.openvinotoolkit.org/repositories/openvino_model_server/packages/2026.3.1/
OpenVINO Model Server 2026.3.0
New features
Improved performance, accuracy, memory consumption, and stability, with particular focus on new models featuring linear attention architectures like qwen3.6.
Added new tool parsers for LFM2.5 and MiniCPM-V5 models. They can be used in agentic scenarios with native function calling.
Added support for the Jinja chat template engine as an alternative to Minja for Vision Language Models (VLMs). By default, ovms package with python integration use Jinja while the package without python can use Minja. Jinja provides support for wider set of chat templates, but most models are supported by Minja. Current exceptions are lfm2, miniCPMv5 and Qwen3-Coder
Introduced automatic detection of key runtime parameters for generative models, including model task category, target device, and tool parsers, reducing required configuration and minimizing deployment errors. Default target device is determined based on the priority: dGPU if present, iGPU if present, CPU. If there are several discrete GPUs, the one with more free VRAM will be selected.
Simplified deployment of classic models by enabling direct deployment without versioning by pointing ‘--model_path’ to a model file, in addition to a folder with model versions.
Extended OVMS CLI support for easier configuration and deployment of local models, including the ability to add classic models to multi-model config file alongside generative models.
Preview support for Kokoro model in OpenAI API ‘/audio/speech’ endpoint, including multi-language support via the optional eSpeek component.
Extended support for Automatic Speech Recognition (ASR) models beyond the Whisper model family, enabling more generic audio endpoint support for models such as Qwen-ASR.
Preview support for Qwen3-OMNI models. This pipeline exposes chat/completions and /responses API. On input, is accepted text, images and audio. Model can generate text and audio. Both unary and streaming responses are covered.
Enabled /v1 REST version for generative endpoints. It can be used just like v3 which remains for compatiblities reasons with previous version. v1 version makes easier integration with some AI frameworks assuming v1 for OpenAI endpoints.
Added option --verbose_response which makes chat/completions endpoint sending raw model prompts and raw model response. It can be used for debugging purposes or testing. Verbose response is compatible with llama.cpp.
Bug fixes
Addressed known security vulnerabilities
Discontinued
TensorFlow Server API is now discontinued. KServe API is addressing the same use cases but has more capabilities and more efficient communication pattern.
Known issues
- Gemma4 models are supported only on CPU. GPU execution might be not reliable yet. Resolution is expected soon.
- Kokoro is supported on CPU only. GPU and NPU execution is to be enabled soon.
- OMNI models are supported as a preview feature with limitations - MOE models are not supported yet, concurrency is not possible - requests will be queued in processing.
- NPU execution on LLM has a limit on max prompt parameter 8k tokens. Model will fail to load with higher parameter. It is expected to be fixed soon.
- NPU plugin for Qwen3-embeddings and Qwen4-reranking fails to load the model when allowed context exceeds 8k tokens. In order to user such models, the workaround is to edit config.json in the model folder and set "max_position_embeddings" to be below 8k. It is expected to be fixed soon.
OpenVINO Model Server 2026.2.1
This is a minor release focused on bug fixes and improvements related to memory management.
The following changes are included:
- updated OpenVINO Runtime to version 2026.2.1
- updated GPU driver in the docker image with ubuntu24 base image
- updated NPU driver in the docker image with ubuntu24 base image
- improvements in input validation for KServer and TFS API
- added new configuration parameters for LLM and VLM models
--cache_interval_multiplier
New parameter cache_interval_multiplier is relevant only for model models with linear attention like Qwen3.6-35B-A3B. It defines how prefix caching algorithm is managing KV cache allocations. The default value 8 is optimized for short context. While using long prompts like over 20k tokens, it is recommended to increase the value to reduce memory consumption. Here is and example for deploying Qwen3.6-35B-A3B model
ovms --model_repository_path c:\models --source_model OpenVINO/Qwen3.6-35B-A3B-int4-ov ^
--task text_generation --target_device GPU --tool_parser qwen3coder --reasoning_parser qwen3 ^
--rest_port 8000 --cache_interval_multiplier 64
You can use OpenVINO Model Server public Docker images based on Ubuntu via the following command:
docker pull openvino/model_server:2026.2.1- CPU device support with image based on Ubuntu 24.04docker pull openvino/model_server:2026.2.1-gpu- GPU, NPU and CPU device support with image based on Ubuntu 24.04
or use provided binary packages. Only packages with suffix _python_on have support for python.
There is also additional distribution channel via https://storage.openvinotoolkit.org/repositories/openvino_model_server/packages/2026.2.1/
Check the instructions how to install the binary package. The prebuilt image is available also on RedHat Ecosystem Catalog
OpenVINO Model Server 2026.2
Performance
- Improved performance on B60 and B70 for Qwen3-30B MoE model family.
- Improved multinomial algorithm performance, reducing latency for generation with temperature > 0.
- Improved model loading and pipeline initialization performance for new inference requests.
New models and hardware support
- Restored support for generative models on hosts with CPUs without AVX2 instruction set when using supported discrete GPUs.
- Added support for Xe GPUs for MoE models, including Intel Arc A770.
- Enabled execution of GPT-OSS-20b with INT8 precision and GPT-OSS-120b with INT4 precision on GPU.
- Enabled models and support for MoE for Qwen3.5, Qwen3.6, Qwen3-Coder-Next Demo
- Fixed chat template rendering for Granite models when processing non-ASCII characters.
- Added tool parsers for Gemma 4 and LFM2 models.
Deployment ease
-
Improved default performance tuning to use resource constraints in Docker containers, with default number of REST workers, OpenVINO inference streams, threads, and CPU pinning configurations avoiding quota and ulimit settings on Linux to prevent overallocation and performance degradation in Docker and Kubernetes environments.
-
Enhanced deployment capabilities with local generative model startup options and runtime parameter configuration through CLI, enabling generative model deployment from read-only filesystems with configurable runtime parameters such as target device and cache size for seamless KServe and OpenShift integration. Demo
-
Improved model pulling recovery mechanisms to resume interrupted Hugging Face model downloads from the previous checkpoint in case of failures or interruptions. Link
New or improved endpoints capabilities
-
Added initial support for
/responsesendpoint. Reference -
Fixed server readiness endpoint behavior -
/v2/health/readynow correctly reports success when all models are fully initialized and returns appropriate errors when models are not loaded. -
Added
min_psampling parameter for enhanced generation control. Link -
Added
skip_special_tokenssampling parameter - when set to False, returns raw model responses including special tokens to users. Link -
Fixed default seed parameter to use random values, ensuring non-deterministic responses from LLM models.
-
Added LoRA adapter support for image generation models. Demo
-
Added support of streaming for audio/transcriptions endpoint. Link
-
Introduced
OVMS_AUDIO_MAX_FILE_SIZE_BYTESenvironment variable that controls the upper bound on memory that a single audio request can allocate for decoded data. Link
Deprecation notice
As a reminder from the announcement in 2025.4, the following features will be dropped from the next release 2026.3:
- TensorFlow Server API - the alternative for classic model is KServe API. While REST API endpoints in /v1 reserved for TFS will be freed, it is planned to enable v1 for generative OpenAI endpoints. /v3 will remain as an alias for compatibility reasons till next year.
- DAG functionality - the alternative is with MediaPipe integration
- Stateful models
Limitations
-
Gemma 4 and LFM2 MoE models supported without Continuous Batching.
-
/responsesendpoint doesn't include built-in tools, audio input and multimodal output. There are also no session management capabilities. -
Using prefix caching with new Linear Attention models such as Qwen3.5/Qwen3.6 consumes exceeding amount of memory which will be addressed shortly.
You can use an OpenVINO Model Server public Docker images based on Ubuntu via the following command:
docker pull openvino/model_server:2026.2- CPU device support with image based on Ubuntu 24.04docker pull openvino/model_server:2026.2-gpu- GPU, NPU and CPU device support with image based on Ubuntu 24.04
or use provided binary packages. Only packages with suffix _python_on have support for python.
There is also additional distribution channel via https://storage.openvinotoolkit.org/repositories/openvino_model_server/packages/2026.2.0/
Check the instructions how to install the binary package. The prebuilt image is available also on RedHat Ecosystem Catalog
OpenVINO Model Server 2026.1
Enhanced support for Qwen3-MOE models and gpt-oss-20b
They deliver now improved performance, accuracy, and robust concurrent request handling with continuous batching capabilities. These models are now available in pre-optimized OpenVINO™ format directly on the Hugging Face hub, making it very easy to deploy them. Check the demos how to used them
- Integration with agentic framework
- Integration with Visual Studio Code
- Integration with OpenWebUI
Added support for Qwen3-VL
This models family gives function calling capabilities, enabling this vision language model in agentic scenarios. Use examples are included in the demos mentioned above.
Extended /image endpoint to support inpainting and outpainting capabilities.
It is now possible to pass the input image along with a mask to edit parts of the image or to extend the input image.
Check how to use those capabilities in the image generation demo
Other improvements and fixes:
- Server logs now report current KV cache allocation alongside current usage metrics. With dynamic cache size (default setting), allocation automatically scales during runtime based on the request’s concurrency and processed context length.
- Generation request cancellation is now supported for NPU devices, where requests from disconnected clients will be cancelled.
- The finish reason now returns
tool_callswhen the model generates a function call, in line with OpenAI API standards. - Corrected tokens usage reporting in the text generation last streaming event with NPU execution
- Added extra streaming event right after the first token is generated, in line with OpenAI API. This will correct TTFT metric benchmarking using tools relying on streaming events.
- Enhanced error handling for Hugging Face Hub model pulling/downloads includes retry and resume capabilities to address network connectivity issues with large model files. Download operations can now recover from previous network connectivity errors or be reported in logs when recovery is not possible.
You can use an OpenVINO Model Server public Docker images based on Ubuntu via the following command:
docker pull openvino/model_server:2026.1- CPU device support with image based on Ubuntu 24.04docker pull openvino/model_server:2026.1-gpu- GPU, NPU and CPU device support with image based on Ubuntu 24.04
or use provided binary packages. Only packages with suffix _python_on have support for python.
There is also additional distribution channel via https://storage.openvinotoolkit.org/repositories/openvino_model_server/packages/2026.1.0/
Check the instructions how to install the binary package. The prebuilt image is available also on RedHat Ecosystem Catalog
OpenVINO Model Server 2026.0
Performance improvements
- Improvements in performance and accuracy for GPT-OSS and Qwen3-MOE models.
- Improvements in execution performance especially on Intel® Core™ Ultra Series 3 built-in GPUs.
- Better accuracy with INT4 precisions especially with long prompts.
- Corrected handling of compilation cache to speed up model loading.
Agentic use case improvements
- Improved chat template examples to fix handling agentic use cases.
- Improvements in tool parsers to be less restrictive for the generated content and improve response reliability.
- Added support for tool parser compatible with
devstralmodel – take advantage of unsloth/Devstral-Small-2507 model or similar for coding tasks. See the code local assistant demo and LLM reference for details.
Audio endpoints improvements
- Improvements in text2speech endpoint:
- Added voice parameter to choose speaker based on provided embeddings vector.
- Improvements in speech2text endpoint:
- Added handling for temperature sampling parameter.
- Support for timestamps in the output.
Check the Audio endpoints demo
VLM pipeline improvements
- New parameters have been added to VLM pipelines to control domain name restrictions for image URLs in requests, with optional URL redirection support. By default, all URLs are blocked. Use
--allowed_media_domainsand--allowed_local_media_pathto configure the allowed sources. See server parameters for details.
Embeddings and reranker improvements
- NPU execution for text embeddings endpoint (preview). Check the embeddings demo for details.
- Exposed tokenizer endpoint for reranker and LLM pipelines.
Deployment improvements
-
Added configurable preprocessing for classic models. Deployed models can include extra preprocessing layers added in at runtime. This can simplify client implementations and enable sending encoded images to models, which are accepted as an array of input. Possible options include:
- Color format change
- Layout change
- Scale changes
- Mean changes
- Precision change
See server parameters and the ONNX model preprocessing demo for details.
-
Optimized file handle usage to reduce the number of open files during high-load operations on Linux deployments.
New or updated demos
- Audio endpoints
- VLM endpoints usage
- Agentic demo
- Visual Studio Code integration for code assistant
- Image classification
Bug fixes
- Optimized file handle usage to reduce the number of open files during high-load operations on Linux deployments.
- Security improvements.
Known issues
- Qwen3-MOE models like Qwen3-Coder-30B-Instruct or Qwen3-30B-A3B in int4 quantization, when deployed on GPU, might have reduced accuracy with long prompts. Temporary workaround it to set environment variable MOE_USE_MICRO_GEMM_PREFILL=0 before starting ovms process. It will deactivate problematic transformation. It will increase slightly TTFT metric. CPU target device or precisions other than int4 on GPU are not impacted.
- gpt-oss model, when deployed on GPU, should be used only with single concurrency. With high concurrency, there is a risk of impacted accuracy results. CPU target device is not impacted by this issue.
Both issues are expected to be fixed soon.
You can use an OpenVINO Model Server public Docker images based on Ubuntu via the following command:
docker pull openvino/model_server:2026.0- CPU device support with image based on Ubuntu 24.04docker pull openvino/model_server:2026.0-gpu- GPU, NPU and CPU device support with image based on Ubuntu 24.04
or use provided binary packages. Only packages with suffix _python_on have support for python.
There is also additional distribution channel via https://storage.openvinotoolkit.org/repositories/openvino_model_server/packages/2026.0.0/
Check the instructions how to install the binary package. The prebuilt image is available also on RedHat Ecosystem Catalog
OpenVINO Model Server 2025.4.1
2025.4.1 is a minor release with bug fixes and improvements based on OpenVINO 2025.4.1.
Preview:
Added preview support for GPT-OSS agentic use case.
As of 2025.4.1, the best accuracy setting is achieved with:
--pipeline_type LM(without continuous batching and concurrency)--target_device GPU(this configuration was validated on Lunar Lake, Arrow Lake-H, and Intel Arc Battlemage dGPU with >=16 GB VRAM)
It is also required to use INT4 precision.
Bug fixes:
- Fixed escaping for whitespace characters in string arguments for qwen3coder tool-call parser.
- Changed requests handling to
chat/completionsendpoint with streaming and usage tracking to LLM pipelines without continuous batching. Such pipelines do not track generated tokens. So far the last chunk wasn't delivered to the client which could result in a missing token in the response. Now the last chunk is delivered with token usage set as 0 which should be ignored. - Minor documentation and demos fixes
You can use an OpenVINO Model Server public Docker images based on Ubuntu via the following command:
docker pull openvino/model_server:2025.4.1- CPU device support with image based on Ubuntu 24.04
docker pull openvino/model_server:2025.4.1-gpu - GPU, NPU and CPU device support with image based on Ubuntu 24.04
or use provided binary packages. Only packages with sufffix _python_on have support for python.
There is also additional distribution channel via https://storage.openvinotoolkit.org/repositories/openvino_model_server/packages/2025.4.1/
OpenVINO Model Server 2025.4
Agentic use case improvements
- Tool parsers for new models Qwen3-Coder-30B and Qwen3-30B-A3B-Instruct have been enabled. These models are supported in OpenVINO Runtime as a preview feature and can be evaluated with “tool calling” capabilities.
- Streaming with “tool calling” for phi-4-mini-instruct and mistral-7B-v0.4 models is supported just like for the rest of supported agentic models.
- Tool parsers for mistral and hermes3 have been improved, resolving multiple issues related to complex generated JSON objects and increasing overall response reliability.
- Guided generation now supports all rules from XGrammar integration. The
response_formatparameter can now accept XGrammar structural tags format (not part of the OpenAI API). Example:{"type": "regex", "pattern": "\\w+\\.\\w+@company\\.com"}.
New or updated demos
- Integration with OpenWebUI
- Integration with Visual Studio Code using the Continue extension
- Agentic client demo
- Audio endpoints
- Windows service usage
- GGUF model pulling
Deployment improvements
-
GGUF model format can now be deployed directly from Hugging Face Hub for several LLM architectures. Architectures such as Qwen2, Qwen2.5, Qwen3 and Llama3 can be deployed with a single command. See Loading GGUF models in OVMS demo for details.
-
OpenVINO Model Server can be deployed as a service in the Windows operating system. It can be managed by service configuration management, shared by all running applications, and controlled using a simplified CLI to pull, configure, enable, and disable models. LINK
-
Pulling the model in IR format has been extended beyond the OpenVINO™ organization in Hugging Face* Hub. While OpenVINO org models are validated by Intel, a rapidly growing ecosystem of IR-format models from other publishers can now also be pulled and deployed via the OVMS CLI. Note: The repository needs to be populated by
optimum-cli export openvinocommand and must include tokenizer model in IR format to be successfully loaded by OpenVINO Model Server. -
CLI simplifications for easier deployment:
--plugin_configparameter can now be applied not only to classic models but also to generative pipelines.
--cache_dirnow enables compilation caching for both classic models and generative pipelines.
--enable_prefix_cachingcan be used the same way for all target devices. -
--add_to_config and --remove_from_config,like –list_models, are now OVMS CLI directives and no longer expect a value. The configuration values should be passed through the following parameters --config_path, --model_repository_path, --model_name or --model_path.
-
When a service is deployed, the CLI can be simplified by setting the environment variable
OVMS_MODEL_REPOSITORY_PATHto point to the models folder. This automatically applies the default parameters for model management, ensuring that --config_path and --model_repository_path are set correctly. For example:
ovms --pull--task text_generation OpenVINO/Qwen3-8B-int4
ovms --list_models
ovms --add_to_models --model_name OpenVINO/Qwen3-8B-int4
ovms --remove_from_models --model_name OpenVINO/Qwen3-8B-int4 -
The
--api_keyparameter is now available, enabling client authorization using an API key. -
Binding parameters are added for both IPv6 and IPv4 addresses for gRPC and REST interfaces.
-
The metrics endpoint is now compatible with Prometheus v3. The output header type has been updated from JSON to plain text.
Performance improvements
-
First-token generation performance has been significantly improved for LLM models with GPU acceleration and prefix caching. This is particularly beneficial for agentic use cases, where repeated chat history creates very long contexts that can now be processed much faster. Prefix caching can be enabled with OVMS CLI parameter
--enable_prefix_caching true -
A new parameter is introduced to increase the allowed prompt length for LLM and VLLM models deployed on NPU. The context can now be extended by adding the CLI parameter
--max_prompt_lenght. The default is 1024 tokens and can be extended up to 10k tokens. Set it to the required value to avoid unnecessary memory usage. For VLM models running on both NPU and CPU, use a device-specific configuration to apply the setting only to the NPU device:--plugin_config '{"DEVICE_PROPERTIES":{"NPU":{"MAX_PROMPT_LEN":2048}}}' -
Model loading time has been reduced through compilation cache, with significant improvements on GPU and NPU devices. Enable caching using the
--cache_dirparameter. -
Improved guided generation performance, including support for tool call guiding.
Audio endpoints added
- Text to speech endpoint compatible with the OpenAI API - /audio/speech
- Speech to text endpoints compatible with the OpenAI API:
/audio/translation - converts provided audio content to English text
/audio/transcription - converts provided audio content to text in the original language
Embeddings endpoints improvements
- A tokenize endpoint has been added to get tokens before sending the input text to embeddings calculation. This helps assess input length to avoid exceeding the model context.
- Embeddings Model now supports three pooling options: CLS, LAST, and MEAN, widening model compatibility. See Text Embeddings Models list for details.
Breaking changes
The old embeddings calculator was removed and replaced by embeddings_ov. The new calculator follows the optimum-cli / Hugging Face model structure and support more features. If you use the old calculator, re-export your models and pull the updated versions from Hugging Face. Demo
Bug fixes
- Fixed model phi-4-mini-instruct generating incorrect responses when context exceeded 4k tokens
- Other minor fixes
Discontinued in 2025
- Deprecated OpenVINO Model Server (OVMS) benchmark client in C++ using TensorFlow Serving API.
Deprecated to be removed in the future
- The dedicated OpenVINO operator for Kubernetes and OpenShift is now deprecated in favor of the recommended KServe operator. The OpenVINO operator will continue to function in upcoming OpenVINO Model Server releases but will no longer be actively developed. Since KServe provides broader capabilities, no loss of functionality is expected. In contrary, more functionalities will be accessible and migration between other serving solutions and OpenVINO Model Server will be much easier.
- TensorFlow Serving (TFS) API support is planned for deprecation. With increasing adoption of the KServe API for classic models and the OpenAI API for generative workloads, usage of the TFS API has significantly declined. Dropping date is to be determined based on the feedback, with a tentative target of mid-2026.
- Support for Stateful models will be deprecated. This capabilities was originally introduced for Kaldi audio models which is no longer relevant. Current audio models support relies on the OpenAI API, and pipelines implemented via OpenVINO GenAI library.
- Directed Acyclic Graph Scheduler will be deprecated in favor of pipelines managed by MediaPipe scheduler and will be removed mid-2026. That approach gives more flexibility, includes wider range of calculators and has support for using processing accelerators.
You can use an OpenVINO Model Server public Docker images based on Ubuntu via the following command:
docker pull openvino/model_server:2025.4- CPU device support with image based on Ubuntu 24.04docker pull openvino/model_server:2025.4-gpu- GPU, NPU and CPU device support with image based on Ubuntu 24.04
or use provided binary packages. Only packages with sufffix _python_on have support for python.
There is also additional distribution channel via https://storage.openvinotoolkit.org/repositories/openvino_model_server/packages/2025.4.0/
Check the instructions how to install the binary package
The prebuilt image is available also on RedHat Ecosystem Catalog