Skip to content

Commit 497a846

Browse files
atobiszeidtrawins
andauthored
Add idle servable management preview docs (#4523)
Documentation to: #4486 Preview: https://openvino-doc.iotg.sclab.intel.com/atobisze_idle_docs/model-server/ovms_docs_idle_servable_management.html Ticket: CVS-194837 --------- Co-authored-by: Trawinski, Dariusz <[email protected]>
1 parent 53d074c commit 497a846

7 files changed

Lines changed: 107 additions & 0 deletions

File tree

‎docs/features.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -16,6 +16,7 @@ ovms_docs_shape_batch_layout
1616
ovms_docs_dynamic_input
1717
ovms_docs_online_config_changes
1818
ovms_docs_model_version_policy
19+
ovms_docs_idle_servable_management
1920
ovms_docs_metrics
2021
ovms_docs_c_api
2122
ovms_docs_advanced
@@ -66,6 +67,11 @@ OpenVINO Model Server regularly checks for changes to the configuration file and
6667

6768
[Learn more](online_config_changes.md)
6869

70+
## Idle Servable Management (Preview)
71+
Unload inactive model and graph groups, then load them with the next inference request. Reduce CPU and GPU memory needed to serve multiple large models.
72+
73+
[Learn more](./idle_servable_management.md)
74+
6975
## Metrics
7076
Use the metrics endpoint compatible with the Prometheus to access performance and utilization statistics.
7177

‎docs/idle_servable_management.md‎

Lines changed: 93 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,93 @@
1+
# Idle Servable Management (Preview) {#ovms_docs_idle_servable_management}
2+
3+
> **Preview feature:** Idle servable management is disabled by default. Its behavior and configuration can change in future releases.
4+
5+
Idle servable management keeps configured models and MediaPipe graphs visible to clients while unloading their heavy runtime resources when they are not needed. A request for an unloaded servable loads it again before inference. This lets one OVMS instance serve several large models with limited CPU or GPU memory.
6+
7+
Optionally set an OpenVINO model cache directory with `--cache_dir` to reduce reload latency. This is recommended because reloading from cache avoids recompiling the model. The first request still bears the latency cost of waiting for the model to wake up.
8+
9+
10+
## Groups
11+
12+
A group is a unit that OVMS loads and unloads together. Add `group_name` to a servable classic model or MediaPipe graph configuration. All servables in the same group load and unload together. Use groups for servables that must be available together, such as a graph and its dependent models.
13+
14+
Without `group_name`, every classic model and MediaPipe graph is assigned to its own group. A group named `permanent` stays loaded and is never unloaded.
15+
16+
For groups that contain only one non-permanent model or graph, setting `group_name` is optional. Keeping it explicit in configuration can improve readability.
17+
18+
OVMS keeps one non-permanent group active. A request for servable from another group waits for current group to finish its requests, then loads requested group. First request after unload includes wake-up latency.
19+
20+
You can create or update `config.json` with OVMS CLI:
21+
22+
```txt
23+
set OVMS_MODEL_REPOSITORY_PATH=c:\models
24+
25+
ovms --pull --source_model OpenVINO/Qwen3-30B-A3B-Instruct-2507-int4-ov
26+
ovms --add_to_config --model_name OpenVINO/Qwen3-30B-A3B-Instruct-2507-int4-ov --group_name permanent
27+
28+
ovms --pull --source_model OpenVINO/bge-base-en-v1.5-int8-ov
29+
ovms --add_to_config --model_name OpenVINO/bge-base-en-v1.5-int8-ov --group_name rag
30+
31+
ovms --pull --source_model OpenVINO/bge-reranker-base-int8-ov
32+
ovms --add_to_config --model_name OpenVINO/bge-reranker-base-int8-ov --group_name rag
33+
34+
ovms --pull --source_model OpenVINO/FLUX.1-schnell-int4-ov
35+
ovms --add_to_config --model_name OpenVINO/FLUX.1-schnell-int4-ov
36+
```
37+
38+
The following resulting configuration keeps one large model loaded and groups retrieval models together. It loads image generation on demand. Download the referenced models before starting OVMS.
39+
40+
```json
41+
{
42+
"model_config_list": [
43+
{
44+
"config": {
45+
"name": "OpenVINO/Qwen3-30B-A3B-Instruct-2507-int4-ov",
46+
"base_path": "./OpenVINO/Qwen3-30B-A3B-Instruct-2507-int4-ov",
47+
"group_name": "permanent"
48+
}
49+
},
50+
{
51+
"config": {
52+
"name": "OpenVINO/bge-base-en-v1.5-int8-ov",
53+
"base_path": "./OpenVINO/bge-base-en-v1.5-int8-ov",
54+
"group_name": "rag"
55+
}
56+
},
57+
{
58+
"config": {
59+
"name": "OpenVINO/bge-reranker-base-int8-ov",
60+
"base_path": "./OpenVINO/bge-reranker-base-int8-ov",
61+
"group_name": "rag"
62+
}
63+
},
64+
{
65+
"config": {
66+
"name": "OpenVINO/FLUX.1-schnell-int4-ov",
67+
"base_path": "./OpenVINO/FLUX.1-schnell-int4-ov"
68+
}
69+
}
70+
]
71+
}
72+
```
73+
74+
> **Important:** Concurrent requests targeting loaded and unloaded groups have no scheduling policy in this preview. While requests run on active group, incoming traffic continues to that group. Do not depend on fair routing or a bounded switch time between groups. Organize groups so that one active non-permanent group fits available host and device memory to avoid out-of-memory conditions during swaps.
75+
76+
## Enable idle management
77+
78+
Set `--idle_unload_timeout_seconds` to a positive number when starting OVMS. `0`, the default, disables idle servable management.
79+
Start OVMS with a cache directory:
80+
81+
```text
82+
ovms --config_path c:\models\config.json --cache_dir c:\models\cache --idle_unload_timeout_seconds 60
83+
```
84+
85+
## Servable status
86+
87+
Idle-unloaded servables remain `AVAILABLE` in readiness and status APIs. They wake on next inference request. Status and metrics requests do not reset idle timeout; only inference activity keeps a group loaded.
88+
89+
## Metrics
90+
91+
Enable metrics as described in [Metrics](./metrics.md). `ovms_graph_loaded` is default gauge for MediaPipe graphs. It has label `name` and reports `1` when graph resources are loaded or `0` when idle-unloaded.
92+
93+
Classic models have no equivalent loaded-state metric in this preview. Monitor request latency to observe model wake-up time.

‎docs/metrics.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -31,6 +31,7 @@ Default metrics
3131
| gauge | ovms_streams | name,version | Number of OpenVINO execution streams. |
3232
| gauge | ovms_current_requests | name,version | Number of requests being currently processed by the model server. |
3333
| gauge | ovms_current_graphs | name | Number of MediaPipe graphs in process. |
34+
| gauge | ovms_graph_loaded | name | Whether MediaPipe graph resources are loaded (`1`) or idle-unloaded (`0`). |
3435
| counter | ovms_requests_success | api,interface,method,name,version | Number of successful requests to a model or a DAG. |
3536
| counter | ovms_requests_fail | api,interface,method,name,version | Number of failed requests to a model or a DAG. |
3637
| counter | ovms_requests_accepted | api,interface,method,name | Number of accepted requests which ended up inserting packet(s) into a MediaPipe graph. |

‎docs/models_repository_classic.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -63,6 +63,8 @@ models/
6363
the version number in parameters, by default, the latest version is served.
6464
- As an alternative for local filesystem only, `model_path` / `base_path` can point directly to a single model file (`.xml`, `.onnx`, `.pdmodel`, `.pdiparams`, `.pb`, `.tflite`). In this mode, Model Server exposes synthetic version `1`.
6565
- Every version folder _must_ include model files, that is, .bin and .xml for IR, .onnx for ONNX, .pdiparams and .pdmodel for Paddlepaddle. The file name can be arbitrary.
66+
67+
> **Note:** [Idle servable management](./idle_servable_management.md) (preview) can unload inactive model groups and reload them on inference request. Set `group_name` in model configuration to load and unload related models together.
6668
- Each model defines input and output tensors in the AI graph. The client passes data to model input tensors by filling appropriate entries in the request input map.
6769
- Prediction results can be read from the response output map. By default, OpenVINO™ Model Server uses model tensor names as input and output names in prediction requests and responses. The client passes the input values to the request and reads the results by referring to the corresponding output names.
6870
- You can optionally add a `mapping_config.json` file to customize input and output names. This file maps tensor names to user-friendly keys, which is particularly useful for models with complex tensor naming. Here is an example:

‎docs/models_repository_graph.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -23,6 +23,7 @@ graph_models
2323

2424
In case the graph includes python nodes, there should be included also a python file with the node implementation.
2525

26+
> **Note:** [Idle servable management](./idle_servable_management.md) (preview) can unload inactive graph groups and reload them on inference request. Set `group_name` in graph configuration to load and unload related servables together.
2627
2728
For more information on how to use MediaPipe graphs, refer to the [article](./mediapipe.md).
2829
Check also the documentation about [python nodes](./python_support/reference.md)

‎docs/prepare_generative_use_cases.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -17,3 +17,5 @@ Prepare models using OVMS [pull mode](./pull_hf_models.md) (recommended).
1717
Prepare models using [python script](../demos/common/export_models/README.md).
1818

1919
Prepare models using OVMS with python [optimum pull mode](./pull_optimum_cli.md).
20+
21+
> **Note:** For several large generative models on limited memory, use [Idle servable management](idle_servable_management.md) (preview). Configure `--cache_dir` to reduce model wake-up latency.

‎extras/llama_swap/README.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,8 @@ In scenario when OVMS is installed on a client platform, it might be common that
77

88
While this tool was implemented for llama-cpp project, it can be easily enabled also for OpenVINO Model Server.
99

10+
For native OVMS unloading and reloading of inactive model and graph groups, see [Idle Servable Management (Preview)](../../docs/idle_servable_management.md).
11+
1012

1113
## Prerequisites
1214

0 commit comments

Comments
 (0)