Skip to content

About

Restart-proof KV cache for llama.cpp: 40K-token TTFT 155s→6s after reboot. llama-server 重启/关机后自动恢复 prompt 前缀缓存,不重算 prefill。

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

10,726 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llama.cpp + Persistent Prompt Disk Cache

Release License: MIT Patch Backend

一句话:关机、重启、换模型之后,4 万 token 的提示词不用重新预填充,首字时间从 155 秒降到 6 秒。

In one line: kill the server, reboot the box, reload the model — a 40K-token prompt no longer re-runs prefill. Time-to-first-token: ~155 s cold → ~6 s warm.

原理是把 KV 缓存的前缀检查点(.centry 快照)持久化到磁盘,服务启动时自动恢复最长 匹配的前缀。客户端不用改、协议不用变,服务自己记得上次算过什么。 The patch persists KV-cache prefix checkpoints to disk (.centry snapshots) and restores the longest matching prefix on boot. No client changes — the server just remembers.

三步跑起来

这是给第一次尝试的用户准备的最短路径。完整参数说明见 server 文档。

**适用范围:**下面的直接下载包是 NVIDIA CUDA 版本,适用于 Windows 和 Linux。 它不包含模型,也不适用于 AMD/Intel/Apple GPU。Windows 用户可以双击压缩包里的 start-llama-server.bat 自动检查显卡、询问模型路径并启动;需要自定义参数时再用下方命令。

1. 下载带缓存功能的 server

  • Windows CUDA(NVIDIA):从 v0.2.3 Release 下载 llama-server-win-cuda-v0.2.3.zip,解压整个目录,不要只拿走 .exe。
  • Linux CUDA(NVIDIA):同一个 Release 里下载 llama-server-linux-cuda-v0.2.3.tar.gz, 该包为 sm_61 / Pascal 编译,因此 GTX 10 系也能跑。
  • AMD / Intel / Apple GPU:本项目只提供 CUDA 预编译包,其他后端请下载源码按 构建说明 编译。注意本项目是 PrismML/llama.cpp 的 prism 分支补丁, 不是 stock ggml-org/llama.cpp 二进制——直接用官方二进制不会有缓存功能。
  • 只想使用 PrismML 官方预编译程序时,可从 Bonsai-demo 下载对应硬件的包; 但请确认包内 server 已包含本项目的 persistent prompt disk cache 功能。

2. 下载模型

准备一个与 server 兼容的 .gguf 模型。Bonsai 系列模型和文件格式说明见 PrismML Bonsai collection。 模型不在压缩包里,需要单独下载;把真实的模型文件路径替换到下面命令的 -m 参数。不要把模型文件提交到 Git 仓库。

3. 启动 server

Windows PowerShell 示例(路径按本机修改):

# 先确认 NVIDIA 驱动可用;至少应能正常显示显卡信息
nvidia-smi

# 在解压后的 llama-server.exe 所在目录打开 PowerShell
.\llama-server.exe `
  -m D:\models\model.gguf `
  --host 127.0.0.1 --port 8080 `
  -c 32768 -np 1 -ngl 99 -fa on `
  --fit `
  --slot-save-path D:\kvstore\model `
  --prompt-cache-disk-budget 100 `
  --checkpoint-min-step 2048 `
  --ctx-checkpoints 64

Linux 示例:

# 确认 NVIDIA 驱动可用
nvidia-smi

tar xzf llama-server-linux-cuda-v0.2.3.tar.gz
./llama-server \
  -m /models/model.gguf \
  --host 0.0.0.0 --port 8080 \
  -c 32768 -np 1 -ngl 99 -fa on \
  --fit \
  --slot-save-path /var/tmp/kvstore/model \
  --prompt-cache-disk-budget 100 \
  --checkpoint-min-step 2048 \
  --ctx-checkpoints 64

如果不想输入命令,使用压缩包里的 start-llama-server.bat。它会询问模型文件路径, 自动使用压缩包目录旁的 cache 文件夹保存缓存。也可以继续手动使用上面的 PowerShell 命令。

到底需要哪些参数

**不传任何缓存路径也能直接工作。**服务默认自动选择空闲空间最大的磁盘,创建 llama-pdcache 目录,并使用 100 GiB 磁盘预算。想固定位置时再手写 --slot-save-path 或 --prompt-cache-disk-path。其余参数都有合理默认值:

参数 必须? 说明
--slot-save-path 否 固定缓存根目录;不指定时自动选择空闲空间最大的磁盘,使用 <磁盘>/llama-pdcache
--prompt-cache-disk 否 磁盘缓存开关,默认已开启,写不写都行。想临时关掉用 --no-prompt-cache-disk
--prompt-cache-disk-budget 否 磁盘占用上限(GiB),默认 100,0 表示不限。超出后自动淘汰旧条目
--prompt-cache-disk-path 否 想把条目放到别的目录时才用;它的优先级高于自动选盘和 --slot-save-path
--prompt-cache-disk-namespace 否 在同一个根目录下按项目隔离子目录,多个项目共用一份缓存时用
--checkpoint-min-step 否 每增长多少 token 存一次快照,默认 8192。对话分支多就调小(2048),更在意磁盘和内存开销就调大(4096)
--ctx-checkpoints 否 内存里最多保留多少个检查点,默认 32

上面的示例命令把可调项都写出来了,方便你照抄后逐项调整;只想开缓存的话,一个缓存路径参数都不用写。 启动日志会打印实际目录,例如:prompt disk cache enabled: D:\llama-pdcache\ (budget: 100 GiB)。

看到下面的结果后,服务已经可以接受请求:

HTTP GET /health -> {"status":"ok"}

若 nvidia-smi 失败,先修复 NVIDIA 驱动。若 server 报显存不足,先降低 -c; 仍然不足时再逐步增加 -ncmoe 1、2、4。 不要直接套用 GTX 1080 的 -ncmoe 23。RTX 30/40/50 系列的完整调参方法见下方指南。

请求端:cache_prompt 默认就是开启

客户端不需要额外改协议。/completion 的 cache_prompt 默认值是 true, 但建议显式写出,排障时一眼能确认:

$body = @{
    prompt = "Explain persistent KV cache in one paragraph."
    cache_prompt = $true
    timings_per_token = $true
    n_predict = 128
    temperature = 0.2
} | ConvertTo-Json
Invoke-RestMethod http://127.0.0.1:8080/completion `
  -Method Post -ContentType "application/json" -Body $body

把这条请求原样再发一次,第二次才有机会看到缓存复用。服务日志中出现 prompt disk: restored N prompt tokens to device 表示从磁盘恢复; prompt cache = total / reused / recomputed 中的 reused 就是本次跳过的 token 数。

连续请求时,必须保持从开头开始的 prompt 前缀完全一致;只有公共前缀能复用, 不同的后缀仍然需要重新计算。响应中的关键字段:

{
  "tokens_cached": 22528,
  "timings": {
    "prompt_n": 4,
    "prompt_per_second": 250.0,
    "predicted_n": 128,
    "predicted_per_second": 24.5
  }
}

tokens_cached 表示本次请求复用的 prompt token 数;prompt_n 是本次重新处理的 prompt token 数;predicted_per_second 是吐字速度。磁盘缓存只负责在重启或分支切换 后恢复状态,仍然需要 cache_prompt: true 才会进入正常的前缀匹配流程。

缓存什么时候会命中,什么时候不会

这块最容易产生误解,先讲清楚:

会命中:

  • 同一个 prompt 原样重发,哪怕中间关了服务、重启了机器
  • 长对话持续增长:每次请求的 prompt = 旧对话 + 新增几条,历史部分直接复用, 只算新增的。上面表格里 40K 提示词只重算 32 个 token 就是这个场景
  • 服务重启后换了进程,但用同一个 --slot-save-path

不会命中:

  • prompt 前缀变了。缓存只认「从第一个 token 开始的公共前缀」。如果你在最前面插了 一句 system prompt,后面全部错位,缓存等于没有。所以别把时间戳、随机 ID 之类 会变化的内容放在 prompt 最前面
  • 换了模型、上下文长度或 KV 类型。检查点带配置签名,不匹配的条目会被安全忽略 (直接当没缓存,不会算错),换回来才会重新命中
  • cache_prompt: false,或者客户端自己在改写 prompt
  • 缓存目录被删了,或者 --slot-save-path 指向了别的位置

**前缀一致性的实际做法:**把稳定内容(system prompt、工具定义、长文档)放在最前面, 把每轮变化的内容(用户新消息、时间戳)放最后。opencode、Claude Code 这类 agent 客户端 天然就是这个结构,所以开箱即用。

**分支回退时更快:**对话分支(同一段历史换个方向继续)本来要从分叉点重算, 有了磁盘检查点可以直接回退到内存里已有的检查点,不用从磁盘重读。

实现位置

核心代码在 tools/server/server-prompt-disk.cpp(约 900 行)加上 slot/task 的接入, 完整参数语义见 tools/server/README.md。 补丁是相对 PrismML prism 分支的单个 commit:

  • 只想用现成的:下 v0.2.3 Release 的预编译包
  • 要改代码:直接 clone 本仓库(main 是当前开发线),或在 prism 分支上 git am v0.1.0 补丁
English summary (what you get)

This repo is an experimental persistent prompt-prefix cache for the PrismML llama.cpp fork. Long agentic sessions re-send the same growing conversation on every request. This patch makes the server persist the KV-cache prefix to disk and restore it automatically, so a restart (or a new process on the same prompt) skips the prefill it has already paid for.

llama-server -m model.gguf -c 32768 \
  --slot-save-path D:/kvstore/mymodel \    # the only REQUIRED flag
  --checkpoint-min-step 4096 \
  --ctx-checkpoints 64 \
  --prompt-cache-disk-budget 100          # --prompt-cache-disk defaults to ON

What you get:

  • Session restore across restarts — on boot the newest checkpoint for the matching prompt is loaded back into the KV cache automatically.
  • Incremental prefix saves — as a long conversation grows past --checkpoint-min-step, a new .centry snapshot is written; the next request reuses the longest cached prefix.
  • Config-signature safety — checkpoints are keyed by model + context + KV quant; entries from a different configuration are ignored instead of corrupted.
  • Disk budget — old entries are evicted once --prompt-cache-disk-budget is exceeded.

Implementation lives in tools/server/server-prompt-disk.cpp (~900 lines) plus slot/task integration; see tools/server/README.md for the full flag reference. Prebuilt CUDA binaries (Windows + Linux) are on the v0.2.3 Release; main is the active development line.

NVIDIA 驱动要求

The v0.2.3 CUDA packages require an NVIDIA driver that supports CUDA 12.0 or newer. Before starting, run nvidia-smi and update the driver if the command fails or reports an older driver. The Windows package needs a Windows R525-class-or-newer driver; the Linux package was built for CUDA sm_61 / Pascal and also requires a compatible Linux NVIDIA driver. A driver error during startup usually appears as CUDA driver version is insufficient for CUDA runtime version; this means the driver must be updated, not that the model or cache is broken.

实测数据

在 GTX 1080 8 GB + Qwen3.6-35B-A3B-NVFP4-Q4_K_M、-c 92160、4 万 token 智能体提示词下测得:

冷算(无缓存) 命中
预填充 40K ~151 s(265 tok/s) 1.35 s(仅重算 32 个 token)
解码 24–28 tok/s 22–28 tok/s(265 个 token 全程稳定)
端到端 ~155 s ~6 s
磁盘开销 1.5 s 写入 2.1 s 读取+解析+恢复(仅重启后首次)

英文对照 / in English:

Cold (no cache) Full prefix hit
Prefill 40K ~151 s (265 tok/s) 1.35 s (32 tokens recomputed)
Decode 24–28 tok/s 22–28 tok/s (sustained over 265 tokens)
End-to-end ~155 s ~6 s
Disk cost 1.5 s write 2.1 s read + parse + restore (first request after restart only)

MoE 分层经验:专家留 GPU,KV 跟层走。保持 -ncmoe 0 + --fit,hybrid 模型不要加 --no-kv-offload(40 层里只有 10 层带 KV)。把 20 层 MoE 搬去 CPU 会直接崩: 预填充掉到 30 tok/s,解码掉到 4 tok/s。混合模型保持 prefix-only 关闭。

MoE layering lesson (same box): keep experts on GPU (-ncmoe 0 + --fit), let KV follow the layers (hybrid models must not add --no-kv-offload — only 10 of 40 layers carry KV). Pushing 20 MoE layers to CPU collapsed prefill to 30 tok/s and decode to 4 tok/s. Keep --prompt-cache-disk-prefix-only off for hybrid models.

NVIDIA RTX 30/40/50 系列配置指南

下面的配置适用于 CUDA 构建和混合 MoE 模型。-ncmoe 没有跨显卡的固定最佳值:它表示放到 CPU 的 MoE 专家层数,数值越大通常越省显存,但会降低速度。先用 GPU KV 跑通,再根据显存余量调整。

推荐起点

RTX 30/40/50 系列优先从下面的配置开始:

-ngl 99
-ncmoe 0
-c 32768                 # 按需要改成 90112 或其他上下文长度
-np 1
-fa on
-ctk q8_0
-ctv q8_0
--fit
--ctx-checkpoints 64
--prompt-cache-disk
--slot-save-path D:\ai\kvstore\apex-ranges

混合模型默认不要加 --no-kv-offload。GPU KV 通常比 CPU KV 快很多;只有显存确实不够启动,或长上下文导致显存不足时,才考虑 CPU KV。

按显存调整 -ncmoe

推荐用阶梯方式测试,而不是直接套用另一台机器的值:

-ncmoe 0  # 首选:专家和 KV 尽量留在 GPU
-ncmoe 1
-ncmoe 2
-ncmoe 4
-ncmoe 8

每次只增加一个档位,并记录启动后的显存和生成速度。出现 OOM、启动失败或 decode 明显下降时,退回上一个档位。当前 GTX 1080 上使用的 -ncmoe 23 是 8 GB 显存的特殊折中值,不应直接用于 RTX 30/40/50 系列。

经验上,显存更大的卡应先尝试 -ncmoe 0;同一显存下,RTX 40/50 系列通常比 RTX 30 系列有更大的速度余量,但最终结果仍取决于模型量化、上下文长度、KV 类型和 batch 参数。

一套可复现的速度测试

保持模型、prompt、-c、-b、-ub 和生成长度不变,只修改 -ncmoe 或是否使用 --no-kv-offload:

$body = @{ prompt = ("Benchmark text: " + (("The quick brown fox tests MoE inference speed. ") * 180)); n_predict = 128; temperature = 0.2; ignore_eos = $true } | ConvertTo-Json -Compress
Invoke-RestMethod http://127.0.0.1:8000/completion -Method Post `
  -Headers @{ Authorization = "Bearer YOUR_KEY"; "Content-Type" = "application/json" } `
  -Body $body

比较日志中的:

prompt eval time ... tokens per second
eval time ... tokens per second
prompt cache = ... reused ... recomputed

选配置时优先看 eval time 的持续 token/s,其次看显存余量和服务稳定性。第一次请求可能是冷算;速度比较至少应使用相同 prompt 重复一次,并区分 prompt prefill 和 decode。

常见问题

现象 处理
启动 OOM 增大 -ncmoe,或降低 -c;不要先关闭 GPU KV
吐字突然变慢 检查是否加了 --no-kv-offload,并比较 GPU 利用率和 eval time
显存还有余量但速度低 从更小的 -ncmoe 重新测试,专家留 CPU 会增加 PCIe/CPU 路径开销
长 prompt 每次都冷算 检查 prompt 前缀是否一致、--slot-save-path 是否相同、缓存目录是否可写
缓存目录被删除 当前版本会在后台写入前自动重建目录;被删除的旧快照无法恢复,需要重新写入
缓存显示配置不匹配 模型、上下文长度或 KV 类型变化会使旧条目被安全忽略

完整的 90K 上下文示例:

-ngl 99 -ncmoe 0 -c 90112 -np 1 -fa on -ctk q8_0 -ctv q8_0
--fit --ctx-checkpoints 64
--checkpoint-range1-end 30000 --checkpoint-range1-step 1000
--checkpoint-range2-end 40000 --checkpoint-range2-step 256
--checkpoint-range3-step 4096
--slot-save-path D:\ai\kvstore\apex-ranges --prompt-cache-disk

如果该配置显存不足,依次尝试 -ncmoe 1、2、4,直到启动稳定;不要把 GTX 1080 的 -ncmoe 23 作为新显卡默认参数。

Important

This is the PrismML fork of llama.cpp (branch prism), the main line behind the Bonsai models. It tracks current mainline llama.cpp and adds the fork's low-bit formats and runtime features on top. New here? Start with Bonsai-demo — it picks the right models and prebuilt binaries for your hardware/backend automatically.

三元模型文件怎么选 / which model file to use(仅本 fork 相关)

Which ternary model file to use:

  • *-PQ2_0.gguf (fork group-128, ggml id 142): preferred on Metal, CUDA, HIP and CPU. About 6% smaller than group-64.
  • *-Q2_0_g64.gguf / 27B *-Q2_g64.gguf (official group-64, ggml id 42): runs on every backend here AND on mainline llama.cpp. If unsure, use this. Newer model releases name this file plain *-Q2_0.gguf.
  • *-Q2_0.gguf on OLDER model repos is the deprecated legacy format (group 128 stored as id 42). It does not load on these builds; the error tells you which file to get instead. If you must run it, use the frozen prism-v5 line and its final release prism-b9601.

Speculative decoding (dspark) is supported via mainline's draft-dspark plus fork patches. Drafters published for older model releases need a one-time conversion with gguf-dspark-to-dflash (see SPECULATIVE.md in Bonsai-demo); newer releases ship ready-to-use drafters.

Do NOT build from prism-v6 (stale mid-migration snapshot) and do NOT mix this fork's ggml-* libraries with a stock llama.cpp build.


llama

Quick start

A few options to get llama.cpp installed on your machine:

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with `llama cli` VLM session with llama cli Built-in web UI against `llama serve` running Qwen 3.6 Built-in web UI against llama serve

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

Backend Target devices
BLAS All
BLIS All
CANN Ascend NPU
CUDA Nvidia GPU
HIP AMD GPU
Hexagon [In Progress] Snapdragon
IBM zDNN IBM Z & LinuxONE
MUSA Moore Threads GPU
Metal Apple Silicon
OpenCL Adreno GPU
OpenVINO [In Progress] Intel CPUs, GPUs, and NPUs
RPC All
SYCL Intel GPU
VirtGPU VirtGPU APIR
Vulkan GPU
WebGPU All
ZenDNN AMD CPU

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

About

Restart-proof KV cache for llama.cpp: 40K-token TTFT 155s→6s after reboot. llama-server 重启/关机后自动恢复 prompt 前缀缓存,不重算 prefill。

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages