Distributed KV Cache User Guide
This guide shows you how to benchmark a KV cache that spans two or more machines: how to prepare the machines, how to fill in each step of the workload form, and how to read the report. To learn what is measured and why, read the Distributed KV Cache Methodology first.
Before you start
You need the following:
- At least two machines of one GPU family. All machines in a pool use the same GPU model, for example all NVIDIA L40S or all AMD MI300X. The number of GPUs per machine may differ. Machines of different vendors never share a pool.
- The Metrum AI Bench agent on every machine. Follow Adding a server for each machine. The agent keeps running after you log out.
- A cluster group that holds the machines. Create it on the Infrastructure page.
- Access to the model. For a gated model, add a Hugging Face token to the project.
- For a shared file tier: one shared filesystem, such as NFS, Lustre or WEKA, mounted at the same path on every machine and writable by the engines.
- For an RDMA stack: RDMA network cards that are cabled and active on every machine. Preflight reads their state before the run.
- Your switches, connected. Machines find the switches they are cabled to on their own. Connect each switch once with a read-only login, so the report shows what the network did. A result cannot be published while a switch the pool's RDMA cards are cabled to is unread. See Network Switches.
Create the workload
-
Open Projects, and then click New project.
-
Add a workload, and choose the KV Cache Offload tool.
-
In Serving framework, choose the engine: vLLM, SGLang or TensorRT-LLM. TensorRT-LLM runs on NVIDIA only.
-
Choose the model.
-
In KV Cache Storage Tier, choose where the engine keeps KV cache on each machine:
- GPU keeps it in GPU memory only. Use it for every stack that does not start an LMCache server.
- CPU RAM and NVMe SSD start an LMCache server on each machine. Use them for the LMCache stacks.
SGLang offers the GPU tier only, because none of its stacks starts an LMCache server.
-
Check the engine settings. For a model that does not fit on one GPU, set the tensor parallel size so the weights fit. For example, a 235B model in BF16 needs about 470 GB, so it needs at least 4 AMD MI300X GPUs or 8 NVIDIA H100 GPUs per engine.
-
Open Distributed setup steps and work through the seven steps below.
The seven setup steps
Step 1: Cluster
Choose the Cluster group whose machines run the pool. The step shows each machine, its GPUs and its network, and recommends a stack the group can run.
Step 2: Strategy
Choose how the machines share KV cache:
| Strategy | Choose it to |
|---|---|
| Each machine keeps its own cache | Measure the baseline every other strategy is compared with. |
| Spill to host memory and disk | See whether RAM or disk holds conversations the GPU cannot. |
| Share one cache pool across machines | Let any machine reuse KV cache any machine computed. |
| Split prefill and decode across machines | Keep long prompts from slowing down answers that are being written. |
| Split and offload | Combine the split with a tier later turns can reuse. |
Step 3: Stack
Choose the KV stack that runs the strategy on your engine. The list shows only stacks your engine version and GPU vendor support. A stack you cannot choose says why.
- Orchestrator. Bare metal starts the engines through each machine's agent. NVIDIA Dynamo, llm-d and vLLM production-stack are shown with a lock and are coming soon.
- Network and Transfer backend. Leave Automatic unless you need a specific path. For an RDMA stack on a machine with several network cards, set RDMA device to the card the pool uses.
Step 4: Sizing
Set the resources each machine gives the pool:
- GPUs per machine. Leave it empty to use every GPU. The platform refuses a value that does not match the tensor parallel size of a stack that runs one engine per machine.
- Host memory tier. The RAM each machine gives to offloaded KV cache, in GB. Keep it well below the machine's RAM.
- Shared store size. The size of the shared Valkey or Mooncake store on the head machine, in GB.
- Disk tier and Disk tier path (or Shared mount path for a shared file tier). The path must be the same on every machine.
- Prefill machines or Prefill GPUs. For a split stack, how much of the pool reads prompts. The rest writes answers.
Step 5: Workload
Choose what the pool serves:
- Agentic data set replay replays recorded coding-agent sessions. It is
the default. Choose the Datasets entry to say which recorded set to
replay. Leave it empty to use the benchmark's own default set,
sammshen. Only this shape reads a dataset. If you switch to another shape, the dataset is removed from the workload, and choosing agentic replay again restores the default. - Synthetic sessions let you set the prefix length, the number of groups, users per group, turns, input and output length, and think time.
- Shared prompt, per-user history and Public trace replay are also available.
Set Arrivals to closed loop, where each user sends the next turn when the last one ends, or to a request rate. Set a Duration to bound each job.
Step 6: Measurement
- Repeats. Run each test point at least 3 times to get confidence intervals. Fewer repeats are marked preliminary.
- Load duration. How long each job keeps its users sending sessions, from 120 to 14400 s. The default is 1200 s. Leave it empty to replay one session per user, which ramps up and drains. The field appears only for the Agentic data set replay shape. The other shapes set their own length, and their results are measured over the whole run.
- Warm-up. Seconds discarded from the start of the load while the pool fills, from 0 to 3600 s. The default is 300 s. It must be less than half the load duration. Published figures come only from requests that start after the warm-up, up to the end of the load. With the defaults, 15 min are measured. If less than 300 s would be measured, the result carries a warning. See the methodology.
- Working set. A ratio above 1 makes the conversations larger than the pool's GPU memory, which is where offload and shared pools can help.
- Routing. The router that spreads requests across the engines. A split stack needs a router that understands prefill and decode.
- Cache state before each test point. Start from an empty cache, or from a warm one.
Step 7: Review
Review the plan: the machines, the engines each one runs, the tiers and the network path. Save the project, and then run it from Monitoring.
Flavors by vendor and engine
The tables list every stack. What to set names the fields that need a value; the other fields can stay at their defaults.
NVIDIA
| Engine | Stack | Strategy | Tier | What to set |
|---|---|---|---|---|
| vLLM | vLLM, prefix cache off | Baseline | GPU | Nothing extra |
| vLLM | vLLM, GPU prefix cache only | Baseline | GPU | Nothing extra |
| vLLM | vLLM, native CPU offload | Offload | GPU | Host memory tier |
| vLLM | vLLM, native offload to a shared filesystem | Shared pool | GPU | Host memory tier, Shared mount path |
| vLLM | vLLM with LMCache, shared Valkey tier | Shared pool | CPU RAM or NVMe SSD | Shared store size |
| vLLM | vLLM with LMCache, shared Mooncake tier | Shared pool | CPU RAM | Shared store size. Every machine needs an LMCache image built with Mooncake support. |
| vLLM | vLLM with Mooncake Store | Shared pool | GPU | Shared store size |
| vLLM | vLLM prefill and decode over NIXL | Split | GPU | Prefill machines |
| vLLM | vLLM prefill and decode over Mooncake | Split | GPU | Prefill machines |
| vLLM | vLLM prefill and decode over NIXL, with native offload | Split and offload | GPU | Prefill machines, Host memory tier |
| vLLM | vLLM prefill and decode over NIXL, with LMCache | Split and offload | CPU RAM or NVMe SSD | Prefill machines, GPUs per machine |
| vLLM | vLLM prefill and decode through the LMCache store | Split | CPU RAM | Prefill machines, Shared store size |
| SGLang | SGLang, radix cache off | Baseline | GPU | Nothing extra |
| SGLang | SGLang, radix cache only | Baseline | GPU | Nothing extra |
| SGLang | SGLang HiCache, host tier | Offload | GPU | Host memory tier |
| SGLang | SGLang HiCache, shared file tier | Shared pool | GPU | Host memory tier, Shared mount path |
| SGLang | SGLang HiCache, Mooncake tier | Shared pool | GPU | Host memory tier, Shared store size |
| SGLang | SGLang prefill and decode over NIXL | Split | GPU | Prefill machines |
| SGLang | SGLang prefill and decode over Mooncake | Split | GPU | Prefill machines |
| SGLang | SGLang prefill and decode with decode-side offload | Split and offload | GPU | Prefill machines, Host memory tier |
| TensorRT-LLM | TensorRT-LLM, block reuse off | Baseline | GPU | Nothing extra |
| TensorRT-LLM | TensorRT-LLM, block reuse only | Baseline | GPU | Nothing extra |
| TensorRT-LLM | TensorRT-LLM, host offload | Offload | GPU | Host memory tier |
| TensorRT-LLM | TensorRT-LLM prefill and decode over NIXL | Split | GPU | Prefill machines |
| TensorRT-LLM | TensorRT-LLM with LMCache, shared Valkey tier | Shared pool | CPU RAM or NVMe SSD | Shared store size |
AMD
AMD pools run vLLM and SGLang on ROCm images. TensorRT-LLM is not available. Every NVIDIA vLLM and SGLang stack above also runs on AMD, except where the engine version is too old. These stacks run on AMD only:
| Engine | Stack | Strategy | What to set | What the machines need |
|---|---|---|---|---|
| vLLM | vLLM prefill and decode over MoRI-IO | Split | Prefill machines | A MoRI-capable RDMA card (AMD Pollara, Broadcom Thor2 or NVIDIA ConnectX-7) and an image with amd-mori |
| SGLang | SGLang HiCache, MoRI tier (AMD) | Shared pool | Host memory tier | A MoRI-capable RDMA card and an SGLang ROCm image with amd-mori |
| SGLang | SGLang prefill and decode over MoRI (AMD) | Split | Prefill machines | A MoRI-capable RDMA card and an SGLang ROCm image with amd-mori |
For a large model on AMD MI300X, set the tensor parallel size in the engine settings so the weights fit, as in Create the workload.
Validated configurations
A validated configuration has at least one valid job on Metrum AI Bench. These were validated on two NVIDIA L40S machines with Qwen 3 8B:
| Engine | Stacks with a valid run |
|---|---|
| vLLM 0.28.0 | Prefix cache off, GPU prefix cache only, native CPU offload, LMCache with a shared Valkey tier, prefill and decode over NIXL, prefill and decode over NIXL with native offload, prefill and decode over NIXL with LMCache, prefill and decode through the LMCache store |
| SGLang 0.5.20 | Radix cache off, radix cache only, HiCache host tier, prefill and decode over NIXL |
| TensorRT-LLM 1.3.0rc28 | Block reuse off, block reuse only, host offload, prefill and decode over NIXL, LMCache with a shared Valkey tier |
AMD validation on MI300X is in progress. Until an AMD stack has a valid run, treat its results as a first measurement and read the report's validity section.
Read the report
Open the run from Monitoring, and then click Report. At the top, Workloads compared shows every workload of the run side by side. A workload with no valid result says why on its own row.
Each run then answers five questions:
- Did it work? Validity, the window the figures were measured over, stability (variation and drift) with any warning, goodput against a latency target, and the accuracy gate.
- How fast was it? Throughput against latency for every job, time to first token by concurrency, and the throughput and latency charts.
- Did distributed KV help? Where prompt tokens came from, the cache hit charts, the KV transfers between machines, the network fabric, and what each switch saw: congestion, pause frames, drops, and whether every network card agrees with its switch port.
- What limited performance? Tier transfer rates, power and utilization, capacity planning answers, and cost.
- What happened underneath? The pool machine by machine, storage, host memory, and every job.
A run that failed before it measured anything shows its reasons at the top and goes straight to what happened underneath.
When the form refuses a setting
| Message | What to do |
|---|---|
| The stack runs one engine per machine, and a machine would run more | Set GPUs per machine to the tensor parallel size. |
| The stack needs an LMCache server on every node | Choose the CPU RAM or NVMe SSD tier. |
| A shared KV pool shares KV between machines, and this group has one machine | Add a machine to the cluster group, or choose an offload or split stack. |
| Round robin cannot route prefill and decode | In Routing, choose a prefill and decode routing policy. |
| An RDMA port is down on a machine | Check that the card the pool uses is cabled and active. Cards that are not in use do not block the run. |
| A node stopped responding | Check that the agent is running on that machine, and then run the workload again. |
| These switches carried the pool's traffic but were not read during it | Connect each switch named on the Network tab, or ignore a switch that carries no benchmark traffic, and then run the workload again. |