AgentX industry impact Β· KV-cache layer
MoRI UMBP
MoRI UMBP (Unified Memory & Bandwidth Pool) is the tiered, distributed KV-cache component of AMDβs MoRI library. It first reached SGLang as a HiCache L3 backend, then became a first-class backend of the KVCache Store Linker. The headline results below were measured on AgentX, serving DeepSeek-V4-Pro disaggregated on MI355X with UMBP as the KV pool.
- stored copies of replicated MLA and DSA KV at TP8
- 8 β 1
- P90 TTFT at concurrency 256 with UMBP on
- β35%
- throughput per GPU on 25% fewer GPUs
- +26β34%
Why agentic serving needs more than a passive byte store
AgentX sessions run about 43 turns and push roughly 1M tokens of context traffic over their life, with a median of 142K input tokens against 444 output tokens per turn. More than 96% of prompt tokens repeat a prefix the server has already seen, so at agentic concurrency the cost of a token is set by where its KV lives rather than by how fast it can be recomputed.
For DeepSeek-V4-Pro on MI355X that raises three problems. The KV cache does not fit in HBM once many long sessions and their subagents are live. It is not sharded across TP ranks, because the MLA latent and the DSA indexer KV are replicated in full on every rank. And a host cache that lives inside the engine process is lost on every restart and every rolling upgrade.
Most L3 KV backends β file stores, Mooncake, NIXL, HF3FS β are passive byte stores: the engine pushes pages down when HBM fills and pulls them back on a hit. UMBP was designed instead around a cluster-wide placement directory spanning HBM, host DRAM, the UMBP DRAM pool, and SSD. Its master tracks per-key access history, node capacity, and fetch latency, so a router can eventually ask not only who holds a prefix but who can serve it fastest, and placement, eviction, and load policies are pluggable interfaces rather than a hard-coded LRU.
The KVCache Store Linker: a direct path from HBM to the pool
UMBP first reached SGLang as a HiCache L3 storage backend, sitting behind a separate L2 host cache on every GPU rank. Running it that way on AgentX exposed six problems, numbered on the left of the figure below. The right side shows, with the same numbers, how the redesign described next fixes each one.
AMD proposed skipping the per-GPU host cache altogether, which matched where the SGLang community was already heading. The MoRI team and SGLang maintainers then built the KVCache Store Linker together, with UMBP supported as a first-class backend alongside HiCache. The linker lets SGLang read cached KV straight from a shared pool in CPU memory and load it into the GPU layer by layer, so the model can start computing before the whole prefix has arrived. The pool can also run as its own process on each machine, so it keeps its contents when the inference engine restarts or is upgraded, and several engine instances can share it.
The linker also stops storing the same data several times. With models like DeepSeek-V4, every GPU in a tensor-parallel group keeps an identical copy of the KV cache, so an 8-GPU group used to write eight copies to CPU memory. Now it writes one: the same memory holds eight times as much cache, and eight times less data moves between CPU memory and the GPUs. A change still under review goes further: instead of every GPU loading the full copy back, each loads a slice and the GPUs exchange the pieces with each other.

Upstream pull requests
What it changed on AgentX
The gain appears even when the DRAM tier is barely used. HBM already served about 95.6% of prompt tokens at concurrency 128β256 and DRAM less than 1%. Turning UMBP on still raised throughput per GPU by 8.3% and cut P90 TTFT by 35% at concurrency 256, and by 2.7% and 34% at concurrency 128. With UMBP off, the host KV pool filled to 72% and 100% while serving at most 0.1% of prompt tokens, so the improvement comes from the direct path rather than from offload.
A large enough cache then changes the topology. Once deduplicated DRAM holds the KV, prefill no longer needs TP8 just for capacity: at concurrency 16β48 the recipe runs a TP4 prefill with a TP8 decode on 12 GPUs instead of 16. UMBP served 34%, 71%, and 75% of prompt tokens at concurrency 16, 32, and 48 while fewer than 3% were recomputed, and throughput per GPU rose 26β34% on 25% fewer GPUs.
Further SGLang optimizations for DeepSeek-V4-Pro on MI355X build on top of UMBP, raising throughput per GPU at concurrency 256 from 55.8k to 62.0k and extending the curve to concurrency 384 and 512. The full configuration is in the October 4 sweep linked below.
Upstream pull requests
Other projects
Inference engine
vLLM
Hybrid-attention prefix retention, CPU KV offload for hybrid models, and a narrowed store and load path.
Read the optimizations βInference engine
SGLang
Sliding-window allocation, HiCache hybrid offload, runtime-scalar context length, and cache-aware DP routing.
Read the optimizations βInference engine
TensorRT-LLM
Boundary-aware incremental tokenization, disaggregated KV descriptor coalescing, and scheduler-lifetime fixes.
Read the optimizations βInference engine
AMD ATOM
Sparse checkpoint retention, recurrent-state checkpoints, CPU offload ownership, and long-prefill parallelism.
Read the optimizations βKernels
ROCm AITER
Context-parallel process groups, 64-bit addressing for large cache pools, and persistent MLA decode kernels.
Read the optimizations βRouter and orchestration
NVIDIA Dynamo
Batched KV matching, request-lease ownership, cheaper router state, and a leaner request plane.
Read the optimizations βKV-cache layer
LMCache
Chunked external-cache loading, hybrid-group storage, AMD Instinct enablement, and DCP-aware offload.
Read the optimizations βTransfer engine
Mooncake
GPU-direct RDMA on ROCm through HIP dmabuf, plus a published ROCm wheel and release path.
Read the optimizations β