Sorry, you need to enable JavaScript to visit this website.

Empirical Analysis of KV Cache Offload for Scalable AI Inference Architectures

Winchester

Wed Sep 30 | 10:35am

Abstract

Recent work across the industry has already established that KV Cache reuse across custom shared and disaggregated cache architectures, RDMA‑enabled data paths, and flash‑backed storage can significantly improve the scalability and efficiency of large‑scale LLM inference. These efforts appropriately focus on new systems that elevate the KV Cache to a first‑class resource, often relying on specialized protocols, purpose‑built services, or custom infrastructure.

This presentation takes a different and complementary approach. We share Micron’s measured evidence that capacity‑centric inference designs, combined with targeted upgrades to lower‑cost tiers of fast, persistent storage, can materially improve near‑term performance while strengthening long‑term returns on CapEx. Drawing from a broad, vendor‑agnostic evaluation of KV Cache offload behavior, analyzed from the perspective of NVMe SSDs and grounded in realistic, production‑inspired workloads, we focus on how KV Cache Reads and Writes actually manifest across today’s widely deployed storage, networking, and accelerator environments. Rather than proposing clean‑slate architectures or fabric‑specific solutions, we examine measured access patterns and drive‑level optimizations that naturally align with both sequential and random KV Cache I/O behavior. Attendees will gain actionable tiering heuristics and architectural insights that enable scalable inference by extending existing systems, balancing TCO across the memory‑storage hierarchy, and extracting greater efficiency without disruptive infrastructure changes.

Finally, we extend these insights to architectural considerations for designing dedicated KV Cache tiers or "KV Cache boxes" built from existing ecosystem components such as custom xPUs, interconnects, and high‑performance NVMe storage. Using measured drive behavior as a first‑order input, we discuss how inference systems may evolve as reliance on KV Cache reuse increases with multi‑turn, multi‑agent interactions and shared resources across future‑proof data center infrastructure.