Sorry, you need to enable JavaScript to visit this website.

How CPUs and GPUs Access Micron SSD Efficiently via NVMe LBA Ranges with Reservations

San Tomas + Lawrence

Wed Sep 30 | 4:35pm

Abstract

GPU‑accelerated workloads—ranging from large‑scale AI pipelines to latency‑sensitive SCADA systems—are pushing storage subsystems into a new regime of extreme, fine‑grained I/O, with per‑device demands exceeding 100M IOPS. Meeting these requirements cannot be achieved through faster media alone; it requires a coordinated redesign of how CPUs, GPUs, and NVMe SSDs interact to deliver high throughput, low latency, and strong isolation in shared environments.

This presentation deep dives into emerging architectural techniques that enable CPUs and GPUs to efficiently and securely share high‑performance NVMe SSDs. Building on this foundation, we demonstrate how proposed NVMe LBA range reservations works. We will examine GPU workloads such linear warp & SOL benchmarks engage with CPU based file system-based workloads. We will analyze performance gains, isolation properties, and scalability trade‑offs, highlighting implications for next‑generation GPU‑centric storage architectures.