San Tomas + Lawrence
Wed Sep 30 | 4:35pm
GPU‑accelerated workloads—ranging from large‑scale AI pipelines to latency‑sensitive SCADA systems—are pushing storage subsystems into a new regime of extreme, fine‑grained I/O, with per‑device demands exceeding 100M IOPS. Meeting these requirements cannot be achieved through faster media alone; it requires a coordinated redesign of how CPUs, GPUs, and NVMe SSDs interact to deliver high throughput, low latency, and strong isolation in shared environments.
This presentation deep dives into emerging architectural techniques that enable CPUs and GPUs to efficiently and securely share high‑performance NVMe SSDs. Building on this foundation, we demonstrate how proposed NVMe LBA range reservations works. We will examine GPU workloads such linear warp & SOL benchmarks engage with CPU based file system-based workloads. We will analyze performance gains, isolation properties, and scalability trade‑offs, highlighting implications for next‑generation GPU‑centric storage architectures.
The evolution of AI use cases is driving new requirements for storage systems and devices and the diversification of these use cases place different competing demands on storage. Endurance use cases (AI Checkpointing), Latency use cases (model off-load), high IOPS (GNN training), and high Bandwidth use cases are frequently looking to SLC media to solve their challenges. However, the change to SLC does not solve all problems simultaneously as there are necessary trade-offs to optimize for a given use case
In this session we discuss the unique requirements for use cases in each of these domains and how NAND can be tuned to satisfy each use case's requirements.
Attendees will leave this session with an understanding of what devices and media are best suited for their use cases and what trade-offs are available to optimize their TCO.
The DNA Data Storage Alliance (DDSA), a SNIA Community with more than 40 member organizations drawn equally from academia and industry, is accelerating its work to build the interoperable ecosystem that will bring DNA-based storage from research labs to data centers. This session provides a co-chair update on the Alliance's recent accomplishments and its active 2026 agenda.
In 2025, the Alliance published six major specifications and white papers SNIA, including the DNA Data Storage Technology Review, a comprehensive assessment of technology maturity, commercial readiness metrics, and the remaining challenges to deployment; a companion white paper on codecs for DNA data storage, covering key metrics and technical attributes; and a formal position paper on biosecurity regulatory policy for DNA data storage in the context of U.S. and EU frameworks, including the forthcoming European Biotech Act.
For 2026, the Alliance is advancing on several fronts. Priorities include the publication of a reference paper for Random Access in DNA data storage and computation, the continuation of the efforts to expand on the DNA data stability method and the write-up of a DNA data storage industry progress report. DDSA is also advancing the approval of a Swordfish template to integrate DNA storage into datacenter management frameworks, as a first step towards enabling greater manageability and interoperability for DNA Data Storage resources. The SNIA Alliance is also actively engaging regulators in both the U.S. and EU to shape proportionate biosecurity policy based on the technical realities of synthetic data DNA — distinct from life-science applications.
While DNA data storage has been proven in research and proof-of-concepts, its path to commercialization is currently hindered by key challenges as with many emerging storage technologies in the past: write/read speeds, capacity limitations, equipment size and complexity, and cost. The work of the Alliance in developing standards and best practices is essential to overcoming these barriers. Attendees of this session will gain a clear understanding of the technology's current status, the Alliance's concrete deliverables to date, and the roadmap for future development.
GPU‑accelerated workloads—ranging from large‑scale AI pipelines to latency‑sensitive SCADA systems—are pushing storage subsystems into a new regime of extreme, fine‑grained I/O, with per‑device demands exceeding 100M IOPS. Meeting these requirements cannot be achieved through faster media alone; it requires a coordinated redesign of how CPUs, GPUs, and NVMe SSDs interact to deliver high throughput, low latency, and strong isolation in shared environments.
This presentation deep dives into emerging architectural techniques that enable CPUs and GPUs to efficiently and securely share high‑performance NVMe SSDs. Building on this foundation, we demonstrate how proposed NVMe LBA range reservations works. We will examine GPU workloads such linear warp & SOL benchmarks engage with CPU based file system-based workloads. We will analyze performance gains, isolation properties, and scalability trade‑offs, highlighting implications for next‑generation GPU‑centric storage architectures.
Sprandom significantly reduces the preconditioning time for a Random write workload, and allows the drive to reach random workload steady state at near sequential write speeds. If sequential testing is required prior to random tests, Sprandom forces a full drive write, meaning any prior sequential write to the drive needs to be rewritten. By modifying the Sprandom algorithm to perform a logical write modification to an already sequentially preconditioned drive, a full drive write can be removed from a performance test cycle.
The drive can be sequentially preconditioned allowing any Sequential Write or Sequential Read tests to be performed. Once that testing is complete, a SSprandom (Short Sprandom) modification write pattern can be applied to the drive which will transition the drive to a random precondition steady state with a minimal set of targeted writes to the already sequentially written drive.
This modification to the Sprandom algorithm will allow further preconditioning test time reductions, with the associated PE cycle reduction to preconditioning, and further enhance the Sprandom method. As drives continue to increase in size, additional precondition improvements will be necessary, and this method continues to improve on the original Sprandom algorithm.
Space is poised to become the next big market for digital electronics. Not only are growing satellite networks bringing Internet access to all corners of the earth today, but the future promises a ballooning number of such systems ranging from large defense networks like the US’ proposed Golden Dome array, which promises to launch thousands of small satellites coordinated to protect against intercontinental ballistic missiles, to current tests that will lead to space-based distributed datacenters. All of these systems must deal with high levels of radiation and challenges with energy consumption and heat dissipation. These issues may require abandoning DRAM and NAND flash, both of which are radiation-sensitive energy hogs, in favor of one of the new memory technologies gaining adoption today. Listen to this presentation, delivered by noted industry analysts Tom Coughlin of Coughlin Associates and Jim Handy of Objective Analysis to learn not only why and how this change will occur, but also to discover how a distributed space system design, and the adoption of low-energy memories, will create changes in software and systems architecture that will ripple through and revolutionize terrestrial datacenter design.
As QLC NAND densities continue to increase and cell geometries shrink, die and block failure rates have become a dominant operational concern at datacenter scale. Traditional SSD firmware manages these failures as a "black box," internally consuming spare blocks until exhaustion leads to abrupt drive retirement—often while 99% of its media remains functional.
This session introduces Host-Managed Over-Provisioning (HM-OP), a cooperative architectural shift that moves lifecycle decisions from firmware to the host filesystem. Unlike Flexible Data Placement (FDP), which optimizes for write ingest, HM-OP focuses on lifecycle optimization by allowing drives to "gracefully shrink" in capacity rather than fail.
We will provide a technical deep dive into the HM-OP workflow, including:
• Health Telemetry: Utilizing 90/30/7-day power-off retention counters and real-time spare block visibility.
• The Cooperative Handshake: How the host orchestrates proactive data evacuation from degrading regions.
• Capacity Shrink Mechanisms: Implementing host-initiated MAX_LBA reduction to reclaim performance and extend drive longevity.
• System Integration: Managing unpredictable die failures through hole consolidation and cluster-level erasure coding.
Attendees will leave with an understanding of the proposed NVMe extensions required to standardize this transparent management model and how shifting to a "managed expense" model for wear-out can significantly improve TCO for modern flash deployments.