Stevens Creek
Wed Sep 30 | 4:35pm
In the age of AI and the ever-important focus on data center efficiency, sharing SSDs across multiple independent tenants is fueling a demand for a wide range of PCI Express® and NVM Express® capabilities. This includes established capabilities such as data placement techniques for mitigating noisy neighbors, and multi-function PCI Express® devices (e.g., SR-IOV) for logical device partitioning. New capabilities are already enabling more efficient data and controller state migration (i.e., live migration), while emerging capabilities such as quality of service for deterministic resource allocation and full NVM Subsystem virtualization promises to bring yet another leap in this space.
Technologies supporting SSD Virtualization have evolved substantially from basic device partitioning at the PCI Express® layer to the device itself providing advanced capabilities directly at the NVM Express® layer. In this session, the evolution of SSD Virtualization technologies and capabilities through the past, present and future are explored in detail.
The session begins with an overview of the past, with the concepts of fully emulated and para-virtualized NVM Express® SSDs. The foundation for understanding the technical topics of the session is constructed by diving into the implementation details of these architectures, including the unavoidable overhead of virtual machine manager (VMM) queue mediation and hypervisor (HV) memory-mapped I/O trapping.
The session then proceeds to investigate current virtualization features such as shadow doorbells, device-assisted data change tracking and controller state migration, and how these capabilities substantially reduce the impact of VMM/HV overhead.
Finally, the session covers the most recent and future advancements of SSD Virtualization technologies, including
controlling tenant resource allocation through quality of service capabilities to guarantee service level agreements,
partitioning the physical SSD into logical SSDs through Exported NVM Subsystems, and
how the concept of NVM Subsystem Templating enables true NVM Subsystem virtualization without relying on the hypervisor and virtual machine manager intercepting commands and mediating the responses.
Combined, these past, present and future capabilities contribute to eliminating the last remaining technical obstacles for unlocking native performance of virtualized SSDs.
Data / Storage Architecture
As datacenter architectures evolve, the question isn’t whether more intelligence moves into the data path- it’s where it belongs. From data reduction to placement to protection, key storage functions now span multiple potential homes: the host, the DPU, and the NVMe SSD itself.
This talk explores the shifting boundary between device-level intelligence and infrastructure-level offload, examining where capabilities naturally align and where overlap creates both opportunity and inefficiency. Through the lens of modern NVMe and disaggregated architectures, we revisit a fundamental question: when does pushing functionality closer to the media win, and when does centralizing it in the data path deliver greater impact?
Rather than prescribing a single answer, we frame the tradeoffs that matter for hyperscale and enterprise environments including performance, efficiency, flexibility, and control, challenging assumptions about where storage functionality should reside.
Data / Storage Architecture
An overview of the ways that AI workloads are different than traditional workloads and how this impacts storage devices, including media choices and SoC design.
Data / Storage Architecture
I'd like to demonstrate using an enterprise storage array provide synchronous replication and namespace migration using a stretched subsystem to interact with a host that does not support dispersed namespaces. I will identify some use cases that are not possible with a stretched subsystem and then demonstrate how using a host and storage array that supports dispersed namespaces enable those use cases. While I will be demonstrating this functionality with specific working products, the intention is more to explore the implications of the stretched subsystem and dispersed namespace design decision, while establishing that either is possible to implement.
Data / Storage Architecture
Despite providing massive parallelism and high-bandwidth local memory, modern accelerators typically access storage through CPU-centered I/O paths. As data-intensive workloads increasingly demand high-rate access to small and irregular I/O, these paths become a bottleneck, leading to underutilization of both compute and storage resources.
Accelerator-integrated Storage I/O (AiSIO) describes a class of system-software architectures that integrate accelerators such as GPUs into the storage I/O path while preserving operating-system managed storage abstractions and semantics, including files and file systems. In contrast, conventional storage access relies on CPU-initiated I/O with host-resident payloads, in which data is transferred through host memory before being consumed by accelerators. AiSIO instead encompasses two accelerator-integrated I/O modes: CPU-initiated I/O with device-resident payloads, where the host retains I/O initiation while directing data transfers directly into accelerator memory via peer-to-peer mechanisms, and device-initiated I/O with device-resident payloads, where accelerators themselves construct and submit storage operations using device memory for I/O payloads.
To make these paths interoperable with existing systems, AiSIO relies on host-coordinated control-plane services for device management, metadata handling, and policy enforcement, while allowing data movement and I/O execution to proceed along accelerator-optimized paths. This separation enables accelerators to participate directly in storage access without relinquishing operating-system control or file-system semantics.
This presentation introduces an open-source AiSIO system software implementation and experimental results comparing AiSIO paths to conventional CPU-initiated I/O with host-resident payloads, as well as contrasting CPU-initiated and device-initiated AiSIO modes. The evaluation quantifies the performance impact of integrating accelerators into the storage I/O path under a common Linux storage environment. Ongoing work focuses on upstreaming required operating-system primitives to the Linux kernel and libraries to support broader adoption.
Data / Storage Architecture
Modern storage observability tools expose metrics from isolated layers of the I/O stack, yet many real-world performance problems emerge from interactions across subsystem boundaries rather than within a single layer. Storage engineers frequently observe large gaps between application-visible latency and device-visible latency, but existing tools rarely explain where latency accumulated, which boundary contributed most to delay, or what diagnostic direction should be taken next.
This session explores applying a Cross-Boundary Latency Attribution framework to correlate latency accumulation across the end-to-end I/O path. Instead of analyzing storage behavior from a single subsystem perspective, the framework models latency across multiple boundaries including application execution, syscall handling, filesystem processing, block-layer queueing, driver overhead, device service time, and completion processing.
The presentation examines how traditional observability tools such as iostat, blktrace, and application benchmarks provide only partial visibility into storage behavior. While these tools expose useful metrics independently, they often fail to explain the relationship between application-visible latency and underlying storage activity. The session demonstrates how hidden latency can accumulate in queue wait time, filesystem overhead, scheduler delays, completion handling, CPU contention, and other host-side behaviors that remain difficult to attribute using existing approaches.
The framework applies progressive telemetry collection and low-overhead instrumentation techniques, including eBPF-based tracing, to correlate events across subsystem boundaries while minimizing workload interference. Particular focus is placed on distinguishing queue wait time from device service time, quantifying unexplained latency gaps, and building diagnostic workflows that guide engineers toward the next investigative step instead of overwhelming them with isolated metrics.
Real-world examples will illustrate how cross-boundary analysis can reveal bottlenecks that remain hidden when observing only application latency or device statistics independently. The session will also discuss emerging observability challenges introduced by AI/ML infrastructure, including vector databases, KV-cache offload, disaggregated NVMe memory tiers, and inference storage pipelines, where small hidden delays can amplify into significant application-visible latency.
Attendees will leave with a practical framework for reasoning about end-to-end latency accumulation, a methodology for correlating behavior across traditionally siloed subsystems, and diagnostic strategies that can be applied to modern storage, AI, and high-performance infrastructure environments.
Data / Storage Architecture
AI hasn’t gone away yet, so we’re still focused on how we push the boundaries of sanity for the performance AI needs. Our opinionated experts from SNIA SFF will use their Natural Intelligence (NI) to expound upon AI’s requirements and the relationships between high bandwidth, signal integrity, power and cooling. The SNIA SFF Technical Working Group is addressing signal integrity for hundreds of Gb/s and advanced cooling (74 6F 20 69 6D 70 72 6F 76 65 20 41 49). Come ask our panel of experts questions to help better understand what is being done and request additional functionality to be developed.