San Tomas + Lawrence
Tue Sep 16 | 2:00pm
In this presentation, we will introduce Storage Multi-Queue on Windows. Storage Multi-Queue is a new storage driver architecture on Windows that is optimized for high performance storage hardware. Whereas the previous storage architecture was based around SCSI concepts, the new architecture uses NVMe concepts and data structures. This allows for a more efficient mapping to modern hardware and eliminates many of the bottlenecks associated with the existing SCSI-based stack. In addition to increased performance, Storage Multi-Queue provides more comprehensive queue management to lessen the implementation burden of the miniport. We will present the architecture of Storage Multi-Queue in Windows and how one can develop a storage miniport driver for this new architecture.
Learn about the architecture of Storage Multi-Queue in Windows Learn about the workflows of a storage driver in Storage Multi-Queue Learn how to develop a storage driver for Storage Multi-Queue
Download PDF
Data / Storage Architecture
As datacenter architectures evolve, the question isn’t whether more intelligence moves into the data path- it’s where it belongs. From data reduction to placement to protection, key storage functions now span multiple potential homes: the host, the DPU, and the NVMe SSD itself.
This talk explores the shifting boundary between device-level intelligence and infrastructure-level offload, examining where capabilities naturally align and where overlap creates both opportunity and inefficiency. Through the lens of modern NVMe and disaggregated architectures, we revisit a fundamental question: when does pushing functionality closer to the media win, and when does centralizing it in the data path deliver greater impact?
Rather than prescribing a single answer, we frame the tradeoffs that matter for hyperscale and enterprise environments including performance, efficiency, flexibility, and control, challenging assumptions about where storage functionality should reside.
Data / Storage Architecture
An overview of the ways that AI workloads are different than traditional workloads and how this impacts storage devices, including media choices and SoC design.
Data / Storage Architecture
I'd like to demonstrate using an enterprise storage array provide synchronous replication and namespace migration using a stretched subsystem to interact with a host that does not support dispersed namespaces. I will identify some use cases that are not possible with a stretched subsystem and then demonstrate how using a host and storage array that supports dispersed namespaces enable those use cases. While I will be demonstrating this functionality with specific working products, the intention is more to explore the implications of the stretched subsystem and dispersed namespace design decision, while establishing that either is possible to implement.
Data / Storage Architecture
Despite providing massive parallelism and high-bandwidth local memory, modern accelerators typically access storage through CPU-centered I/O paths. As data-intensive workloads increasingly demand high-rate access to small and irregular I/O, these paths become a bottleneck, leading to underutilization of both compute and storage resources.
Accelerator-integrated Storage I/O (AiSIO) describes a class of system-software architectures that integrate accelerators such as GPUs into the storage I/O path while preserving operating-system managed storage abstractions and semantics, including files and file systems. In contrast, conventional storage access relies on CPU-initiated I/O with host-resident payloads, in which data is transferred through host memory before being consumed by accelerators. AiSIO instead encompasses two accelerator-integrated I/O modes: CPU-initiated I/O with device-resident payloads, where the host retains I/O initiation while directing data transfers directly into accelerator memory via peer-to-peer mechanisms, and device-initiated I/O with device-resident payloads, where accelerators themselves construct and submit storage operations using device memory for I/O payloads.
To make these paths interoperable with existing systems, AiSIO relies on host-coordinated control-plane services for device management, metadata handling, and policy enforcement, while allowing data movement and I/O execution to proceed along accelerator-optimized paths. This separation enables accelerators to participate directly in storage access without relinquishing operating-system control or file-system semantics.
This presentation introduces an open-source AiSIO system software implementation and experimental results comparing AiSIO paths to conventional CPU-initiated I/O with host-resident payloads, as well as contrasting CPU-initiated and device-initiated AiSIO modes. The evaluation quantifies the performance impact of integrating accelerators into the storage I/O path under a common Linux storage environment. Ongoing work focuses on upstreaming required operating-system primitives to the Linux kernel and libraries to support broader adoption.
Data / Storage Architecture
Modern storage observability tools expose metrics from isolated layers of the I/O stack, yet many real-world performance problems emerge from interactions across subsystem boundaries rather than within a single layer. Storage engineers frequently observe large gaps between application-visible latency and device-visible latency, but existing tools rarely explain where latency accumulated, which boundary contributed most to delay, or what diagnostic direction should be taken next.
This session explores applying a Cross-Boundary Latency Attribution framework to correlate latency accumulation across the end-to-end I/O path. Instead of analyzing storage behavior from a single subsystem perspective, the framework models latency across multiple boundaries including application execution, syscall handling, filesystem processing, block-layer queueing, driver overhead, device service time, and completion processing.
The presentation examines how traditional observability tools such as iostat, blktrace, and application benchmarks provide only partial visibility into storage behavior. While these tools expose useful metrics independently, they often fail to explain the relationship between application-visible latency and underlying storage activity. The session demonstrates how hidden latency can accumulate in queue wait time, filesystem overhead, scheduler delays, completion handling, CPU contention, and other host-side behaviors that remain difficult to attribute using existing approaches.
The framework applies progressive telemetry collection and low-overhead instrumentation techniques, including eBPF-based tracing, to correlate events across subsystem boundaries while minimizing workload interference. Particular focus is placed on distinguishing queue wait time from device service time, quantifying unexplained latency gaps, and building diagnostic workflows that guide engineers toward the next investigative step instead of overwhelming them with isolated metrics.
Real-world examples will illustrate how cross-boundary analysis can reveal bottlenecks that remain hidden when observing only application latency or device statistics independently. The session will also discuss emerging observability challenges introduced by AI/ML infrastructure, including vector databases, KV-cache offload, disaggregated NVMe memory tiers, and inference storage pipelines, where small hidden delays can amplify into significant application-visible latency.
Attendees will leave with a practical framework for reasoning about end-to-end latency accumulation, a methodology for correlating behavior across traditionally siloed subsystems, and diagnostic strategies that can be applied to modern storage, AI, and high-performance infrastructure environments.
Data / Storage Architecture
AI hasn’t gone away yet, so we’re still focused on how we push the boundaries of sanity for the performance AI needs. Our opinionated experts from SNIA SFF will use their Natural Intelligence (NI) to expound upon AI’s requirements and the relationships between high bandwidth, signal integrity, power and cooling. The SNIA SFF Technical Working Group is addressing signal integrity for hundreds of Gb/s and advanced cooling (74 6F 20 69 6D 70 72 6F 76 65 20 41 49). Come ask our panel of experts questions to help better understand what is being done and request additional functionality to be developed.