Sorry, you need to enable JavaScript to visit this website.

Enabling Ultra-High-Scale RAG with Cost-Efficient, All-in-Storage ANNS with High-Capacity SSDs

Winchester

Tue Sep 29 | 2:35pm

Abstract

As semantic search and Retrieval-Augmented Generation (RAG) systems scale to billions or even trillions of vectors, the traditional DRAM-based vector search solutions become prohibitively expensive and challenging to scale. The demand for high-scale RAG continues to grow as enterprises seek to index, search, and reason over ever-larger volumes of information.

This talk will explore scaling cost, index build time, and high-capacity SSD-based architectures to efficiently serve large scale semantic search workloads. We will demonstrate that high-density SSD systems can meet RAG latency and throughput requirements at the scale of 10s/100s of billions of vectors offering significantly lower cost, scalable, and a sustainable alternative to memory-based approaches. This session will share the architectural insights of an SSD-friendly vector search engine technology, performance results, and highlight the SSDs as a practical foundation for next generation, web scale semantic search RAG deployments.