
12/15/2023 · Burcin Bozkaya
What this post added
This post demonstrates the integration of Milvus with NVIDIA Merlin for efficient vector similarity search in recommender workflows. It highlights Milvus's GPU acceleration capabilities, leveraging NVIDIA RAFT, to achieve significant speedups (37x to 91x) in ANN search for user and item embeddings generated by Merlin Models. The technical details cover the use of NVTabular for data preprocessing, Merlin Models for training Two-Tower deep learning models, Milvus for building GPU-accelerated indexes and performing similarity searches, and NVIDIA Triton Inference Server for the inference stage. The post also touches upon Milvus's system design, separating compute and storage for scalability, and its reliance on underlying libraries like cuDF and RAFT for GPU acceleration.