
12/17/2024 · James Luan
What this post added
This post introduces native lexical search support through Sparse-BM25 in Milvus 2.5, building upon existing hybrid search capabilities. It details the implementation using Tantivy for tokenization and preprocessing, distributed vocabulary and term frequency management, sparse vector generation via TF and TF-IDF, and inverted index support with the WAND algorithm. The post highlights advantages over Elasticsearch, including algorithm flexibility, cost efficiency through compression and quantization, and superior performance in long query optimization by combining sparse embeddings with graph indices.