AI Data Storage Engine
Bring Vector Compression to the Extreme: How Milvus Serves 3× More Queries with RaBitQ

Bring Vector Compression to the Extreme: How Milvus Serves 3× More Queries with RaBitQ

5/13/2025 · Alexandr Guzhva, Li Liu, Jiang Chen

What this post added

Introduces and details the integration of RaBitQ, a novel 1-bit vector compression technique, into Milvus. Explains the engineering challenges and tradeoffs in implementing RaBitQ within Milvus's Knowhere search engine, including pre-computation of auxiliary data and hardware acceleration using CPU instructions (VPOPCNTDQ for AVX512). Describes query optimization techniques like scalar quantization on query vectors and refinement. Introduces the new `IVF_RABITQ` index type in Milvus 2.6, which combines RaBitQ with IVF clustering. Provides usage examples and benchmarking results showing a 3x QPS increase with comparable accuracy.

Read the original post ↗