BlogsQdrantCPU Performance for Vector Search

CPU Performance for Vector Search

CPU Performance for Vector Search

4
posts
2024–2026

Qdrant continues to optimize performance for vector search, demonstrating significant speed improvements for disk-based retrieval. This post details benchmarks showing Qdrant achieving 2x higher throughput and 50% lower latency than Elastic's DiskBBQ at the same recall target, using significantly less CPU and RAM per node. This is achieved through advanced quantization (TurboQuant 4-bit), async disk scoring, and a two-stage retrieval process. The results highlight Qdrant's efficiency and performance. Version 1.15 introduces new quantization modes including 1.5-bit and 2-bit quantization for improved compression and accuracy, as well as asymmetric quantization for combining binary storage with scalar queries. Text indexing is enhanced with multilingual tokenization, stop words, stemming, and phrase matching. MMR reranking is introduced for diversifying search results. Optimizations include HNSW healing and migration to Gridstore for faster ingestion.

2026

Qdrant 1.19 - TurboQuant Datatype & Memory Tiers - Qdrant

8/5/2026

Introduced Turbo4 datatype for 4-bit vector storage, reducing storage by up to nine times without a full-precision copy. Unified memory configuration with Memory Tiers (pinned, cached, cold). Enabled Per-Tenant IDF Statistics for improved BM25 scoring in multi-tenant deployments. Added prefix matching to keyword indexes and a slice filter condition for deterministic subset partitioning. Enhanced Web UI with live resharding progress, an overhauled Collection Visualizer, and interactive payload index configuration.

Qdrant Beats Elastic’s DiskBBQ at 2x Throughput, Half the Latency, and 1/3 the Compute - Qdrant

7/8/2026

This post presents a benchmark comparing Qdrant's disk-based retrieval performance against Elastic's DiskBBQ. It details Qdrant's optimized configuration using TurboQuant 4-bit, async disk scoring, and a two-stage retrieval plan, demonstrating superior throughput and lower latency compared to Elastic's proprietary solution. The benchmark setup, methodology, and results are provided, along with instructions for reproducing the benchmark. The post emphasizes Qdrant's open-source nature and efficiency gains.

2025

Qdrant 1.15 - Smarter Quantization & better Text Filtering - Qdrant

7/18/2025

Introduced 1.5-bit and 2-bit quantization for improved compression and accuracy, and asymmetric quantization for combining binary storage with scalar queries. Enhanced text indexing with multilingual tokenization, stop words, stemming, and phrase matching. Added Maximal Marginal Relevance (MMR) reranking for diversifying search results. Implemented HNSW healing for more efficient index updates and migration to Gridstore for faster ingestion.

2024

Intel’s New CPU Powers Faster Vector Search - Qdrant

5/10/2024

This post presents benchmark results from Qdrant's R&D division, demonstrating that Intel's 5th gen Xeon processors (Emerald Rapids) offer up to 38% faster vector search compared to the 4th gen Sapphire Rapids. Tests on machines with 32 cores showed 1.38x faster query execution and a 2.79x reduction in latency compared to Sapphire Rapids. The findings suggest that CPUs are a strong choice for vector search, especially for enterprise-scale AI/ML applications, and Qdrant recommends specific core ranges for optimal performance.