Qdrant FineWeb-10B: benchmarking four vector databases

Learn why vector benchmarks need exact ground truth, how to test a dataset slice, and how four databases compare on one million vectors.

Player not loading? Watch on YouTube

This video examines Qdrant's FineWeb-10B dataset and uses a smaller slice to compare Qdrant, Milvus, Elasticsearch and pgvector. The speaker explains why small datasets can hide memory pressure and disk access costs, and why recall measurements need exact nearest-neighbor answers rather than another approximate search.

FineWeb-10B contains 10 billion web-document vectors, with dense and sparse embeddings generated using GTE multilingual base. The speaker describes 120,000 queries, including filtered searches, and notes that Qdrant's choices of corpus, model and queries still influence the benchmark. Supernova is the open source toolchain used for embedding generation, database loading and load testing, released under Apache 2.0.

The practical example uses Docker containers and one million vectors. Its script computes fresh brute-force ground truth for that subset, then tests each database separately at five search-effort settings. The full dataset's answers do not apply to a slice because many nearest neighbors may lie outside it.

In the speaker's test, Elasticsearch takes 7.5 milliseconds at 94% recall, but tops out around 95%; Qdrant and Milvus reach 98%. These results describe the tested subset. The speaker explicitly cautions that performance at 10 billion vectors could differ.