
Qdrant is a self-hosted vector database for developers building semantic search, retrieval-augmented generation (RAG), recommendations and AI agent memory. It stores embeddings alongside JSON metadata, so applications can find similar content while restricting results by attributes such as location, text or numeric ranges. The Rust engine is open source under Apache 2.0 and runs locally in Docker or on your own servers.
Hybrid search combines semantic similarity with keyword matching in one query. Qdrant supports dense and sparse vectors, including BM25, SPLADE++ and miniCOIL, plus multiple embeddings per object for models such as ColBERT. Metadata filters apply during search. Relevance controls let developers boost scores with business rules or use Maximum Marginal Relevance to reduce repetitive results.
For larger datasets, quantization and on-disk storage reduce RAM use, while sharding and replication spread data across servers. New vectors become searchable without a full index rebuild. Qdrant uses CPU acceleration and supports NVIDIA and AMD GPUs for indexing. REST and gRPC APIs connect it to applications, and a built-in web interface lets developers inspect collections and test queries.
Self-hosted deployments keep stored vectors and metadata on your infrastructure. Qdrant Cloud is a separate managed service on AWS, GCP or Azure; Cloud Inference generates text and image embeddings there. Hybrid Cloud runs the data side on your Kubernetes infrastructure, while Private Cloud supports air-gapped deployments.
Claim this page with an email at qdrant.tech. Qdrant gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Qdrant?Promote it
Something wrong or outdated on this page?
23.2KUpdated 1 day ago
macOS · Windows · Linux · Docker#Semantic search
pgvector adds vector storage and similarity search to Postgres, so developers can keep embeddings alongside application records in a self-hosted database. It suits applications that need to find similar items while retaining SQL queries, joins and transactional guarantees. It runs on Linux, macOS and Windows, with Docker also supported.
33.1KUpdated 4 weeks ago
Web#Hybrid search#Knowledge graphs#MCP
13KUpdated 21 hours agoApache-2.0
Docker#Agent Skills#Hugging Face integration#Knowledge graphs
26.6KUpdated 6 days agoGPL-3.0
macOS · Linux · Docker#Hybrid search#Multimodal input#RAG
16.9KUpdated 1 day ago
Docker#Hybrid search#RAG#Reranking
4.7KUpdated 1 week agoApache-2.0
macOS · Windows · Linux · Docker#Hybrid search#Reranking#Semantic search
SurrealDB is a self-hosted database for developers building AI agents, knowledge graphs and applications that need several kinds of data together. It stores documents, relationships, vectors and time-series data in one engine, so an application's records and its AI retrieval layer can share the same database.
txtai is a Python framework for developers building search applications, chat with their data, and AI agents on their own hardware or servers. Its embeddings database combines sparse and dense vector search with graphs and relational data, so the same system can find related content and supply context to language models. It's open source under Apache 2.0.
Typesense combines typo-tolerant site search with vector and semantic search in a self-hosted engine. It's for developers building searchable apps, product catalogs or AI search over their own data. The C++ engine uses an in-memory architecture for low-latency results as users type.
Weaviate is a self-hosted vector database for developers building search applications, RAG systems, recommendation engines, and chatbots. It stores data objects alongside their vector embeddings, so applications can search by meaning and filter results using structured data. You can run the database locally with Docker, deploy it on Kubernetes, or use the hosted Weaviate Cloud service.
Infinity is a self-hosted database for developers building search and retrieval-augmented generation (RAG) into LLM applications. It combines embedding search with full-text search and structured filters, so an application can retrieve relevant records through both meaning and exact terms.