
Vespa is a self-hosted AI search platform for developers building search, RAG, and recommendation systems over large, changing datasets. It combines retrieval with machine-learned ranking, so an application can find candidate results and evaluate their relevance in the same platform. The code is open source under Apache 2.0. You can run it on your own servers or use the managed Vespa Cloud service, where applications run in the cloud.
Vespa searches text, vectors and structured data together. For RAG applications, it supports hybrid search, relevance models and multi-vector representations, giving developers ways to select context beyond vector similarity alone. It also works with tensors and evaluates machine-learned models during queries.
Its distributed architecture spreads data and model evaluation across multiple nodes. It's built for workloads that need fast responses while their underlying content changes continuously, with support for organizing and aggregating results as part of serving a query. That makes it relevant to teams whose search and ranking requirements extend beyond storing embeddings.
Recommendation and personalization applications can retrieve eligible content and score it with models. E-commerce applications can combine search and recommendations with structured filters across text and image content. For personal or private search, Vespa has a streaming search mode that avoids building indexes when each query accesses only a small portion of the total dataset.
Claim this page with an email at vespa.ai. Vespa gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Vespa?Promote it
Something wrong or outdated on this page?
22.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Hybrid search#Ollama integration
localGPT is a self-hosted AI document chat app for people who want to question and summarise files on their own hardware. Its local Ollama setup keeps documents and conversations on your machine. Answers include source passages, so you can check what the model used.
26.6KUpdated 6 days agoGPL-3.0
macOS · Linux · Docker#Hybrid search#Multimodal input#RAG
1.1KUpdated 3 weeks agoMPL-2.0
Linux#Ollama integration#RAG#Semantic search
4.7KUpdated 1 week agoApache-2.0
macOS · Windows · Linux · Docker#Hybrid search#Reranking#Semantic search
38.7KUpdated 11 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Multimodal input#Ollama integration#OpenAI-compatible API
46.3KUpdated 1 day agoApache-2.0
macOS · Linux#Hybrid search#Semantic search
Typesense combines typo-tolerant site search with vector and semantic search in a self-hosted engine. It's for developers building searchable apps, product catalogs or AI search over their own data. The C++ engine uses an in-memory architecture for low-latency results as users type.
chromem-go is a vector database that runs inside your Go application, so developers can add semantic search or retrieval augmented generation (RAG) without maintaining a separate database server. It stores text alongside embeddings and retrieves related documents for use in LLM answers. Its focus is ordinary application workloads rather than collections containing millions of documents.
Infinity is a self-hosted database for developers building search and retrieval-augmented generation (RAG) into LLM applications. It combines embedding search with full-text search and structured filters, so an application can retrieve relevant records through both meaning and exact terms.
Langchain-Chatchat is a self-hosted application for asking questions about your own documents and using AI agents. It focuses on Chinese-language use and open models, with a fully offline setup that can keep documents and model processing on your hardware. Its code is open source under Apache 2.0.
Milvus is an open-source vector database for developers building RAG applications, image search and recommendation systems. It stores embeddings alongside metadata so applications can retrieve related text, images or multimodal data. You can run it on your own hardware, from a laptop prototype to a distributed production cluster.