
FastEmbed is a Python library that generates embeddings on your own hardware for semantic search and retrieval-augmented generation (RAG). It's for developers who need to turn text into searchable vectors without relying on a cloud embedding API. It can run on a CPU or use GPU acceleration, and its Apache 2.0 license makes it open source.
The library uses ONNX Runtime and quantized model weights, with few external dependencies and no requirement for PyTorch. That smaller dependency footprint also makes it suitable for serverless environments such as AWS Lambda, where the code runs in the cloud rather than on your machine. Data parallelism supports encoding large datasets.
FastEmbed covers several parts of a retrieval system:
Developers can extend its dense embedding and cross encoder support with models outside the supplied selection. This gives projects room to choose a model suited to their text and retrieval task.
Qdrant maintains the library and integrates it into the Qdrant Python client. That integration connects document embedding with vector storage and text queries, including use with an in-process Qdrant database.
Claim this page and we'll verify you by hand. FastEmbed gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find FastEmbed?Promote it
Something wrong or outdated on this page?
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
1.6KUpdated 9 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
19.1KUpdated 1 week agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.
Sentence Transformers is an open-source Python library for developers building semantic search and document retrieval on their own hardware. It runs embedding and reranker models locally, turning content into numerical representations for similarity comparisons and scoring results against a query. The library uses the Apache 2.0 license.
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.