
Sentence Transformers is an open-source Python library for developers building semantic search and document retrieval on their own hardware. It runs embedding and reranker models locally, turning content into numerical representations for similarity comparisons and scoring results against a query. The library uses the Apache 2.0 license.
Its scope covers several approaches to search. Dense embeddings support matching by meaning, while sparse encoders work with vocabulary-based representations. Cross-Encoder rerankers score candidate documents for relevance. Multi-Vector Encoder models retain representations of individual tokens for ColBERT-style retrieval, and support ColPali-style search over document images.
The library also handles multimodal content. Supported models can encode text, images, audio or video. Qwen3-VL embedding and reranker models support comparisons between text and images, including ranking mixed text and image documents. Visual document models can match a text query against page images without an OCR step.
Pretrained models are available through Hugging Face, including all-MiniLM-L6-v2, SPLADE and ColBERT-compatible models. The selection includes multilingual and task-specific models. Developers can train, fine-tune and evaluate embedding models, rerankers, sparse encoders and multi-vector encoders on their own data. PyTorch underpins the library, with CUDA support and ONNX and OpenVINO options.
Claim this page with an email at sbert.net. Sentence Transformers gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Sentence Transformers?Promote it
Something wrong or outdated on this page?
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
3.2KUpdated 23 hours agoApache-2.0
#Batch processing#Multilingual#ONNX
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
1.6KUpdated 9 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
FastEmbed is a Python library that generates embeddings on your own hardware for semantic search and retrieval-augmented generation (RAG). It's for developers who need to turn text into searchable vectors without relying on a cloud embedding API. It can run on a CPU or use GPU acceleration, and its Apache 2.0 license makes it open source.
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.