Favicon of Sentence Transformers

Sentence Transformers

Python library for running and training embedding and reranker models locally, with Apache 2.0 licensing and pretrained models on Hugging Face.

Screenshot of Sentence Transformers website

Sentence Transformers is an open-source Python library for developers building semantic search and document retrieval on their own hardware. It runs embedding and reranker models locally, turning content into numerical representations for similarity comparisons and scoring results against a query. The library uses the Apache 2.0 license.

Its scope covers several approaches to search. Dense embeddings support matching by meaning, while sparse encoders work with vocabulary-based representations. Cross-Encoder rerankers score candidate documents for relevance. Multi-Vector Encoder models retain representations of individual tokens for ColBERT-style retrieval, and support ColPali-style search over document images.

The library also handles multimodal content. Supported models can encode text, images, audio or video. Qwen3-VL embedding and reranker models support comparisons between text and images, including ranking mixed text and image documents. Visual document models can match a text query against page images without an OCR step.

Pretrained models are available through Hugging Face, including all-MiniLM-L6-v2, SPLADE and ColBERT-compatible models. The selection includes multilingual and task-specific models. Developers can train, fine-tune and evaluate embedding models, rerankers, sparse encoders and multi-vector encoders on their own data. PyTorch underpins the library, with CUDA support and ONNX and OpenVINO options.

Similar to Sentence Transformers