rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.
Local choices include SentenceTransformer and Transformers cross-encoders, T5 models such as InRanker and MonoT5, and BGE models based on Gemma and MiniCPM. ColBERT support includes PyLate. For CPU inference, FlashRank provides ONNX-based models. The core package has no dependencies; model-specific backends bring their own requirements.
The same interface also connects to Cohere, Jina, Voyage, MixedBread, Pinecone and Isaacus APIs. These routes send queries and documents to the chosen service, while local model backends perform inference on your hardware. RankGPT and RankLLM support LLM-based document ranking, including GPT models. Hugging Face Text Embeddings Inference is another supported server backend.
The library also handles image reranking through MonoVLMRanker, including MonoQwen2-VL. Results keep document IDs and metadata alongside ranks and scores where the model provides them, so applications can retain source information when selecting the most relevant documents. English and multilingual model choices are available for supported backends.
The project describes this as a beta release. Some RankLLM-backed model integrations are untested, so check the supported-model notes for the backend you choose.
Claim this page and we'll verify you by hand. rerankers gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find rerankers?Promote it
Something wrong or outdated on this page?
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
3.2KUpdated 23 hours agoApache-2.0
#Batch processing#Multilingual#ONNX
19.1KUpdated 1 week agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
16.3KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Image-to-image#Multilingual
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
FastEmbed is a Python library that generates embeddings on your own hardware for semantic search and retrieval-augmented generation (RAG). It's for developers who need to turn text into searchable vectors without relying on a cloud embedding API. It can run on a CPU or use GPU acceleration, and its Apache 2.0 license makes it open source.
Sentence Transformers is an open-source Python library for developers building semantic search and document retrieval on their own hardware. It runs embedding and reranker models locally, turning content into numerical representations for similarity comparisons and scoring results against a query. The library uses the Apache 2.0 license.
Transformers.js is a JavaScript library for developers building web apps that run AI models on the user's device. Inference happens in the browser, so an app doesn't need a separate model server to process its inputs. The library is open source under Apache 2.0.
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.