Favicon of rerankers

rerankers

A Python reranking library with a shared interface for local models and cloud APIs, CPU inference through FlashRank, and an Apache 2.0 license.

rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.

Local choices include SentenceTransformer and Transformers cross-encoders, T5 models such as InRanker and MonoT5, and BGE models based on Gemma and MiniCPM. ColBERT support includes PyLate. For CPU inference, FlashRank provides ONNX-based models. The core package has no dependencies; model-specific backends bring their own requirements.

The same interface also connects to Cohere, Jina, Voyage, MixedBread, Pinecone and Isaacus APIs. These routes send queries and documents to the chosen service, while local model backends perform inference on your hardware. RankGPT and RankLLM support LLM-based document ranking, including GPT models. Hugging Face Text Embeddings Inference is another supported server backend.

The library also handles image reranking through MonoVLMRanker, including MonoQwen2-VL. Results keep document IDs and metadata alongside ranks and scores where the model provides them, so applications can retain source information when selecting the most relevant documents. English and multilingual model choices are available for supported backends.

The project describes this as a beta release. Some RankLLM-backed model integrations are untested, so check the supported-model notes for the backend you choose.

Similar to rerankers