ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
The original model combines PaliGemma with ColBERT-style retrieval. It represents different parts of a page separately and compares them with a text query, so visual details can contribute to the match. This approach suits documents where extracting text alone would lose information about the page.
The Python library includes training and inference code for ColPali and related retrievers, including ColQwen2, ColQwen2.5 and ColSmol. Qwen-based models support dynamic image resolution. The code runs with PyTorch and supports NVIDIA CUDA GPUs and Apple Silicon through MPS.
For larger document collections, fast-plaid provides faster matching. Similarity maps show which regions of a page correspond to individual query terms, giving researchers a way to inspect retrieval results. Optional scoring kernels reduce memory use during scoring and training on supported GPUs.
The repository's code is open source under MIT. Model licenses differ: ColPali checkpoints use the Gemma license, while the listed ColQwen2 and ColSmol checkpoints use Apache 2.0. Sentence Transformers supports the model family through MultiVectorEncoder, and colpali-engine remains available for existing projects and research reproduction.
Claim this page and we'll verify you by hand. ColPali gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find ColPali?Promote it
Something wrong or outdated on this page?
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
3.2KUpdated 23 hours agoApache-2.0
#Batch processing#Multilingual#ONNX
1.6KUpdated 9 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
19.1KUpdated 1 week agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
2.2KUpdated 1 day agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
FastEmbed is a Python library that generates embeddings on your own hardware for semantic search and retrieval-augmented generation (RAG). It's for developers who need to turn text into searchable vectors without relying on a cloud embedding API. It can run on a CPU or use GPU acceleration, and its Apache 2.0 license makes it open source.
rerankers is a Python library for developers building search and retrieval systems who want to compare reranking models without rewriting their integration each time. It takes a query and candidate documents, then ranks their relevance through a shared interface across local models and hosted services. It's open source under Apache 2.0.
Sentence Transformers is an open-source Python library for developers building semantic search and document retrieval on their own hardware. It runs embedding and reranker models locally, turning content into numerical representations for similarity comparisons and scoring results against a query. The library uses the Apache 2.0 license.