
ModernBERT is a family of open-source text encoder models for developers building document search, classification and code retrieval on their own hardware. Its longer context lets it process documents and code passages that exceed the limits of older BERT models. The code and models use the Apache 2.0 license.
The base and large checkpoints work with Hugging Face Transformers. They predict masked words out of the box; retrieval, question answering and classification require task-specific fine-tuning. That distinction matters if you're choosing a ready-made embedding model versus a foundation for your own application.
ModernBERT accepts up to 8,192 tokens, compared with the 512-token limit common in earlier encoders. Its training includes substantial code data, making it relevant to search across source code as well as prose. Retrieval training examples cover dense embeddings through Sentence Transformers and ColBERT through PyLate.
Efficiency is a central design choice. Alternating local and global attention reduces the cost of longer inputs, while unpadding avoids computation on empty padding tokens. Flash Attention 2 provides further acceleration on compatible GPUs. The models were evaluated on an NVIDIA RTX 4090, with attention to throughput for inputs of mixed lengths and memory use on smaller GPUs.
The research code builds on MosaicBERT and includes FlexBERT, a modular system for constructing encoders. It provides pre-training and evaluation code, including GLUE tests and retrieval evaluations, for teams that need to train or compare their own models.
Claim this page and we'll verify you by hand. ModernBERT gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find ModernBERT?Promote it
Something wrong or outdated on this page?
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
22.2KUpdated 1 week agoMIT
#Hugging Face integration#Multilingual
5.8KUpdated 2 days agoApache-2.0
#Hugging Face integration#Multilingual#Quantization
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
huggingface.coEmbedding and Reranker Models
#Batch processing#Hugging Face integration#Multilingual
2.2KUpdated 1 day agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
E5 Embeddings is a family of text embedding models for developers building search and retrieval systems on their own hardware. It converts text into numerical representations for matching queries with relevant passages. The family includes English and multilingual models, plus instruction-based variants for task-specific embeddings.
EmbeddingGemma is a text embedding model for developers building search and document features that run on phones, laptops or tablets. Based on Gemma 3, it converts text into numerical representations so applications can find related passages by meaning. Embeddings stay on your hardware, and the model works without an internet connection.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
GTE (General Text Embedding) is Alibaba’s family of downloadable models for representing text as vectors. Developers use these representations to compare queries with documents, cluster related text or supply retrieval components for larger applications.