Favicon of ModernBERT

ModernBERT

Open-source text encoder models under Apache 2.0 for local retrieval and classification, with long context and Hugging Face Transformers support.

Screenshot of ModernBERT website

ModernBERT is a family of open-source text encoder models for developers building document search, classification and code retrieval on their own hardware. Its longer context lets it process documents and code passages that exceed the limits of older BERT models. The code and models use the Apache 2.0 license.

The base and large checkpoints work with Hugging Face Transformers. They predict masked words out of the box; retrieval, question answering and classification require task-specific fine-tuning. That distinction matters if you're choosing a ready-made embedding model versus a foundation for your own application.

ModernBERT accepts up to 8,192 tokens, compared with the 512-token limit common in earlier encoders. Its training includes substantial code data, making it relevant to search across source code as well as prose. Retrieval training examples cover dense embeddings through Sentence Transformers and ColBERT through PyLate.

Efficiency is a central design choice. Alternating local and global attention reduces the cost of longer inputs, while unpadding avoids computation on empty padding tokens. Flash Attention 2 provides further acceleration on compatible GPUs. The models were evaluated on an NVIDIA RTX 4090, with attention to throughput for inputs of mixed lengths and memory use on smaller GPUs.

The research code builds on MosaicBERT and includes FlexBERT, a modular system for constructing encoders. It provides pre-training and evaluation code, including GLUE tests and retrieval evaluations, for teams that need to train or compare their own models.

Similar to ModernBERT