Favicon of E5 Embeddings

E5 Embeddings

Text embedding models for local retrieval, with English and multilingual variants, Hugging Face checkpoints, and MIT-licensed Python code.

E5 Embeddings is a family of text embedding models for developers building search and retrieval systems on their own hardware. It converts text into numerical representations for matching queries with relevant passages. The family includes English and multilingual models, plus instruction-based variants for task-specific embeddings.

The English models come in small, base, and large sizes, including E5-small-v2, E5-base-v2, and E5-large-v2. The multilingual family offers the same size choices and includes multilingual-e5-large-instruct. These choices let developers compare model size and retrieval performance within one family rather than commit to a single checkpoint.

E5-mistral-7b-instruct provides a separate option based on Mistral, with support through Hugging Face Transformers. The checkpoints are available on Hugging Face under the intfloat account. The Python project is open source under the MIT license and includes evaluation code that runs on local GPU machines.

For teams comparing embedding models, the project provides BEIR retrieval evaluations and MTEB evaluations for other embedding tasks, including multilingual evaluation. It also distinguishes models trained only on unlabeled datasets through the -unsupervised suffix. The evaluation scripts can use all available GPUs, and corpus encoding can take substantial time, especially with E5-mistral-7b-instruct.

Similar to E5 Embeddings