
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
Embedding models turn text into numerical representations that let an application compare meaning rather than rely only on matching words. Rerankers score query-document pairs to reorder retrieved results. BGE provides both, so developers can choose models for finding candidate passages and for assessing their relevance before passing them to an LLM.
The models cover different retrieval needs. BAAI/bge-m3 supports multilingual search with dense, sparse and ColBERT-style multi-vector retrieval, plus longer text inputs. BAAI/bge-en-icl is an English embedding model that accepts task instructions and example pairs through in-context learning. BAAI/bge-multilingual-gemma2 provides another multilingual option with task-based instructions.
For reranking, BAAI/bge-reranker-v2-m3 is a lightweight multilingual cross-encoder. BAAI/bge-reranker-v2-gemma also handles multilingual query-document scoring, while BAAI/bge-reranker-v2-minicpm-layerwise lets developers select output layers to reduce inference work. FlagEmbedding includes examples for embedding and reranker inference, evaluation and fine-tuning.
Claim this page with an email at bge-model.com. BGE Embeddings gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find BGE Embeddings?Promote it
Something wrong or outdated on this page?
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
2.2KUpdated 1 day agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
3.2KUpdated 23 hours agoApache-2.0
#Batch processing#Multilingual#ONNX
19.1KUpdated 1 week agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
7.7KUpdated 2 years agoMIT
Docker · Web#Multilingual
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
FastEmbed is a Python library that generates embeddings on your own hardware for semantic search and retrieval-augmented generation (RAG). It's for developers who need to turn text into searchable vectors without relying on a cloud embedding API. It can run on a CPU or use GPU acceleration, and its Apache 2.0 license makes it open source.
Sentence Transformers is an open-source Python library for developers building semantic search and document retrieval on their own hardware. It runs embedding and reranker models locally, turning content into numerical representations for similarity comparisons and scoring results against a query. The library uses the Apache 2.0 license.
MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.