
GTE (General Text Embedding) is Alibaba’s family of downloadable models for representing text as vectors. Developers use these representations to compare queries with documents, cluster related text or supply retrieval components for larger applications.
The gte-multilingual-base model supports more than 70 languages, dense and sparse representations, and inputs up to 8,192 tokens. Its dense embeddings can be shortened to trade vector size against retrieval quality. These features are specific to that variant; other GTE releases have different architectures and limits.
The official model card provides local Transformers and SentenceTransformers examples, including batched encoding and document similarity scoring. It also documents serving through Infinity or Hugging Face Text Embeddings Inference, with Docker examples for CPU and GPU execution. The card links instructions for using its custom code offline after download.
The gte-multilingual-base weights are released under Apache-2.0. Check the card for the exact variant you choose, including its hardware needs and any custom model code. Alibaba’s hosted embedding APIs are separate services, and their underlying models are not necessarily identical to the downloadable GTE models.
Claim this page and we'll verify you by hand. GTE Embeddings gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find GTE Embeddings?Promote it
Something wrong or outdated on this page?
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
2KUpdated 1 year ago
Docker#Batch processing#Hugging Face integration#Multilingual
Qwen3-Embedding is a family of text embedding models for developers building search and document analysis on their own hardware or servers. It turns text into numerical representations that applications can compare by meaning. Its multilingual support covers languages including English, Chinese, Arabic and Ukrainian, as well as programming languages.
22.2KUpdated 1 week agoMIT
#Hugging Face integration#Multilingual
5.8KUpdated 2 days agoApache-2.0
#Hugging Face integration#Multilingual#Quantization
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
2.2KUpdated 1 day agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
E5 Embeddings is a family of text embedding models for developers building search and retrieval systems on their own hardware. It converts text into numerical representations for matching queries with relevant passages. The family includes English and multilingual models, plus instruction-based variants for task-specific embeddings.
EmbeddingGemma is a text embedding model for developers building search and document features that run on phones, laptops or tablets. Based on Gemma 3, it converts text into numerical representations so applications can find related passages by meaning. Embeddings stay on your hardware, and the model works without an internet connection.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.