Favicon of GTE Embeddings

GTE Embeddings

Alibaba’s GTE models turn text into vectors for retrieval and similarity matching, with downloadable weights and local Python inference.

Screenshot of GTE Embeddings website

GTE (General Text Embedding) is Alibaba’s family of downloadable models for representing text as vectors. Developers use these representations to compare queries with documents, cluster related text or supply retrieval components for larger applications.

The gte-multilingual-base model supports more than 70 languages, dense and sparse representations, and inputs up to 8,192 tokens. Its dense embeddings can be shortened to trade vector size against retrieval quality. These features are specific to that variant; other GTE releases have different architectures and limits.

The official model card provides local Transformers and SentenceTransformers examples, including batched encoding and document similarity scoring. It also documents serving through Infinity or Hugging Face Text Embeddings Inference, with Docker examples for CPU and GPU execution. The card links instructions for using its custom code offline after download.

The gte-multilingual-base weights are released under Apache-2.0. Check the card for the exact variant you choose, including its hardware needs and any custom model code. Alibaba’s hosted embedding APIs are separate services, and their underlying models are not necessarily identical to the downloadable GTE models.

Similar to GTE Embeddings