Favicon of EmbeddingGemma

EmbeddingGemma

A multilingual text embedding model that runs offline on phones, laptops and tablets, with open weights and a quantized memory footprint under 200MB.

Screenshot of EmbeddingGemma website

EmbeddingGemma is a text embedding model for developers building search and document features that run on phones, laptops or tablets. Based on Gemma 3, it converts text into numerical representations so applications can find related passages by meaning. Embeddings stay on your hardware, and the model works without an internet connection.

Its uses include semantic search, information retrieval, text classification and clustering. For a chatbot that needs to consult local documents, it can supply the retrieval part of a RAG pipeline, with Gemma 3n handling the conversational response. EmbeddingGemma itself produces embeddings rather than chat replies.

The model trains on text in over 100 languages. Its adjustable embedding size lets developers trade some representation detail for lower storage needs and faster processing. Matryoshka Representation Learning supports dimensions from 768 down to 128, while the input context accommodates up to 2,000 tokens.

A quantized model can run with less than 200MB of RAM, making its memory footprint relevant to applications on everyday devices. Local processing also lets those applications embed sensitive documents without sending their text to a cloud service.

Google provides open weights for fine-tuning and deployment, including responsible commercial use. The separate Gemma Python library uses JAX and carries an Apache 2.0 license; that license applies to the library rather than establishing the model weights' terms. The weights use the Gemma terms rather than Apache-2.0. Hugging Face access is gated and requires accepting the model terms. Downloads are also available through Kaggle.

Similar to EmbeddingGemma