
EmbeddingGemma is a text embedding model for developers building search and document features that run on phones, laptops or tablets. Based on Gemma 3, it converts text into numerical representations so applications can find related passages by meaning. Embeddings stay on your hardware, and the model works without an internet connection.
Its uses include semantic search, information retrieval, text classification and clustering. For a chatbot that needs to consult local documents, it can supply the retrieval part of a RAG pipeline, with Gemma 3n handling the conversational response. EmbeddingGemma itself produces embeddings rather than chat replies.
The model trains on text in over 100 languages. Its adjustable embedding size lets developers trade some representation detail for lower storage needs and faster processing. Matryoshka Representation Learning supports dimensions from 768 down to 128, while the input context accommodates up to 2,000 tokens.
A quantized model can run with less than 200MB of RAM, making its memory footprint relevant to applications on everyday devices. Local processing also lets those applications embed sensitive documents without sending their text to a cloud service.
Google provides open weights for fine-tuning and deployment, including responsible commercial use. The separate Gemma Python library uses JAX and carries an Apache 2.0 license; that license applies to the library rather than establishing the model weights' terms. The weights use the Gemma terms rather than Apache-2.0. Hugging Face access is gated and requires accepting the model terms. Downloads are also available through Kaggle.
Claim this page with an email at ai.google.dev. EmbeddingGemma gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find EmbeddingGemma?Promote it
Something wrong or outdated on this page?
jina.aiEmbedding and Reranker Models
Docker#GGUF#LoRA#MLX
Jina Embeddings is a family of models that converts content into vectors for retrieval, similarity matching, classification and clustering. It includes multilingual text models and multimodal variants for searching across different media.
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
22.2KUpdated 1 week agoMIT
#Hugging Face integration#Multilingual
5.8KUpdated 2 days agoApache-2.0
Android#LM Studio integration#LoRA#Multilingual
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
E5 Embeddings is a family of text embedding models for developers building search and retrieval systems on their own hardware. It converts text into numerical representations for matching queries with relevant passages. The family includes English and multilingual models, plus instruction-based variants for task-specific embeddings.
Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.