Favicon of Gemma

Gemma

An open-weight local LLM family from Google DeepMind for phones, PCs and servers, with Ollama and LM Studio support and an Apache 2.0 JAX library.

Screenshot of Gemma website

Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.

The models cover different jobs. Gemma 4 supports reasoning and multimodal applications, while specialized variants include MedGemma for medical text and images, TranslateGemma for translation, EmbeddingGemma for on-device embeddings, and ShieldGemma 2 for content classification. DiffusionGemma is designed for fast text generation. This gives developers a choice of model suited to the task and the hardware where it will run.

For local use, Gemma works with Ollama, LM Studio, Gemma.cpp and Google AI Edge. The accompanying JAX library supports conversations, multimodal use and fine-tuning, including LoRA. That library is open source under Apache License 2.0 and can use a CPU, GPU or TPU. Gemma models also have integrations with PyTorch and Keras.

Similar to Gemma