Favicon of PaliGemma

PaliGemma

Vision-language models that answer image questions, read text in images and detect objects, with downloadable weights and GPU research code.

Screenshot of PaliGemma website

PaliGemma is a family of downloadable vision-language models for developers and researchers building applications that work with images and text. It combines SigLIP's image processing with Gemma's language capabilities to answer questions about visual content. Its main appeal is task-specific fine-tuning: you can adapt a base model to your own image data and intended use.

The models can generate captions for images and short videos, detect objects, and read text within images. They accept an image alongside a text prompt, so an application can ask for a description or a more focused answer about what the image contains. Model variants offer different input resolutions and sizes for different workloads.

The checkpoint type matters when choosing a starting point. PT models need fine-tuning before they produce useful results. FT models target specific research datasets, while mix models have training across several tasks and can handle common use cases without further tuning. These choices let researchers compare a general base with task-trained checkpoints.

Weights are available through Kaggle and Hugging Face. The associated big_vision research code uses JAX/Flax and supports training on GPU machines you control or on Google Cloud TPU VMs; Colab notebooks provide cloud-based examples for detection, generation and fine-tuning. The research code carries the Apache 2.0 license.

PaliGemma weights use the Gemma terms of use, which are separate from the Apache-licensed research code.

Similar to PaliGemma