
PaliGemma is a family of downloadable vision-language models for developers and researchers building applications that work with images and text. It combines SigLIP's image processing with Gemma's language capabilities to answer questions about visual content. Its main appeal is task-specific fine-tuning: you can adapt a base model to your own image data and intended use.
The models can generate captions for images and short videos, detect objects, and read text within images. They accept an image alongside a text prompt, so an application can ask for a description or a more focused answer about what the image contains. Model variants offer different input resolutions and sizes for different workloads.
The checkpoint type matters when choosing a starting point. PT models need fine-tuning before they produce useful results. FT models target specific research datasets, while mix models have training across several tasks and can handle common use cases without further tuning. These choices let researchers compare a general base with task-trained checkpoints.
Weights are available through Kaggle and Hugging Face. The associated big_vision research code uses JAX/Flax and supports training on GPU machines you control or on Google Cloud TPU VMs; Colab notebooks provide cloud-based examples for detection, generation and fine-tuning. The research code carries the Apache 2.0 license.
PaliGemma weights use the Gemma terms of use, which are separate from the Apache-licensed research code.
Claim this page with an email at ai.google.dev. PaliGemma gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PaliGemma?Promote it
Something wrong or outdated on this page?
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
8.7KUpdated 2 years agoMIT
Linux#Hugging Face integration#Multimodal input#ONNX
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
Hallo turns a single portrait and a speech recording into an animated talking video on your own hardware. It's a local AI tool for creators working with talking portraits and researchers who want access to both generation and training code. The Python code uses the MIT license; required pretrained models and dependencies have their own terms.