Open Vision-Language Models (VLMs)

Models that read images, screenshots and scanned documents and answer questions about them, such as Qwen2.5-VL and MiniCPM-V.

26 tools
An open-weight local LLM family from Google DeepMind for phones, PCs and servers, with Ollama and LM Studio support and an Apache 2.0 JAX library.

5.8KUpdated 2 days agoApache-2.0

Android#LM Studio integration#LoRA#Multilingual

AI content moderation models check text and images in LLM inputs and outputs. Use downloadable models or Meta's hosted Llama API.

4.4KUpdated 20 hours ago

#Guardrails#Multilingual#Multimodal input

A local vision-language model for image analysis, document extraction and video understanding, with Apache 2.0 weights and Hugging Face Transformers support.

huggingface.coOCR and Document Scanning

Linux#Batch processing#Hugging Face integration#Multimodal input

Multimodal AI models you can run locally with Mistral's GPU inference library, which is open source under Apache 2.0 and archived.

10.8KUpdated 3 months agoApache-2.0

Docker#Hugging Face integration#Multimodal input#Tool calling

Open-source vision language model for local image and text tasks, with Apache 2.0 licensing, Transformers support, and a small GPU memory footprint.

3.9KUpdated 1 week agoApache-2.0

#Hugging Face integration#LoRA#Multimodal input

A family of AI models you can run offline with Ollama, llama.cpp or LM Studio, with open weights and training data for building specialized agents.

2.1KUpdated 3 weeks agoApache-2.0

Linux#GGUF#Guardrails#Hugging Face integration

LLM models for local or cloud inference, with downloadable weights and function calling. The Apache 2.0 inference library is archived.

10.8KUpdated 3 months agoApache-2.0

Docker#Hugging Face integration#Multimodal input#Tool calling

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

26.5KUpdated 3 weeks agoApache-2.0

macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend

Vision-language models that answer image questions, read text in images and detect objects, with downloadable weights and GPU research code.

3.5KUpdated 1 year agoApache-2.0

#Hugging Face integration#Multimodal input

A self-hosted vision-language model project for image chat, with a local Gradio interface, GPU inference and Apache 2.0 code.

25KUpdated 2 years agoApache-2.0

macOS · Web#LoRA#Multimodal input#Quantization

An open-source AI model family for on-premises deployment, licensed under Apache 2.0, with language, speech, vision and guardrail models.

273Updated 2 years agoApache-2.0

Linux#Guardrails#Hugging Face integration#LM Studio integration

An open-source image and text model for local image classification without task-specific training. Runs through PyTorch on CPU or CUDA GPUs under MIT.

34.4KUpdated 6 months agoMIT

#Batch processing#Multimodal input

A local image captioning model with open weights, Apache 2.0 code, SFW and NSFW coverage, and support for ComfyUI and vLLM.

1.3KUpdated 7 months agoApache-2.0

Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API

An open-source object detection model that finds objects using text prompts. Run it locally with PyTorch on a CPU or NVIDIA GPU. Apache 2.0 licensed.

10.6KUpdated 2 years agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input

A Python library for running and training CLIP image-text models on your own hardware, with local checkpoints and Hugging Face model support.

14.2KUpdated 5 days ago

#Hugging Face integration#Multimodal input

A local multimodal AI model for answering image questions and generating pictures, with downloadable weights and a Gradio interface.

17.8KUpdated 2 years agoMIT

Web#Batch processing#Hugging Face integration#Multimodal input

Local OCR model that extracts plain or formatted text from images, supports multi-page recognition, and works with Hugging Face Transformers.

8.2KUpdated 2 years ago

#Batch processing#GGUF#Hugging Face integration

Local OCR model that converts images and PDFs to text or Markdown. Runs on NVIDIA GPUs with vLLM or Transformers under the MIT license.

23.9KUpdated 8 months agoMIT

Linux#Batch processing#Hugging Face integration#Multimodal input

A vision-language model family with local Python code for training and evaluation, Apache 2.0 licensing, and variants built on OLMo and Qwen2.

937Updated 2 years agoApache-2.0

#Hugging Face integration#Multimodal input#Works offline

An open-source vision model that runs locally with PyTorch and Hugging Face Transformers, supports CPU or CUDA GPUs, and uses the MIT license.

huggingface.coComputer Vision Models

#Hugging Face integration#Multimodal input#Structured output

An open-weight vision model for image questions, captions and object detection. Run it locally or use hosted inference and fine-tuning.

10.1KUpdated 5 months agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline

A local LLM family for developers and researchers, with downloadable weights, text and vision models, and custom licensing for research and commercial use.

7.7KUpdated 12 months ago

#Hugging Face integration#Multimodal input#Quantization

Open-source vision-language models for visual chat, document questions and image retrieval, with downloadable weights and Hugging Face Transformers support.

10.2KUpdated 1 year agoMIT

#Hugging Face integration#Multimodal input

A self-hosted OCR model that extracts text and page structure from PDFs and images, with vLLM, Hugging Face Transformers and CPU inference support.

9.2KUpdated 6 months agoMIT

Docker#Hugging Face integration#Multilingual#Multimodal input

Local OCR software for PDFs and images, with reading order, tables and math. Runs on CPU, Apple Silicon or NVIDIA GPUs; code uses Apache 2.0.

21.4KUpdated 3 weeks agoApache-2.0

macOS · Web#Batch processing#llama.cpp backend#Multilingual

Favicon of PaddleOCR

PaddleOCR

2 videos
An open source OCR toolkit that runs locally, converts PDFs and images to Markdown or JSON, and supports multilingual text under Apache 2.0.

90.4KUpdated 2 weeks agoApache-2.0

Web#Multilingual#ONNX#Structured output

More in Open Models