
InternVL is a family of downloadable vision-language models for developers and researchers building AI that can interpret images and discuss them in text. It combines visual recognition with language models, supporting both multimodal chat and tasks such as image classification and image-text retrieval.
Its capabilities extend beyond describing photos. Document and chart question answering, visual mathematics, and image or video classification are among the tasks it targets. The underlying vision models also support pixel-level recognition, which makes the project relevant to computer vision research as well as conversational applications.
The model family includes smaller and larger variants rather than a single fixed model. InternViT supplies the vision component, paired with language models including Qwen2.5 and InternLM. This gives developers a choice of model sizes and language backbones when selecting a model for their own deployment.
Downloadable checkpoints are available through Hugging Face and ModelScope. The project provides its own checkpoint format alongside a format compatible with Hugging Face Transformers. Its Python repository uses the MIT license.
A hosted chat demo and API are also available. Those services run remotely; the downloadable models are the route for developers who want to host inference themselves. The project includes standalone vision components as well as models intended for multimodal dialogue.
Claim this page and we'll verify you by hand. InternVL gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find InternVL?Promote it
Something wrong or outdated on this page?
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
34.4KUpdated 6 months agoMIT
#Batch processing#Multimodal input
CLIP is an open-source image and text model that lets developers and researchers classify images using labels written in natural language, without collecting training examples for each task. It runs locally through PyTorch on a CPU or CUDA GPU. The code and model weights use the MIT license.
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.