Favicon of InternVL

InternVL

Open-source vision-language models for visual chat, document questions and image retrieval, with downloadable weights and Hugging Face Transformers support.

Screenshot of InternVL website

InternVL is a family of downloadable vision-language models for developers and researchers building AI that can interpret images and discuss them in text. It combines visual recognition with language models, supporting both multimodal chat and tasks such as image classification and image-text retrieval.

Its capabilities extend beyond describing photos. Document and chart question answering, visual mathematics, and image or video classification are among the tasks it targets. The underlying vision models also support pixel-level recognition, which makes the project relevant to computer vision research as well as conversational applications.

The model family includes smaller and larger variants rather than a single fixed model. InternViT supplies the vision component, paired with language models including Qwen2.5 and InternLM. This gives developers a choice of model sizes and language backbones when selecting a model for their own deployment.

Downloadable checkpoints are available through Hugging Face and ModelScope. The project provides its own checkpoint format alongside a format compatible with Hugging Face Transformers. Its Python repository uses the MIT license.

A hosted chat demo and API are also available. Those services run remotely; the downloadable models are the route for developers who want to host inference themselves. The project includes standalone vision components as well as models intended for multimodal dialogue.

Similar to InternVL