Favicon of Molmo

Molmo

A vision-language model family with local Python code for training and evaluation, Apache 2.0 licensing, and variants built on OLMo and Qwen2.

Screenshot of Molmo website

Molmo is Ai2's family of vision-language models, with code for running and training models on your own hardware. It's for developers and researchers who need to work with images and text, adapt a model, or evaluate it against visual tasks. The Python codebase is open source under Apache 2.0 and builds on OLMo, adding image encoding and generative evaluation.

The family includes MolmoE-1B, based on OLMoE, and Molmo-7B-O, based on OLMo. Molmo-7B-D and Molmo-72B use Qwen2. These models pair their language backbones with an OpenAI CLIP vision encoder. For custom training, the code also supports SigLIP, MetaCLIP and DINOv2, so researchers can compare different approaches to image processing.

Ai2 provides both pretrained and fine-tuned weights, including checkpoints compatible with this codebase and models on Hugging Face. The repository covers pretraining, multitask training and evaluation, rather than only inference. It also supports evaluation through VLMEvalkit, though that integration doesn't cover every dataset. A vLLM implementation offers faster inference than the repository's evaluation path.

Local evaluation supports high-resolution images and distributed processing across machines. Evaluating the 72B model with this codebase requires multiple nodes. Ai2 also offers a public web demo of Molmo-7B-D, separate from running the models on your own hardware.

Similar to Molmo