Favicon of SmolVLM

SmolVLM

Open-source vision language model for local image and text tasks, with Apache 2.0 licensing, Transformers support, and a small GPU memory footprint.

Screenshot of SmolVLM website

SmolVLM is a compact vision language model from Hugging Face for developers building local AI applications that work with images and text. It can describe pictures, answer questions about diagrams, and read information from documents such as invoices. Its small memory footprint makes on-device use practical on laptops and smaller local setups.

You can include multiple images in one conversation and mix them with text prompts. SmolVLM Instruct is the ready-to-use variant for interactive applications, while SmolVLM-Base provides a starting point for custom training. SmolVLM-Synthetic is a separate variant fine-tuned on synthetic data.

The model works with Hugging Face Transformers and supports CPU or CUDA execution. The published GPU inference tests used about 5 GB of memory. Its architecture follows Idefics3, pairing SmolLM2 with a SigLIP vision encoder and compressing image information to reduce memory use. This also keeps memory growth more moderate when a prompt contains multiple images.

SmolVLM is open source under Apache 2.0 and permits commercial use. The release includes model weights, training datasets, recipes, and fine-tuning tools. Developers can adapt it through Transformers, use LoRA or QLoRA for lighter fine-tuning, and apply preference tuning through TRL. VLMEvalKit integration supports evaluation of both the original model and customized versions.

Similar to SmolVLM