
SmolVLM is a compact vision language model from Hugging Face for developers building local AI applications that work with images and text. It can describe pictures, answer questions about diagrams, and read information from documents such as invoices. Its small memory footprint makes on-device use practical on laptops and smaller local setups.
You can include multiple images in one conversation and mix them with text prompts. SmolVLM Instruct is the ready-to-use variant for interactive applications, while SmolVLM-Base provides a starting point for custom training. SmolVLM-Synthetic is a separate variant fine-tuned on synthetic data.
The model works with Hugging Face Transformers and supports CPU or CUDA execution. The published GPU inference tests used about 5 GB of memory. Its architecture follows Idefics3, pairing SmolLM2 with a SigLIP vision encoder and compressing image information to reduce memory use. This also keeps memory growth more moderate when a prompt contains multiple images.
SmolVLM is open source under Apache 2.0 and permits commercial use. The release includes model weights, training datasets, recipes, and fine-tuning tools. Developers can adapt it through Transformers, use LoRA or QLoRA for lighter fine-tuning, and apply preference tuning through TRL. VLMEvalKit integration supports evaluation of both the original model and customized versions.
Claim this page and we'll verify you by hand. SmolVLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find SmolVLM?Promote it
Something wrong or outdated on this page?
26.5KUpdated 3 weeks agoApache-2.0
macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend
MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
1.3KUpdated 7 months agoApache-2.0
Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
7.7KUpdated 12 months ago
#Hugging Face integration#Multimodal input#Quantization
Llama is Meta's family of large language models for developers, researchers and businesses that want to run models on their own hardware or servers. Its downloadable weights let you build generative AI applications with local inference. Access requires license acceptance and approval, and the weights use custom licensing for research and commercial use.
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.