MLX VLM runs vision-language and audio/video models locally on Macs through MLX. It's an open-source Python package under the MIT license, aimed at developers and researchers who want to build multimodal apps or fine-tune models on their own hardware. It supports Apple Silicon GPUs through Metal.
Model support includes Qwen2-VL, Qwen2.5-VL, LLaVA-OneVision and Gemma 4, alongside document-focused models such as DeepSeek-OCR and PaddleOCR-VL. Depending on the model, it can answer questions about images, compare several images in one conversation, understand video or interpret audio. MiniCPM-o supports speech generation with reference audio.
You can work through a Python API, a command-line interface or a Gradio chat interface. Its self-hosted server exposes OpenAI-compatible chat and Responses APIs, so applications can send text and media to models running on your Mac. The server also provides embeddings and document reranking, plus speech transcription and synthesis through mlx-audio. It can discover models in local folders and the Hugging Face cache.
For repeated conversations about an image, the server reuses cached visual features instead of processing the image again on every turn. Batching lets multiple requests share GPU compute, while speculative decoding can accelerate generation with supported model pairings. Model conversion and quantization help reduce memory demands, and TurboQuant compresses the attention cache with custom Metal kernels.
Claim this page and we'll verify you by hand. MLX VLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MLX VLM?Promote it
Something wrong or outdated on this page?
7.7KUpdated 18 hours agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
7.2KUpdated 9 hours agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
23.2KUpdated 14 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
6.1KUpdated 6 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
130KUpdated 20 hours agoMIT
Web#Code execution#GGUF#Hugging Face integration
15.8KUpdated 1 day agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.