Favicon of MLX VLM

MLX VLM

A local AI Python package for running and fine-tuning multimodal models on Macs, with an MIT license and an OpenAI-compatible server.

MLX VLM runs vision-language and audio/video models locally on Macs through MLX. It's an open-source Python package under the MIT license, aimed at developers and researchers who want to build multimodal apps or fine-tune models on their own hardware. It supports Apple Silicon GPUs through Metal.

Model support includes Qwen2-VL, Qwen2.5-VL, LLaVA-OneVision and Gemma 4, alongside document-focused models such as DeepSeek-OCR and PaddleOCR-VL. Depending on the model, it can answer questions about images, compare several images in one conversation, understand video or interpret audio. MiniCPM-o supports speech generation with reference audio.

You can work through a Python API, a command-line interface or a Gradio chat interface. Its self-hosted server exposes OpenAI-compatible chat and Responses APIs, so applications can send text and media to models running on your Mac. The server also provides embeddings and document reranking, plus speech transcription and synthesis through mlx-audio. It can discover models in local folders and the Hugging Face cache.

For repeated conversations about an image, the server reuses cached visual features instead of processing the image again on every turn. Batching lets multiple requests share GPU compute, while speculative decoding can accelerate generation with supported model pairings. Model conversion and quantization help reduce memory demands, and TurboQuant compresses the attention cache with custom Metal kernels.

Similar to MLX VLM