Favicon of MiniCPM-V

MiniCPM-V

Local vision-language models for image and video understanding, with Apache 2.0 code, mobile deployment and support for Ollama and llama.cpp.

MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.

The models answer questions about individual images, work with multiple images together, read text within pictures and describe video content. Their emphasis on mobile deployment makes them relevant to projects where visual understanding needs to run on-device. A hosted API is also available; that service runs remotely rather than on your hardware.

MiniCPM-V 4.6 combines SigLIP2 with a Qwen-based language model and compresses visual information early to reduce processing work. Developers can trade finer image detail for faster inference. Quantized models are available in GGUF, BNB, AWQ and GPTQ formats, with support for Ollama, llama.cpp, vLLM, SGLang and Hugging Face Transformers. SWIFT and LLaMA-Factory support fine-tuning for specialized tasks on consumer GPUs.

The related MiniCPM-o family adds live audio and video conversation with speech and text responses. It can listen, watch and speak at the same time, and supports voice imitation from reference audio. Its self-hosted PyTorch web demo requires an NVIDIA GPU with at least 28 GB of GPU memory and can serve as an API backend for other applications.

Similar to MiniCPM-V