MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.
The models answer questions about individual images, work with multiple images together, read text within pictures and describe video content. Their emphasis on mobile deployment makes them relevant to projects where visual understanding needs to run on-device. A hosted API is also available; that service runs remotely rather than on your hardware.
MiniCPM-V 4.6 combines SigLIP2 with a Qwen-based language model and compresses visual information early to reduce processing work. Developers can trade finer image detail for faster inference. Quantized models are available in GGUF, BNB, AWQ and GPTQ formats, with support for Ollama, llama.cpp, vLLM, SGLang and Hugging Face Transformers. SWIFT and LLaMA-Factory support fine-tuning for specialized tasks on consumer GPUs.
The related MiniCPM-o family adds live audio and video conversation with speech and text responses. It can listen, watch and speak at the same time, and supports voice imitation from reference audio. Its self-hosted PyTorch web demo requires an NVIDIA GPU with at least 28 GB of GPU memory and can serve as an API backend for other applications.
Claim this page and we'll verify you by hand. MiniCPM-V gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MiniCPM-V?Promote it
Something wrong or outdated on this page?
3.9KUpdated 1 week agoApache-2.0
#Hugging Face integration#LoRA#Multimodal input
SmolVLM is a compact vision language model from Hugging Face for developers building local AI applications that work with images and text. It can describe pictures, answer questions about diagrams, and read information from documents such as invoices. Its small memory footprint makes on-device use practical on laptops and smaller local setups.
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
3.2KUpdated 4 days agoMIT
macOS · Windows · Linux · iOS · Android#GGUF#Human approval#LM Studio integration
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
11.2KUpdated 5 months agoApache-2.0
macOS · iOS · Web#Hugging Face integration#MLX#Quantization
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
Off Grid AI runs language models on iOS, Android, macOS, Windows and Linux. You can chat, analyze documents and generate images on your own hardware. The mobile app uses the MIT license, while the desktop app uses AGPL.
Moshi is a voice AI model and dialogue framework that can listen while it speaks. It processes speech directly, retaining information such as emotion and non-verbal cues that a text transcription can miss. It's aimed at researchers and developers building spoken AI applications, with local inference and self-hosted server options.