
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
There are models for different hardware needs. Moondream 3.1 uses a mixture-of-experts design for visual reasoning and longer queries. Moondream 2 is a smaller choice that runs on CPUs, GPUs and edge devices, while Moondream 2 0.5B is aimed at more constrained hardware and fine-tuning. Photon, Moondream’s inference engine, runs models locally on Mac, Windows and CUDA hardware, including NVIDIA Jetson devices. You can also self-host inference on a server.
Moondream has cloud services alongside the local models. Moondream Cloud provides hosted inference through an OpenAI-compatible API, so requests sent there are processed off your machine. Lens is a hosted fine-tuning service for adapting a model to images from a specific use case; it doesn’t require dataset uploads, and you keep the resulting weights. A tuned model can run through Moondream Cloud or locally with Photon. That gives teams working with specialized images, such as warehouse or manufacturing footage, a path from an open model to a model trained for their own visual tasks.
Claim this page with an email at moondream.ai. Moondream gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Moondream?Promote it
Something wrong or outdated on this page?
62.1KUpdated 23 hours agoAGPL-3.0
#ONNX
Ultralytics YOLO is an open-source Python computer vision library for developers building applications that analyze images and video on their own hardware. It supports local and edge deployment, including NVIDIA Jetson, Raspberry Pi and mobile phones. A separate hosted platform provides browser-based annotation, cloud GPU training and managed prediction endpoints.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
29.9KUpdated 3 weeks ago
macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization
InsightFace is a face analysis toolkit for developers and teams building identity verification, access control, or face editing software. The code uses the MIT license. Its Python tools and self-hosted recognition server run inference on your own hardware. It also offers commercial models and API access for face swapping and deepfake detection.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.