MNN is a lightweight C++ framework for developers who want AI models to run on phones, PCs and embedded devices. It handles inference and training on the device, with a focus on small application footprints and hardware acceleration. The project is open source under Apache 2.0, and Alibaba uses it in apps including Taobao, Youku and DingTalk.
MNN-LLM runs language models locally, including Qwen, Baichuan, Zhipu and Llama. It also supports Qwen3.5 and Qwen3-VL. MNN-Diffusion brings Stable Diffusion image generation to local hardware. Companion Android and iOS apps provide multimodal chat; the Android app covers text generation, image understanding, audio transcription and image generation. The TaoAvatar app supports offline conversations with a 3D avatar, with language, speech and avatar models all running on the device.
The engine supports iOS, Android and POSIX embedded devices. It can use ARM and x64 CPUs, mobile GPUs through Metal, OpenCL and Vulkan, and NVIDIA GPUs through CUDA. Qualcomm Hexagon DSP acceleration is also available. FP16 and Int8 quantization reduce model size, which matters when fitting models into mobile apps or constrained devices.
Developers can bring models from TensorFlow, TensorFlow Lite, Caffe, ONNX and TorchScript into MNN. Its tools cover model conversion, compression and training, while a Python API provides inference and image processing without requiring C++ application code. MNN Workbench includes pretrained models, visual training tools and deployment to devices.
Claim this page and we'll verify you by hand. MNN gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MNN?Promote it
Something wrong or outdated on this page?
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
1.7KUpdated 2 days ago
Linux#Multimodal input#Quantization
RKLLM is a software stack for developers building local AI applications on Rockchip hardware. It uses the chip's neural processing unit (NPU) to run language and multimodal models on development boards, with support for the RK3588, RK3576, RK3562 and RV1126B series.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.