Favicon of MNN

MNN

An open source on-device AI engine for local LLMs and image models, with iOS, Android, CPU and GPU support under Apache 2.0.

MNN is a lightweight C++ framework for developers who want AI models to run on phones, PCs and embedded devices. It handles inference and training on the device, with a focus on small application footprints and hardware acceleration. The project is open source under Apache 2.0, and Alibaba uses it in apps including Taobao, Youku and DingTalk.

MNN-LLM runs language models locally, including Qwen, Baichuan, Zhipu and Llama. It also supports Qwen3.5 and Qwen3-VL. MNN-Diffusion brings Stable Diffusion image generation to local hardware. Companion Android and iOS apps provide multimodal chat; the Android app covers text generation, image understanding, audio transcription and image generation. The TaoAvatar app supports offline conversations with a 3D avatar, with language, speech and avatar models all running on the device.

The engine supports iOS, Android and POSIX embedded devices. It can use ARM and x64 CPUs, mobile GPUs through Metal, OpenCL and Vulkan, and NVIDIA GPUs through CUDA. Qualcomm Hexagon DSP acceleration is also available. FP16 and Int8 quantization reduce model size, which matters when fitting models into mobile apps or constrained devices.

Developers can bring models from TensorFlow, TensorFlow Lite, Caffe, ONNX and TorchScript into MNN. Its tools cover model conversion, compression and training, while a Python API provides inference and image processing without requiring C++ application code. MNN Workbench includes pretrained models, visual training tools and deployment to devices.

Similar to MNN