Paddle Lite is an on-device AI inference engine for developers building mobile apps, embedded systems and edge devices. It focuses on running neural networks with smaller model files, a smaller runtime and lower memory use. It's open source under Apache 2.0.
It runs on Android and iOS, embedded Linux devices, and Windows, macOS and Linux computers. Applications can integrate it through C++, Java or Python APIs. The supplied examples cover image classification, object detection, face landmarks and portrait segmentation.
PaddlePaddle is its native model format. X2Paddle provides a conversion path for models built with PyTorch, TensorFlow or Caffe, so teams can use models developed outside the PaddlePaddle ecosystem.
Hardware support extends beyond CPUs. Paddle Lite supports OpenCL and Metal acceleration, alongside backends such as Android NNAPI, Qualcomm QNN, Huawei Kirin and Ascend NPUs, and Intel OpenVINO. A single model can split execution across different types of hardware, which matters when a device combines a CPU with a dedicated AI accelerator.
Its optimization tools use quantization and combine model operations to reduce resource demands and improve execution speed. They also check which operations a target platform supports. The execution runtime can ship separately from the model analysis tools to keep the deployed application smaller, and developers have debugging and performance profiling tools for investigating inference behavior.
Claim this page with an email at paddlepaddle.org.cn. Paddle Lite gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Paddle Lite?Promote it
Something wrong or outdated on this page?
6.1KUpdated 6 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
5.1KUpdated 2 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
16.2KUpdated 2 days agoApache-2.0
Windows · iOS · Android#Image-to-image#Multimodal input#ONNX
23.9KUpdated 11 hours ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
10.9KUpdated 11 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
MNN is a lightweight C++ framework for developers who want AI models to run on phones, PCs and embedded devices. It handles inference and training on the device, with a focus on small application footprints and hardware acceleration. The project is open source under Apache 2.0, and Alibaba uses it in apps including Taobao, Youku and DingTalk.
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.