Favicon of Paddle Lite

Paddle Lite

An open-source AI inference engine that runs PaddlePaddle models on Android, iOS, Linux, Windows and macOS, with CPU and hardware acceleration support.

Paddle Lite is an on-device AI inference engine for developers building mobile apps, embedded systems and edge devices. It focuses on running neural networks with smaller model files, a smaller runtime and lower memory use. It's open source under Apache 2.0.

It runs on Android and iOS, embedded Linux devices, and Windows, macOS and Linux computers. Applications can integrate it through C++, Java or Python APIs. The supplied examples cover image classification, object detection, face landmarks and portrait segmentation.

PaddlePaddle is its native model format. X2Paddle provides a conversion path for models built with PyTorch, TensorFlow or Caffe, so teams can use models developed outside the PaddlePaddle ecosystem.

Hardware support extends beyond CPUs. Paddle Lite supports OpenCL and Metal acceleration, alongside backends such as Android NNAPI, Qualcomm QNN, Huawei Kirin and Ascend NPUs, and Intel OpenVINO. A single model can split execution across different types of hardware, which matters when a device combines a CPU with a dedicated AI accelerator.

Its optimization tools use quantization and combine model operations to reduce resource demands and improve execution speed. They also check which operations a target platform supports. The execution runtime can ship separately from the model analysis tools to keep the deployed application smaller, and developers have debugging and performance profiling tools for investigating inference behavior.

Similar to Paddle Lite