Favicon of LiteRT

LiteRT

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

Screenshot of LiteRT website

LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.

The framework runs across Android, iOS, Linux, macOS, Windows and the web, with support for embedded devices too. It can use CPUs, GPUs and supported NPUs, so apps can take advantage of the hardware available on a device. LiteRT-LM supports Apple Silicon GPU acceleration on iOS and macOS; its browser runtime requires WebGPU.

LiteRT converts models from PyTorch, TensorFlow and JAX into .tflite format and supports quantization to reduce model size and computation needs. Preconverted models include Gemma, Llama 3.2, Qwen 3.5 and MobileNet. Its shared interface for supported NPUs reduces the need for separate integrations with each chipset provider.

For local LLM apps, LiteRT-LM builds on the runtime and handles conversations, streaming responses, multimodal input and tool calling. Models can run entirely on the device without sending inference requests to a server. Developers can also use LiteRT for diffusion models, camera image segmentation and speech recognition with Whisper, Moonshine or Parakeet.

Speech synthesis works fully offline and can begin playback while the rest of the output is still being generated. FastVLM supports on-device vision-language tasks with NPU acceleration on Qualcomm devices.

Similar to LiteRT