Favicon of ExecuTorch

ExecuTorch

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

Screenshot of ExecuTorch website

ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.

Its workloads include local LLM inference, computer vision and voice processing. Text models support quantization, long context, tool calling and speculative decoding. Voice capabilities include streaming and offline speech recognition, speech synthesis, voice activity detection and speaker diarization. Vision workloads cover image classification, object detection and segmentation, alongside multimodal models that work with text, images and audio.

The runtime can use CPU, GPU, NPU and DSP acceleration. Supported backends include XNNPACK for CPUs, Core ML for Apple platforms, OpenVINO on Linux and Qualcomm accelerators on Android. It delegates supported parts of a model to hardware backends and retains portable CPU kernels for fallback. Developers can keep the runtime small by including only the operators and hardware components their application needs.

Application APIs cover C++, Python, Swift/Objective-C and Kotlin/Java. Profiling and numeric debugging tools help developers inspect model performance and output, and the runtime supports custom operators and backends. ExecuTorch also runs AI features on Meta Quest and Ray-Ban Meta devices.

The README tracks the main branch. MLX, browser/WebGPU and the OpenAI-compatible local server are experimental paths; use release-matched documentation for deployment.

Similar to ExecuTorch