Favicon of ONNX Runtime

ONNX Runtime

An open source engine for running ONNX models on Windows, macOS, Linux, mobile devices and the web, with CPU, GPU and NPU support.

Screenshot of ONNX Runtime website

ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.

The runtime supports models originating in PyTorch, TensorFlow/Keras, scikit-learn, LightGBM and XGBoost. It can run generative AI workloads such as text generation and image synthesis, alongside other machine learning models. Built-in model optimizations and hardware acceleration help reduce latency, memory use and compute demands. It runs on CPUs, GPUs and NPUs, so developers can use the hardware available to their application.

ONNX Runtime has interfaces for Python, C#, C++, Java, JavaScript and Rust. ONNX Runtime Web brings inference into the browser, while ONNX Runtime Mobile targets Android and iOS apps. These options make it useful for teams deploying a model across several products or devices without adopting a separate inference engine for each one.

It also supports training. Developers can train on device to personalize an app locally, or accelerate transformer training on multi-node NVIDIA GPU systems. ONNX Runtime is licensed under MIT. The project says it may collect usage data and send it to Microsoft.

Similar to ONNX Runtime