
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
The runtime supports models originating in PyTorch, TensorFlow/Keras, scikit-learn, LightGBM and XGBoost. It can run generative AI workloads such as text generation and image synthesis, alongside other machine learning models. Built-in model optimizations and hardware acceleration help reduce latency, memory use and compute demands. It runs on CPUs, GPUs and NPUs, so developers can use the hardware available to their application.
ONNX Runtime has interfaces for Python, C#, C++, Java, JavaScript and Rust. ONNX Runtime Web brings inference into the browser, while ONNX Runtime Mobile targets Android and iOS apps. These options make it useful for teams deploying a model across several products or devices without adopting a separate inference engine for each one.
It also supports training. Developers can train on device to personalize an app locally, or accelerate transformer training on multi-node NVIDIA GPU systems. ONNX Runtime is licensed under MIT. The project says it may collect usage data and send it to Microsoft.
Claim this page with an email at onnxruntime.ai. ONNX Runtime gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find ONNX Runtime?Promote it
Something wrong or outdated on this page?
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
3.5KUpdated 6 days agoApache-2.0
#Hugging Face integration#ONNX#Quantization
16.2KUpdated 1 day agoApache-2.0
Windows · iOS · Android#Image-to-image#Multimodal input#ONNX
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.
MNN is a lightweight C++ framework for developers who want AI models to run on phones, PCs and embedded devices. It handles inference and training on the device, with a focus on small application footprints and hardware acceleration. The project is open source under Apache 2.0, and Alibaba uses it in apps including Taobao, Youku and DingTalk.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.