Player not loading? Watch on YouTube
This tutorial explains ONNX (Open Neural Network Exchange) through Python export and inference examples. The speaker compares the format to PDF: models defined in different frameworks can share a runtime for inference without those training frameworks installed.
The first PyTorch example adds one to its input rather than training on a dataset. It uses sample input with torch.onnx.export and enables the modern exporter with dynamo=True. A separate script loads the file through ONNX Runtime, reads input and output names from metadata, and supplies a NumPy float32 array. The TensorFlow example uses TF2ONNX and an explicit input signature. The presenter reuses the inference script by changing the model path.
A fuller example trains an MNIST classifier for five epochs and exports it with dummy input shaped for 28-by-28 images. ONNX Runtime handles prediction, though this demonstration still uses a torch dependency to load test data.
The final example downloads a quantized ONNX model described as optimized for CPU and mobile, then uses ONNX Runtime GenAI to run a local LLM with a tokenizer and chat template. The repository must provide the required ONNX format. The presenter reports that unrestricted generation crashed his computer; setting the maximum sequence length to the input token count plus 20 allows his arithmetic example to finish.