ONNX tutorial: export PyTorch and TensorFlow models

Learn to export models for ONNX Runtime inference, test an MNIST classifier, and limit generation length in a downloaded Hugging Face model.

Player not loading? Watch on YouTube

This tutorial explains ONNX (Open Neural Network Exchange) through Python export and inference examples. The speaker compares the format to PDF: models defined in different frameworks can share a runtime for inference without those training frameworks installed.

The first PyTorch example adds one to its input rather than training on a dataset. It uses sample input with torch.onnx.export and enables the modern exporter with dynamo=True. A separate script loads the file through ONNX Runtime, reads input and output names from metadata, and supplies a NumPy float32 array. The TensorFlow example uses TF2ONNX and an explicit input signature. The presenter reuses the inference script by changing the model path.

A fuller example trains an MNIST classifier for five epochs and exports it with dummy input shaped for 28-by-28 images. ONNX Runtime handles prediction, though this demonstration still uses a torch dependency to load test data.

The final example downloads a quantized ONNX model described as optimized for CPU and mobile, then uses ONNX Runtime GenAI to run a local LLM with a tokenizer and chat template. The repository must provide the required ONNX format. The presenter reports that unrestricted generation crashed his computer; setting the maximum sequence length to the input token count plus 20 allows his arithmetic example to finish.