Favicon of Optimum

Optimum

Open-source Python tools optimize Hugging Face models for local, edge and cloud hardware, with ONNX Runtime, OpenVINO and TensorRT-LLM integrations.

Screenshot of Optimum website

Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.

Its main role is model optimization: exporting models to other formats, applying quantization and optimizing model graphs. These capabilities let developers adapt existing models to their deployment hardware while staying within the Hugging Face ecosystem.

  • ONNX export and ONNX Runtime support cover Transformers, Diffusers, Sentence Transformers and TIMM models. The ONNX integration lives in the separate optimum-onnx package.
  • OpenVINO provides optimization, quantization and deployment for Intel hardware. NVIDIA TensorRT-LLM supports accelerated inference on NVIDIA hardware.
  • ExecuTorch supports exporting Transformers models for on-device inference on mobile and edge devices within the PyTorch ecosystem.
  • Training integrations include Intel Gaudi, AWS Trainium and ONNX Runtime. Optimum also has packages for Google TPUs, AWS Inferentia and FuriosaAI WARBOY.

The hardware-specific packages make Optimum relevant to teams targeting several deployment environments. Local and edge integrations run models on the chosen devices; AWS Trainium and Inferentia integrations target AWS infrastructure. For developers building their own optimization workflows, Torch FX support allows custom transformations of PyTorch Transformers model graphs.

Similar to Optimum