Favicon of TensorRT Model Optimizer

TensorRT Model Optimizer

Open-source model optimization library under Apache 2.0. Compress Hugging Face, PyTorch and ONNX models for TensorRT-LLM, vLLM and SGLang.

TensorRT Model Optimizer, called NVIDIA Model Optimizer or ModelOpt, is a Python library for developers preparing models for local or self-hosted inference. It reduces model size and memory use and can speed up inference through compression and other optimization techniques. It's open source under Apache 2.0.

ModelOpt accepts Hugging Face, PyTorch and ONNX models, with support for language models, vision-language models and diffusion models. Its Python APIs let developers combine techniques within one library, then export optimized checkpoints to TensorRT-LLM, TensorRT, vLLM or SGLang. Hugging Face export covers both transformers and diffusers models.

Quantization lowers the precision used to store and compute model values. ModelOpt supports post-training quantization as well as quantization-aware training and distillation to recover accuracy lost during compression, including workflows for FP8 and NVFP4. Pruning removes unnecessary weights, while distillation trains smaller models to reproduce the behavior of larger ones. It also supports neural architecture search, sparsity and speculative decoding with trained draft modules.

The library integrates with NVIDIA Megatron-Bridge, Megatron-LM and Hugging Face Accelerate for techniques that require training. Windows quantization workflows and NVIDIA container images are available. NVIDIA also publishes pre-quantized checkpoints on Hugging Face for deployment with TensorRT-LLM, vLLM and SGLang.

Install the Python package with the dependencies for the intended workflow. Linux CUDA examples target NVIDIA GPUs; supported model formats, quantization schemes and deployment runtimes depend on the hardware and support matrix.

Similar to TensorRT Model Optimizer