TensorRT Model Optimizer, called NVIDIA Model Optimizer or ModelOpt, is a Python library for developers preparing models for local or self-hosted inference. It reduces model size and memory use and can speed up inference through compression and other optimization techniques. It's open source under Apache 2.0.
ModelOpt accepts Hugging Face, PyTorch and ONNX models, with support for language models, vision-language models and diffusion models. Its Python APIs let developers combine techniques within one library, then export optimized checkpoints to TensorRT-LLM, TensorRT, vLLM or SGLang. Hugging Face export covers both transformers and diffusers models.
Quantization lowers the precision used to store and compute model values. ModelOpt supports post-training quantization as well as quantization-aware training and distillation to recover accuracy lost during compression, including workflows for FP8 and NVFP4. Pruning removes unnecessary weights, while distillation trains smaller models to reproduce the behavior of larger ones. It also supports neural architecture search, sparsity and speculative decoding with trained draft modules.
The library integrates with NVIDIA Megatron-Bridge, Megatron-LM and Hugging Face Accelerate for techniques that require training. Windows quantization workflows and NVIDIA container images are available. NVIDIA also publishes pre-quantized checkpoints on Hugging Face for deployment with TensorRT-LLM, vLLM and SGLang.
Install the Python package with the dependencies for the intended workflow. Linux CUDA examples target NVIDIA GPUs; supported model formats, quantization schemes and deployment runtimes depend on the hardware and support matrix.
Claim this page and we'll verify you by hand. TensorRT Model Optimizer gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find TensorRT Model Optimizer?Promote it
Something wrong or outdated on this page?
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
2.7KUpdated 1 week agoApache-2.0
Linux · Docker#Hugging Face integration#Quantization
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
Intel Neural Compressor is a Python library for developers compressing AI models for deployment on their own hardware or servers. It supports local LLM work as well as other deep learning models, with particular attention to Intel CPUs, GPUs and Gaudi accelerators. It's open source under the Apache 2.0 license.