Favicon of TensorRT

TensorRT

AI inference SDK that optimizes models for NVIDIA GPUs on local PCs, edge devices and servers, with ONNX, PyTorch and Hugging Face support.

Screenshot of TensorRT website

TensorRT is NVIDIA's AI inference SDK for developers deploying models on their own NVIDIA GPUs. It combines a model compiler and runtime to reduce response latency and increase throughput on workstations, laptops, edge devices and data center servers. It's intended for application developers and teams serving models, rather than people looking for a ready-made chat app.

The SDK accepts ONNX models and integrates with PyTorch and Hugging Face. It supports applications such as image generation, speech processing and video analysis alongside language models. TensorRT-LLM handles LLM inference on NVIDIA data center and workstation GPUs, while TensorRT for RTX targets GeForce RTX and RTX Pro PCs with engines portable across operating systems and GPUs.

Model Optimizer compresses models through quantization, pruning and distillation. It supports deployment through TensorRT, TensorRT-LLM, vLLM and SGLang. For self-hosted serving, NVIDIA Dynamo Triton can run TensorRT models with dynamic batching and concurrent model execution.

Inference runs on the target GPU. The separate TensorRT Cloud service builds optimized engines in NVIDIA's cloud for a chosen GPU and performance target. Linux and Docker are supported environments, and deployment targets include Jetson and NVIDIA DRIVE devices.

The open source components use Apache 2.0, but cover only part of the full SDK. They include the ONNX parser, plugins and sample applications.

Similar to TensorRT