
TensorRT is NVIDIA's AI inference SDK for developers deploying models on their own NVIDIA GPUs. It combines a model compiler and runtime to reduce response latency and increase throughput on workstations, laptops, edge devices and data center servers. It's intended for application developers and teams serving models, rather than people looking for a ready-made chat app.
The SDK accepts ONNX models and integrates with PyTorch and Hugging Face. It supports applications such as image generation, speech processing and video analysis alongside language models. TensorRT-LLM handles LLM inference on NVIDIA data center and workstation GPUs, while TensorRT for RTX targets GeForce RTX and RTX Pro PCs with engines portable across operating systems and GPUs.
Model Optimizer compresses models through quantization, pruning and distillation. It supports deployment through TensorRT, TensorRT-LLM, vLLM and SGLang. For self-hosted serving, NVIDIA Dynamo Triton can run TensorRT models with dynamic batching and concurrent model execution.
Inference runs on the target GPU. The separate TensorRT Cloud service builds optimized engines in NVIDIA's cloud for a chosen GPU and performance target. Linux and Docker are supported environments, and deployment targets include Jetson and NVIDIA DRIVE devices.
The open source components use Apache 2.0, but cover only part of the full SDK. They include the ONNX parser, plugins and sample applications.
Claim this page with an email at developer.nvidia.com. TensorRT gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find TensorRT?Promote it
Something wrong or outdated on this page?
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
23.9KUpdated 10 hours ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
10.9KUpdated 10 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
5.1KUpdated 1 day ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
7.3KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · iOS · Android#Quantization
Paddle Lite is an on-device AI inference engine for developers building mobile apps, embedded systems and edge devices. It focuses on running neural networks with smaller model files, a smaller runtime and lower memory use. It's open source under Apache 2.0.
1.7KUpdated 3 days ago
Linux#Multimodal input#Quantization
RKLLM is a software stack for developers building local AI applications on Rockchip hardware. It uses the chip's neural processing unit (NPU) to run language and multimodal models on development boards, with support for the RK3588, RK3576, RK3562 and RV1126B series.
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.