Favicon of TensorRT-LLM

TensorRT-LLM

A self-hosted LLM inference library built on PyTorch for NVIDIA GPUs, with a Python API, OpenAI-compatible serving, and multi-node support.

Screenshot of TensorRT-LLM website

TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.

The Python LLM API handles inference within applications, while trtllm-serve exposes models through APIs compatible with OpenAI clients. Supported workloads include text generation, multimodal requests and embeddings. Model-specific support includes DeepSeek-R1, Llama and GPT-OSS-120B. It also integrates with NVIDIA Dynamo and Triton Inference Server for larger serving deployments.

Its performance features address different bottlenecks: quantization reduces model precision, in-flight batching groups active requests, and cache reuse avoids repeating some work across requests. Speculative decoding and parallel execution provide further optimization options. For large deployments, it can separate prompt processing from token generation. LoRA support and guided decoding cover adapted models and structured output, including JSON schemas.

Inference runs on your hardware. Anonymous usage telemetry is enabled by default and can be disabled. It reports deployment and hardware metadata, but excludes prompts, generated outputs, model weights and local model paths. Benchmarking and evaluation tools help teams measure throughput and assess their deployments.

Similar to TensorRT-LLM