Favicon of Triton Inference Server

Triton Inference Server

Self-hosted AI inference server runs TensorRT, PyTorch and ONNX models on GPUs or CPUs, with dynamic batching and a BSD-3-Clause license.

Screenshot of Triton Inference Server website

Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.

Supported backends include TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL. Teams can run multiple models concurrently and combine them into pipelines with preprocessing and postprocessing. Custom backends let developers extend the server for workloads that need their own execution logic.

Dynamic batching groups requests for processing, while sequence batching supports workloads that carry state between requests. Triton handles real-time requests, batch jobs and audio/video streams. Its Model Analyzer helps teams profile model configurations and compare performance before choosing how to serve them.

GPU hardware is optional. Triton runs on x86 and ARM CPUs, NVIDIA GPUs and non-NVIDIA accelerators, including AWS Inferentia. Linux Docker containers are available for x86 and ARM, alongside Windows and NVIDIA Jetson releases.

Applications connect through HTTP or gRPC, with Python and C++ client libraries. Kubernetes integration supports scaling, and Prometheus collects monitoring metrics. NVIDIA also offers commercial support through NVIDIA AI Enterprise. The separate NVIDIA Dynamo project complements Triton with LLM-specific capabilities such as prefix caching and disaggregated serving.

Similar to Triton Inference Server