
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.
Supported backends include TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL. Teams can run multiple models concurrently and combine them into pipelines with preprocessing and postprocessing. Custom backends let developers extend the server for workloads that need their own execution logic.
Dynamic batching groups requests for processing, while sequence batching supports workloads that carry state between requests. Triton handles real-time requests, batch jobs and audio/video streams. Its Model Analyzer helps teams profile model configurations and compare performance before choosing how to serve them.
GPU hardware is optional. Triton runs on x86 and ARM CPUs, NVIDIA GPUs and non-NVIDIA accelerators, including AWS Inferentia. Linux Docker containers are available for x86 and ARM, alongside Windows and NVIDIA Jetson releases.
Applications connect through HTTP or gRPC, with Python and C++ client libraries. Kubernetes integration supports scaling, and Prometheus collects monitoring metrics. NVIDIA also offers commercial support through NVIDIA AI Enterprise. The separate NVIDIA Dynamo project complements Triton with LLM-specific capabilities such as prefix caching and disaggregated serving.
Claim this page with an email at developer.nvidia.com. Triton Inference Server gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Triton Inference Server?Promote it
Something wrong or outdated on this page?
6KUpdated 22 hours agoApache-2.0
#Hugging Face integration#ONNX#OpenAI-compatible API
KServe is an open source platform for teams serving LLMs and predictive machine learning models on their own Kubernetes infrastructure. It puts both kinds of workloads under a common serving API, so teams can manage different model frameworks through the same platform. It uses the Apache 2.0 license.
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
2.5KUpdated 1 day ago
macOS · Windows · Linux · Docker#Batch processing#Code execution#Multimodal input
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
818Updated 4 years agoApache-2.0
macOS · Windows · Linux · Docker#Home Assistant integration#Works offline
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
Roboflow Inference is a self-hosted computer vision server for teams building camera and image analysis systems. It runs on your own computer, server, or edge device and combines model predictions with workflows for tracking, counting, measuring, and responding to events. Roboflow also offers hosted servers and a Serverless Cloud API, where processing runs on its infrastructure.
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
DeepStack is a self-hosted computer vision API for developers adding image analysis to camera systems, home automation or other applications. It runs prebuilt and custom models on your own hardware and works fully offline. Image processing stays on the device or server where you host it, with no cloud service required.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.