
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
The Python LLM API handles inference within applications, while trtllm-serve exposes models through APIs compatible with OpenAI clients. Supported workloads include text generation, multimodal requests and embeddings. Model-specific support includes DeepSeek-R1, Llama and GPT-OSS-120B. It also integrates with NVIDIA Dynamo and Triton Inference Server for larger serving deployments.
Its performance features address different bottlenecks: quantization reduces model precision, in-flight batching groups active requests, and cache reuse avoids repeating some work across requests. Speculative decoding and parallel execution provide further optimization options. For large deployments, it can separate prompt processing from token generation. LoRA support and guided decoding cover adapted models and structured output, including JSON schemas.
Inference runs on your hardware. Anonymous usage telemetry is enabled by default and can be disabled. It reports deployment and hardware metadata, but excludes prompts, generated outputs, model weights and local model paths. Benchmarking and evaluation tools help teams measure throughput and assess their deployments.
Claim this page and we'll verify you by hand. TensorRT-LLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find TensorRT-LLM?Promote it
Something wrong or outdated on this page?
93KUpdated 1 hour agoApache-2.0
macOS · Docker#Batch processing#Distributed execution#GGUF
vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.
130KUpdated 37 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
36.7KUpdated 1 hour agoApache-2.0
#Batch processing#Distributed execution#LoRA
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
19.5KUpdated 1 week agoApache-2.0
Docker#LoRA#Multimodal input#Prompt caching
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
SGLang is a self-hosted inference framework for teams that need to serve language and multimodal models on their own hardware. It runs on a single GPU or across distributed clusters and exposes an OpenAI-compatible API. The project is open source under the Apache 2.0 license.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
KTransformers is an open-source framework for running and fine-tuning large language models on your own hardware. It focuses on mixture-of-experts (MoE) models, distributing work between CPU memory and GPU resources to reduce the GPU memory needed. It's aimed at researchers and developers who want to serve or adapt models such as DeepSeek-V3 and DeepSeek-R1.