
SGLang is a self-hosted inference framework for teams that need to serve language and multimodal models on their own hardware. It runs on a single GPU or across distributed clusters and exposes an OpenAI-compatible API. The project is open source under the Apache 2.0 license.
Model support covers Llama, Qwen, DeepSeek, Kimi, GLM, Gemma and Mistral, along with many Hugging Face models. SGLang also serves embedding and reward models, plus diffusion models such as WAN and Qwen-Image. That range makes it relevant to teams running different AI workloads through the same serving framework.
Its runtime is built for low latency and high request volume. Prefix caching reuses work when requests share text, while continuous batching and speculative decoding help the server process requests efficiently. It supports structured outputs, quantized models and multiple LoRA adapters, giving teams ways to serve different workloads without treating each model as a separate project.
SGLang supports NVIDIA and AMD GPUs, Google TPUs, Ascend NPUs and Moore Threads MUSA accelerators. Intel Xeon CPUs are supported. For larger deployments, it can split model work across multiple devices and servers. It also acts as a rollout backend for reinforcement learning workflows, with integrations including AReaL, Miles, slime, Tunix and verl.
Claim this page with an email at docs.sglang.ai. SGLang gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find SGLang?Promote it
Something wrong or outdated on this page?
93KUpdated 1 hour agoApache-2.0
macOS · Docker#Batch processing#Distributed execution#GGUF
vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.
182KUpdated 16 hours agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input
14.7KUpdated 21 hours ago
Docker#Batch processing#Distributed execution#LoRA
130KUpdated 36 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
47.7KUpdated 1 month agoApache-2.0
macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX
3.1KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization
Ollama runs language models on your own computer or server. It provides a command-line runner and a local API for people building AI applications or connecting existing tools to models they host themselves. The software is distributed under the MIT license.
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
exo is a local LLM runner that combines your devices into a cluster, letting you use models too large for one machine's memory. It's for people who want to run large models on their own hardware and developers connecting existing AI clients to local inference. It runs on macOS and Linux under the Apache 2.0 license.
Distributed Llama runs a local LLM across several computers, sharing both the computation and the model's memory use. It's for people who want to use their own networked hardware for inference rather than keep the entire workload on one machine. The C++ project is open source under the MIT license.