Favicon of SGLang

SGLang

An open-source inference framework for serving language and multimodal models on your own hardware, with an OpenAI-compatible API.

Screenshot of SGLang website

SGLang is a self-hosted inference framework for teams that need to serve language and multimodal models on their own hardware. It runs on a single GPU or across distributed clusters and exposes an OpenAI-compatible API. The project is open source under the Apache 2.0 license.

Model support covers Llama, Qwen, DeepSeek, Kimi, GLM, Gemma and Mistral, along with many Hugging Face models. SGLang also serves embedding and reward models, plus diffusion models such as WAN and Qwen-Image. That range makes it relevant to teams running different AI workloads through the same serving framework.

Its runtime is built for low latency and high request volume. Prefix caching reuses work when requests share text, while continuous batching and speculative decoding help the server process requests efficiently. It supports structured outputs, quantized models and multiple LoRA adapters, giving teams ways to serve different workloads without treating each model as a separate project.

SGLang supports NVIDIA and AMD GPUs, Google TPUs, Ascend NPUs and Moore Threads MUSA accelerators. Intel Xeon CPUs are supported. For larger deployments, it can split model work across multiple devices and servers. It also acts as a rollout backend for reinforcement learning workflows, with integrations including AReaL, Miles, slime, Tunix and verl.

Similar to SGLang