Player not loading? Watch on YouTube
This comparison examines how SGLang, vLLM and llama.cpp serve a local LLM as requests overlap and concurrency rises. The speaker presents a single-GPU benchmark with Qwen3.5-35B-A3B, SGLang 0.5.15, vLLM 0.20.1 and an unspecified latest llama.cpp build. The reported results describe these workloads, rather than establishing one engine as universally fastest.
For unique prompts at 32 concurrent requests, the speaker reports 1,743 output tokens per second for vLLM and 1,661 for SGLang. With a shared system prefix, SGLang reportedly pulls ahead as reusable context grows. The explanation centers on RadixAttention: a radix tree stores cached prompt prefixes so later requests can reuse shared branches. The video also discusses structured JSON output through finite-state machines.
The speaker recommends llama.cpp for personal use with GGUF files on laptops, Macs using Metal and CPU systems, while describing throughput and queueing limits under heavier concurrent traffic. vLLM's PagedAttention and continuous batching receive attention as mechanisms for serving more requests. For a self-hosted AI agent workload with repeated tool definitions or chat history, the speaker favors SGLang. Quantization formats and an OpenAI-compatible localhost endpoint round out the discussion; hardware, prompt overlap and context length remain relevant to the choice.