vLLM tutorial: Hugging Face comparison and API setup

Compare Hugging Face inference with vLLM, launch an OpenAI-compatible server, and test up to 20 concurrent users with a Gradio dashboard.

Player not loading? Watch on YouTube

This walkthrough uses a 135M-parameter model to compare a basic Hugging Face inference script with vLLM, then builds a self-hosted inference API. The lab checks its environment before running the same model and prompt through both engines. The speaker reports higher tokens per second with vLLM in this example.

The memory exercises explain KV cache allocation and PagedAttention through an operating system paging analogy. In the demonstration, memory utilization rises from about 20% with contiguous allocation to about 95% with pages allocated on demand.

The API task uses the OpenAI SDK with a localhost endpoint on port 8000. Load tests cover 1, 5, 10 and 20 concurrent users. The speaker describes higher aggregate throughput alongside a slight increase in per-request latency. Later tasks vary context length and concurrent sequence limits, then display performance comparisons and live metrics in Gradio.

For hardware selection, the speaker favors vLLM for GPU-based multi-user serving and llama.cpp for CPU/RAM inference on consumer hardware. The tutorial focuses on serving and measurement rather than a desktop chat interface.