llama.cpp source build and quantization tutorial

Learn to build llama.cpp with CUDA on Linux Mint and choose quantized models, with tests showing how context consumes a 24 GB GPU's memory.

Player not loading? Watch on YouTube

This tutorial builds llama.cpp from source on Linux Mint with Nvidia drivers installed. The demonstrated setup uses build 9156 and Nvidia toolkit version 13.2.1-595, providing CUDA 13.2. Prerequisites include Git, CMake and OpenBLAS. The speaker configures CUDA support, runs a release build and checks the resulting CLI version before testing the server through a browser on port 8080.

The comparison with LM Studio focuses on convenience versus build control. The speaker favors compiling for access to recent code, experimental branches and hardware-specific options. An NCCL warning concerns coordination across multiple CUDA GPUs; the demonstrated machine has one GPU, so the speaker proceeds without it.

For local LLM model selection, the tutorial explains quantization as a tradeoff between precision and memory use. The speaker recommends Q4 for general use, rather than establishing it as a universal quality threshold. A roughly 17 GB Gemma 4 31B model uses about 22 GB of VRAM with 64,000 tokens of context, illustrating why model file size alone cannot determine fit.

After a slow dense-model test, the speaker switches to a 26B A4B MoE model at Q4. With 6,400 tokens of context, the reported result is 17.9 GB of VRAM use and about 92 tokens per second. These are results from this setup. Connecting OpenCode as a coding assistant is planned for a later episode.