Player not loading? Watch on YouTube
This explainer covers the loading stage of local LLM inference: how an engine takes model files on disk and prepares their weights for computation. The speaker introduces weight and configuration artifacts, then describes how llama.cpp can divide model weights between RAM and an accelerator.
Memory mapping, or mmap, lets the operating system load file-backed memory as needed and reload pages after eviction. The speaker reports getting a first token in under 10 seconds with llama.cpp in one example, while describing vLLM startup as potentially taking minutes because of compilation and heavier initialization. These are examples from the presentation, rather than a controlled benchmark or a general speed guarantee.
The quantization discussion explains rounding, group scales, and symmetric versus asymmetric ranges. It then covers GGUF-related quantization choices and K-quants, followed by AWQ's use of calibration data to identify important weights. EXL2 receives a separate explanation of mixed precision and sensitivity estimates. A Llama 2 13B comparison favors EXL2 in the example shown.
For people who run models locally, the hardware discussion matters: the speaker identifies Hopper support for native FP8 and Blackwell support for NVFP4. Prefill, decoding, scheduling, and KV cache management remain outside this video's main scope.