Player not loading? Watch on YouTube
This comparison separates the local LLM stack into model formats, execution runtimes, serving engines, model managers and desktop applications. The speaker argues that a tokens-per-second screenshot says little without the hardware, format and workload behind it.
llama.cpp is presented as the option for GGUF portability and detailed CPU/GPU offloading control. Ollama handles model downloads and exposes an API on port 11434, with less access to runtime settings. LM Studio adds a graphical workbench, document retrieval and MCP support over bundled llama.cpp and MLX runtimes. The speaker distinguishes those open source runtimes from the proprietary desktop application and discusses restrictions on redistribution and SaaS use.
For concurrent serving, the comparison turns to vLLM and SGLang. It describes their cache management and batching approaches, while noting that vLLM's GGUF support is experimental and unoptimized. TensorRT-LLM requires a more rigid NVIDIA build setup; ExLlamaV3 uses converted EXL3 models, and MLX LM targets Apple hardware.
The final section examines FreeToken and Colibrì for large mixture-of-experts models. The speaker states that FreeToken's CLI requires Linux x86-64, an NVIDIA GPU and CUDA 13, with curated model support and text-only multimodal operation. Colibrì streams cold experts from SSD storage, but cache misses can sharply limit response speed.