Player not loading? Watch on YouTube
This tutorial explains how to run models locally through four approaches: LM Studio, Ollama, Docker Model Runner and Python code that invokes llama.cpp. The speaker first describes model weights, parameter counts, quantization and the inference engine that loads a model and generates responses.
Hardware selection centers on GPU VRAM or unified memory on M-series Macs. The speaker explains that longer context windows consume additional memory and discusses the tradeoff between model size and generation speed. His memory ranges and speed examples are rules of thumb and results from his own hardware, rather than guarantees for every model.
The LM Studio walkthrough covers downloading a model, choosing quantization, adjusting context size and loading it for chat or API access. Ollama's demo uses list, pull and run commands, then explains how applications can call its local API. The speaker presents Docker Model Runner as an experimental Docker Desktop feature, discusses Linux with Nvidia hardware, and introduces deployment through Docker and Compose files. The Python example loads a model directly and also calls an Ollama backend.