Player not loading? Watch on YouTube
This tutorial explains how to run models locally with llama.cpp, then expose a model through a server API and connect OpenCode. The presenter uses Docker on WSL and selects a Gemma 4 26B Q4 GGUF file from Hugging Face for the demonstration.
Quantization is the starting point: smaller model files reduce memory demands, with a tradeoff in output quality. The presenter recommends Q4 as a starting choice and discusses CPU execution alongside GPU acceleration. He lists Linux, Mac and Windows support and cites 4 GB of RAM plus a modern CPU as a baseline, rather than a requirement for the larger model used here.
The setup covers Docker image variants and hardware-specific tags, a model directory mounted at /models, and NVIDIA Container Toolkit installation on the host. In llama-cli, the tutorial uses the NGL setting to control GPU layer offload and checks logs and GPU memory usage before adjusting performance.
The server example binds to 0.0.0.0 and exposes a chat interface and API on port 8080. Context size affects memory use and how much conversation the model can retain. The final section configures OpenCode with the server URL, a model reference and token limits. The presenter describes offline use with local models, though the final OpenCode example performs a web fetch.