Pi and llama.cpp: local GGUF model setup

Learn to connect Pi to llama.cpp, choose GGUF quantization for your hardware, and load a model with /llama. The example uses an M4 Max with 36 GB.

Player not loading? Watch on YouTube

This tutorial shows how to run a local LLM in Pi through llama.cpp. The presenter copies the installer command from llama.app, starts a local server with llama serve, and connects Pi to it. He describes the demonstrated inference as staying on his computer, with no prompts or code sent to external APIs.

Model selection centers on GGUF files and hardware capacity. On Hugging Face, he adds a Mac with an M4 Max and 36 GB to his profile's hardware settings, then checks the model's compatibility recommendations. For that configuration, the page recommends 4-bit quantization. This is a recommendation for his example machine, rather than a requirement for every setup.

In an updated version of Pi, he opens /llama, chooses the download option, pastes a model ID, and selects Q4. The model is already downloaded in his demonstration, so he skips transferring it again. He then loads it and uses /model to select llama.cpp before sending a greeting.

The closing discussion suggests using a cloud model for planning and a local coding assistant for implementation; the presenter's privacy claims concern the local portion of that workflow.