Player not loading? Watch on YouTube
This tutorial builds a local AI coding assistant with llama.cpp and Pi Coding Agent on an RTX 3060 with 12GB VRAM, DDR4 RAM, and a four-core CPU. The speaker uses Q4 REAP variants of Qwen3.6 and GLM-4.7 Flash, offloading part of each model to system RAM. He presents them as capable options for coding work, while cautioning that they do not match Claude or GPT-5.
The tuning focuses on prefill: processing the instructions, files, and tool details an agent needs before generating a response. Using llama.bench, the speaker finds that three threads outperform four on his machine. He reports Qwen prompt processing rising from 300 to 1,142 tokens per second when the microbatch increases from 256 to 2,048. For regular use, he chooses 1,024 to leave more VRAM available, reporting 870 tokens per second. These are results from his setup, and he recommends benchmarking each machine. The demonstrated TurboQuant fork of llama.cpp adds KV cache compression, with different tradeoffs for the two models.
The setup connects Pi through the pi-llama-cpp extension and a server URL in settings.json. An INI preset file lets Pi discover models and ask llama-server to unload one and load another, though large models still take time to switch. Optional Tailscale access connects a laptop to the home server through its Tailscale IP.