llama.cpp setup: Qwen3.6-35B-A3B on a 6GB GPU

Learn CPU expert offloading and memory tuning in llama.cpp. The speaker reports 17 tokens per second with a GTX 1060 and 32GB of RAM.

Player not loading? Watch on YouTube

This tutorial explains how to run models locally when the model file exceeds GPU memory. The speaker reports running a 22.36GB, four-bit Qwen3.6-35B-A3B model at 17 tokens per second on a 6GB GTX 1060 with 32GB of system RAM. Its mixture of experts architecture activates only part of the model for each token, which makes selective CPU offloading practical.

The llama.cpp setup uses --n-cpu-moe to keep routed expert weights in system RAM, then adjusts GPU layers to use the remaining VRAM. The speaker recommends reducing the number of CPU expert layers until memory runs out, then stepping back. Loading weights into RAM and keeping pages resident can help, but memory mapping may work better when the model exceeds available RAM. The tutorial also describes the consolidated --load-mode option and explains why larger micro-batches can improve prompt processing separately from answer generation.

System memory bandwidth remains a limit for this local LLM configuration. Long context also consumes KV cache memory. The speaker discusses TurboQuant's three-bit cache compression, but says it requires a fork rather than upstream llama.cpp and can reduce decoding speed. Speculative decoding can increase expert memory traffic in offloaded models. The closing assessment treats this setup as a local assistant, with reported throughput too slow for an unattended coding agent, and recommends testing quantized model quality on your own repository.