llama.cpp and Qwen 3.6: 35B setup on 6GB VRAM

Learn MoE offloading and memory flags for llama.cpp. The speaker reports 17 tokens/s with Qwen 3.6 35B-A3B on a GTX 1060 and 24GB RAM.

Player not loading? Watch on YouTube

This tutorial explains a local LLM setup using llama.cpp and Qwen 3.6 35B-A3B on a GTX 1060 with 6GB VRAM, an Intel i3-8100 and 24GB DDR4 RAM. The supplied setup uses Proxmox, LXC and Docker. The speaker reports about 3 tokens per second with a basic split between GPU and CPU layers.

The main adjustment keeps expert weights in CPU memory while placing other model components on the GPU. The speaker reports 10 tokens per second with --n-cpu-moe 41, then 13.5 after --no-mmap loads the roughly 20GB model into RAM upfront. Reducing the CPU expert layer count to 35 reportedly reaches 17 tokens per second, but leaves less VRAM for context, reducing it to about 64,000 tokens.

The context experiment uses TurboQuant cache compression; the description links a llama.cpp fork. The speaker reports fitting 128,000 tokens, then 256,000 after moving another layer's experts to the CPU, with 5.9GB VRAM usage. These are reported results on this rig, including the claimed preservation of speed and answer quality.

For sustained operation, the guide covers memory locking permissions in LXC, Docker's IPC lock capability and --mlock. It also documents a failed speculative decoding test: a Qwen 3.5 800M draft model reduced reported speed from 17 to 11 tokens per second despite roughly 65% acceptance.