Player not loading? Watch on YouTube
This guide compares Qwen3.8-27B configurations for Apple Silicon Macs with 16GB to 48GB of unified memory. The speaker explains how its hybrid attention layout reduces cache growth, while model weights, prefill buffers and macOS still compete for memory. The guide excludes 8GB machines as impractical and treats Intel Macs separately.
For 16GB, the speaker recommends a text-only 2-bit MLX build for speed, reporting about 11 tokens per second on an M4 with a 10-core GPU. The tested 3-bit alternative needs system tuning and runs at 3.7 tokens per second; maintainer testing reportedly found that agentic loops failed at this memory tier. At 24GB, standard 4-bit builds become practical for shorter contexts. A 2-bit build with a 4-bit cache reportedly reached 16,000 tokens on a base M5.
The speaker recommends 32GB for a more flexible local LLM setup: 6-bit for quality or 4-bit for speed and context. An M2 Max test reached 32,000 tokens at 15.8 tokens per second. At 48GB, reported tests reached 128,000 tokens, with separate community testing reaching 195,000. The advertised 262,000-token window remains a configuration to validate, rather than a demonstrated result here.
The guide also reports a llama.cpp issue beyond roughly 130,000 tokens and recommends validated MLX or oMLX environments for longer contexts. It advises disabling experimental Apple Neural Engine prefill in oMLX because of reported timeouts.