Slotstream: Qwen 3.8 Flash Next on a 48 GB Mac explained

Learn how SlotStream streams Qwen experts from SSD, why its memory target caps at 33 GB, and what the reported 12 tokens per second leaves out.

Player not loading? Watch on YouTube

SlotStream is an open source engine for running Qwen 3.8 Flash Next as a local LLM on Apple silicon. The video examines a reported result: a 104 GB model running on an M5 Pro Mac with 48 GB of memory at about 12 tokens per second during warm decode. The speaker attributes the figures to the project's README and developer reports.

The explanation centers on expert streaming. Each layer has 512 experts, but a token uses only 10. SlotStream keeps the 3.8 GB dense trunk in memory and reads expert byte ranges from SSD into one fixed slot pool shared across all 48 layers. According to the account, ordinary memory mapping fails here because MLX evaluates the whole expert tensor when gathering selected experts.

The reported memory target caps at 33 GB. The developer's sweep found no speed gain above that point, so 48 GB and 128 GB Macs receive the same memory plan. That finding comes from one measured machine. Smaller Mac results are estimates, including roughly four tokens per second for a 16 GB Mac. Prompt processing also takes time: the speaker reports about a minute before the first token for an 8,000-token document on the measured Mac.

The experiment requires Apple silicon, macOS 14 or later, and about 110 GB of free disk. It uses Swift and MLX, supports only Flash Next at 4-bit quantization, and has no Linux or CUDA support. The developer describes it as an experiment rather than a commercial product.