Player not loading? Watch on YouTube
This tutorial tests EXO on three base M4 Mac Minis, each with 16 GB of unified memory. The presenter connects them in a triangle with Thunderbolt 4 cables, then assigns each macOS Thunderbolt bridge a static IP on the same subnet. EXO discovers the nodes and provides a browser interface on port 52415 for loading models and choosing how many nodes to use.
The local LLM tests show the tradeoff between memory capacity and generation speed. The presenter reports about 21 tokens per second for Qwen 3.5 9B on one Mini, compared with 19 across two or three. He attributes the slowdown to Thunderbolt 4 bandwidth and network overhead. Qwen 3.5 27B at 8-bit quantization requires all three machines in this setup and produces about 3.7 tokens per second. He considers that too slow for interactive chat, but suggests background jobs such as overnight code checks as possible uses.
The walkthrough also connects Open WebUI through EXO's Ollama endpoint so the interface can address the cluster as one service. For hardware selection, the presenter compares the base Mini with the M4 Pro Mini and M3 Ultra and M5 Max Mac Studios. He favors the base model for smaller workloads or adding memory to an existing setup, while pointing to Thunderbolt 5 and RDMA on the alternatives for larger clustered workloads. These recommendations reflect his tests and the comparison benchmarks he gathered.