Favicon of Slotstream

Slotstream

Local LLM runner for Apple Silicon Macs that runs Qwen3.8-Flash-Next from SSD. Works offline after download and connects to coding agents and chat apps.

Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.

It requires macOS 14 or later, at least 16 GB of RAM and about 110 GB of free SSD space. Windows, Linux and Intel Macs aren't supported. After the model download, inference works offline without a cloud account. Connected agents can still contact internet services for searches or other tools, depending on their settings.

Claude Code, Codex, Pi, opencode and Hermes can use it as their model backend. Chat apps such as Open WebUI can connect through its Ollama-, OpenAI- or Anthropic-compatible APIs. Developers also have a Swift library. Slotstream supplies the inference engine; connected apps supply their own interfaces and tools.

The native Swift engine uses Apple's MLX and Metal without a Python runtime. It adjusts its memory cache as other apps need room and reuses conversation history to reduce repeated prompt processing. Long prompts can still take minutes before a reply starts, and connected apps share one generation at a time.

Slotstream supports only Qwen3.8-Flash-Next. It also runs on Macs with 96 GB or more, but engines that keep the model entirely in RAM may reply faster there. The engine is open source under MIT; the model weights use a separate Qwen community license.

Similar to Slotstream