Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.
It requires macOS 14 or later, at least 16 GB of RAM and about 110 GB of free SSD space. Windows, Linux and Intel Macs aren't supported. After the model download, inference works offline without a cloud account. Connected agents can still contact internet services for searches or other tools, depending on their settings.
Claude Code, Codex, Pi, opencode and Hermes can use it as their model backend. Chat apps such as Open WebUI can connect through its Ollama-, OpenAI- or Anthropic-compatible APIs. Developers also have a Swift library. Slotstream supplies the inference engine; connected apps supply their own interfaces and tools.
The native Swift engine uses Apple's MLX and Metal without a Python runtime. It adjusts its memory cache as other apps need room and reuses conversation history to reduce repeated prompt processing. Long prompts can still take minutes before a reply starts, and connected apps share one generation at a time.
Slotstream supports only Qwen3.8-Flash-Next. It also runs on Macs with 96 GB or more, but engines that keep the model entirely in RAM may reply faster there. The engine is open source under MIT; the model weights use a separate Qwen community license.
Claim this page and we'll verify you by hand. Slotstream gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Slotstream?Promote it
Something wrong or outdated on this page?
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
23.2KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.