Player not loading? Watch on YouTube
Locally AI developer Adrien Grondin explains how to run models locally on iPhone with MLX, Apple's framework for Apple Silicon. He introduces Locally AI as a native chatbot available on iPhone, iPad and macOS, with support for MLX-compatible models and Apple Foundation.
For developers, the talk outlines a setup using MLX Swift LM. Choose compatible weights from the MLX community on Hugging Face, then pass the model ID to the framework to download and load them. This is an introduction rather than a code walkthrough. Grondin recommends 4-bit to 8-bit quantization for phones, depending on model size, and says lower precision can hurt output quality.
The Gemma 4 demonstration illustrates streaming output that he describes as running offline at about 40 tokens per second on a recent iPhone. He says older iPhones can also run models, but viewers should not expect the same speed. The narration gives inconsistent bit-depth details for that demo, so the reported speed is best read as a result from the speaker's example. Model downloads, which he puts at roughly 1 to 3 GB, remain a practical constraint.
The closing discussion describes LM Studio's model downloads, local server and MLX and Llama CPP engines. In the Q&A, Grondin says MLX Swift LM supports tool calling but not yet structured generation. He also explains that Locally AI limits its model selection to options he has checked work on iPhone.