Player not loading? Watch on YouTube
This beginner tutorial walks through Locally AI on Mac, iPad and iPhone, from App Store installation to downloading a local LLM and starting a chat. The speaker attributes the claim that the app collects no information or uploads conversations to its developer. Downloaded models can run offline, according to the tutorial. Apple Intelligence is optional: the app can select the Apple Foundation model when it is installed, or use downloaded alternatives.
The model selection walkthrough covers Gemma 4 E2B and E4B. The walkthrough shows the app recommending the 4B option for M3 iPads and newer, and the 2B option for M2 and newer; Bonsai is presented as an option for M1 iPads. The speaker explains parameter counts and speed tradeoffs, then enables thinking mode, which is off by default and adds response time. Image tests extract text from a document cover but fail to identify Notre Dame in a photograph.
Personalization examples show how custom instructions and temperature affect replies, including a reusable email response prompt. For coding, the speaker switches to a downloaded 9B coding model identified in the narration as Code Llama 3.5 and asks it to write a system prompt of at most 1,000 characters. A calendar reference image then guides a frontend mockup. The generated HTML is previewed in CodePen; dates overflow, so the speaker suggests further corrections through chat or screenshots.