Player not loading? Watch on YouTube
This tutorial explains how to run Gemma 4 as a local LLM on a computer or phone. The speaker describes it as an open weight model and outlines four variants: E2B for lightweight devices, E4B for consumer computers, a 26B mixture-of-experts model for high-end GPUs, and a 31B dense model for enterprise or multiple GPUs.
The desktop walkthrough uses Ollama on Windows and Mac. It covers model downloads with ollama pull gemma4:e4b or ollama pull gemma4:e2b, then shows Windows graphics preferences for both Ollama executables without measuring GPU offload. The phone walkthrough separately covers its accelerator setting. Examples include chat and PDF summarization with the internet disconnected. He recommends starting context length at 16K or 32K rather than the advertised maximum.
On phones, Google AI Edge Gallery provides separate workspaces for chat, image questions and audio. The tutorial explains temperature, TopK, TopP, response length, GPU acceleration and thinking mode. These workspaces do not smoothly combine camera, audio and text in one conversation. Agent skills can contact internet services even when reasoning happens on-device; desktop web search also requires a connection.
The closing assessment weighs hardware demands and slower responses against local processing. The speaker gives January 2025 as the knowledge cutoff and warns that VRAM can constrain usable context. He treats Gemma 4 as a supplement for sensitive documents or disconnected work and advises checking company compliance rules.