Player not loading? Watch on YouTube
The presenter tests Gemma 4 12B as a local LLM on an M4 Mac mini with 16GB of RAM, using LM Studio for chat and a local server. He describes the model as open source and discusses Google's claimed performance relative to Gemma 4 26B. Those benchmark claims are separate from the hands-on tests shown here.
Memory is the main setup constraint. The selected download is 7.56GB, while LM Studio estimates about 8GB to load it with a 4,000-token context. Increasing context length raises memory use: the presenter selects roughly 31,000 tokens at about 12GB, while the displayed estimate for 131,000 tokens exceeds 26GB. After loading, he reports total system memory use of 14.9GB. These figures describe his configuration, with limited room left for other applications. He also warns that large initial prompts in coding harnesses can exceed the selected context budget.
The practical test asks Gemma to extract details from an invoice image, then generate an HTML invoice parser that calls LM Studio's server. The presenter previews the generated app in VS Code and sends the invoice through it. The connection works, but both extraction attempts return an incorrect balance due. The demo shows a working application flow with an accuracy problem that would need validation before use for office tasks. The video also gives the presenter's high-level explanation of multimodal processing and multi-token prediction, which he describes as accelerating decoding rather than prompt processing.