Gemma 4 with Ollama: local setup and GPU checks

Learn how to install Gemma 4 with Ollama, check GPU memory on Windows, test image prompts, and run the 31B model on a rented RTX 5090 server.

Player not loading? Watch on YouTube

This beginner tutorial explains how to run models locally with Ollama, using Google's Gemma 4 family. The speaker describes four sizes: effective 2B and 4B models, a 26B mixture-of-experts model with 3.8B active parameters, and a 31B dense model. He also states that the family uses an Apache 2.0 license.

The setup starts with Ollama's download options for Windows, Mac and Linux, then follows a Windows installation. Gemma 4 is absent from the app's model selector during the walkthrough, so the speaker uses the terminal to download a smaller variant. After installation, it appears in the chat interface. He notes that the first response takes time while the model loads into VRAM.

GPU memory is a practical focus. The tutorial checks an RTX 5070 Ti with 16 GB of VRAM using NVIDIA's SMI utility and Windows Task Manager. The speaker warns that larger models can fall back to the CPU and respond slowly when GPU memory is insufficient. His hardware guidance is approximate, rather than a benchmark of each variant.

For the local LLM demonstration, the speaker submits a photo of a yellow McLaren. The model describes the scene and returns a license plate reading. He then starts Ollama on a rented server with an RTX 5090 and runs the 31B model in the terminal. That section uses remote hardware; the speaker's claims about cost and privacy are not independently established. The tutorial ends with commands for listing and removing installed models.