Player not loading? Watch on YouTube
The video introduces Gemma 4 and compares its 31B dense model with Gemma 3 27B before walking through an Ollama and Open WebUI setup. The speaker describes four variants: E2B, E4B, 31B, and a 26B mixture of experts model with roughly 4B active parameters. The architecture comparison covers attention heads, normalization and a reported 256K-token context window for the larger models.
The walkthrough uses a rented RunPod GPU because the speaker cannot demonstrate the 31B model on his PC. He describes a download of about 20 GB and says GPU memory must accommodate it plus inference overhead; the demonstrated pod has 48 GB VRAM. On Linux, he installs Zstandard and Ollama, starts Ollama in the background, then installs Open WebUI and connects it through the Ollama base URL. Open WebUI serves on port 8005. After creating an admin account, he downloads Gemma 4 31B and refreshes the interface to select it.
Tests include an options-trading explanation, thumbnail descriptions and locating an eagle in a logo. The speaker also reports successful OCR and speech-to-text tests, and says smaller variants accept video with audio while larger ones accept video without audio. These are reported capabilities, not a guarantee for every runtime. For viewers planning to run models locally, the hardware discussion is useful, though the demonstrated self-hosted deployment runs on a remote server.