Gemma 4: Ollama and Open WebUI setup and Gemma 3 comparison

Learn to serve Gemma 4 31B with Ollama and Open WebUI on RunPod, expose port 8005, and test text prompts and image inputs.

Player not loading? Watch on YouTube

The video introduces Gemma 4 and compares its 31B dense model with Gemma 3 27B before walking through an Ollama and Open WebUI setup. The speaker describes four variants: E2B, E4B, 31B, and a 26B mixture of experts model with roughly 4B active parameters. The architecture comparison covers attention heads, normalization and a reported 256K-token context window for the larger models.

The walkthrough uses a rented RunPod GPU because the speaker cannot demonstrate the 31B model on his PC. He describes a download of about 20 GB and says GPU memory must accommodate it plus inference overhead; the demonstrated pod has 48 GB VRAM. On Linux, he installs Zstandard and Ollama, starts Ollama in the background, then installs Open WebUI and connects it through the Ollama base URL. Open WebUI serves on port 8005. After creating an admin account, he downloads Gemma 4 31B and refreshes the interface to select it.

Tests include an options-trading explanation, thumbnail descriptions and locating an eagle in a logo. The speaker also reports successful OCR and speech-to-text tests, and says smaller variants accept video with audio while larger ones accept video without audio. These are reported capabilities, not a guarantee for every runtime. For viewers planning to run models locally, the hardware discussion is useful, though the demonstrated self-hosted deployment runs on a remote server.