Gemma 4 in LM Studio: setup and 12B vs E4B

Learn to run Gemma 4 in LM Studio on a 16 GB Ryzen 5600G PC, adjust thinking mode, and compare 4-bit 12B and E4B models.

Player not loading? Watch on YouTube

This tutorial covers Gemma 4 setup in LM Studio on Windows, using a Ryzen 5600G PC with 16 GB of RAM and integrated graphics. The presenter also notes that LM Studio is available for Mac and Linux. He downloads the 7.15 GB Gemma 4 12B QAT model through the model search interface and starts a chat.

Model selection depends on available memory. The presenter recommends allowing room beyond the base weights for the operating system and context window. He explains quantization-aware training and suggests trying a post-training quantized version if QAT causes problems with text analysis or coding.

Turning off thinking mode reduces the delay before simple answers in his tests. With the computer disconnected from the internet, he demonstrates offline writing and French translation, though he does not verify the translation's accuracy. Other examples include image identification, a roughly 200-word summary of a 3,900-word chapter, and a Python BMI calculator that he runs in an IDE. The demonstrated context setting is 8192 tokens; he cautions that increasing it uses more memory and reduces performance.

The final comparison tests a 4-bit E4B model against 12B. E4B reaches 11.75 tokens per second on a greeting, versus 6.0 for the earlier 12B run, but fails to identify an animal that 12B recognized. These examples show the speed and accuracy tradeoff on this particular local LLM setup.