Player not loading? Watch on YouTube
This tutorial walks through Gemma 4 setup in LM Studio, then connects the model to Continue in VS Code. The speaker uses Windows 11 on an HP laptop with 24 GB of RAM and an RTX 3050 with 4 GB of VRAM. LM Studio also supports macOS and Linux, according to the tutorial.
The hardware guidance recommends efficient 2B or 4B models for 8 GB RAM, 26B for 16 to 24 GB, and 31B dense for more than 24 GB. These are the speaker's recommendations, rather than guaranteed requirements. The walkthrough selects Q4 downloads and disables GPU acceleration because the laptop lacks enough VRAM for full acceleration of larger models.
The 4B model uses about 7 GB RAM in this demonstration. Although the speaker lists context limits of 128K for 2B and 4B and 256K for 26B and 31B, both tested models use a 32K setting. Tests cover HTML generation, reasoning and math. The speaker reports roughly 11 tokens per second for 26B. An image test identifies a Raspberry Pi but mistakes a Pi 5 for a Pi 4.
For a local coding assistant, the tutorial enables LM Studio's server, installs Continue and selects LM Studio as its provider. The speaker sets Continue's tool options to automatic before requesting a to-do app. Generating its HTML, CSS and JavaScript takes about 20 minutes on the demonstrated hardware.