Gemma 4, Ollama and Continue: VS Code setup guide

Learn to connect Gemma 4 through Ollama to Continue in VS Code on Windows, then generate a to-do app and stopwatch with automatic file editing.

Player not loading? Watch on YouTube

This tutorial sets up Gemma 4 as a local coding assistant in VS Code using Ollama and Continue. The presenter works on Windows with an HP gaming laptop containing 24 GB of RAM. He recommends 16 to 24 GB of RAM, a four-core CPU and an NVIDIA RTX GPU with at least 4 GB of VRAM for partial GPU offloading. These are his suggested requirements for the demonstration, rather than verified requirements for every model size.

The walkthrough installs Ollama through PowerShell, checks its version and confirms that its local service responds in a browser. The presenter selects what he calls the four-billion model, downloads it and tests a Rust prompt in the terminal. He reports roughly 7 GB of RAM usage when loading it. His Ollama session shows a default 4K context; he says it can increase to 64K but keeps the default for the examples.

In VS Code, he installs Continue, connects it to Ollama and sets every built-in tool to automatic so the AI agent can create and modify workspace files. After restarting the editor and selecting Gemma 4, he requests an HTML to-do app. He reports about three minutes to generate its HTML, CSS and JavaScript files. He then deletes those files before requesting a stopwatch, which he says takes about two minutes. The examples illustrate his local AI workflow; they do not establish performance on other hardware.