Player not loading? Watch on YouTube
This tutorial connects Gemma 4 to Roo Code in VS Code through LM Studio's local server. The speaker presents it as a possible replacement for his GitHub Copilot workflow, though he later clarifies that his subscription is still active. The demonstration covers chat and a local coding assistant rather than establishing equivalent performance to cloud services.
The Windows setup uses a 3060 GPU with 12 GB of dedicated memory and an i7 CPU. The downloaded model is about 5.89 GB. These are the demonstration's specifications, not minimum requirements. The speaker reports that the 31B model ran slowly on his machine and suggests trying 2B on weaker hardware.
After loading the model, he starts LM Studio's server and selects Roo Code's OpenAI-compatible provider. He enters the server's base URL, a dummy API key and the model identifier. An initial request fails because the prompt needs 8,539 tokens while the loaded context is 4,096. Increasing the context and reloading lets the request proceed. He also reports trouble with Continue's release version and better results with its prerelease.
The AI agent generates files for a FastAPI product-management project, but the video ends without confirming that the application runs. The speaker notes that downloaded weights do not update themselves; adding recent information through tools or fine-tuning requires further work.