Player not loading? Watch on YouTube
Nick demonstrates how he connects Claude Code to a local LLM through LM Studio's API server. The setup covers terminal installation, a base URL environment variable pointing to the server, an API access token, and a model name passed when launching Claude Code. He uses local models alongside paid services, including handing implementation tasks to them after Claude helps prepare a plan.
The coding assistant first uses Gemma 4 with 7.5 billion parameters on a MacBook with an M4 Pro and 24 GB of RAM. In the demonstration, it creates a basic HTML to-do page in about 90 seconds. Adding an input field takes 2 minutes and 18 seconds, but the result fails in the browser. Repeated repair attempts leave the error unresolved; Nick manually adds a missing HTML block, after which adding tasks works but the appearance remains poor.
He then connects to Gemma 4's 26B variant through LM Studio on another machine on his local network. That desktop has an AMD Ryzen 7, 128 GB of RAM, and an Nvidia GeForce RTX 4060 Ti with 16 GB of video RAM. The larger model takes about three minutes for the initial page and eight minutes to add working task entry. It also completes the task completion toggle. These are results from a small demonstration, which Nick explicitly distinguishes from a broad model evaluation. He explains that total task time includes multiple model requests and iterations.