Player not loading? Watch on YouTube
Kyle walks through a local LLM coding setup built around LM Studio. He explains model parameters, context length and quantization, then uses GPU memory estimates to choose and load models. The hardware discussion covers dedicated VRAM on Windows and unified memory on newer Macs. His examples use a GPU with 16 GB of memory; this is his setup, not a stated minimum requirement.
The configuration section shows how GPU offload and context length affect memory use. Kyle reports roughly 124 tokens per second with GPT-OSS-20B fitting in GPU memory, compared with about 24 after increasing context and spilling into system memory. He also experiments with moving mixture-of-experts weights to the CPU. These results describe his hardware and settings.
LM Studio's server connects the coding assistant to other applications through OpenAI-compatible endpoints. In VS Code, Kyle configures Continue with Qwen 2.5 Coder 1.5B for autocomplete and a separate model for agent tasks. He covers model IDs, the /v1 API base, tool capabilities and logs for troubleshooting delayed suggestions.
The Copilot setup uses VS Code Insiders at recording time. Kyle says that route still requires internet access, while his Pi terminal workflow can operate offline. He closes with Qwen 3.6 and Claude Sonnet 4.6 examples: both Sudoku apps took about nine minutes, while an identical bug fix took roughly 2.5 minutes locally versus 45 seconds with Sonnet.