LM Studio and VS Code: local coding agent setup

Learn to connect LM Studio to VS Code for coding agents and configure Continue autocomplete, with model selection based on VRAM and unified memory.

Player not loading? Watch on YouTube

This tutorial sets up a local coding assistant with LM Studio and Visual Studio Code, then adds Continue for autocomplete. The speaker explains how Windows GPU VRAM and Apple M-series unified memory affect model selection. His hardware tables are rough guidelines: quantization, context length and other running applications also affect memory use and speed.

The workflow uses a small model for completions and a larger model with tool calling for chat and file edits. On an M5 Max with 64 GB of unified memory, the speaker tests a 35-billion-parameter model in Q4 quantization and a larger Coder Next model. He reports slower responses from the latter and uses the smaller option for the coding demonstration.

Setup covers downloading models, adjusting GPU offload and context length, and starting LM Studio's development server. In VS Code, the speaker adds a custom chat completions endpoint through Manage Language Models and fills in the model ID, URL and capabilities. Continue gets a separate autocomplete configuration. The intended result is offline coding after the software and models are installed.

The demonstration also shows limits. The agent creates chess-game files but does not use the requested React framework, encounters an integrated-browser error, and produces a game that fails to load. The speaker recommends smaller edits and functions, while acknowledging that results vary and fall short of stronger cloud models.