Player not loading? Watch on YouTube
This tutorial connects Claude Code to a local LLM in LM Studio through the free-claude-code proxy. The speaker demonstrates the setup on Windows with PowerShell after reaching a Claude subscription limit. The model handles inference locally; using the Claude Code interface does not mean the setup runs Anthropic's Claude models.
The walkthrough starts with downloading and loading a GGUF model in LM Studio. The speaker identifies it as Qwen 3.5 9B and uses an 8GB GPU with 16GB RAM. They choose a 70K context length and adjust GPU offload to fit available memory. Although the speaker suggests about 8GB RAM and says a GPU is optional, this demonstration does not establish those requirements for every model or context setting.
LM Studio serves requests on localhost port 1234. The speaker clones the proxy repository, installs uv and Python 3.14, then configures the provider, model name and base URL with the /v1 suffix. The proxy runs on port 8082, which becomes Claude Code's endpoint.
The final steps install Claude Code and select the LM Studio model through /model. The speaker warns that choosing the default model would use their subscription credits. A prompt then appears in LM Studio's processing view and returns an answer. Response speed is slow in this example, so the video is primarily a setup guide for a local coding assistant.