Player not loading? Watch on YouTube
This Chinese-language tutorial shows a Windows setup that connects the desktop client the speaker calls Claude Code to a local LLM through Ollama and CC Switch. Ollama runs the models, while CC Switch forwards the client's requests to them. The speaker presents this as a way to use the coding assistant offline without Claude's official API.
The walkthrough starts with Git and the three applications. It then covers model downloads, a custom provider in CC Switch, and the Ollama endpoint at 127.0.0.1:11434/v1. The configuration uses OpenAI compatibility mode and a placeholder API key. Later steps enable local routing, Windows virtualization features and WSL, followed by developer mode in the desktop client and an exported registry configuration.
For model selection, he downloads Gemma 4 26B, roughly 17 GB, on a GPU with 24 GB of VRAM. He also installs Qwen 3.5 9B, described as a 6.6 GB download suitable for about 8 GB of VRAM; this is his hardware guidance, not a verified minimum.
The client’s Haiku and Opus labels do not identify the underlying model selected in Ollama. The final tests request a cyberpunk homepage and a 3D running game. The speaker reports nearly full GPU utilization and says Gemma's game works, while the smaller Qwen model produces faulty gameplay. These examples compare his results, without establishing general performance or equivalence to Claude Sonnet.