Player not loading? Watch on YouTube
This tutorial demonstrates two ways to connect Claude Desktop to other models through its developer-mode third-party inference settings. The Windows walkthrough uses Bifrost with OpenRouter for cloud inference, then connects directly to LM Studio to run models locally. The speaker demonstrates both routes without signing into an Anthropic account.
For the cloud setup, the presenter starts Bifrost with NPX, adds an OpenRouter API key, and configures a routing rule targeting Nvidia Nemotron's free tier. The tutorial stresses the model name's :free suffix and notes rate limits. It also warns that the selected provider logs prompts and outputs for model improvement. Bifrost's request logs show which model actually responds, even when Claude Desktop displays an Anthropic model label. That label does not establish that Claude generated the answer.
The local setup serves a Google 4-billion-parameter model through LM Studio on port 1234, using a Claude-style alias for discovery. A request exceeds 36,000 tokens, too large for the initial 8,000-token context. The presenter reloads with a 40,000-token context and reports substantial memory use and slow processing on a CPU-only machine. The demonstration retains Cowork and Code features, but the speaker cautions that reasoning and code quality depend on the replacement model.