Claude Code with NVIDIA NIM: LiteLLM Proxy setup

Learn to route Claude Code through LiteLLM Proxy to NVIDIA NIM, configure localhost port 4000, and check requests in the proxy logs.

Player not loading? Watch on YouTube

This tutorial connects Claude Code to NVIDIA NIM through LiteLLM Proxy. The speaker explains that the proxy translates Claude Code's requests into the provider's expected format. The proxy runs on the user's machine, while NVIDIA's servers handle model inference. This setup does not run models locally or provide offline use.

The walkthrough installs LiteLLM with pip, creates config.yaml, and starts the server on localhost port 4000. The configuration maps a Claude model alias to an NVIDIA NIM model and separates the proxy's master key from the NVIDIA API key. The supplied example maps claude-3-5-sonnet-20240620 to nvidia_nim/meta/llama-3.1-405b-instruct; the alias does not mean Claude performs the inference. The speaker also notes that credentials can go in environment variables.

Claude Code's settings.json points to http://127.0.0.1:4000 and uses the proxy master key for authentication. Installation guidance covers Windows, macOS, Linux and WSL. The speaker tests the connection with a question, checks the proxy logs, then asks the coding assistant to create app.js.

The speaker reports using the API without payment or a credit card in this demonstration. They warn that busy models can queue requests and recommend checking response times before choosing one. For other providers, the tutorial advises checking LiteLLM's documentation for the correct model prefix.