LM Studio setup: Gemma 4, Obsidian and Claude Code

Learn to load Gemma 4 E2B in LM Studio, chat with files and images, and connect Obsidian and Claude Code on a MacBook Pro with 16GB RAM.

Player not loading? Watch on YouTube

This tutorial explains how to run models locally with LM Studio and use its server endpoint in other applications. The presenter downloads Gemma 4 E2B, describes it as using about 4GB of memory, and demonstrates it on a MacBook Pro with 16GB of unified memory. He explains the difference between storing a model on disk and loading it into RAM. Increasing the context length also raises memory use.

The chat examples cover image descriptions and questions about uploaded documents through RAG. The upload interface shown allows up to five files at a time and displays a 30MB limit. Model settings include context length and a thinking mode that the presenter disables when he wants a quicker answer.

For Obsidian, the walkthrough uses the Copilot plugin, an exact model identifier, and the LM Studio server URL with /v1 appended. The presenter then adds Firecrawl to mcp.json, supplies an API key, enables the tool, and approves a website scraping request. This step uses a website and an external service; local inference alone does not establish that the whole workflow is offline.

The final section points Claude Code at LM Studio through environment variables and selects the local LLM. The demonstrated response takes longer than the direct chat example. The presenter suggests that the project's prompts, skills and MCP configuration may account for the delay.