Player not loading? Watch on YouTube
This LM Studio tutorial explains how to download models, load them into memory and use them through chat or a local inference server. The speaker directs viewers to lmstudio.ai/download for the standard application rather than LM Studio Bionic.
Model selection starts with the hardware panel. The speaker recommends checking GPU VRAM on Windows and unified memory on modern Macs, then using LM Studio's model recommendations. A small LFM model provides the initial example: its Q8 download is about 1.25 GB. The guide explains weight quantization and shows how increasing context length raises memory usage. It also covers K and V cache quantization, loading multiple models and leaving memory headroom for other applications.
The chat walkthrough includes a system prompt that requests Indonesian replies, saved presets and file-question plugins. For a larger model, the speaker recommends Q4 cache settings and speculative decoding in MTP mode; those recommendations concern the demonstrated setup.
The coding assistant example connects Cline in VS Code to an OpenAI-compatible endpoint using LM Studio's base URL and model identifier. Slow prompt processing leads the speaker to reduce the context window, after which the retry returns a response. The final example generates and runs a Python script that calls localhost:1234/v1, showing another way to run models locally outside the chat interface.