Player not loading? Watch on YouTube
This tutorial walks through fine-tuning a local LLM with Unsloth Studio, from installation to comparing a trained model with its base version. The speaker describes Studio as a free, open source interface that runs on your machine without an account and can work offline. Installation and launch use terminal commands; the browser interface opens at localhost port 8888, and the terminal must stay running.
The example uses Llama 3.1 8B Instruct in Unsloth's BNB 4-bit variant with QLoRA. The speaker explains how adapters change responses while leaving the base weights untouched, and contrasts this with RAG's retrieval of information into a prompt. He recommends an Nvidia GPU with at least 12 GB of VRAM for Windows, or a modern M-series Mac with ample unified memory. Available memory limits model size. For this workflow, he specifies Safetensors rather than GGUF training files.
The dataset targets an Indonesian tutor that uses vocabulary the speaker already knows. He shows Claude-generated examples and a Studio recipe that turns seed vocabulary into structured training data. The walkthrough covers instruction, input and output fields, column mapping, epoch settings and training loss. He stops the live training run and uses a previously trained model for the comparison, reporting simpler replies in his examples. Export options include GGUF for use with Ollama, a full 16-bit model and a LoRA adapter.