Player not loading? Watch on YouTube
This tutorial explains LoRA and QLoRA, then applies them to Qwen 2.5 7B using Sherlock Holmes prose. The speaker trains on a 12GB RTX 3060 and compares the result with the base 7B Instruct model and a fully fine-tuned 0.5B Qwen model from the previous episode.
LoRA freezes the base weights and trains two smaller matrices to approximate weight updates. The explanation covers rank, alpha, target layers and initialization: B starts at zero so the adapter initially leaves the model's behavior unchanged. QLoRA adds a frozen 4-bit base, with discussion of NF4, double quantization and paged optimizers.
The code walkthrough uses Hugging Face PEFT and bitsandbytes. Its adapter configuration targets all linear layers with rank 16, alpha 32 and dropout 0.05. Training uses gradient checkpointing and eight accumulation passes per optimizer step. The speaker reports 5.18GB of VRAM after loading the quantized base, 7.36GB after attaching adapters and usage slightly below 8GB during early training. The saved adapter is 165MB.
The comparison is a small prose demonstration, not a general benchmark for a local LLM. The speaker favors the adapted 7B model's Holmes style, but notes repetitive details and other quirks. Validation loss changes little in the second epoch, which he flags as possible overfitting.