Player not loading? Watch on YouTube
This tutorial explains how to run models locally on Windows with llamafile and a fine-tuned Qwen 3 4B thinking model stored on a USB drive. The speaker uses llamafile 0.10.5 and a quantized GGUF file of about 2.5 GB. The demonstrated setup works offline after the initial downloads; the speaker says prompts stay on the computer.
The stated requirements are at least 8 GB of free drive space and about 4 GB of free RAM. The walkthrough recommends USB 3.0 or an external SSD. It starts by formatting the drive as exFAT, which erases its contents and avoids FAT32's 4 GB file limit. On Windows, the llamafile binary needs an .exe extension. The executable and model go in the same folder, alongside a four-line launcher saved as a .bat file.
The launcher runs from its own folder so a changed USB drive letter does not break the paths. After loading the model, the server provides a browser chat at 127.0.0.1:8080. The terminal must remain open because closing it stops the server. The linked article supplies the exact launcher text.
The speaker describes the CPU setup as useful for drafting and summarizing, but below large cloud models in capability. USB 2.0 can load slowly; moving the model to internal storage is the suggested workaround. Answers can be confidently wrong and need checking when accuracy matters.