Unsloth and Outlines: fine-tune a tiny local model

Learn synthetic data generation, fine-tuning and deployment for a 135M model, using Outlines, Unsloth and 300,000 research-paper training examples.

Player not loading? Watch on YouTube

This tutorial follows supervised fine-tuning of a 135-million-parameter SmolLM model for narrow research-paper tasks. The workflow turns passages from 1,500 papers into synthetic instructions, trains the model with Unsloth, evaluates its answers and packages it behind a task-specific API.

For data generation, the presenter uses a 4-bit Qwen 3.5 4B model through MLX LM on a MacBook. He names llama.cpp, vLLM and Transformers as alternatives for Windows or Linux. Outlines constrains decoding to produce structured responses. The distinction matters: a valid schema does not guarantee a correct answer. Generated question-answer pairs become Alpaca-format examples, alongside summaries, knowledge-graph triplets and passage comparisons.

The training walkthrough covers chat templates, LoRA and loss on assistant responses only. The presenter reports generating 300,000 rows locally and evaluates 500 examples with GPT-5.4 Nano as the judge. These results concern the demonstrated tasks, rather than general chat ability. He warns that small prompt changes reduce performance because training used fixed instructions.

The deployment harness preserves those prompts and asks users for the relevant passage and task inputs. Neural TXT packages this local LLM workflow as a library. The presenter reports up to 350 tokens per second after 8-bit quantization on his two-year-old MacBook; that speed is specific to his demonstration.