Unsloth and TRL: DPO tutorial with ORPO comparison

Learn to build preference pairs and train a 135M-parameter model with Unsloth and TRL, then compare DPO, ORPO and the SFT baseline.

Player not loading? Watch on YouTube

This tutorial applies direct preference optimization (DPO) to a 135-million-parameter language model previously trained with supervised fine-tuning (SFT). It focuses on passage-based tasks such as extracting bullet points and question-answer pairs. The speaker explains preference tuning as a way to favor better responses the model can already produce.

The workflow starts with response analysis: 100 test prompts yield 1,600 responses, which the speaker assesses for diversity, coherence and correctness. For training, four responses per example across 40,000 examples receive rankings from a local 4-billion-parameter judge. These rankings become roughly 150,000 preference examples with prompt, accepted and rejected fields.

Unsloth and Hugging Face TRL handle training. The explanation covers the frozen reference model, trainable policy, sequence log probabilities and beta parameter, followed by LoRA fine-tuning through PEFT and trainer metrics. ORPO provides a comparison that combines an SFT loss with a preference penalty. The speaker links the preference dataset and training code for readers who want to repeat the experiment.

In the speaker's judge-based evaluation, DPO raises the score from 3.2 to 3.3 out of 5 and wins 67% of comparisons against SFT. DPO also leads ORPO by 55% to 45%. These are results from this experiment. Response diversity falls, and free-form continuation remains error-prone: preference tuning has not supplied the model with new external knowledge.