Transformers tutorial: tokenizers and Qwen 2.5 models

Learn to load Qwen 2.5 3B, compare its tokenizer with GPT-2, and test base versus instruct outputs on an RTX 3060 with 12 GB of memory.

Player not loading? Watch on YouTube

This tutorial loads pretrained models with the Hugging Face Transformers library. It introduces model names, parameter counts, base and instruct variants, and quantization before working through Python scripts. The hardware examples use the creator's RTX 3060 with 12 GB; weight-size estimates are rough examples rather than complete runtime memory requirements.

The practical section uses AutoTokenizer and from_pretrained to compare Qwen 2.5 and GPT-2 tokenizers on ordinary text, cryptocurrency terms and code. It inspects token IDs, attention masks, special tokens and Qwen's chat template. Matching the tokenizer to its model keeps the token IDs aligned with the intended embeddings.

The creator then loads Qwen 2.5 3B, prints its configuration and parameter counts, and relates its attention and feed-forward layers to the smaller transformer built in an earlier episode. The same generation prompts produce coherent text where the mini GPT produced gibberish.

A Bitcoin question shows the difference between base and instruct outputs: the base model continues with another question, while the instruct model answers it. A general answer about Kalshi sets up a future fine-tuning lesson. Fine-tuning and RAG are discussed as later topics.