Favicon of nanoGPT

nanoGPT

An open-source local LLM training toolkit in Python for training GPTs from scratch or fine-tuning GPT-2, with CPU, NVIDIA and Apple Silicon support.

nanoGPT is a Python toolkit for developers and researchers who want to train GPT models on their own hardware or fine-tune existing GPT-2 checkpoints. Its author has deprecated the project and points readers to nanochat. The MIT-licensed code remains available for study and modification.

A rewrite of minGPT, it puts practical training at the center of a small, readable PyTorch codebase. The training loop and model definition are short enough to inspect and adapt, which matters if you want to change the model itself rather than work through a larger training framework.

It supports training from scratch, fine-tuning on your own text, and generating text from saved models. You can also load OpenAI's GPT-2 checkpoints, including GPT-2 Medium, Large and XL, through Hugging Face Transformers. Included examples cover character-level Shakespeare models and GPT-2 training on OpenWebText. Benchmarking and profiling tools help assess training performance.

Hardware needs depend on the model. Small experiments can run on a CPU, including a MacBook, while Apple Silicon GPUs can accelerate training through PyTorch's MPS backend. NVIDIA GPUs support larger runs, and distributed training can span multiple GPUs or server nodes. Reproducing the included GPT-2 training result requires an eight-GPU A100 server with 40 GB per GPU.

Training and text generation run on the hardware you choose, with checkpoints saved there. Dataset and pretrained-weight downloads use external services, and Weights & Biases logging is optional.

The documented Windows workaround disables torch.compile with --compile=False.

Similar to nanoGPT