nanoGPT is a Python toolkit for developers and researchers who want to train GPT models on their own hardware or fine-tune existing GPT-2 checkpoints. Its author has deprecated the project and points readers to nanochat. The MIT-licensed code remains available for study and modification.
A rewrite of minGPT, it puts practical training at the center of a small, readable PyTorch codebase. The training loop and model definition are short enough to inspect and adapt, which matters if you want to change the model itself rather than work through a larger training framework.
It supports training from scratch, fine-tuning on your own text, and generating text from saved models. You can also load OpenAI's GPT-2 checkpoints, including GPT-2 Medium, Large and XL, through Hugging Face Transformers. Included examples cover character-level Shakespeare models and GPT-2 training on OpenWebText. Benchmarking and profiling tools help assess training performance.
Hardware needs depend on the model. Small experiments can run on a CPU, including a MacBook, while Apple Silicon GPUs can accelerate training through PyTorch's MPS backend. NVIDIA GPUs support larger runs, and distributed training can span multiple GPUs or server nodes. Reproducing the included GPT-2 training result requires an eight-GPU A100 server with 40 GB per GPU.
Training and text generation run on the hardware you choose, with checkpoints saved there. Dataset and pretrained-weight downloads use external services, and Weights & Biases logging is optional.
The documented Windows workaround disables torch.compile with --compile=False.
Claim this page and we'll verify you by hand. nanoGPT gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find nanoGPT?Promote it
Something wrong or outdated on this page?
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
21.7KUpdated 1 day agoApache-2.0
macOS#Distributed execution#Hugging Face integration#LoRA
58.3KUpdated 3 months agoMIT
macOS#Code execution#Tool calling
4.6KUpdated 1 week agoApache-2.0
Web#Hugging Face integration#Multilingual
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
1.1KUpdated 2 years agoMIT
#Hugging Face integration#LoRA#Quantization
PEFT is an open-source Python library for developers who want to adapt pretrained models on their own hardware with less compute and storage than full fine-tuning requires. It trains a small subset of parameters, often through adapters, while leaving the base model intact. It's licensed under Apache 2.0.
nanochat is an MIT-licensed toolkit for training your own LLM and chatting with it on hardware you control. It's aimed at researchers and developers who want to study or modify the full training pipeline, with a small Python codebase built on PyTorch.
AutoTrain trains custom machine learning models from your own data through a no-code interface. It's for people who need to fine-tune an LLM or build a classifier without writing a training pipeline. The local AutoTrain Advanced project is no longer maintained, so it won't receive bug fixes or new features.
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.