Favicon of nanochat

nanochat

Open-source LLM training toolkit built on PyTorch. Train and chat with your own models on NVIDIA GPUs, with smaller CPU and Apple Silicon examples.

nanochat is an MIT-licensed toolkit for training your own LLM and chatting with it on hardware you control. It's aimed at researchers and developers who want to study or modify the full training pipeline, with a small Python codebase built on PyTorch.

The pipeline covers tokenizer training, pretraining, supervised finetuning, reinforcement learning, evaluation and inference. A command-line chat interface lets you talk to the resulting model, and the model can execute Python code as a tool. Evaluation includes science, math and coding tasks, alongside benchmarks for inference speed and memory use.

Its distinctive approach is to tie model size and training settings to transformer depth. Choosing a smaller or larger model automatically adjusts its width, attention heads and training hyperparameters. That makes it useful for comparing models at different scales without managing a large collection of separate settings. The repository includes experiments for studying scaling laws and measuring training time against GPT-2 capability on the DCLM CORE benchmark.

Training runs on a single GPU node, either on your own hardware or a rented server. The reference setup uses NVIDIA H100 GPUs, and A100 GPUs also work. A single GPU can run the pipeline more slowly; GPUs with less memory need adjusted batch sizes. CPU and Apple Silicon examples train much smaller models with limited capabilities. Research runs can use Weights & Biases to track training metrics.

Similar to nanochat