Favicon of Nanotron

Nanotron

Open-source LLM pretraining library built on PyTorch for custom datasets, with distributed NVIDIA GPU training. Licensed under Apache 2.0.

Nanotron is a Python library for researchers and developers who want to pretrain language models on their own datasets and GPU infrastructure. Built on PyTorch, it supports NVIDIA CUDA GPUs and training across multiple servers with Slurm. It's open source under Apache 2.0.

Its main focus is distributing training across GPUs. Data, tensor and pipeline parallelism let teams divide both the training workload and the model, while expert parallelism supports mixture-of-experts models. Explicit interfaces for tensor and pipeline parallelism make distributed training easier to inspect and debug.

Nanotron includes Llama training examples, plus examples for Mamba and mixture-of-experts architectures. Custom data loaders and Datatrove integration let teams use their own data pipelines. Spectral µTransfer supports scaling up neural networks, and DoReMi examples cover an approach to improving training efficiency through data selection.

For larger training jobs, it supports the ZeRO-1 optimizer, parameter sharding and custom module checkpointing, alongside FP32 gradient accumulation. Teams can generate text from saved checkpoints, including across multiple GPUs. Checkpoints can stay on the training infrastructure or upload automatically to S3; Hugging Face Hub integration handles model and dataset transfers, and Weights & Biases provides external experiment tracking.

CUDA event timing measures GPU performance. Published benchmarks and the Ultrascale Playbook cover training configurations, memory use and scaling across model sizes and node counts.

Similar to Nanotron