Nanotron is a Python library for researchers and developers who want to pretrain language models on their own datasets and GPU infrastructure. Built on PyTorch, it supports NVIDIA CUDA GPUs and training across multiple servers with Slurm. It's open source under Apache 2.0.
Its main focus is distributing training across GPUs. Data, tensor and pipeline parallelism let teams divide both the training workload and the model, while expert parallelism supports mixture-of-experts models. Explicit interfaces for tensor and pipeline parallelism make distributed training easier to inspect and debug.
Nanotron includes Llama training examples, plus examples for Mamba and mixture-of-experts architectures. Custom data loaders and Datatrove integration let teams use their own data pipelines. Spectral µTransfer supports scaling up neural networks, and DoReMi examples cover an approach to improving training efficiency through data selection.
For larger training jobs, it supports the ZeRO-1 optimizer, parameter sharding and custom module checkpointing, alongside FP32 gradient accumulation. Teams can generate text from saved checkpoints, including across multiple GPUs. Checkpoints can stay on the training infrastructure or upload automatically to S3; Hugging Face Hub integration handles model and dataset transfers, and Weights & Biases provides external experiment tracking.
CUDA event timing measures GPU performance. Published benchmarks and the Ultrascale Playbook cover training configurations, memory use and scaling across model sizes and node counts.
Claim this page and we'll verify you by hand. Nanotron gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Nanotron?Promote it
Something wrong or outdated on this page?
5.8KUpdated 19 hours agoBSD-3-Clause
Linux#Distributed execution#Hugging Face integration
torchtitan is an open-source training platform for researchers and developers building generative AI models on their own GPU machines or server clusters. It uses PyTorch's distributed training tools and keeps the model code relatively simple when spreading work across GPUs. The Python codebase has extension points and replaceable components for experiments with model architectures and training infrastructure.
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
2KUpdated 2 days agoGPL-3.0
Windows · Linux#Distributed execution#LoRA#Quantization
28.6KUpdated 23 hours agoMIT
macOS · Linux#Distributed execution#LoRA
21.9KUpdated 1 day agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
3KUpdated 5 days ago
Linux · iOS#Hugging Face integration#LoRA#Quantization
TorchAO is a PyTorch library for developers who want to train or run models on their own hardware with less memory and faster computation. It reduces the precision of model weights and activations, with options for language models and image or video generation. Its PyTorch integration works with torch.compile and FSDP2 across most Hugging Face PyTorch models.
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
diffusion-pipe is a local diffusion model training tool for people fine-tuning image and video models on their own GPU hardware. Its main distinction is that it can divide a model across several GPUs when it won't fit on one, while also distributing training work across GPUs. The Python project is open source under GPL-3.0 and uses DeepSpeed.
MLX is a machine learning array framework for researchers and developers building models on their own hardware. Its distinctive feature on Apple silicon is shared CPU and GPU memory: both processors can work on the same arrays without copying data between them. It's open source under the MIT license.
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.