Favicon of Megatron-LM

Megatron-LM

LLM training framework with ready-made research scripts, NVIDIA GPU parallelism, and Hugging Face checkpoint conversion through Megatron Bridge.

Megatron-LM is a Python framework for research teams training large language models on NVIDIA GPU infrastructure. It pairs ready-made training scripts with Megatron Core, a library developers can use to build their own training systems. Its focus is distributed training, with benchmarks on H100 clusters spanning thousands of GPUs.

The two components suit different needs. Researchers can use Megatron-LM's training examples to study model architectures and distributed training. Framework developers and ML engineers can use Core's transformer components within custom pipelines, rather than adopting the whole reference training setup.

Core supports tensor, pipeline, data, expert, and context parallelism, so teams can distribute models and training workloads across GPUs in several ways. Mixed precision includes FP16, BF16, FP8, and FP4. The training pipeline also provides checkpointing and fault tolerance, while communication optimizations overlap GPU computation with data transfers.

Model support includes mixture-of-experts architectures, with training configurations for DeepSeek-V3, Mixtral, and Qwen3. Falcon-H1 combines transformer and Mamba layers, and Core also supports BitNet ternary quantization. Megatron Bridge converts checkpoints in both directions between Hugging Face and Megatron and supplies model training recipes.

The repository also includes reinforcement learning code with RLHF, inference engines and a server, plus post-training tools for quantization, distillation, and pruning. Models can be exported to TensorRT-LLM.

Similar to Megatron-LM