Megatron-LM is a Python framework for research teams training large language models on NVIDIA GPU infrastructure. It pairs ready-made training scripts with Megatron Core, a library developers can use to build their own training systems. Its focus is distributed training, with benchmarks on H100 clusters spanning thousands of GPUs.
The two components suit different needs. Researchers can use Megatron-LM's training examples to study model architectures and distributed training. Framework developers and ML engineers can use Core's transformer components within custom pipelines, rather than adopting the whole reference training setup.
Core supports tensor, pipeline, data, expert, and context parallelism, so teams can distribute models and training workloads across GPUs in several ways. Mixed precision includes FP16, BF16, FP8, and FP4. The training pipeline also provides checkpointing and fault tolerance, while communication optimizations overlap GPU computation with data transfers.
Model support includes mixture-of-experts architectures, with training configurations for DeepSeek-V3, Mixtral, and Qwen3. Falcon-H1 combines transformer and Mamba layers, and Core also supports BitNet ternary quantization. Megatron Bridge converts checkpoints in both directions between Hugging Face and Megatron and supplies model training recipes.
The repository also includes reinforcement learning code with RLHF, inference engines and a server, plus post-training tools for quantization, distillation, and pruning. Models can be exported to TensorRT-LLM.
Claim this page and we'll verify you by hand. Megatron-LM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Megatron-LM?Promote it
Something wrong or outdated on this page?
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
1.1KUpdated 2 years agoMIT
#Hugging Face integration#LoRA#Quantization
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
3.8KUpdated 23 hours agoApache-2.0
#Hugging Face integration#Quantization
3.5KUpdated 6 days agoApache-2.0
#Hugging Face integration#ONNX#Quantization
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
LLM Compressor is an open-source Python library for developers preparing models to run on their own hardware with vLLM. It reduces model size and memory requirements through quantization, and accepts local checkpoints or models from Hugging Face repositories. It's licensed under Apache 2.0.
Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.