
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
You can fully fine-tune models or use LoRA and QLoRA to reduce training memory needs. Supported families include GPT-OSS, Llama, Mistral, Mixtral, Qwen, and Gemma, alongside models available through Hugging Face Transformers. Training also covers vision-language and audio models such as Qwen2-VL, Pixtral, LLaVA, and Voxtral, with support for image, video, and audio data.
Beyond fine-tuning, Axolotl supports preference training with DPO, reinforcement learning with GRPO, and reward modelling. A shared configuration covers dataset preparation, training, evaluation, quantization, and inference. For larger workloads, it supports multiple GPUs and multiple machines through FSDP, DeepSpeed, Torchrun, and Ray. Attention optimizations and sequence packing help make better use of training hardware.
Axolotl requires an NVIDIA or AMD GPU. It accepts local datasets as well as data from Hugging Face and cloud storage, so training data doesn't have to go to an external AI service. Telemetry is enabled by default and can be disabled; it collects basic system information, model types, and error rates, excluding personal data and file paths.
Claim this page with an email at axolotl.ai. Axolotl gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Axolotl?Promote it
Something wrong or outdated on this page?
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
5.8KUpdated 5 months agoBSD-3-Clause
#Hugging Face integration#LoRA#Quantization
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
18KUpdated 1 day ago
Docker#Distributed execution#Hugging Face integration#Quantization
10.1KUpdated 2 weeks agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
torchtune is a Python library for developers and researchers who want to adapt LLMs on their own GPU hardware using PyTorch. Its editable training recipes suit work that needs control over the training code and model implementations. The project is no longer actively maintained.
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.
Megatron-LM is a Python framework for research teams training large language models on NVIDIA GPU infrastructure. It pairs ready-made training scripts with Megatron Core, a library developers can use to build their own training systems. Its focus is distributed training, with benchmarks on H100 clusters spanning thousands of GPUs.
OpenRLHF is a self-hosted Python framework for researchers and teams training language models with human feedback or custom rewards. It runs on your own NVIDIA GPU hardware, with Docker support and distributed training across servers. It's open source under Apache 2.0.
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.