ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
Training can update all model parameters or use methods such as LoRA and QLoRA to reduce resource needs. Beyond instruction fine-tuning, it supports pre-training, preference alignment with DPO, and reinforcement learning with GRPO. You can also train embedding and reranker models, use custom datasets, and mix text, images, video and audio in multimodal training. Agent training includes multi-turn dialogue and tool calling.
Hardware support covers NVIDIA GPUs through CUDA, AMD GPUs through ROCm, Apple MPS and CPUs, as well as Ascend NPUs and MetaX GPUs. For larger workloads, DeepSpeed, FSDP and Megatron support distributed training across machines, including mixture-of-experts models.
A Gradio web interface covers training, inference, evaluation and quantization alongside Python and command-line access. ModelScope and Hugging Face supply model and dataset downloads. For serving, ms-swift works with Transformers, vLLM, SGLang and LMDeploy and exposes OpenAI-compatible interfaces. EvalScope handles model evaluation, while quantization export supports AWQ, GPTQ, FP8 and BNB.
Claim this page and we'll verify you by hand. ms-swift gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find ms-swift?Promote it
Something wrong or outdated on this page?
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
5.8KUpdated 5 months agoBSD-3-Clause
#Hugging Face integration#LoRA#Quantization
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
58.3KUpdated 3 months agoMIT
macOS#Code execution#Tool calling
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.
torchtune is a Python library for developers and researchers who want to adapt LLMs on their own GPU hardware using PyTorch. Its editable training recipes suit work that needs control over the training code and model implementations. The project is no longer actively maintained.
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
nanochat is an MIT-licensed toolkit for training your own LLM and chatting with it on hardware you control. It's aimed at researchers and developers who want to study or modify the full training pipeline, with a small Python codebase built on PyTorch.