AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
Its model support covers FLUX, SDXL and Stable Diffusion 1.5, alongside Qwen-Image and image-editing models such as FLUX.1-Kontext-dev. Video training includes Wan, LTX and MiniMax-H3, with support for reference-image-to-video training. It also supports audio models such as ACE-Step and the multimodal Qwen2.5-Omni model.
You can train through a browser interface or the command line. The web UI lets you start, stop and monitor jobs, and training continues when the interface isn't running. Token authentication can restrict access to a server's UI. Checkpoints let interrupted training resume rather than start over.
For fine-tuning, it supports LoRA and LoKr adapters, including control over which model layers train. Its image loader handles different aspect ratios and resizes images for batching, so datasets don't need uniform crops. Training outputs include checkpoints and sample images.
Local training writes outputs to your machine. GPU rental through Ostris Cloud, RunPod or Modal moves training to cloud hardware; Modal stores models and samples in cloud storage. Access to FLUX.1-dev through Hugging Face requires an account, an access token and model approval.
Claim this page and we'll verify you by hand. AI Toolkit (ostris) gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find AI Toolkit (ostris)?Promote it
Something wrong or outdated on this page?
21.7KUpdated 1 day agoApache-2.0
macOS#Distributed execution#Hugging Face integration#LoRA
PEFT is an open-source Python library for developers who want to adapt pretrained models on their own hardware with less compute and storage than full fine-tuning requires. It trains a small subset of parameters, often through adapters, while leaving the base model intact. It's licensed under Apache 2.0.
2.9KUpdated 19 hours agoAGPL-3.0
macOS · Docker · Web#ControlNet#Distributed execution#Human approval
2KUpdated 2 days agoGPL-3.0
Windows · Linux#Distributed execution#LoRA#Quantization
63.5KUpdated 11 months agoMIT
macOS · Windows#Hugging Face integration
nanoGPT is a Python toolkit for developers and researchers who want to train GPT models on their own hardware or fine-tune existing GPT-2 checkpoints. Its author has deprecated the project and points readers to nanochat. The MIT-licensed code remains available for study and modification.
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
75.2KUpdated 2 days agoApache-2.0
Web#LoRA#Multimodal input#OpenAI-compatible API
SimpleTuner is an open-source toolkit for fine-tuning image, video and audio generation models on your own hardware or GPU servers. It's for creators and researchers adapting models to their datasets, and teams sharing training infrastructure. A web dashboard manages training jobs.
diffusion-pipe is a local diffusion model training tool for people fine-tuning image and video models on their own GPU hardware. Its main distinction is that it can divide a model across several GPUs when it won't fit on one, while also distributing training work across GPUs. The Python project is open source under GPL-3.0 and uses DeepSpeed.
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
LLaMA-Factory is an open-source framework for developers and researchers who want to adapt language and multimodal models on their own hardware. It brings training and inference into one toolkit, with support for LLaMA, Qwen3, Qwen3-VL, DeepSeek, Gemma, Mistral and LLaVA. Its license is Apache 2.0.