Favicon of Augmentoolkit

Augmentoolkit

An open-source LLM training toolkit that turns documents into specialist datasets, with offline generation on macOS and Linux and optional cloud compute.

Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.

Dataset generation can run offline on consumer hardware without an external API key, using a model built for that task. It also works with DeepSeek and Llama, and can use services such as DeepInfra for faster generation. Local generation keeps that work on your computer; API generation sends it to an external service. Training needs a powerful machine or rented compute, which the toolkit can arrange automatically.

Its factual training workflow combines question answering, ways of expressing the same information, mistake correction and examples that teach a model to acknowledge gaps in its knowledge. It balances subject-specific material with general training data and can carry the process through training, model download and preparation for local inference. You can stop at datasets.

Other workflows create multi-turn roleplaying data or build text classifiers from unlabelled material. It also prepares datasets for retrieval-augmented generation and can serve the resulting model. A graphical interface and CLI support both approaches, interrupted runs resume automatically, and developers can add their own generation pipelines in Python. Custom models can also run as Discord bots with the hosting code on your computer.

Similar to Augmentoolkit