Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
Dataset generation can run offline on consumer hardware without an external API key, using a model built for that task. It also works with DeepSeek and Llama, and can use services such as DeepInfra for faster generation. Local generation keeps that work on your computer; API generation sends it to an external service. Training needs a powerful machine or rented compute, which the toolkit can arrange automatically.
Its factual training workflow combines question answering, ways of expressing the same information, mistake correction and examples that teach a model to acknowledge gaps in its knowledge. It balances subject-specific material with general training data and can carry the process through training, model download and preparation for local inference. You can stop at datasets.
Other workflows create multi-turn roleplaying data or build text classifiers from unlabelled material. It also prepares datasets for retrieval-augmented generation and can serve the resulting model. A graphical interface and CLI support both approaches, interrupted runs resume automatically, and developers can add their own generation pipelines in Python. Custom models can also run as Discord bots with the hosting code on your computer.
Claim this page and we'll verify you by hand. Augmentoolkit gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Augmentoolkit?Promote it
Something wrong or outdated on this page?
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
10.6KUpdated 2 years agoMIT
macOS · Windows · Linux · Docker#Hugging Face integration
Petals lets developers and researchers use large language models that won't fit on a single consumer GPU by sharing the work across a network of machines. It supports text generation and fine-tuning from a desktop computer or Google Colab. Each participant holds part of the model, while other computers handle the remaining parts.
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
25KUpdated 2 years agoApache-2.0
macOS · Web#LoRA#Multimodal input#Quantization
7.4KUpdated 12 hours agoApache-2.0
Linux#Agent Skills#OpenAI-compatible API#Prompt versioning
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
LLaVA is a family of vision-language models for researchers and developers who want to ask questions about images on their own hardware. It pairs a CLIP vision encoder with a language model to support image descriptions, visual reasoning and reading text in pictures. Its Python code is open source under Apache 2.0; the project places research-use restrictions on its data and checkpoints, with additional terms from the underlying models.
Reef is self-hosted infrastructure for developers who want AI agents to improve through feedback on actual interactions. It connects inference and learning with versioned deployment, so an agent can update its model weights or its prompts, rules, and skills while continuing to serve requests. It's open source under Apache 2.0.