
DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.
Workflows can use local models or API-based LLMs such as OpenAI's GPT-4. Local model work runs on your own hardware; calls to OpenAI use its cloud service. You can combine multiple prompting steps to create examples for a new task, augment an existing dataset, or clean data before training. Its examples include generating research abstracts and summaries, then using those pairs to train a Hugging Face T5 model.
Training covers instruction tuning, distillation, and alignment with human preferences, using existing or generated data. LoRA lets you train a subset of model parameters, while quantization provides another way to reduce resource demands. DataDreamer also supports running models and training across multiple GPUs, as well as executing workflow steps in parallel.
Caching and resumability help avoid repeating completed work when an experiment stops or needs another run. Shareable workflows support reproducibility across experiments. For publishing results, DataDreamer can send datasets and trained models to Hugging Face Hub and generate data cards, model cards, metadata, and required citation lists.
Claim this page with an email at datadreamer.dev. DataDreamer gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find DataDreamer?Promote it
Something wrong or outdated on this page?
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
10.1KUpdated 2 weeks agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
OpenRLHF is a self-hosted Python framework for researchers and teams training language models with human feedback or custom rewards. It runs on your own NVIDIA GPU hardware, with Docker support and distributed training across servers. It's open source under Apache 2.0.