Favicon of DataDreamer

DataDreamer

Open-source Python library for synthetic data generation and LLM training, with local models, API-based models, caching, and resumable workflows.

Screenshot of DataDreamer website

DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.

Workflows can use local models or API-based LLMs such as OpenAI's GPT-4. Local model work runs on your own hardware; calls to OpenAI use its cloud service. You can combine multiple prompting steps to create examples for a new task, augment an existing dataset, or clean data before training. Its examples include generating research abstracts and summaries, then using those pairs to train a Hugging Face T5 model.

Training covers instruction tuning, distillation, and alignment with human preferences, using existing or generated data. LoRA lets you train a subset of model parameters, while quantization provides another way to reduce resource demands. DataDreamer also supports running models and training across multiple GPUs, as well as executing workflow steps in parallel.

Caching and resumability help avoid repeating completed work when an experiment stops or needs another run. Shareable workflows support reproducibility across experiments. For publishing results, DataDreamer can send datasets and trained models to Hugging Face Hub and generate data cards, model cards, metadata, and required citation lists.

Similar to DataDreamer