Favicon of Distilabel

Distilabel

A Python framework for generating and evaluating LLM datasets, with Apache 2.0 licensing and integrations for Anthropic, Cohere and Argilla.

Screenshot of Distilabel website

Distilabel is an open-source Python framework for engineers building datasets to train or evaluate AI models. It pairs synthetic data generation with LLM feedback, so a pipeline can create examples and judge their quality. It uses the Apache 2.0 license.

Its tasks cover classification and information extraction as well as instruction following, dialogue generation and judging model responses. For fine-tuning work, it can produce preference data and use AI judgments to filter examples. The generation and evaluation methods draw on published research, giving teams established approaches to work from when developing their own datasets.

Distilabel brings LLM providers behind a shared API, with integrations for Anthropic and Cohere's cloud services. Those integrations send model requests to the provider. It also supports exporting generated datasets to Argilla for further review, while Outlines and Instructor integrations support structured generation.

The framework suits teams that need repeatable data pipelines and control over how examples are produced and assessed. Ray support lets them distribute pipeline work, and fault tolerance supports larger jobs. Data processing capabilities include sentence embeddings with FAISS, text clustering with UMAP and Scikit-learn, and duplicate detection with MinHash. Community datasets built with Distilabel include OpenHermesPreference and Intel Orca DPO.

The original authors have moved on. Community collaborators maintain the project, with current fixes and improvements available on its development branch.

Similar to Distilabel