Favicon of DataTrove

DataTrove

An open-source Python library for preparing LLM training data locally or on Slurm and Ray clusters, with filtering, deduplication and generation.

DataTrove is an open-source Python library for teams preparing large text datasets, including LLM training corpora. It runs on your own machine or on Slurm and Ray clusters, with processing steps that carry across those environments. It uses the Apache 2.0 license.

Its reusable blocks cover reading and writing data, extracting text from HTML, filtering documents, tokenization and dataset statistics. Deduplication includes MinHash and exact sentence matching. You can add custom processing logic alongside the supplied blocks, so a dataset's cleanup rules don't have to fit a fixed workflow. Supported inputs include web archives and Parquet, and storage access through fsspec covers local files, S3 and Hugging Face datasets or buckets.

The library targets workloads where memory use and recovery from failed jobs matter. Its staged processing approach keeps memory use relatively low, and parallel execution can use multiple CPU cores or machines. It tracks completed tasks. Restarting an interrupted job reruns only unfinished work, while dataset statistics record how much data each stage processed or removed.

DataTrove also generates synthetic data through vLLM, SGLang and OpenAI-compatible endpoints. Generation can use servers running on your own hardware or send requests to an external endpoint. Custom generation logic can make several model calls per document, split long documents into chunks and combine the outputs.

Similar to DataTrove