DataTrove is an open-source Python library for teams preparing large text datasets, including LLM training corpora. It runs on your own machine or on Slurm and Ray clusters, with processing steps that carry across those environments. It uses the Apache 2.0 license.
Its reusable blocks cover reading and writing data, extracting text from HTML, filtering documents, tokenization and dataset statistics. Deduplication includes MinHash and exact sentence matching. You can add custom processing logic alongside the supplied blocks, so a dataset's cleanup rules don't have to fit a fixed workflow. Supported inputs include web archives and Parquet, and storage access through fsspec covers local files, S3 and Hugging Face datasets or buckets.
The library targets workloads where memory use and recovery from failed jobs matter. Its staged processing approach keeps memory use relatively low, and parallel execution can use multiple CPU cores or machines. It tracks completed tasks. Restarting an interrupted job reruns only unfinished work, while dataset statistics record how much data each stage processed or removed.
DataTrove also generates synthetic data through vLLM, SGLang and OpenAI-compatible endpoints. Generation can use servers running on your own hardware or send requests to an external endpoint. Custom generation logic can make several model calls per document, split long documents into chunks and combine the outputs.
Claim this page and we'll verify you by hand. DataTrove gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find DataTrove?Promote it
Something wrong or outdated on this page?
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
1.8KUpdated 19 hours agoApache-2.0
Linux · Docker#Distributed execution#Hugging Face integration#Multilingual
1.7KUpdated 2 days agoApache-2.0
#Batch processing#Code execution#Multimodal input
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
NeMo Curator is an open source Python toolkit for ML engineers and data teams preparing AI training datasets on their own hardware. It handles text, images, video and audio, with reusable pipelines that can run on a laptop or scale across a multi-node Ray cluster. NVIDIA uses it to prepare data for Nemotron models.
Curator is a Python library for developers preparing LLM training datasets or extracting structured records from existing data. It supports local inference through Ollama and vLLM alongside cloud model APIs, so the same data pipeline can use models on your hardware or a hosted provider. It's open source under Apache 2.0.
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.