Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
Its reusable processing steps cover text, images, audio and video. You can combine them into reproducible pipelines for cleaning, filtering, deduplication and synthetic data generation, then share or version those pipelines as YAML recipes. Custom operators extend the built-in collection.
For document retrieval, it supports extraction, normalization and semantic chunking. Agent-focused processing can remove identifying information, assess interaction quality and produce HTML reports of problematic examples. Dataset analysis and tracing help you inspect data quality and see which samples a pipeline changed.
Processing can scale beyond a laptop through Ray, with adaptive parallelism and CUDA acceleration. Automatic fusion combines compatible processing stages to reduce overhead. The framework supports local storage as well as S3 and HDFS for dataset input and export.
The Juicer model handles text cleaning, filtering and semantic labeling from natural-language instructions and can run locally. Model-backed operations also support OpenAI-compatible APIs and LiteLLM, so those calls go to the selected provider. Alibaba Cloud PAI offers a cloud integration for running Data-Juicer jobs.
Claim this page and we'll verify you by hand. Data-Juicer gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Data-Juicer?Promote it
Something wrong or outdated on this page?
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
5.1KUpdated 1 year agoApache-2.0
Web#Multi-user access#Semantic search
Argilla is an open-source data annotation and feedback tool for AI engineers and domain experts who build training and evaluation datasets. You can run your own Argilla server or deploy it on Hugging Face Spaces. It's licensed under Apache 2.0.
3.4KUpdated 10 months agoApache-2.0
#Structured output
Distilabel is an open-source Python framework for engineers building datasets to train or evaluate AI models. It pairs synthetic data generation with LLM feedback, so a pipeline can create examples and judge their quality. It uses the Apache 2.0 license.
5.8KUpdated 4 years agoApache-2.0
LayoutParser is an open-source Python library for developers and researchers who need to detect page structure in document images and turn OCR output into structured data. Its pretrained deep learning models share a common interface, so you can work with models trained on different document datasets without rewriting the surrounding pipeline.
19.7KUpdated 6 months agoApache-2.0
Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API
1.3KUpdated 7 months agoApache-2.0
Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API
olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.
JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.