Favicon of Data-Juicer

Data-Juicer

Open-source AI data processing framework for local machines and Ray clusters, with multimodal cleaning, deduplication and Apache 2.0 licensing.

Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.

Its reusable processing steps cover text, images, audio and video. You can combine them into reproducible pipelines for cleaning, filtering, deduplication and synthetic data generation, then share or version those pipelines as YAML recipes. Custom operators extend the built-in collection.

For document retrieval, it supports extraction, normalization and semantic chunking. Agent-focused processing can remove identifying information, assess interaction quality and produce HTML reports of problematic examples. Dataset analysis and tracing help you inspect data quality and see which samples a pipeline changed.

Processing can scale beyond a laptop through Ray, with adaptive parallelism and CUDA acceleration. Automatic fusion combines compatible processing stages to reduce overhead. The framework supports local storage as well as S3 and HDFS for dataset input and export.

The Juicer model handles text cleaning, filtering and semantic labeling from natural-language instructions and can run locally. Model-backed operations also support OpenAI-compatible APIs and LiteLLM, so those calls go to the selected provider. Alibaba Cloud PAI offers a cloud integration for running Data-Juicer jobs.

Similar to Data-Juicer