Favicon of NeMo Curator

NeMo Curator

A self-hosted Python toolkit for curating AI training data across text, images, video and audio, with NVIDIA GPU support and an Apache 2.0 license.

NeMo Curator is an open source Python toolkit for ML engineers and data teams preparing AI training datasets on their own hardware. It handles text, images, video and audio, with reusable pipelines that can run on a laptop or scale across a multi-node Ray cluster. NVIDIA uses it to prepare data for Nemotron models.

Text processing includes language detection, quality filtering and classification, plus exact, fuzzy and substring deduplication. Image tools generate CLIP embeddings, detect NSFW content and filter by aesthetics. Video processing finds scenes, extracts clips and filters by motion. Audio tools cover speech transcription, quality assessment, audio tagging and speaker diarization, with word error rate filtering for speech datasets.

Its GPU processing uses NVIDIA RAPIDS alongside Ray for distributed workloads. Pipelines overlap CPU and GPU work and automatically adjust workers across stages. Deployment options include Docker, Slurm and Kubernetes. CPU-only text processing is available; the GPU text quickstart targets Linux x86_64 with CUDA and about 16 GB of GPU memory, and needs internet access to download a Hugging Face model.

Curator also supports an OpenAI-compatible LLM endpoint within a pipeline for classification and synthetic data generation. Its Nemotron-CC recipes cover Common Crawl extraction through filtering, deduplication and synthetic data generation. The toolkit uses the Apache 2.0 license.

Similar to NeMo Curator