NeMo Curator is an open source Python toolkit for ML engineers and data teams preparing AI training datasets on their own hardware. It handles text, images, video and audio, with reusable pipelines that can run on a laptop or scale across a multi-node Ray cluster. NVIDIA uses it to prepare data for Nemotron models.
Text processing includes language detection, quality filtering and classification, plus exact, fuzzy and substring deduplication. Image tools generate CLIP embeddings, detect NSFW content and filter by aesthetics. Video processing finds scenes, extracts clips and filters by motion. Audio tools cover speech transcription, quality assessment, audio tagging and speaker diarization, with word error rate filtering for speech datasets.
Its GPU processing uses NVIDIA RAPIDS alongside Ray for distributed workloads. Pipelines overlap CPU and GPU work and automatically adjust workers across stages. Deployment options include Docker, Slurm and Kubernetes. CPU-only text processing is available; the GPU text quickstart targets Linux x86_64 with CUDA and about 16 GB of GPU memory, and needs internet access to download a Hugging Face model.
Curator also supports an OpenAI-compatible LLM endpoint within a pipeline for classification and synthetic data generation. Its Nemotron-CC recipes cover Common Crawl extraction through filtering, deduplication and synthetic data generation. The toolkit uses the Apache 2.0 license.
Claim this page and we'll verify you by hand. NeMo Curator gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find NeMo Curator?Promote it
Something wrong or outdated on this page?
7.1KUpdated 2 days agoApache-2.0
Docker#Batch processing#Distributed execution#Multimodal input
Data-Juicer is a Python framework for preparing AI datasets on your own machine or a distributed Ray cluster. It's for researchers and teams curating model training data, agent interaction records or documents for retrieval. The project is open source under Apache 2.0.
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
3.4KUpdated 1 day agoApache-2.0
#Batch processing#Distributed execution#Multilingual
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.
DataTrove is an open-source Python library for teams preparing large text datasets, including LLM training corpora. It runs on your own machine or on Slurm and Ray clusters, with processing steps that carry across those environments. It uses the Apache 2.0 license.