
Hugging Face Datasets is an open source Python library for preparing data for AI training and evaluation on your own machine. It's for developers and researchers working with local files or datasets from the Hugging Face Hub. The library runs locally; downloading, streaming or sharing data through the Hub uses Hugging Face's hosted service.
Apache Arrow storage lets it work with large datasets through memory-mapped, zero-copy reads, reducing the need to load everything into RAM. Streaming lets you process data as it arrives without downloading the full dataset first. The library caches processed results and reuses them, so repeated work doesn't always require another preprocessing pass.
It handles text, images, audio and video, plus PDFs, 3D medical images in NIfTI format, and AI agent traces containing prompts, tool calls and responses. Local file support includes CSV, JSON, JSONL and Parquet. You can process your own data with the same library you use for public datasets.
For existing training workflows, it exchanges data with PyTorch, TensorFlow and JAX, as well as NumPy, Pandas, Polars and Spark. Parallel processing helps prepare larger collections, while indexing, slicing and random access support dataset inspection. FAISS and Elasticsearch integrations provide similarity search over datasets. The library uses the Apache 2.0 license.
Claim this page and we'll verify you by hand. Hugging Face Datasets gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Hugging Face Datasets?Promote it
Something wrong or outdated on this page?
5.1KUpdated 1 year agoApache-2.0
Web#Multi-user access#Semantic search
Argilla is an open-source data annotation and feedback tool for AI engineers and domain experts who build training and evaluation datasets. You can run your own Argilla server or deploy it on Hugging Face Spaces. It's licensed under Apache 2.0.
17.8KUpdated 2 weeks agoApache-2.0
#Code execution#Hugging Face integration#Human approval
1.1KUpdated 2 years agoMIT
#Hugging Face integration#LoRA#Quantization
1.8KUpdated 19 hours agoApache-2.0
Linux · Docker#Distributed execution#Hugging Face integration#Multilingual
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
CAMEL is an open-source Python framework for developers and researchers building systems where AI agents work together. Its focus is on agent roles, communication, and behavior across extended tasks, with applications in synthetic training data, task automation, and simulated societies. It uses the Apache 2.0 license.
DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.
NeMo Curator is an open source Python toolkit for ML engineers and data teams preparing AI training datasets on their own hardware. It handles text, images, video and audio, with reusable pipelines that can run on a laptop or scale across a multi-node Ray cluster. NVIDIA uses it to prepare data for Nemotron models.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.