Favicon of Hugging Face Datasets

Hugging Face Datasets

Python dataset library for AI training and evaluation. Process local files or load data from the Hugging Face Hub, with Apache Arrow storage and streaming.

Screenshot of Hugging Face Datasets website

Hugging Face Datasets is an open source Python library for preparing data for AI training and evaluation on your own machine. It's for developers and researchers working with local files or datasets from the Hugging Face Hub. The library runs locally; downloading, streaming or sharing data through the Hub uses Hugging Face's hosted service.

Apache Arrow storage lets it work with large datasets through memory-mapped, zero-copy reads, reducing the need to load everything into RAM. Streaming lets you process data as it arrives without downloading the full dataset first. The library caches processed results and reuses them, so repeated work doesn't always require another preprocessing pass.

It handles text, images, audio and video, plus PDFs, 3D medical images in NIfTI format, and AI agent traces containing prompts, tool calls and responses. Local file support includes CSV, JSON, JSONL and Parquet. You can process your own data with the same library you use for public datasets.

For existing training workflows, it exchanges data with PyTorch, TensorFlow and JAX, as well as NumPy, Pandas, Polars and Spark. Parallel processing helps prepare larger collections, while indexing, slicing and random access support dataset inspection. FAISS and Elasticsearch integrations provide similarity search over datasets. The library uses the Apache 2.0 license.

Similar to Hugging Face Datasets