Hugging Face Datasets
Library for accessing and sharing ML datasets
About
Hugging Face Datasets is a library for loading, processing, and sharing datasets across NLP, vision, audio, and multimodal tasks. Storage is backed by Apache Arrow and Parquet files with memory-mapped reads, so most datasets stream and process out of core. It integrates with the Hugging Face Hub for download and upload, supports streaming mode for datasets larger than disk, and ships a unified load_dataset() entry point.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Datasets & Training
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- Apache-2.0
- Added
- Jan 29, 2026
Related Tools
Fine-tune LLMs 2x faster with 80% less memory
Tool for fine-tuning LLMs with various configurations
Open source data labeling platform for ML projects
Hugging Face library for parameter-efficient fine-tuning
Python library for synthetic data generation and curation pipelines with batch LLM inference.
Detects label errors, outliers, and duplicates in machine learning datasets using confident learning.