Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on HuggingFace
Xinyu Yang, Weixin Liang, James Zou
Abstract
Advances in machine learning are closely tied to the creation of datasets. While data documentation is widely recognized as essential to the reliability, reproducibility, and transparency of ML, we lack a systematic empirical understanding of current dataset documentation practices. To shed light on this question, here we take Hugging Face -one of the largest platforms for sharing and collaborating on ML models and datasets -as a prominent case study. By analyzing all 7,433 dataset documentation on Hugging Face, our investigation provides an overview of the Hugging Face dataset ecosystem and insights into dataset documentation practices, yielding 5 main findings: (1) The dataset card completion rate shows marked heterogeneity correlated with dataset popularity: While 86.0% of the top 100 downloaded dataset cards fill out all sections suggested by Hugging Face community, only 7.9% of dataset cards with no downloads complete all these sections. (2) A granular examination of each section within the dataset card reveals that the practitioners seem to prioritize Dataset Description and Dataset Structure sections, accounting for 36.2% and 33.6% of the total card length, respectively, for the most downloaded datasets. In contrast, the Considerations for Using the Data section receives the lowest proportion of content, accounting for just 2.1% of the text. (3) By analyzing the subsections within each section and utilizing topic modeling to identify key topics, we uncover what is discussed in each section, and underscore significant themes encompassing both technical and social impacts, as well as limitations within the Considerations for Using the Data section. (4) Our findings also highlight the need for improved accessibility and reproducibility of datasets in the Usage sections. (5) In addition, our human annotation evaluation emphasizes the pivotal role of comprehensive dataset content in shaping individuals' perceptions of a dataset card's overall quality. Overall, our study offers a unique perspective on analyzing dataset documentation through large-scale data science analysis and underlines the need for more thorough dataset documentation in machine learning research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72b695ab-fd5b-49e0-ad71-295315bb45f5Cited by top-tier papers6
- RiskRAG: A Data-Driven Solution for Improved AI Model Risk ReportingPooja S. B. Rao, Sanja Scepanovic, Ke Zhou, Edyta Paulina Bogucka et al.CHI 2025 · 7 citations
- Navigating Uncertainties: How GenAI Developers Document Their Models on Open-Source PlatformsNingjing Tang, Megan Li, Amy A. Winecoff, Michael Madaio et al.CHI 2026 · 3 citations
- Talking About the Assumption in the RoomRamaravind Kommiya Mothilal, Faisal M. Lalani, Syed Ishtiaque Ahmed, Shion Guha et al.CHI 2025 · 1 citation
- From Reflection to Repair: A Scoping Review of Dataset Documentation ToolsPedro Reynolds-Cuéllar, Marisol Wong-Villacres, Adriana Alvarado Garcia, Heila PrecelCHI 2026 · 1 citation
- Bridging the Data Provenance Gap Across Text, Speech, and VideoShayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary et al.ICLR 2025
Builds on1
Related papers
- Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and DesiderataAmy Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach et al.CSCW 2022 · 58 citations
- Answering User Questions About Machine Learning Models Through Standardized Model CardsTajkia Rahman Toma, Balreet Grewal, Cor-Paul BezemerICSE 2025 · 1 citation
- An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryWenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer et al.ICSE 2023 · 62 citations
- Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and TraceabilityAvinash Bhat, Austin Coursey, Grace Hu, Sixian Li et al.CHI 2023 · 29 citations
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset DevelopmentMorgan Klaus Scheuerman, Alex Hanna, Emily DentonCSCW 2021 · 169 citations
