Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on HuggingFace
Xinyu Yang, Weixin Liang, James Zou
摘要
Advances in machine learning are closely tied to the creation of datasets. While data documentation is widely recognized as essential to the reliability, reproducibility, and transparency of ML, we lack a systematic empirical understanding of current dataset documentation practices. To shed light on this question, here we take Hugging Face -one of the largest platforms for sharing and collaborating on ML models and datasets -as a prominent case study. By analyzing all 7,433 dataset documentation on Hugging Face, our investigation provides an overview of the Hugging Face dataset ecosystem and insights into dataset documentation practices, yielding 5 main findings: (1) The dataset card completion rate shows marked heterogeneity correlated with dataset popularity: While 86.0% of the top 100 downloaded dataset cards fill out all sections suggested by Hugging Face community, only 7.9% of dataset cards with no downloads complete all these sections. (2) A granular examination of each section within the dataset card reveals that the practitioners seem to prioritize Dataset Description and Dataset Structure sections, accounting for 36.2% and 33.6% of the total card length, respectively, for the most downloaded datasets. In contrast, the Considerations for Using the Data section receives the lowest proportion of content, accounting for just 2.1% of the text. (3) By analyzing the subsections within each section and utilizing topic modeling to identify key topics, we uncover what is discussed in each section, and underscore significant themes encompassing both technical and social impacts, as well as limitations within the Considerations for Using the Data section. (4) Our findings also highlight the need for improved accessibility and reproducibility of datasets in the Usage sections. (5) In addition, our human annotation evaluation emphasizes the pivotal role of comprehensive dataset content in shaping individuals' perceptions of a dataset card's overall quality. Overall, our study offers a unique perspective on analyzing dataset documentation through large-scale data science analysis and underlines the need for more thorough dataset documentation in machine learning research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- RiskRAG: A Data-Driven Solution for Improved AI Model Risk ReportingPooja S. B. Rao, Sanja Scepanovic, Ke Zhou, Edyta Paulina Bogucka 等CHI 2025 · 被引用 7 次
- Navigating Uncertainties: How GenAI Developers Document Their Models on Open-Source PlatformsNingjing Tang, Megan Li, Amy A. Winecoff, Michael Madaio 等CHI 2026 · 被引用 3 次
- Talking About the Assumption in the RoomRamaravind Kommiya Mothilal, Faisal M. Lalani, Syed Ishtiaque Ahmed, Shion Guha 等CHI 2025 · 被引用 1 次
- From Reflection to Repair: A Scoping Review of Dataset Documentation ToolsPedro Reynolds-Cuéllar, Marisol Wong-Villacres, Adriana Alvarado Garcia, Heila PrecelCHI 2026 · 被引用 1 次
- Bridging the Data Provenance Gap Across Text, Speech, and VideoShayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary 等ICLR 2025
它引用的顶会 Paper1
相关 Paper
- Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and DesiderataAmy Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach 等CSCW 2022 · 被引用 58 次
- Answering User Questions About Machine Learning Models Through Standardized Model CardsTajkia Rahman Toma, Balreet Grewal, Cor-Paul BezemerICSE 2025 · 被引用 1 次
- An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryWenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer 等ICSE 2023 · 被引用 62 次
- Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and TraceabilityAvinash Bhat, Austin Coursey, Grace Hu, Sixian Li 等CHI 2023 · 被引用 29 次
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset DevelopmentMorgan Klaus Scheuerman, Alex Hanna, Emily DentonCSCW 2021 · 被引用 169 次
