Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and Desiderata
Amy Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach, Jennifer Wortman Vaughan
摘要
Data is central to the development and evaluation of machine learning (ML) models. However, the use of problematic or inappropriate datasets can result in harms when the resulting models are deployed. To encourage responsible AI practice through more deliberate reflection on datasets and transparency around the processes by which they are created, researchers and practitioners have begun to advocate for increased data documentation and have proposed several data documentation frameworks. However, there is little research on whether these data documentation frameworks meet the needs of ML practitioners, who both create and consume datasets. To address this gap, we set out to understand ML practitioners' data documentation perceptions, needs, challenges, and desiderata, with the ultimate goal of deriving design requirements that can inform future data documentation frameworks. We conducted a series of semi-structured interviews with 14 ML practitioners at a single large, international technology company. We had them answer a list of questions taken from datasheets for datasets 2018datasheets. Our findings show that current approaches to data documentation are largely ad hoc and myopic in nature. Participants expressed needs for data documentation frameworks to be adaptable to their contexts, integrated into their existing tools and workflows, and automated wherever possible. Despite the fact that data documentation frameworks are often motivated from the perspective of responsible AI, participants did not make the connection between the questions that they were asked to answer and their responsible AI implications. In addition, participants often had difficulties prioritizing the needs of dataset consumers and providing information that someone unfamiliar with their datasets might need to know. Based on these findings, we derive seven design requirements for future data documentation frameworks such as more actionable guidance on how the characteristics of datasets might result in harms and how these harms might be mitigated, more explicit prompts for reflection, automated adaptation to different contexts, and integration into ML practitioners' existing tools and workflows.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Seamful XAI: Operationalizing Seamful Design in Explainable AIUpol Ehsan, Q. Vera Liao, Samir Passi, Mark O. Riedl 等CSCW 2024 · 被引用 41 次
- Out of Context: Investigating the Bias and Fairness Concerns of "Artificial Intelligence as a Service"Kornel Lewicki, Michelle Seng Ah Lee, Jennifer Cobbe, Jatinder SinghCHI 2023 · 被引用 31 次
- Let's Talk Futures: A Literature Review of HCI's Future OrientationCamilo Sanchez, Sui Wang, Kaisa Savolainen, Felix Anand Epp 等CHI 2025 · 被引用 30 次
- Angler: Helping Machine Translation Practitioners Prioritize Model ImprovementsSamantha Robertson, Zijie J. Wang, Dominik Moritz, Mary Beth Kery 等CHI 2023 · 被引用 20 次
- Tinker, Tailor, Configure, Customize: The Articulation Work of Contextualizing an AI Fairness ChecklistMichael A. Madaio, Jingya Chen, Hanna M. Wallach, Jennifer Wortman VaughanCSCW 2024 · 被引用 13 次
它引用的顶会 Paper9
- Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIMichael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. WallachCHI 2020 · 被引用 428 次
- Where Responsible AI meets Reality: Practitioner Perspectives on Enablers for Shifting Organizational PracticesBogdana Rakova, Jingying Yang, Henriette Cramer, Rumman ChowdhuryCSCW 2021 · 被引用 326 次
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 被引用 260 次
- Between Subjectivity and Imposition: Power Dynamics in Data Annotation for Computer VisionMilagros Miceli, Martin Schuessler, Tianling YangCSCW 2020 · 被引用 148 次
- Who Is Included in Human Perceptions of AI?: Trust and Perceived Fairness around Healthcare AI and Cultural MistrustMin Kyung Lee, Katherine RichCHI 2021 · 被引用 135 次
相关 Paper
- From Reflection to Repair: A Scoping Review of Dataset Documentation ToolsPedro Reynolds-Cuéllar, Marisol Wong-Villacres, Adriana Alvarado Garcia, Heila PrecelCHI 2026 · 被引用 1 次
- Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on HuggingFaceXinyu Yang, Weixin Liang, James ZouICLR 2024 · 被引用 41 次
- Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and TraceabilityAvinash Bhat, Austin Coursey, Grace Hu, Sixian Li 等CHI 2023 · 被引用 29 次
- Documenting Data Production Processes: A Participatory Approach for Data WorkMilagros Miceli, Tianling Yang, Adriana Alvarado Garcia, Julian Posada 等CSCW 2022 · 被引用 28 次
- Datasheets for Datasets help ML Engineers Notice and Understand Ethical Issues in Training DataKaren L. BoydCSCW 2021 · 被引用 68 次
