"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen K. Paritosh, Lora Aroyo
摘要
AI models are increasingly applied in high-stakes domains like health and conservation. Data quality carries an elevated significance in high-stakes AI due to its heightened downstream impact, impacting predictions like cancer detection, wildlife poaching, and loan allocations. Paradoxically, data is the most under-valued and de-glamorised aspect of AI. In this paper, we report on data practices in high-stakes AI, from interviews with 53 AI practitioners in India, East and West African countries, and USA. We define, identify, and present empirical evidence on Data Cascades—compounding events causing negative, downstream effects from data issues—triggered by conventional AI/ML practices that undervalue data quality. Data cascades are pervasive (92% prevalence), invisible, delayed, but often avoidable. We discuss HCI opportunities in designing and incentivizing data excellence as a first-class citizen of AI, resulting in safer and more robust systems for all.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper108
- Data Filtering NetworksAlex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt 等ICLR 2024 · 被引用 251 次
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset DevelopmentMorgan Klaus Scheuerman, Alex Hanna, Emily DentonCSCW 2021 · 被引用 169 次
- What's In My Big Data?Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander 等ICLR 2024 · 被引用 135 次
- Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and ProcessNadia Nahar, Shurui Zhou, Grace A. Lewis, Christian KästnerICSE 2022 · 被引用 122 次
- The Data-Production DispositifMilagros Miceli, Julian PosadaCSCW 2022 · 被引用 117 次
它引用的顶会 Paper3
- A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic RetinopathyEmma Beede, Elizabeth Elliott Baylor, Fred Hersch, Anna Iurchenko 等CHI 2020 · 被引用 589 次
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 被引用 260 次
- Understanding and Visualizing Data Iteration in Machine LearningFred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, Kayur PatelCHI 2020 · 被引用 114 次
相关 Paper
- Whose AI Dream? In search of the aspiration in data annotationDing Wang, Shantanu Prabhat, Nithya SambasivanCHI 2022 · 被引用 66 次
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein 等CHI 2023 · 被引用 51 次
- Analyzing Collaborative Challenges and Needs of UX Practitioners when Designing with AI/MLMeena Devii Muralikumar, David W. McDonaldCSCW 2024 · 被引用 6 次
- Stress-Testing ML Pipelines with Adversarial Data CorruptionJiongli Zhu, Geyang Xu, Felipe Lorenzi, Boris Glavic 等VLDB 2025 · 被引用 2 次
- "I Don't Think RAI Applies to My Model" - Engaging Non-champions with Sticky Stories for Responsible AI WorkNadia Nahar, Chenyang Yang, Yanxin Chen, Wesley Hanwen Deng 等CHI 2026 · 被引用 2 次
