A hunt for the Snark: Annotator Diversity in Data Practices
Shivani Kapania, Alex S. Taylor, Ding Wang
摘要
Diversity in datasets is a key component to building responsible AI/ML. Despite this recognition, we know little about the diversity among the annotators involved in data production. We investigated the approaches to annotator diversity through 16 semi-structured interviews and a survey with 44 AI/ML practitioners. While practitioners described nuanced understandings of annotator diversity, they rarely designed dataset production to account for diversity in the annotation process. The lack of action was explained through operational barriers: from the lack of visibility in the annotator hiring process, to the conceptual difficulty in incorporating worker diversity. We argue that such operational barriers and the widespread resistance to accommodating annotator diversity surface a prevailing logic in data practices-where neutrality, objectivity and 'representationalist thinking' dominate. By understanding this logic to be part of a regime of existence, we explore alternative ways of accounting for annotator subjectivity and diversity in data practices.
• Human-centered computing → Empirical studies in HCI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- FACET: Fairness in Computer Vision Evaluation BenchmarkLaura Gustafson, Chloé Rolland, Nikhila Ravi, Quentin Duval 等ICCV 2023 · 被引用 74 次
- Everyday Uncertainty: How Blind People Use GenAI Tools for Information AccessXinru Tang, Ali Abdolrahmani, Darren Gergle, Anne Marie PiperCHI 2025 · 被引用 26 次
- Making Data Work CountSrravya Chandhiramowuli, Alex S. Taylor, Sara Heitlinger, Ding WangCSCW 2024 · 被引用 25 次
- Wikibench: Community-Driven Data Curation for AI Evaluation on WikipediaTzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim 等CHI 2024 · 被引用 19 次
- Better Little People Pictures: Generative Creation of Demographically Diverse AnthropographicsPriya Dhawka, Lauren Perera, Wesley WillettCHI 2024 · 被引用 8 次
它引用的顶会 Paper15
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 被引用 195 次
- Between Subjectivity and Imposition: Power Dynamics in Data Annotation for Computer VisionMilagros Miceli, Martin Schuessler, Tianling YangCSCW 2020 · 被引用 148 次
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel 等CHI 2022 · 被引用 134 次
- The Landscape and Gaps in Open Source Fairness ToolkitsMichelle Seng Ah Lee, Jatinder SinghCHI 2021 · 被引用 117 次
相关 Paper
- "It is currently hodgepodge": Examining AI/ML Practitioners' Challenges during Co-production of Responsible AI ValuesRama Adithya Varanasi, Nitesh GoyalCHI 2023 · 被引用 52 次
- Towards a Non-Ideal Methodological Framework for Responsible MLRamaravind Kommiya Mothilal, Shion Guha, Syed Ishtiaque AhmedCHI 2024 · 被引用 7 次
- Whose AI Dream? In search of the aspiration in data annotationDing Wang, Shantanu Prabhat, Nithya SambasivanCHI 2022 · 被引用 66 次
- Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and DesiderataAmy Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach 等CSCW 2022 · 被引用 58 次
- Documenting Data Production Processes: A Participatory Approach for Data WorkMilagros Miceli, Tianling Yang, Adriana Alvarado Garcia, Julian Posada 等CSCW 2022 · 被引用 28 次
