Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development
Morgan Klaus Scheuerman, Alex Hanna, Emily Denton
摘要
Data is a crucial component of machine learning. The field is reliant on data to train, validate, and test models. With increased technical capabilities, machine learning research has boomed in both academic and industry settings, and one major focus has been on computer vision. Computer vision is a popular domain of machine learning increasingly pertinent to real-world applications, from facial recognition in policing to object detection for autonomous vehicles. Given computer vision's propensity to shape machine learning research and impact human life, we seek to understand disciplinary practices around dataset documentation - how data is collected, curated, annotated, and packaged into datasets for computer vision researchers and practitioners to use for model tuning and development. Specifically, we examine what dataset documentation communicates about the underlying values of vision data and the larger practices and goals of computer vision as a field. To conduct this study, we collected a corpus of about 500 computer vision datasets, from which we sampled 114 dataset publications across different vision tasks. Through both a structured and thematic content analysis, we document a number of values around accepted data practices, what makes desirable data, and the treatment of humans in the dataset construction process. We discuss how computer vision datasets authors value efficiency at the expense of care; universality at the expense of contextuality; impartiality at the expense of positionality; and model work at the expense of data work. Many of the silenced values we identify sit in opposition with social computing practices. We conclude with suggestions on how to better incorporate silenced values into the dataset creation and curation process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot 等ICML 2023 · 被引用 751 次
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel 等CHI 2022 · 被引用 134 次
- The Data-Production DispositifMilagros Miceli, Julian PosadaCSCW 2022 · 被引用 117 次
- Towards Transparency in Dermatology Image Datasets with Skin Tone Annotations by Experts, Crowds, and an AlgorithmMatthew Groh, Caleb Harris, Roxana Daneshjou, Omar Badri 等CSCW 2022 · 被引用 62 次
它引用的顶会 Paper6
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Critical Race Theory for HCIIhudiya Finda Ogbonnaya-Ogburu, Angela D. R. Smith, Alexandra To, Kentaro ToyamaCHI 2020 · 被引用 397 次
- For You, or For"You"?: Everyday LGBTQ+ Encounters with TikTokEllen Simpson, Bryan C. SemaanCSCW 2020 · 被引用 228 次
- How We've Taught Algorithms to See Identity: Constructing Race and Gender in Image Databases for Facial AnalysisMorgan Klaus Scheuerman, Kandrea Wade, Caitlin Lustig, Jed R. BrubakerCSCW 2020 · 被引用 198 次
- Relational, Flexible, Everyday: Learning from Ethics in Dementia ResearchJames Hodge, Sarah Foley, Rens Brankaert, Gail Kenning 等CHI 2020 · 被引用 44 次
相关 Paper
- From Human to Data to Dataset: Mapping the Traceability of Human Subjects in Computer Vision DatasetsMorgan Klaus Scheuerman, Katy Weathington, Tarun Mugunthan, Emily Denton 等CSCW 2023 · 被引用 16 次
- Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and DesiderataAmy Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach 等CSCW 2022 · 被引用 58 次
- From Reflection to Repair: A Scoping Review of Dataset Documentation ToolsPedro Reynolds-Cuéllar, Marisol Wong-Villacres, Adriana Alvarado Garcia, Heila PrecelCHI 2026 · 被引用 1 次
- Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on HuggingFaceXinyu Yang, Weixin Liang, James ZouICLR 2024 · 被引用 41 次
- Documenting Data Production Processes: A Participatory Approach for Data WorkMilagros Miceli, Tianling Yang, Adriana Alvarado Garcia, Julian Posada 等CSCW 2022 · 被引用 28 次
