Datasheets for Datasets help ML Engineers Notice and Understand Ethical Issues in Training Data
Karen L. Boyd
Abstract
The social computing community has demonstrated interest in the ethical issues sometimes produced by machine learning (ML) models, like violations of privacy, fairness, and accountability. This paper discovers what kinds of ethical considerations machine learning engineers recognize, how they build understanding, and what decisions they make when working with a real-world dataset. In particular, it illustrates ways in which Datasheets for Datasets, an accountability intervention designed to help engineers explore unfamiliar training data, scaffolds the process of issue discovery, understanding, and ethical decision-making. Participants were asked to review an intentionally ethically problematic dataset and asked to think aloud as they used it to solve a given ML problem. Out of 23 participants, 11 were given a Datasheet they could use while completing the task. Participants were ethically sensitive enough to identify concerns in the dataset; participants who had a Datasheet did open and refer to it; and those with Datasheets mentioned ethical issues during the think-aloud earlier and more often than than those without. The think-aloud protocol offered a grounded description of how participants recognized, understood, and made a decision about ethical problems in an unfamiliar dataset. The method used in this study can test other interventions that claim to encourage recognition, promote understanding, and support decision-making among technologists.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f1b8a118-2af9-4c9b-994a-fa04429f8644Cited by top-tier papers15
- Seeing Like a Toolkit: How Toolkits Envision the Work of AI EthicsRichmond Y. Wong, Michael A. Madaio, Nick MerrillCSCW 2023 · 108 citations
- Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User ExperienceQ. Vera Liao, Hariharan Subramonyam, Jennifer Wang, Jennifer Wortman VaughanCHI 2023 · 81 citations
- Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and DesiderataAmy Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach et al.CSCW 2022 · 58 citations
- A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness EvaluationsGlen Berman, Nitesh Goyal, Michael MadaioCHI 2024 · 40 citations
- Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and TraceabilityAvinash Bhat, Austin Coursey, Grace Hu, Sixian Li et al.CHI 2023 · 29 citations
Related papers
- Ethics Sheets for AI TasksSaif M. MohammadACL 2022 · 38 citations
- Documenting Data Production Processes: A Participatory Approach for Data WorkMilagros Miceli, Tianling Yang, Adriana Alvarado Garcia, Julian Posada et al.CSCW 2022 · 28 citations
- Data Ethics Emergency Drill: A Toolbox for Discussing Responsible AI for Industry TeamsVanessa Aisyahsari Hanschke, Dylan Rees, Merve Alanyali, David Hopkinson et al.CHI 2024 · 7 citations
- D-BIAS: A Causality-Based Human-in-the-Loop System for Tackling Algorithmic BiasBhavya Ghai, Klaus MuellerIEEE VIS 2022 · 45 citations
- Is a Seat at the Table Enough? Engaging Teachers and Students in Dataset Specification for ML in EducationMei Tan, Hansol Lee, Dakuo Wang, Hari SubramonyamCSCW 2024 · 9 citations
