Preventing Harmful Data Practices by using Participatory Input to Navigate the Machine Learning Multiverse
Jan Simson, Fiona Draxler, Samuel Mehr, Christoph Kern
Abstract
In light of inherent trade-offs regarding fairness, privacy, interpretability and performance, as well as normative questions, the machine learning (ML) pipeline needs to be made accessible for public input, critical reflection and engagement of diverse stakeholders.
In this work, we introduce a participatory approach to gather input from the general public on the design of an ML pipeline. We show how people's input can be used to navigate and constrain the multiverse of decisions during both model development and evaluation. We highlight that central design decisions should be democratized rather than "optimized" to acknowledge their critical impact on the system's output downstream. We describe the iterative development of our approach and its exemplary implementation on a citizen science platform. Our results demonstrate how public participation can inform critical design decisions along the model-building pipeline and combat widespread lazy data practices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f58d521a-f099-42f5-8aae-8a2903531720Builds on13
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIMichael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. WallachCHI 2020 · 428 citations
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel et al.CHI 2022 · 134 citations
- Exploring the Whole Rashomon Set of Sparse Decision TreesRui Xin, Chudi Zhong, Zhi Chen, Takuya Takagi et al.NeurIPS 2022 · 117 citations
- ORES: Lowering Barriers with Participatory Machine Learning in WikipediaAaron Halfaker, R. Stuart GeigerCSCW 2020 · 84 citations
Related papers
- Learning Representations by Humans, for HumansSophie Hilgard, Nir Rosenfeld, Mahzarin R. Banaji, Jack Cao et al.ICML 2021 · 33 citations
- Computational Notebooks as Co-Design Tools: Engaging Young Adults Living with Diabetes, Family Carers, and Clinicians with Machine Learning ModelsAmid Ayobi, Jacob Hughes, Christopher J. Duckworth, Jakub J. Dylag et al.CHI 2023 · 26 citations
- Experimental Analysis of Multi-Step Pipelines for Fair Classifications - More than the Sum of Their Parts?Nico Lässig, Melanie HerschelICDE 2025
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas et al.ICML 2020 · 146 citations
- Towards Fairness in Practice: A Practitioner-Oriented Rubric for Evaluating Fair ML ToolkitsBrianna Richardson, Jean Garcia-Gathright, Samuel F. Way, Jennifer Thom et al.CHI 2021 · 56 citations
