Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback
Maria Lerner, Florian E. Dorner, Elliott Ash, Naman Goel
摘要
There is a growing body of work on learning from human feedback to align various aspects of machine learning systems with human values and preferences.We consider the setting of fairness in content moderation, in which human feedback is used to determine how two comments -referencing different sensitive attribute groups -should be treated in comparison to one another.With a novel dataset collected from Prolific and MTurk, we find significant gaps in fairness preferences depending on the race, age, political stance, educational level, and LGBTQ+ identity of annotators.We also demonstrate that demographics mentioned in text have a strong influence on how users perceive individual fairness in moderation.Further, we find that differences also exist in downstream classifiers trained to predict human preferences.Finally, we observe that an ensemble, giving equal weight to classifiers trained on annotations from different demographics, performs better for different demographic intersections; compared to a single classifier that gives equal weight to each annotation.Warning: This paper discusses examples of content that may be offensive or disturbing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel 等CHI 2022 · 被引用 134 次
- Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity AnnotationNitesh Goyal, Ian D. Kivlichan, Rachel Rosen, Lucy VassermanCSCW 2022 · 被引用 74 次
- FuzzE: Fuzzy Fairness Evaluation of Offensive Language Classifiers on African-American EnglishAnthony RiosAAAI 2020 · 被引用 26 次
- To Aggregate or Not? Learning with Separate Noisy LabelsJiaheng Wei, Zhaowei Zhu, Tianyi Luo, Ehsan Amid 等KDD 2023 · 被引用 21 次
相关 Paper
- Impact of Annotator Demographics on Sentiment Dataset LabelingYi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs 等CSCW 2022 · 被引用 20 次
- Which Demographics do LLMs Default to During Annotation?Johannes Schäfer, Aidan Combs, Christopher Bagdon, Jiahui Li 等ACL 2025 · 被引用 11 次
- Being Right for Whose Right Reasons?Terne Sasha Thorn Jakobsen, Laura Cabello, Anders SøgaardACL 2023 · 被引用 8 次
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 被引用 44 次
- Perturbation Augmentation for Fairer NLPRebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith 等EMNLP 2022 · 被引用 54 次
