Being Right for Whose Right Reasons?
Terne Sasha Thorn Jakobsen, Laura Cabello, Anders Søgaard
摘要
Explainability methods are used to benchmark the extent to which model predictions align with human rationales i.e., are ‘right for the right reasons’. Previous work has failed to acknowledge, however, that what counts as a rationale is sometimes subjective. This paper presents what we think is a first of its kind, a collection of human rationale annotations augmented with the annotators demographic information. We cover three datasets spanning sentiment analysis and common-sense reasoning, and six demographic groups (balanced across age and ethnicity). Such data enables us to ask both what demographics our predictions align with and whose reasoning patterns our models’ rationales align with. We find systematic inter-group annotator disagreement and show how 16 Transformer-based models align better with rationales provided by certain demographic groups: We find that models are biased towards aligning best with older and/or white annotators. We zoom in on the effects of model size and model distillation, finding –contrary to our expectations– negative correlations between model size and rationale agreement as well as no evidence that either model size or model distillation improves fairness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon 等ICML 2022 · 被引用 144 次
- Fairness without Demographics through Knowledge DistillationJunyi Chai, Taeuk Jang, Xiaoqian WangNeurIPS 2022 · 被引用 57 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- Evaluating model performance under worst-case subpopulationsMike Li, Hongseok Namkoong, Shangzhou XiaNeurIPS 2021 · 被引用 19 次
相关 Paper
- Fair Dataset Distillation via Cross-Group Barycenter AlignmentMohammad Hossein Moslemi, Nima Hosseini Dashtbayaz, Zhimin Mei, Bissan Ghaddar 等ICML 2026
- Impact of Annotator Demographics on Sentiment Dataset LabelingYi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs 等CSCW 2022 · 被引用 20 次
- Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human FeedbackMaria Lerner, Florian E. Dorner, Elliott Ash, Naman GoelACL 2024
- Learning to Rationalize for Nonmonotonic Reasoning with Distant SupervisionFaeze Brahman, Vered Shwartz, Rachel Rudinger, Yejin ChoiAAAI 2021 · 被引用 46 次
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
