"Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
Karina Halevy, Anna Sotnikova, Badr AlKhamissi, Syrielle Montariol, Antoine Bosselut
摘要
Weight-based model editing methods update the parametric knowledge of language models post-training. However, these methods can unintentionally alter unrelated parametric knowledge representations, potentially increasing the risk of harm. In this work, we investigate how weight editing methods unexpectedly amplify model biases after edits. We introduce a novel benchmark dataset, SEESAW-CF, for measuring bias amplification of model editing methods for demographic traits such as race, geographic origin, and gender. We use SEESAW-CF to examine the impact of model editing on bias in five large language models. Our results demonstrate that edited models exhibit, to various degrees, more biased behavior for certain demographic groups than before they were edited, specifically becoming less confident in properties for Asian and African subjects. Additionally, editing facts about place of birth, country of citizenship, or gender has particularly negative effects on the model's knowledge about unrelated properties, such as field of work, a pattern observed across multiple models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 被引用 195 次
- Editing Large Language Models: Problems, Methods, and OpportunitiesYunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng 等EMNLP 2023 · 被引用 83 次
- Prompting GPT-3 To Be ReliableChenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang 等ICLR 2023 · 被引用 68 次
- Mass-Editing Memory in a TransformerKevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov 等ICLR 2023 · 被引用 52 次
相关 Paper
- Unveiling the Pitfalls of Knowledge Editing for Large Language ModelsZhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang 等ICLR 2024 · 被引用 47 次
- Can We Debias Multimodal Large Language Models via Model Editing?Zecheng Wang, Xinye Li, Zhanyue Qin, Chunshan Li 等ACM MM 2024 · 被引用 2 次
- Tracing and Reversing Edits in LLMsPaul Youssef, Zhixue Zhao, Christin Seifert, Jörg SchlöttererICLR 2026 · 被引用 7 次
- FAME: Towards Factual Multi-Task Model EditingZeng Li, Yingyu Shan, Zeming Liu, Jiashu Yao 等EMNLP 2024 · 被引用 1 次
- Should We Really Edit Language Models? On the Evaluation of Edited Language ModelsQi Li, Xiang Liu, Zhenheng Tang, Peijie Dong 等NeurIPS 2024 · 被引用 25 次
