Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language Models
Ryan Steed, Swetasudha Panda, Ari Kobren, Michael L. Wick
摘要
A few large, homogenous, pre-trained models undergird many machine learning systems — and often, these models contain harmful stereotypes learned from the internet. We investigate the bias transfer hypothesis: the theory that social biases (such as stereotypes) internalized by large language models during pre-training transfer into harmful task-specific behavior after fine-tuning. For two classification tasks, we find that reducing intrinsic bias with controlled interventions before fine-tuning does little to mitigate the classifier’s discriminatory behavior after fine-tuning. Regression analysis suggests that downstream disparities are better explained by biases in the fine-tuning dataset. Still, pre-training plays a role: simple alterations to co-occurrence rates in the fine-tuning dataset are ineffective when the model has been pre-trained. Our results encourage practitioners to focus more on dataset quality and context-specific harms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 被引用 117 次
- Overwriting Pretrained Bias with Finetuning DataAngelina Wang, Olga RussakovskyICCV 2023 · 被引用 50 次
- Equi-Tuning: Group Equivariant Fine-Tuning of Pretrained ModelsSourya Basu, Prasanna Sattigeri, Karthikeyan Natesan Ramamurthy, Vijil Chenthamarakshan 等AAAI 2023 · 被引用 25 次
- Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream TasksMiaomiao Li, Hao Chen, Yang Wang, Tingyuan Zhu 等ACL 2026 · 被引用 16 次
- Inductive biases of multi-task learning and finetuning: multiple regimes of feature reuseSamuel Lippl, Jack W. LindseyNeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper3
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset DevelopmentMorgan Klaus Scheuerman, Alex Hanna, Emily DentonCSCW 2021 · 被引用 169 次
- Towards Debiasing Sentence RepresentationsPaul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim 等ACL 2020 · 被引用 149 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
相关 Paper
- Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningFan Zhou, Yuzhou Mao, Liu Yu, Yi Yang 等ACL 2023 · 被引用 21 次
- Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language ModelsLaura Cabello, Emanuele Bugliarello, Stephanie Brandl, Desmond ElliottEMNLP 2023 · 被引用 3 次
- Debiasing Pretrained Text Encoders by Paying Attention to Paying AttentionYacine Gaci, Boualem Benatallah, Fabio Casati, Khalid BenabdeslemEMNLP 2022 · 被引用 12 次
- Probing Toxic Content in Large Pre-Trained Language ModelsNedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song 等ACL 2021
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen 等ICLR 2026 · 被引用 6 次
