Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language Models
Ryan Steed, Swetasudha Panda, Ari Kobren, Michael L. Wick
Abstract
A few large, homogenous, pre-trained models undergird many machine learning systems — and often, these models contain harmful stereotypes learned from the internet. We investigate the bias transfer hypothesis: the theory that social biases (such as stereotypes) internalized by large language models during pre-training transfer into harmful task-specific behavior after fine-tuning. For two classification tasks, we find that reducing intrinsic bias with controlled interventions before fine-tuning does little to mitigate the classifier’s discriminatory behavior after fine-tuning. Regression analysis suggests that downstream disparities are better explained by biases in the fine-tuning dataset. Still, pre-training plays a role: simple alterations to co-occurrence rates in the fine-tuning dataset are ineffective when the model has been pre-trained. Our results encourage practitioners to focus more on dataset quality and context-specific harms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50a59a07-dbdc-4923-99bd-c23b2cbff51fCited by top-tier papers15
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 117 citations
- Overwriting Pretrained Bias with Finetuning DataAngelina Wang, Olga RussakovskyICCV 2023 · 50 citations
- Equi-Tuning: Group Equivariant Fine-Tuning of Pretrained ModelsSourya Basu, Prasanna Sattigeri, Karthikeyan Natesan Ramamurthy, Vijil Chenthamarakshan et al.AAAI 2023 · 25 citations
- Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream TasksMiaomiao Li, Hao Chen, Yang Wang, Tingyuan Zhu et al.ACL 2026 · 16 citations
- Inductive biases of multi-task learning and finetuning: multiple regimes of feature reuseSamuel Lippl, Jack W. LindseyNeurIPS 2024 · 13 citations
Builds on3
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset DevelopmentMorgan Klaus Scheuerman, Alex Hanna, Emily DentonCSCW 2021 · 169 citations
- Towards Debiasing Sentence RepresentationsPaul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim et al.ACL 2020 · 149 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
Related papers
- Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningFan Zhou, Yuzhou Mao, Liu Yu, Yi Yang et al.ACL 2023 · 21 citations
- Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language ModelsLaura Cabello, Emanuele Bugliarello, Stephanie Brandl, Desmond ElliottEMNLP 2023 · 3 citations
- Debiasing Pretrained Text Encoders by Paying Attention to Paying AttentionYacine Gaci, Boualem Benatallah, Fabio Casati, Khalid BenabdeslemEMNLP 2022 · 12 citations
- Probing Toxic Content in Large Pre-Trained Language ModelsNedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song et al.ACL 2021
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen et al.ICLR 2026 · 6 citations
