Debiasing Algorithm through Model Adaptation
Tomasz Limisiewicz, David Marecek, Tomás Musil
Abstract
Large language models are becoming the go-to solution for the ever-growing number of tasks. However, with growing capacity, models are prone to rely on spurious correlations stemming from biases and stereotypes present in the training data. This work proposes a novel method for detecting and mitigating gender bias in language models. We perform causal analysis to identify problematic model components and discover that mid-upper feed-forward layers are most prone to convey bias. Based on the analysis results, we intervene in the model by applying a linear projection to the weight matrices of these layers. Our titular method, DAMA, significantly decreases bias as measured by diverse metrics while maintaining the model's performance on downstream tasks. We release code for our method and models, which retrain LLaMA's state-of-the-art performance while being significantly less biased.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEditQizhou Chen, Taolin Zhang, Chengyu Wang, Xiaofeng He et al.AAAI 2025 · 9 citations
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 4 citations
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model ResponsesXin Xu, Xunzhi He, Churan Zhi, Ruizhe Chen et al.ICLR 2026 · 4 citations
- Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion ModelsKatarzyna Zaleska, Lukasz Popek, Monika Wysoczanska, Kamil DejaCVPR 2026 · 2 citations
- Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language ModelsYisong Xiao, Aishan Liu, Siyuan Liang, Xianglong Liu et al.ISSTA 2025 · 2 citations
Builds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell et al.NeurIPS 2023 · 305 citations
Related papers
- Unsupervised Concept Vector Extraction for Bias Control in LLMsHannah Cyberey, Yangfeng Ji, David EvansEMNLP 2025 · 4 citations
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen et al.ICLR 2026 · 6 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- A Causal Inference Method for Reducing Gender Bias in Word Embedding RelationsZekun Yang, Juan FengAAAI 2020 · 40 citations
- Attention Speaks Volumes: Localizing and Mitigating Bias in Language ModelsRishabh Adiga, Besmira Nushi, Varun ChandrasekaranACL 2025
