The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models
Yan Liu, Yu Liu, Xiaokang Chen, Pin-Yu Chen, Daoguang Zan, Min-Yen Kan, Tsung-Yi Ho
Abstract
Pre-trained Language models (PLMs) have been acknowledged to contain harmful information, such as social biases, which may cause negative social impacts or even bring catastrophic results in application. Previous works on this problem mainly focused on using black-box methods such as probing to detect and quantify social biases in PLMs by observing model outputs. As a result, previous debiasing methods mainly finetune or even pre-train PLMs on newly constructed anti-stereotypical datasets, which are high-cost. In this work, we try to unveil the mystery of social bias inside language models by introducing the concept of SOCIAL BIAS NEURONS. Specifically, we propose INTEGRATED GAP GRADIENTS (IG 2 ) to accurately pinpoint units (i.e., neurons) in a language model that can be attributed to undesirable behavior, such as social bias. By formalizing undesirable behavior as a distributional property of language, we employ sentiment-bearing prompts to elicit classes of sensitive words (demographics) correlated with such sentiments. Our IG 2 thus attributes the uneven distribution for different demographics to specific Social Bias Neurons, which track the trail of unwanted behavior inside PLM units to achieve interpretability. Moreover, derived from our interpretable technique, BIAS NEURON SUPPRESSION (BNS) is further proposed to mitigate social biases. By studying BERT, RoBERTa, and their attributable differences from debiased FairBERTa, IG 2 allows us to locate and suppress identified neurons, and further mitigate undesired behaviors. As measured by prior metrics from StereoSet, our model achieves a higher degree of fairness while maintaining language modeling ability with low cost 12 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 172b88ff-bb14-42aa-b286-48b301874823Cited by top-tier papers16
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen et al.ICLR 2026 · 6 citations
- LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-DisjointQianli Ma, Dongrui Liu, Qian Chen, Linfeng Zhang et al.ACL 2025 · 5 citations
- The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language ModelsChen Qian, Dongrui Liu, Jie Zhang, Yong Liu et al.ACL 2025 · 3 citations
- MAVias: Mitigate any Visual BiasIoannis Sarridis, Christos Koutlis, Symeon Papadopoulos, Christos DiouICCV 2025 · 1 citation
- RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as PriorJunyao Yang, Jianwei Wang, Huiping Zhuang, Cen Chen et al.AAAI 2026 · 1 citation
Builds on12
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language ModelsNicholas Meade, Elinor Poole-Dayan, Siva ReddyACL 2022 · 160 citations
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 117 citations
- On Second Thought, Let's Not Think Step by Step! Bias and Toxicity in Zero-Shot ReasoningOmar Shaikh, Hongxin Zhang, William Barr Held, Michael S. Bernstein et al.ACL 2023 · 61 citations
- Perturbation Augmentation for Fairer NLPRebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith et al.EMNLP 2022 · 54 citations
Related papers
- Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsYue Guo, Yi Yang, Ahmed AbbasiACL 2022
- Interpretable Debiasing of Vision-Language Models for Social FairnessNa Min An, Yoonna Jang, Yusuke Hirota, Ryo Hachiuma et al.CVPR 2026 · 7 citations
- Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron EnhancementJinhao Pan, Chahat Raj, Anjishnu Mukherjee, Sina Mansouri et al.ICML 2026
- BiasWipe: Mitigating Unintended Bias in Text Classifiers through Model InterpretabilityMamta Mamta, Rishikant Chigrupaatii, Asif EkbalEMNLP 2024 · 4 citations
- Mitigating Biases for Instruction-following Language Models via Bias Neurons EliminationNakyeong Yang, Taegwan Kang, Stanley Jungkyu Choi, Honglak Lee et al.ACL 2024
