Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
Nakyeong Yang, Taegwan Kang, Stanley Jungkyu Choi, Honglak Lee, Kyomin Jung
Abstract
Instruction-following language models often show undesirable biases. These undesirable biases are accelerated in the real-world usage of language models, where a wide range of instructions is used through zero-shot example prompting. To solve this problem, we first define bias neuron, which significantly affects biased outputs, and prove its existence empirically. Furthermore, we propose a novel and practical bias mitigation method, CRISPR, to eliminate bias neurons of language models in instruction-following settings. CRISPR automatically determines biased outputs and categorizes neurons that affect the biased outputs as bias neurons using an attribution, an explainability method. Experimental results demonstrate the effectiveness of our method in mitigating biases under zero-shot instructionfollowing settings without losing the model's task performance and existing knowledge. The experimental results reveal the generalizability of our method as it shows robustness under various instructions and datasets. Surprisingly, our method can mitigate the bias in language models by eliminating only a few neurons (e.g., three neurons). guage models typically arise from the relationship 042 between labels (e.g., "poor people") and tokens 043 (e.g., "drugs") within data instances (Zhao et al., 044 2021; Fei et al., 2023). 045 However, the association between labels and in-046 structions also causes a critical bias since various 047 instructions affect language models to behave in-048 consistently. Figure 2 shows the inconsistent be-049 havior of the Flan-T5-base in various synonymous 050 instructions on four datasets (Wang et al., 2018; 051 Parrish et al., 2021). These results indicate that 052 a language model is easily distracted by varying 053 instructions despite given semantically the same 054 meaning. These phenomena suggest that language 055 models exhibit significant cognitive biases in in-056 struction settings, and these are some of the most 057
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdee9ccd-dbfa-45e8-8219-9ac24ee04147Cited by top-tier papers10
- Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust UnlearningNakyeong Yang, Dong-Kyum Kim, Jea Kwon, Minsung Kim et al.ICLR 2026 · 7 citations
- Neuron-Level Analysis of Cultural Understanding in Large Language ModelsTaisei Yamamoto, Ryoma Kumon, Danushka Bollegala, Hitomi YanakaICLR 2026 · 1 citation
- Debiasing the Fine-Grained Classification Task in LLMs with Bias-Aware PEFTDaiying Zhao, Xinyu Yang, Hang ChenACL 2025 · 1 citation
- CachePrune: Teaching LLMs What Not to Follow via KV-Cache EditingRui Wang, Junda Wu, Yu Xia, Tong Yu et al.ACL 2026
- Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron EnhancementJinhao Pan, Chahat Raj, Anjishnu Mukherjee, Sina Mansouri et al.ICML 2026
Builds on3
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- An Information-theoretic Approach to Prompt Engineering Without Ground Truth LabelsTaylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw et al.ACL 2022 · 142 citations
- Task-Specific Skill Localization in Fine-tuned Language ModelsAbhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, Sanjeev AroraICML 2023 · 100 citations
Related papers
- The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language ModelsYan Liu, Yu Liu, Xiaokang Chen, Pin-Yu Chen et al.ICLR 2024 · 32 citations
- BiasWipe: Mitigating Unintended Bias in Text Classifiers through Model InterpretabilityMamta Mamta, Rishikant Chigrupaatii, Asif EkbalEMNLP 2024 · 4 citations
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen et al.ICLR 2026 · 6 citations
- PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding ProjectionMahdiyar Molahasani, Azadeh Motamedi, Michael A. Greenspan, Il-Min Kim et al.ICCV 2025 · 5 citations
- SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal BiasWenqian Ye, Di Wang, Guangtao Zheng, Bohan Liu et al.AAAI 2026
