Attention Pruning: Automated Fairness Repair of Language Models via Surrogate Simulated Annealing
Vishnu Asutosh Dasu, Md Rafi Ur Rashid, Vipul Gupta, Saeid Tizpaz-Niari, Gang Tan
Abstract
This paper explores pruning attention heads as a post-processing bias mitigation method for large language models (LLMs). LLMs have been expanding into sensitive social contexts and socio-economic decision-making where fairness concerns become especially crucial. Since LLMs develop their decision-making patterns by training on massive datasets of human-generated content, they naturally encode and perpetuate societal biases. While modifying training datasets and algorithms is prohibitive, post-processing techniquessuch as pruning attention heads in pre-trained LLMs-can provide feasible and effective approaches to improve fairness. However, identifying the optimal subset of parameters to prune presents a combinatorial challenge within the immense parameter space of LLMs, requiring efficient solutions that balance competing objectives across the frontiers of model fairness and utility.
We explore a search-based program repair approach via simulated annealing to address the computational challenges. Given the prohibitive evaluation costs in billion-parameter LLMs, we develop surrogate deep neural networks that efficiently model the relationship between attention head states (active/inactive) and their corresponding fairness/utility metrics. This allows us to perform optimization over the surrogate models and efficiently identify optimal subsets of attention heads for pruning rather than directly searching through the LLM parameter space. This paper introduces Attention Pruning, a fairness-aware surrogate simulated annealing approach to prune attention heads in LLMs that disproportionately contribute to bias while minimally impacting overall model utility. Our experimental evaluation shows that Attention Pruning achieves a reduction of up to 40% in gender bias and outperforms state-of-the-art bias mitigation strategies.
Warning: This paper contains content that some readers may find offensive and harmful.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 470659be-b2f7-4383-b4a8-dadbd78a6771Cited by top-tier papers1
Ask how each one uses itBuilds on21
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
- Fairway: a way to build fair ML softwareJoymallya Chakraborty, Suvodeep Majumder, Zhe Yu, Tim MenziesFSE 2020 · 131 citations
- White-box fairness testing through adversarial samplingPeixin Zhang, Jingyi Wang, Jun Sun, Guoliang Dong et al.ICSE 2020 · 127 citations
- Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipelineSumon Biswas, Hridesh RajanFSE 2021 · 101 citations
Related papers
- Deciphering Stereotypes in Pre-Trained Language ModelsWeicheng Ma, Henry Scheible, Brian Wang, Goutham Veeramachaneni et al.EMNLP 2023 · 7 citations
- Attention Speaks Volumes: Localizing and Mitigating Bias in Language ModelsRishabh Adiga, Besmira Nushi, Varun ChandrasekaranACL 2025
- Multi-Feature Quantized Self-Attention for Fair Large Language ModelsJaeil Park, Sung-Bae ChoICLR 2026
- KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language ModelsSeorin Kim, Dongyoung Lee, Jaejin LeeEMNLP 2025
- Beyond Linear Approximations: A Novel Pruning Approach for Attention MatrixYingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song et al.ICLR 2025
