Lune

ICSE2026Top-tier venue

Attention Pruning: Automated Fairness Repair of Language Models via Surrogate Simulated Annealing

Vishnu Asutosh Dasu, Md Rafi Ur Rashid, Vipul Gupta, Saeid Tizpaz-Niari, Gang Tan

2026Year
1Top-tier citations

Abstract

This paper explores pruning attention heads as a post-processing bias mitigation method for large language models (LLMs). LLMs have been expanding into sensitive social contexts and socio-economic decision-making where fairness concerns become especially crucial. Since LLMs develop their decision-making patterns by training on massive datasets of human-generated content, they naturally encode and perpetuate societal biases. While modifying training datasets and algorithms is prohibitive, post-processing techniquessuch as pruning attention heads in pre-trained LLMs-can provide feasible and effective approaches to improve fairness. However, identifying the optimal subset of parameters to prune presents a combinatorial challenge within the immense parameter space of LLMs, requiring efficient solutions that balance competing objectives across the frontiers of model fairness and utility.

We explore a search-based program repair approach via simulated annealing to address the computational challenges. Given the prohibitive evaluation costs in billion-parameter LLMs, we develop surrogate deep neural networks that efficiently model the relationship between attention head states (active/inactive) and their corresponding fairness/utility metrics. This allows us to perform optimization over the surrogate models and efficiently identify optimal subsets of attention heads for pruning rather than directly searching through the LLM parameter space. This paper introduces Attention Pruning, a fairness-aware surrogate simulated annealing approach to prune attention heads in LLMs that disproportionately contribute to bias while minimally impacting overall model utility. Our experimental evaluation shows that Attention Pruning achieves a reduction of up to 40% in gender bias and outperforms state-of-the-art bias mitigation strategies.

Warning: This paper contains content that some readers may find offensive and harmful.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 470659be-b2f7-4383-b4a8-dadbd78a6771

Cited by top-tier papers1

Ask how each one uses it

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines