Multi-Feature Quantized Self-Attention for Fair Large Language Models
Jaeil Park, Sung-Bae Cho
摘要
Large language models (LLMs) often encode social biases tied to sensitive features such as race and gender, undermining fairness in downstream tasks even after instruction tuning. Conventional debiasing methods require expensive fine-tuning, are tied to specific architectures, or operate only at the input or decoding stage while neglecting attention-level representations, which can result in compromised task performance. Moreover, most approaches are tailored to single-attribute settings and do not explicitly address scenarios with multiple, overlapping protected attributes and their intersections. This paper proposes a novel method of multi-feature quantized attention regularization (MQAR) to mitigate multi-feature bias by injecting a structured quantization into frozen self-attention layers. MQAR disentangles attribute-specific activations through vector-quantized regularization and uses a discriminator-guided autoencoding regularizer to adversarially suppress protected-attribute information while preserving task-relevant semantics. Crucially, the proposed method operates without modifying the backbone parameters or accessing pre-training data, ensuring architecture-agnostic applicability and minimizing representation distortion. MQAR is evaluated on five diverse LLMs (BERT, T5, GPT-Neo, Mixtral, and LLaMA 3.2) using three standard bias benchmarks (WinoBias, StereoSet, and CrowS-Pairs). Across these models, MQAR consistently reduces bias for multiple protected attributes and their intersections while maintaining downstream accuracy within at most 0.4 %, on average, of non-debiased baselines on sentiment analysis, abusive language detection, and text generation tasks. These findings highlight quantized attention regularization as a scalable and effective method for mitigating social bias in modern language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 被引用 495 次
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He 等ICLR 2020 · 被引用 391 次
- Towards Debiasing Sentence RepresentationsPaul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim 等ACL 2020 · 被引用 149 次
- PromptBERT: Improving BERT Sentence Embeddings with PromptsTing Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang 等EMNLP 2022 · 被引用 148 次
相关 Paper
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen 等ICLR 2026 · 被引用 6 次
- KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language ModelsSeorin Kim, Dongyoung Lee, Jaejin LeeEMNLP 2025
- Debiasing Pretrained Text Encoders by Paying Attention to Paying AttentionYacine Gaci, Boualem Benatallah, Fabio Casati, Khalid BenabdeslemEMNLP 2022 · 被引用 12 次
- Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsYue Guo, Yi Yang, Ahmed AbbasiACL 2022
- Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language ModelsYisong Xiao, Aishan Liu, Siyuan Liang, Xianglong Liu 等ISSTA 2025 · 被引用 2 次
