LIDAO: Towards Limited Interventions for Debiasing (Large) Language Models
Tianci Liu, Haoyu Wang, Shiyang Wang, Yu Cheng, Jing Gao
摘要
Warning: this paper contains model outputs exhibiting offensiveness and biases. Large language models (LLMs) have achieved impressive performance on various natural language generation tasks. Nonetheless, they suffer from generating negative and harmful contents that are biased against certain demographic groups (e.g., female), raising severe fairness concerns. As remedies, prior works intervened the generation by removing attitude or demographic information, inevitably degrading the generation quality and resulting in notable fairness-fluency trade-offs. However, it is still under-explored to what extent the fluency has to be affected in order to achieve a desired level of fairness. In this work, we conduct the first formal study from an information-theoretic perspective. We show that previous approaches are excessive for debiasing and propose LIDAO, a general framework to debias a (L)LM at a better fluency provably. We further robustify LIDAO in adversarial scenarios, where a carefully-crafted prompt may stimulate LLMs exhibiting instruction-following abilities to generate texts with fairness issue appears only when the prompt is also taken into account. Experiments on three LMs ranging from 0.7B to 7B parameters demonstrate the superiority of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
相关 Paper
- KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language ModelsSeorin Kim, Dongyoung Lee, Jaejin LeeEMNLP 2025
- Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron EnhancementJinhao Pan, Chahat Raj, Anjishnu Mukherjee, Sina Mansouri 等ICML 2026
- Prompting Fairness: Integrating Causality to Debias Large Language ModelsJingling Li, Zeyu Tang, Xiaoyu Liu, Peter Spirtes 等ICLR 2025
- Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsYue Guo, Yi Yang, Ahmed AbbasiACL 2022
- Mitigate Extrinsic Social Bias in Pre-trained Language Models via Continuous Prompts AdjustmentYiwei Dai, Hengrui Gu, Ying Wang, Xin WangEMNLP 2024 · 被引用 1 次
