LLM Bias Detection and Mitigation through the Lens of Desired Distributions
Ingroj Shrestha, Padmini Srinivasan
摘要
Although prior work on bias mitigation has focused on promoting social equality and demographic parity, less attention has been given to aligning LLM's outputs to desired distributions. For example, we might want to align a model with real-world distributions to support factual grounding. Thus, we define bias as deviation from a desired distribution, which may be an equal or real-world distribution, depending on application goals. We propose a weighted adaptive loss 1 based fine-tuning method that aligns LLM's gender-profession output distribution with the desired distribution, while preserving language modeling capability. Using 3 profession sets-male-dominated, female-dominated, and gender-balanced-derived from U.S. labor statistics (2024), we assess both our adaptive method for reflecting reality and a non-adaptive variant for equality. Across three masked language models, bias is observed under both distributions. We achieve near-complete mitigation under equality and 30-75% reduction under real-world settings. Autoregressive LLMs show no bias under equality but notable bias under real-world settings, with the Llama Instruct models (3.2-3B, 3.1-8B) achieving a 50-62% reduction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Identity-Robust Language Model Generation via Content Integrity PreservationMiao Zhang, Kelly Chen, Md Mehrab Tanjim, Rumi ChunaraACL 2026 · 被引用 1 次
- CHOIR: Harmonizing Structured Persona Diversity for Robust Collaborative LLM ReasoningXiangjue Dong, Cong Wang, Maria Teleki, Millennium Bismay 等ACL 2026
- Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-TuningYanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau 等ACL 2026
它引用的顶会 Paper9
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
- ADEPT: A DEbiasing PrompT FrameworkKe Yang, Charles Yu, Yi Ren Fung, Manling Li 等AAAI 2023 · 被引用 40 次
- Debiasing Pretrained Text Encoders by Paying Attention to Paying AttentionYacine Gaci, Boualem Benatallah, Fabio Casati, Khalid BenabdeslemEMNLP 2022 · 被引用 12 次
相关 Paper
- Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsYue Guo, Yi Yang, Ahmed AbbasiACL 2022
- KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language ModelsSeorin Kim, Dongyoung Lee, Jaejin LeeEMNLP 2025
- Finetuning Text-to-Image Diffusion Models for FairnessXudong Shen, Chao Du, Tianyu Pang, Min Lin 等ICLR 2024 · 被引用 97 次
- GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language ModelsTao Zhang, Ziqian Zeng, Yuxiang Xiao, Huiping Zhuang 等ACL 2025 · 被引用 18 次
- Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMsCamilla Casula, Sebastiano Vecellio Salto, Elisa Leonardelli, Sara TonelliEMNLP 2025
