Unmasking Style Sensitivity: A Causal Analysis of Bias Evaluation Instability in Large Language Models
Jiaxu Zhao, Meng Fang, Kun Zhang, Mykola Pechenizkiy
摘要
Warning: This paper contains content that may be offensive or upsetting. Natural language processing applications are increasingly prevalent, but social biases in their outputs remain a critical challenge. While various bias evaluation methods have been proposed, these assessments show unexpected instability when input texts undergo minor stylistic changes. This paper conducts a comprehensive analysis of how different style transformations impact bias evaluation results across multiple language models and bias types using causal inference techniques. Our findings reveal that formality transformations significantly affect bias scores, with informal style showing substantial bias reductions (up to 8.33% in LLaMA-2-13B). We identify appearance bias, sexual orientation bias, and religious bias as most susceptible to style changes, with variations exceeding 20%. Larger models demonstrate greater sensitivity to stylistic variations, with bias measurements fluctuating up to 3.1% more than in smaller models. These results highlight critical limitations in current bias evaluation methods and emphasize the need for reliable and fair assessments of language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
相关 Paper
- Understanding Large Language Model Vulnerabilities to Social Bias AttacksJiaxu Zhao, Meng Fang, Fanghua Ye, Ke Xu 等ACL 2025
- Debiasing Algorithm through Model AdaptationTomasz Limisiewicz, David Marecek, Tomás MusilICLR 2024 · 被引用 24 次
- Social Bias Probing: Fairness Benchmarking for Language ModelsMarta Marchiori Manerba, Karolina Stanczak, Riccardo Guidotti, Isabelle AugensteinEMNLP 2024 · 被引用 4 次
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen 等ICLR 2026 · 被引用 6 次
- B-score: Detecting biases in large language models using response historyAn Vo, Mohammad Reza Taesiri, Daeyoung Kim, Anh Totti NguyenICML 2025
