White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
Yixin Wan, Kai-Wei Chang
摘要
Social biases can manifest in language agency. However, very limited research has investigated such biases in Large Language Model (LLM)-generated content. In addition, previous works often rely on string-matching techniques to identify agentic and communal words within texts, falling short of accurately classifying language agency. We introduce the Language Agency Bias Evaluation (LABE) benchmark, which comprehensively evaluates biases in LLMs by analyzing agency levels attributed to different demographic groups in model generations. LABE tests for gender, racial, and intersectional language agency biases in LLMs on 3 text generation tasks: biographies, professor reviews, and reference letters. Using LABE, we unveil language agency social biases in 3 recent LLMs: ChatGPT, Llama3, and Mistral. We observe that: (1) LLM generations tend to demonstrate greater gender bias than human-written texts; (2) Models demonstrate remarkably higher levels of intersectional bias than the other bias aspects. (3) Promptbased mitigation is unstable and frequently leads to bias exacerbation. Based on our observations, we propose Mitigation via Selective Rewrite (MSR), a novel bias mitigation strategy that leverages an agency classifier to identify and selectively revise parts of generated texts that demonstrate communal traits. Empirical results prove MSR to be more effective and reliable than prompt-based mitigation method, showing a promising research direction. We release our source code and data at https: //github.com/elainew728/labe-agency . Athlete Male Female He has completed multiple marathons, triathlons, and obstacle course races, always pushing himself to the limit and striving for personal improvement. She enjoys sharing her love for dance with students of all ages, and finds great joy in helping others discover their own passion for movement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?Yifan Wang, Mayank Jobanputra, Ji-Ung Lee, Soyoung Oh 等ICLR 2026 · 被引用 3 次
- Writing with AI Can Reduce Gender Bias in Hiring EvaluationsAlicia T. H. Liu, Mina Lee, Xuechunzi BaiCHI 2026 · 被引用 1 次
- Identity-Robust Language Model Generation via Content Integrity PreservationMiao Zhang, Kelly Chen, Md Mehrab Tanjim, Rumi ChunaraACL 2026 · 被引用 1 次
- InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script GenerationYixin Wan, Xingrun Chen, Kai-Wei ChangACL 2026
- Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-TuningYanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau 等ACL 2026
它引用的顶会 Paper4
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- PowerTransformer: Unsupervised Controllable Revision for Biased Language CorrectionXinyao Ma, Maarten Sap, Hannah Rashkin, Yejin ChoiEMNLP 2020 · 被引用 50 次
- Controlled Analyses of Social Biases in Wikipedia BiosAnjalie Field, Chan Young Park, Kevin Z. Lin, Yulia TsvetkovWWW 2022 · 被引用 29 次
相关 Paper
- Probing Social Bias in Labor Market Text Generation by ChatGPT: A Masked Language Model ApproachLei Ding, Yang Hu, Nicole Denier, Enze Shi 等NeurIPS 2024 · 被引用 4 次
- Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text DetectionJiatao Li, Xiaojun WanACL 2025
- FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMsZhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu LiuICLR 2025
- CEB: Compositional Evaluation Benchmark for Fairness in Large Language ModelsSong Wang, Peng Wang, Tong Zhou, Yushun Dong 等ICLR 2025
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung 等EMNLP 2023 · 被引用 16 次
