White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
Yixin Wan, Kai-Wei Chang
Abstract
Social biases can manifest in language agency. However, very limited research has investigated such biases in Large Language Model (LLM)-generated content. In addition, previous works often rely on string-matching techniques to identify agentic and communal words within texts, falling short of accurately classifying language agency. We introduce the Language Agency Bias Evaluation (LABE) benchmark, which comprehensively evaluates biases in LLMs by analyzing agency levels attributed to different demographic groups in model generations. LABE tests for gender, racial, and intersectional language agency biases in LLMs on 3 text generation tasks: biographies, professor reviews, and reference letters. Using LABE, we unveil language agency social biases in 3 recent LLMs: ChatGPT, Llama3, and Mistral. We observe that: (1) LLM generations tend to demonstrate greater gender bias than human-written texts; (2) Models demonstrate remarkably higher levels of intersectional bias than the other bias aspects. (3) Promptbased mitigation is unstable and frequently leads to bias exacerbation. Based on our observations, we propose Mitigation via Selective Rewrite (MSR), a novel bias mitigation strategy that leverages an agency classifier to identify and selectively revise parts of generated texts that demonstrate communal traits. Empirical results prove MSR to be more effective and reliable than prompt-based mitigation method, showing a promising research direction. We release our source code and data at https: //github.com/elainew728/labe-agency . Athlete Male Female He has completed multiple marathons, triathlons, and obstacle course races, always pushing himself to the limit and striving for personal improvement. She enjoys sharing her love for dance with students of all ages, and finds great joy in helping others discover their own passion for movement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?Yifan Wang, Mayank Jobanputra, Ji-Ung Lee, Soyoung Oh et al.ICLR 2026 · 3 citations
- Writing with AI Can Reduce Gender Bias in Hiring EvaluationsAlicia T. H. Liu, Mina Lee, Xuechunzi BaiCHI 2026 · 1 citation
- Identity-Robust Language Model Generation via Content Integrity PreservationMiao Zhang, Kelly Chen, Md Mehrab Tanjim, Rumi ChunaraACL 2026 · 1 citation
- InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script GenerationYixin Wan, Xingrun Chen, Kai-Wei ChangACL 2026
- Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-TuningYanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau et al.ACL 2026
Builds on4
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- PowerTransformer: Unsupervised Controllable Revision for Biased Language CorrectionXinyao Ma, Maarten Sap, Hannah Rashkin, Yejin ChoiEMNLP 2020 · 50 citations
- Controlled Analyses of Social Biases in Wikipedia BiosAnjalie Field, Chan Young Park, Kevin Z. Lin, Yulia TsvetkovWWW 2022 · 29 citations
Related papers
- Probing Social Bias in Labor Market Text Generation by ChatGPT: A Masked Language Model ApproachLei Ding, Yang Hu, Nicole Denier, Enze Shi et al.NeurIPS 2024 · 4 citations
- Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text DetectionJiatao Li, Xiaojun WanACL 2025
- FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMsZhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu LiuICLR 2025
- CEB: Compositional Evaluation Benchmark for Fairness in Large Language ModelsSong Wang, Peng Wang, Tong Zhou, Yushun Dong et al.ICLR 2025
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung et al.EMNLP 2023 · 16 citations
