CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
Song Wang, Peng Wang, Tong Zhou, Yushun Dong, Zhen Tan, Jundong Li
摘要
As Large Language Models (LLMs) are increasingly deployed to handle various natural language processing (NLP) tasks, concerns regarding the potential negative societal impacts of LLM-generated content have also arisen. To evaluate the biases exhibited by LLMs, researchers have recently proposed a variety of datasets. However, existing bias evaluation efforts often focus on only a particular type of bias and employ inconsistent evaluation metrics, leading to difficulties in comparison across different datasets and LLMs. To address these limitations, we collect a variety of datasets designed for the bias evaluation of LLMs, and further propose CEB, a Compositional Evaluation Benchmark with 11,004 samples that cover different types of bias across different social groups and tasks. The curation of CEB is based on our newly proposed compositional taxonomy, which characterizes each dataset from three dimensions: bias types, social groups, and tasks. By combining the three dimensions, we develop a comprehensive evaluation strategy for the bias in LLMs. Our experiments demonstrate that the levels of bias vary across these dimensions, thereby providing guidance for the development of specific bias mitigation methods. Our code is provided at https://github.com/SongW-SW/CEB .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent ApproachYuchen Wu, Edward Sun, Kaijie Zhu, Jianxun Lian 等NeurIPS 2025 · 被引用 20 次
- Large Language Models Can Be Contextual Privacy Protection LearnersYijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu 等EMNLP 2024 · 被引用 18 次
- AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language ModelsKai Li, Can Shen, Yile Liu, Jirui Han 等ICLR 2026 · 被引用 17 次
- FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and StereotypesJanki Atul Nawale, Mohammed Safi Ur Rahman Khan, Janani D, Mansi Gupta 等ACL 2025 · 被引用 5 次
- Mitigating Biases in Language Models via Bias UnlearningDianqing Liu, Yi Liu, Guoqing Jin, Zhendong MaoEMNLP 2025 · 被引用 4 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- Training individually fair ML models with sensitive subspace robustnessMikhail Yurochkin, Amanda Bower, Yuekai SunICLR 2020 · 被引用 123 次
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani 等EMNLP 2022 · 被引用 56 次
相关 Paper
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model ResponsesXin Xu, Xunzhi He, Churan Zhi, Ruizhe Chen 等ICLR 2026 · 被引用 4 次
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung 等EMNLP 2023 · 被引用 16 次
- RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language ModelsSoumya Barikeri, Anne Lauscher, Ivan Vulic, Goran GlavasACL 2021
- Adaptive Generation of Bias-Eliciting Questions for LLMsRobin Staab, Jasper Dekoninck, Maximilian Baader, Martin VechevICML 2026
- MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-JudgeSua Lee, Sanghee Park, Jinbae ImACL 2026 · 被引用 1 次
