A Multilingual Social Bias Benchmark Incorporating Thinking Processes
Masahiro Kaneko, Danushka Bollegala, Timothy Baldwin
摘要
Large Language Models (LLMs) can learn both useful knowledge and harmful stereotypes, making bias evaluation essential. Existing frameworks fall into two types: those considering reasoning steps (Thinking Process-Aware Evaluation, TPAE) and those focusing only on final outputs (Straight-to-the-Answer Evaluation, SAE). Prior TPAE studies showed effectiveness in assessing gender bias but relied on template-based, word-counting prompts, limiting generalization to other bias types, languages, and reasoning-based methods. In this study, we introduce MBTP 1 , a multilingual social bias benchmark that incorporates humangenerated pro-and anti-stereotype reasoning as part of the thinking process, and propose a few-shot meta-evaluation method that enables scalable bias assessment without model finetuning. From experiments evaluating 13 social bias categories across 8 languages, we find that human-generated thinking consistently yields higher-quality evaluations than LLM-generated or template-based approaches. Furthermore, TPAE demonstrates superior performance over SAE, highlighting the importance of considering reasoning processes in bias evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
相关 Paper
- EuroGEST: Investigating gender stereotypes in multilingual language modelsJacqueline Rowe, Mateusz Klimaszewski, Liane Guillou, Shannon Vallor 等EMNLP 2025
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 被引用 495 次
- Bias in Language Models: Beyond Trick Tests and Towards RUTEd EvaluationKristian Lum, Jacy Reese Anthis, Kevin Robinson, Chirag Nagpal 等ACL 2025 · 被引用 41 次
- Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language ModelsZara Siddique, Liam D. Turner, Luis Espinosa AnkeEMNLP 2024 · 被引用 2 次
- Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning TasksFangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann 等ACL 2025 · 被引用 14 次
