Adaptive Generation of Bias-Eliciting Questions for LLMs
Robin Staab, Jasper Dekoninck, Maximilian Baader, Martin Vechev
Abstract
Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions of users worldwide. Despite their widespread adoption, growing reliance on their outputs raises significant concerns, particularly as users may be exposed to model-inherent biases that disadvantage or stereotype certain groups. However, existing bias benchmarks commonly rely on simple templated prompts or restrictive multiple-choice questions that fail to capture the complexity of real-world user interactions. In this work, we address this gap by introducing a counterfactual framework that automatically generates realistic, open-ended questions for LLM bias evaluation. Through iterative question mutation, our approach systematically explores areas where models are most likely to exhibit biased behavior. Beyond just detecting harmful biases, we also capture increasingly relevant response dimensions, such as asymmetric refusals and explicit bias acknowledgment. Building on this, we construct CAB, a diverse and human-verified benchmark for realistic and nuanced bias evaluations on current frontier LLMs. Our evaluation using CAB highlights the continued need for fairness research by showing that all examined models exhibit persistent biases across certain scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d4e88e8-d70b-48c6-992f-b0c068b81b2aBuilds on9
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani et al.EMNLP 2022 · 56 citations
- Perturbation Augmentation for Fairer NLPRebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith et al.EMNLP 2022 · 54 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung et al.EMNLP 2023 · 16 citations
- "You Gotta be a Doctor, Lin" : An Investigation of Name-Based Bias of Large Language Models in Employment RecommendationsHuy Nghiem, John Prindle, Jieyu Zhao, Hal Daumé IIIEMNLP 2024 · 4 citations
Related papers
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model ResponsesXin Xu, Xunzhi He, Churan Zhi, Ruizhe Chen et al.ICLR 2026 · 4 citations
- The Impossibility of Fair LLMsJacy Reese Anthis, Kristian Lum, Michael D. Ekstrand, Avi Feller et al.ACL 2025
- Certifying Counterfactual Bias in LLMsIsha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi et al.ICLR 2025 · 3 citations
- FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMsZhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu LiuICLR 2025
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu et al.ICML 2026
