F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations
Tian Lan, Jiang Li, Yemin Wang, Xu Liu, Xiangdong Su, Guanglai Gao
Abstract
Warning: This paper contains content that may be offensive or harmful With the growing adoption of large language models (LLMs) in NLP tasks, concerns about their fairness have intensified. Yet, most existing fairness benchmarks rely on closed-ended evaluation formats, which diverge from realworld open-ended interactions. These formats are prone to position bias and introduce a "minimum score" effect, where models can earn partial credit simply by guessing. Moreover, such benchmarks often overlook factuality considerations rooted in historical, social, physiological, and cultural contexts, and rarely account for intersectional biases. To address these limitations, we propose F 2 Bench: an openended fairness evaluation benchmark for LLMs that explicitly incorporates factuality considerations. F 2 Bench comprises 2,568 instances across 10 demographic groups and two openended tasks. By integrating text generation, multi-turn reasoning, and factual grounding, F 2 Bench aims to more accurately reflect the complexities of real-world model usage. We conduct a comprehensive evaluation of several LLMs across different series and parameter sizes. Our results reveal that all models exhibit varying degrees of fairness issues. We further compare open-ended and closedended evaluations, analyze model-specific disparities, and provide actionable recommendations for future model development. Our code and dataset are publicly available at https: //github.com/VelikayaScarlet/F2Bench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e079151c-68fc-40d9-95d5-ced2d36162caBuilds on9
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2024 · 424 citations
- French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than EnglishAurélie Névéol, Yoann Dupont, Julien Bezançon, Karën FortACL 2022 · 61 citations
- WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language ModelsVirginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang, Jonathan MayACL 2023 · 46 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
- RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language ModelsSoumya Barikeri, Anne Lauscher, Ivan Vulic, Goran GlavasACL 2021
Related papers
- FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMsZhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu LiuICLR 2025
- Adaptive Generation of Bias-Eliciting Questions for LLMsRobin Staab, Jasper Dekoninck, Maximilian Baader, Martin VechevICML 2026
- The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented InterventionYixin Wan, Di Wu, Haoran Wang, Kai-Wei ChangEMNLP 2024 · 3 citations
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model ResponsesXin Xu, Xunzhi He, Churan Zhi, Ruizhe Chen et al.ICLR 2026 · 4 citations
- FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality EvaluationFarima Fatahi Bayat, Lechen Zhang, Sheza Munir, Lu WangACL 2025
