Lune

EMNLP2025顶会

F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations

Tian Lan, Jiang Li, Yemin Wang, Xu Liu, Xiangdong Su, Guanglai Gao

2025年份
3被引次数

摘要

Warning: This paper contains content that may be offensive or harmful With the growing adoption of large language models (LLMs) in NLP tasks, concerns about their fairness have intensified. Yet, most existing fairness benchmarks rely on closed-ended evaluation formats, which diverge from realworld open-ended interactions. These formats are prone to position bias and introduce a "minimum score" effect, where models can earn partial credit simply by guessing. Moreover, such benchmarks often overlook factuality considerations rooted in historical, social, physiological, and cultural contexts, and rarely account for intersectional biases. To address these limitations, we propose F 2 Bench: an openended fairness evaluation benchmark for LLMs that explicitly incorporates factuality considerations. F 2 Bench comprises 2,568 instances across 10 demographic groups and two openended tasks. By integrating text generation, multi-turn reasoning, and factual grounding, F 2 Bench aims to more accurately reflect the complexities of real-world model usage. We conduct a comprehensive evaluation of several LLMs across different series and parameter sizes. Our results reveal that all models exhibit varying degrees of fairness issues. We further compare open-ended and closedended evaluations, analyze model-specific disparities, and provide actionable recommendations for future model development. Our code and dataset are publicly available at https: //github.com/VelikayaScarlet/F2Bench .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext e079151c-68fc-40d9-95d5-ced2d36162ca

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖