Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
Yinong Oliver Wang, Nivedha Sivakumar, Falaah Arif Khan, Katherine Metcalf, Adam Golinski, Natalie Mackraz, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff
Abstract
The recent rapid adoption of large language models (LLMs) highlights the critical need for benchmarking their fairness. Conventional fairness metrics, which focus on discrete accuracy-based evaluations (i.e., prediction correctness), fail to capture the implicit impact of model uncertainty (e.g., higher model confidence about one group over another despite similar accuracy). To address this limitation, we propose an uncertainty-aware fairness metric, UCerF, to enable a fine-grained evaluation of model fairness that is more reflective of the internal bias in model decisions compared to conventional fairness measures. Furthermore, observing data size, diversity, and clarity issues in current datasets, we introduce a new genderoccupation fairness evaluation dataset with 31,756 samples for co-reference resolution, offering a more diverse and suitable dataset for evaluating modern LLMs. We establish a benchmark, using our metric and dataset, and apply it to evaluate the behavior of ten open-source LLMs. For example, Mistral-7B exhibits suboptimal fairness due to high confidence in incorrect predictions, a detail overlooked by Equalized Odds but captured by UCerF. Overall, our proposed LLM benchmark, which evaluates fairness with uncertainty awareness, paves the way for developing more transparent and accountable AI systems. Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs Uncertain Certain Certain Correct Incorrect 𝐷(𝑥 ! ) (a) Fair(and performant) Female Group (Pro-stereotypical) Male Group (Anti-stereotypical) Model Behavior Case Study Uncertain Certain Certain Correct Incorrect 𝐷(𝑥 ! ) (b) Biased(moderately) Uncertain Certain Certain Correct Incorrect 𝐷(𝑥 ! ) (d) Fair(though not performant) Uncertain Certain Certain Correct Incorrect 𝐷(𝑥 ! )
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09a18409-80ad-46da-ad71-73866ee4dfc9Cited by top-tier papers3
- DSO: Direct Steering Optimization for Bias MitigationLucas Monteiro Paes, Nivedha Sivakumar, Yinong Oliver Wang, Masha Fedzechkina et al.CVPR 2026 · 3 citations
- Quantifying Metric and Model Agreement in Bias Evaluation of Large Language ModelsArash Asgari, Huan Wu, Amirreza Naziri, Mojtaba Kolahdouzi et al.ACL 2026
- A Game-Theoretic Framework for Measuring and Explaining Metric Compatibility in Fair Machine LearningLingfeng Zhang, Jingran Yang, Zhaohui Wang, Min Zhang et al.ICML 2026
Builds on12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMsHariharan Subramonyam, Roy Pea, Christopher Lawrence Pondoc, Maneesh Agrawala et al.CHI 2024 · 137 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
Related papers
- Fair in Mind, Fair in Action? A Synchronous Benchmark for Understanding and Generation in UMLLMsYiran Zhao, Lu Zhou, Xiaogang Xu, Zhe Liu et al.ICLR 2026 · 1 citation
- Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language ModelsZhengyang Shan, Emily Diana, Jiawei ZhouACL 2025 · 3 citations
- UNCLE: Benchmarking Uncertainty Expressions in Long-Form GenerationRuihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang et al.EMNLP 2025
- LLMS ON TRIAL: Evaluating Judicial Fairness For Large Language ModelsYiran Hu, Zongyue Xue, Haitao Li, Siyuan Zheng et al.ICLR 2026 · 4 citations
- Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning TasksFangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann et al.ACL 2025 · 14 citations
