Quantifying Metric and Model Agreement in Bias Evaluation of Large Language Models
Arash Asgari, Huan Wu, Amirreza Naziri, Mojtaba Kolahdouzi, Laleh Seyyed-Kalantari
Abstract
Bias evaluation in large language models (LLMs) uses many metrics and benchmarks, but lacks a systematic way to measure agreement across bias metrics and models. As a result, improvements observed under one metric may contradict another, and model rankings may reflect benchmark-specific artifacts rather than stable bias profiles. In this work, we introduce Metric Agreement Score (MeAS) and Model Agreement Score (MoAS), which quantify cross-metric and cross-model agreement in bias rankings, respectively. We apply these measures to ten LLMs, seven bias metrics, and nine corpora. Our results reveal disagreement among both metrics and models: Contrary to expectations, we find that metrics within the same category (generationbased and probabilistic) often behave independently of each other. For instance, HONEST shows independence with toxicity metrics, and the CAT score shows no correlation with Language Modeling Bias metric. At the model level, DeepSeek-family models invert bias rankings relative to most others, indicating that the model family strongly shapes specific bias profiles. These findings challenge the assumption that bias mitigation is universally transferable and highlight the need for agreement-aware evaluation. Our code can be accessed at the following: https://github.com/arashasg/ LLM-Bias-Agreement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ab8090a-1918-4a80-9c4e-29e60c7dbb9fBuilds on5
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani et al.EMNLP 2022 · 56 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
- Social Bias Probing: Fairness Benchmarking for Language ModelsMarta Marchiori Manerba, Karolina Stanczak, Riccardo Guidotti, Isabelle AugensteinEMNLP 2024 · 4 citations
- Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMsYinong Oliver Wang, Nivedha Sivakumar, Falaah Arif Khan, Katherine Metcalf et al.ICML 2025
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
Related papers
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model ResponsesXin Xu, Xunzhi He, Churan Zhi, Ruizhe Chen et al.ICLR 2026 · 4 citations
- Quantifying Biases in LLM-as-a-Judge EvaluationsMagda Dubois, Harry Coppock, Mario Giulianelli, Ole Jorgensen et al.ICML 2026
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung et al.EMNLP 2023 · 16 citations
- Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMsHyejun Jeong, Shiqing Ma, Amir HoumansadrICLR 2026 · 1 citation
- Benchmarking Overton Pluralism in LLMsElinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei et al.ICLR 2026 · 9 citations
