Quantifying Biases in LLM-as-a-Judge Evaluations
Magda Dubois, Harry Coppock, Mario Giulianelli, Ole Jorgensen, Timo Flesch, Lennart Luettgau, Cozmin Ududec
Abstract
The evaluation of large language models (LLMs) is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation, they are not free from biases (e.g., favouring longer outputs or generations from their own model family). Here we propose a statistical framework based on Bayesian generalised linear models (GLMs) that enables researchers to address their primary research questions (e.g., LLM capability or risk assessment), while simultaneously identifying, quantifying and mitigating various biases in their autograders. Our approach can be applied to various evaluation formats (e.g., absolute scores or pairwise preferences) and augments traditional metrics (e.g., inter-rater agreement) by providing precise uncertainty estimates and clarifying sources of disagreement between graders. This framework also enables efficient counterfactual simulations without costly re-evaluation (e.g., assessing agreement after removing systematic biases). We demonstrate these capabilities through simulated examples, with all methods available in an open-source software package. Overall, we introduce a novel framework for autograder evaluation which allows researchers to detect, quantify and correct for various biases in a systematic way.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias DetectorHaoyan Yang, Runxue Bao, Cao (Danica) Xiao, Jun Ma et al.NeurIPS 2025 · 15 citations
- ELSPR: Evaluator LLM Training Data Self-Purification on Non-Transitive Preferences via Tournament Graph ReconstructionYan Yu, Yilun Liu, Minggui He, Shimin Tao et al.AAAI 2026 · 2 citations
- Large Language Models are not Fair EvaluatorsPeiyi Wang, Lei Li, Liang Chen, Zefan Cai et al.ACL 2024
Related papers
- Justice or Prejudice? Quantifying Biases in LLM-as-a-JudgeJiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen et al.ICLR 2025
- Bridging Human and LLM Judgments: Understanding and Narrowing the GapFelipe Maia Polo, Xinhe Wang, Mikhail Yurochkin, Gongjun Xu et al.NeurIPS 2025 · 7 citations
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn et al.ICML 2026 · 24 citations
- LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language ModelsMinsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu et al.IEEE VIS 2024 · 23 citations
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
