Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
Felipe Maia Polo, Xinhe Wang, Mikhail Yurochkin, Gongjun Xu, Moulinath Banerjee, Yuekai Sun
Abstract
Large language models are increasingly used as judges (LLM-as-a-judge) to evaluate model outputs at scale, but their assessments often diverge systematically from human judgments. We present Bridge 1 , a unified statistical framework that explicitly bridges human and LLM evaluations under both absolute scoring and pairwise comparison paradigms. Bridge posits a latent human preference score for each prompt-response pair and models LLM deviations as linear transformations of covariates that capture sources of discrepancies. This offers a simple and principled framework for refining LLM ratings and characterizing systematic discrepancies between humans and LLMs. We provide an efficient fitting algorithm with asymptotic guarantees for statistical inference. Using six LLM judges and two benchmarks (BigGen Bench and Chatbot Arena), Bridge achieves higher agreement with human ratings (accuracy, calibration, and KL divergence) and exposes systematic human-LLM gaps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74c96171-2647-43ab-82d7-1f65281159d8Cited by top-tier papers2
- Margin-Adaptive Confidence Ranking for Reliable LLM JudgementGaojie Jin, Yong Tao, Lijia Yu, Tianjin HuangICML 2026
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response TheoryJunhyuk Choi, Sohhyung Park, chanhee cho, Hyeonchu Park et al.ICML 2026
Builds on8
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsSeungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin et al.EMNLP 2024 · 38 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal ModelYebin Lee, Imseong Park, Myungjoo KangACL 2024 · 3 citations
- Large Language Models are not Fair EvaluatorsPeiyi Wang, Lei Li, Liang Chen, Zefan Cai et al.ACL 2024
Related papers
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
- Prediction-Powered Ranking of Large Language ModelsIvi Chatzi, Eleni Straitouri, Suhas Thejaswi, Manuel Gomez RodriguezNeurIPS 2024 · 34 citations
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate ThemYidong Wang, Yunze Song, Tingyuan Zhu, Xuanwang Zhang et al.ICLR 2026 · 29 citations
- Quantifying Biases in LLM-as-a-Judge EvaluationsMagda Dubois, Harry Coppock, Mario Giulianelli, Ole Jorgensen et al.ICML 2026
- Who can we trust? LLM-as-a-jury for Comparative AssessmentMengjie Qian, Guangzhi Sun, Mark Gales, Kate KnillICML 2026 · 5 citations
