How to Correctly Report LLM-as-a-Judge Evaluations
Chungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn, Kangwook Lee
Abstract
Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensitivity and specificity of the LLM judges induce bias in naive evaluation scores. We propose a simple plug-in framework that corrects this bias and enables statistically principled uncertainty quantification. Our framework constructs confidence intervals that account for uncertainty from both the test dataset and a human-labeled calibration dataset. Additionally, it uses an adaptive strategy to allocate calibration samples for tighter intervals. Importantly, we characterize parameter regimes defined by the true evaluation score and the LLM judge’s sensitivity and specificity in which our LLM-based evaluation yields more reliable estimates than human-only evaluation. Moreover, we show that our framework remains unbiased under distribution shift between the test and calibration datasets, in contrast to existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87cd5e21-59b1-4896-8df8-4c10fef0d02cCited by top-tier papers5
- Efficient Inference for Noisy LLM-as-a-Judge EvaluationYiqun Chen, Sizhu Lu, Sijia Li, Moran Guo et al.ICML 2026 · 5 citations
- ReEfBench: Quantifying the Reasoning Efficiency of LLMsZhizhang Fu, Yuancheng Gu, Chenkai Hu, Hanmeng Liu et al.ACL 2026 · 1 citation
- POLCA: Stochastic Generative Optimization with LLMXuanfei Ren, Allen Nie, Tengyang Xie, Ching-An ChengICML 2026 · 1 citation
- Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document CollectionsLukasz Borchmann, Jordy Van Landeghem, Michał Turski, Shreyansh Padarha et al.ICML 2026
- Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising ModelsKrishna Balasubramanian, Aleksandr Podkopaev, Shiva KasiviswanathanICML 2026
Builds on8
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Efficient Inference for Noisy LLM-as-a-Judge EvaluationYiqun Chen, Sizhu Lu, Sijia Li, Moran Guo et al.ICML 2026 · 5 citations
- Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimalityArpan Mukherjee, Marcello Bullo, Debabrota Basu, Deniz GunduzICLR 2026 · 3 citations
Related papers
- Uncertainty Quantification for LLM-Based Survey SimulationsChengpiao Huang, Yuhang Wu, Kaizheng WangICML 2025
- Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal PredictionHuanxin Sheng, Xinyi Liu, Hangfeng He, Jieyu Zhao et al.EMNLP 2025 · 1 citation
- Multicalibration for Confidence Scoring in LLMsGianluca Detommaso, Martin Bertran Lopez, Riccardo Fogliato, Aaron RothICML 2024 · 39 citations
- Quantifying Biases in LLM-as-a-Judge EvaluationsMagda Dubois, Harry Coppock, Mario Giulianelli, Ole Jorgensen et al.ICML 2026
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
