How to Correctly Report LLM-as-a-Judge Evaluations
Chungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn, Kangwook Lee
摘要
Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensitivity and specificity of the LLM judges induce bias in naive evaluation scores. We propose a simple plug-in framework that corrects this bias and enables statistically principled uncertainty quantification. Our framework constructs confidence intervals that account for uncertainty from both the test dataset and a human-labeled calibration dataset. Additionally, it uses an adaptive strategy to allocate calibration samples for tighter intervals. Importantly, we characterize parameter regimes defined by the true evaluation score and the LLM judge’s sensitivity and specificity in which our LLM-based evaluation yields more reliable estimates than human-only evaluation. Moreover, we show that our framework remains unbiased under distribution shift between the test and calibration datasets, in contrast to existing approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Efficient Inference for Noisy LLM-as-a-Judge EvaluationYiqun Chen, Sizhu Lu, Sijia Li, Moran Guo 等ICML 2026 · 被引用 5 次
- ReEfBench: Quantifying the Reasoning Efficiency of LLMsZhizhang Fu, Yuancheng Gu, Chenkai Hu, Hanmeng Liu 等ACL 2026 · 被引用 1 次
- POLCA: Stochastic Generative Optimization with LLMXuanfei Ren, Allen Nie, Tengyang Xie, Ching-An ChengICML 2026 · 被引用 1 次
- Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document CollectionsLukasz Borchmann, Jordy Van Landeghem, Michał Turski, Shreyansh Padarha 等ICML 2026
- Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising ModelsKrishna Balasubramanian, Aleksandr Podkopaev, Shiva KasiviswanathanICML 2026
它引用的顶会 Paper8
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
- Efficient Inference for Noisy LLM-as-a-Judge EvaluationYiqun Chen, Sizhu Lu, Sijia Li, Moran Guo 等ICML 2026 · 被引用 5 次
- Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimalityArpan Mukherjee, Marcello Bullo, Debabrota Basu, Deniz GunduzICLR 2026 · 被引用 3 次
相关 Paper
- Uncertainty Quantification for LLM-Based Survey SimulationsChengpiao Huang, Yuhang Wu, Kaizheng WangICML 2025
- Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal PredictionHuanxin Sheng, Xinyi Liu, Hangfeng He, Jieyu Zhao 等EMNLP 2025 · 被引用 1 次
- Multicalibration for Confidence Scoring in LLMsGianluca Detommaso, Martin Bertran Lopez, Riccardo Fogliato, Aaron RothICML 2024 · 被引用 39 次
- Quantifying Biases in LLM-as-a-Judge EvaluationsMagda Dubois, Harry Coppock, Mario Giulianelli, Ole Jorgensen 等ICML 2026
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
