Efficient Inference for Noisy LLM-as-a-Judge Evaluation
Yiqun Chen, Sizhu Lu, Sijia Li, Moran Guo, Shengyi Li
Abstract
Large language models (LLMs) are increasingly used as automatic evaluators of generative AI outputs, a paradigm often referred to as "LLM-as-a-judge." In practice, LLM judges are imperfect predictions for the underlying truth and can exhibit systematic, non-random errors. Two main approaches have recently been proposed to address this issue: (i) direct measurementerror correction based on misclassification models such as Rogan-Gladen-style estimators, and (ii) surrogate-outcome approaches such as prediction-powered inference (PPI), which correct bias by calibrating prediction residuals on a small set of gold-standard human labels. In this paper, we systematically study the performance of these two approaches for estimating mean parameters (e.g., average benchmark scores or pairwise win rates). Leveraging tools from semiparametric efficiency theory, we unify the two classes of estimators by deriving explicit forms of efficient influence function (EIF)-based efficient estimators and characterize conditions under which PPIstyle estimators attain strictly smaller asymptotic variance than measurement-error corrections. We verify our theoretical results in simulations and demonstrate the methods on a real-data example. We provide an implementation of the benchmarked methods and comparison utilities at https://github.com/yiqunchen/debias-llm-as-a-judge .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b36a8fe4-7720-4f1e-81aa-1e7b49d87dd2Cited by top-tier papers2
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn et al.ICML 2026 · 24 citations
- Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising ModelsKrishna Balasubramanian, Aleksandr Podkopaev, Shiva KasiviswanathanICML 2026
Builds on4
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language ModelsNaoki Egami, Musashi Hinck, Brandon M. Stewart, Hanying WeiNeurIPS 2023 · 74 citations
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn et al.ICML 2026 · 24 citations
- CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-JudgesHaitao Li, Junjie Chen, Qingyao Ai, Zhumin Chu et al.ACL 2025
Related papers
- Adaptive Prediction-Powered AutoEval with Reliability and Efficiency GuaranteesSangwoo Park, Matteo Zecchin, Osvaldo SimeoneNeurIPS 2025 · 10 citations
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle et al.ICLR 2026 · 24 citations
- Multiple-Prediction-Powered InferenceCharlie Cowen-Breen, Alekh Agarwal, Stephen Bates, William W. Cohen et al.ICLR 2026 · 11 citations
- Generative Augmented InferenceCheng Lu, Mengxin Wang, Dennis Zhang, Heng ZhangICML 2026 · 1 citation
- How efficient is LLM-generated code? A rigorous & high-standard benchmarkRuizhong Qiu, Weiliang Will Zeng, James Ezick, Christopher Lott et al.ICLR 2025
