Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction
Huanxin Sheng, Xinyi Liu, Hangfeng He, Jieyu Zhao, Jian Kang
Abstract
LLM-as-a-judge has become a promising paradigm for using large language models (LLMs) to evaluate natural language generation (NLG), but the uncertainty of its evaluation remains underexplored.This lack of reliability may limit its deployment in many applications.This work presents the first framework to analyze the uncertainty by offering a prediction interval of LLM-based scoring via conformal prediction.Conformal prediction constructs continuous prediction intervals from a single evaluation run, and we design an ordinal boundary adjustment for discrete rating tasks.We also suggest a midpoint-based score within the interval as a low-bias alternative to raw model score and weighted average.We perform extensive experiments and analysis, which show that conformal prediction can provide valid prediction interval with coverage guarantees.We also explore the usefulness of interval midpoint and judge reprompting for better judgment. 1 LLM-as-a-judge: Rate an evaluation and output logits 1. You'll be provided with a task to evaluate.These are the introduction, criteria and evaluation steps: ... ... 2. The task for you to evaluate is as follows: ... ..
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70db3f50-fb41-4878-857e-e5fbee564071Cited by top-tier papers4
- WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation MetricsChenxu Liu, Yingjie Fu, Wei Yang, Ying Zhang et al.ACL 2026 · 10 citations
- Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-JudgeXin Sun, Di Wu, Sijing Qin, Isao Echizen et al.ACL 2026 · 2 citations
- Multimodal Learning on Low-Quality Data with Conformal Predictive Self-CalibrationXun Jiang, Yufan Gu, Disen Hu, Yuqing Hou et al.CVPR 2026
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response TheoryJunhyuk Choi, Sohhyung Park, chanhee cho, Hyeonchu Park et al.ICML 2026
Builds on12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala et al.ICLR 2024 · 132 citations
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 107 citations
- Conformal Prediction using Conditional HistogramsMatteo Sesia, Yaniv RomanoNeurIPS 2021 · 106 citations
Related papers
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn et al.ICML 2026 · 24 citations
- Conf-Gen: Conformal Uncertainty Quantification for Generative ModelsGabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc Law et al.ICML 2026 · 1 citation
- Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language GenerationMykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp HochreiterICLR 2026 · 10 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Prune 'n Predict: Optimizing LLM Decision-making with Conformal PredictionHarit Vishwakarma, Alan Mishler, Thomas Cook, Niccolò Dalmasso et al.ICML 2025
