Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction
Huanxin Sheng, Xinyi Liu, Hangfeng He, Jieyu Zhao, Jian Kang
摘要
LLM-as-a-judge has become a promising paradigm for using large language models (LLMs) to evaluate natural language generation (NLG), but the uncertainty of its evaluation remains underexplored.This lack of reliability may limit its deployment in many applications.This work presents the first framework to analyze the uncertainty by offering a prediction interval of LLM-based scoring via conformal prediction.Conformal prediction constructs continuous prediction intervals from a single evaluation run, and we design an ordinal boundary adjustment for discrete rating tasks.We also suggest a midpoint-based score within the interval as a low-bias alternative to raw model score and weighted average.We perform extensive experiments and analysis, which show that conformal prediction can provide valid prediction interval with coverage guarantees.We also explore the usefulness of interval midpoint and judge reprompting for better judgment. 1 LLM-as-a-judge: Rate an evaluation and output logits 1. You'll be provided with a task to evaluate.These are the introduction, criteria and evaluation steps: ... ... 2. The task for you to evaluate is as follows: ... ..
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation MetricsChenxu Liu, Yingjie Fu, Wei Yang, Ying Zhang 等ACL 2026 · 被引用 10 次
- Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-JudgeXin Sun, Di Wu, Sijing Qin, Isao Echizen 等ACL 2026 · 被引用 2 次
- Multimodal Learning on Low-Quality Data with Conformal Predictive Self-CalibrationXun Jiang, Yufan Gu, Disen Hu, Yuqing Hou 等CVPR 2026
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response TheoryJunhyuk Choi, Sohhyung Park, chanhee cho, Hyeonchu Park 等ICML 2026
它引用的顶会 Paper12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
- Language Models with Conformal Factuality GuaranteesChristopher Mohri, Tatsunori HashimotoICML 2024 · 被引用 107 次
- Conformal Prediction using Conditional HistogramsMatteo Sesia, Yaniv RomanoNeurIPS 2021 · 被引用 106 次
相关 Paper
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn 等ICML 2026 · 被引用 24 次
- Conf-Gen: Conformal Uncertainty Quantification for Generative ModelsGabriel Loaiza-Ganem, Kevin Zhang, Wei Cui, Marc Law 等ICML 2026 · 被引用 1 次
- Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language GenerationMykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp HochreiterICLR 2026 · 被引用 10 次
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
- Prune 'n Predict: Optimizing LLM Decision-making with Conformal PredictionHarit Vishwakarma, Alan Mishler, Thomas Cook, Niccolò Dalmasso 等ICML 2025
