Lune

EMNLP2025顶会

Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction

Huanxin Sheng, Xinyi Liu, Hangfeng He, Jieyu Zhao, Jian Kang

2025年份
1被引次数
4顶会引用

摘要

LLM-as-a-judge has become a promising paradigm for using large language models (LLMs) to evaluate natural language generation (NLG), but the uncertainty of its evaluation remains underexplored.This lack of reliability may limit its deployment in many applications.This work presents the first framework to analyze the uncertainty by offering a prediction interval of LLM-based scoring via conformal prediction.Conformal prediction constructs continuous prediction intervals from a single evaluation run, and we design an ordinal boundary adjustment for discrete rating tasks.We also suggest a midpoint-based score within the interval as a low-bias alternative to raw model score and weighted average.We perform extensive experiments and analysis, which show that conformal prediction can provide valid prediction interval with coverage guarantees.We also explore the usefulness of interval midpoint and judge reprompting for better judgment. 1 LLM-as-a-judge: Rate an evaluation and output logits 1. You'll be provided with a task to evaluate.These are the introduction, criteria and evaluation steps: ... ... 2. The task for you to evaluate is as follows: ... ..

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖