Benchmarking LLMs' Judgments with No Gold Standard
Shengwei Xu, Yuxuan Lu, Grant Schoenebeck, Yuqing Kong
摘要
We introduce the GEM (Generative Estimator for Mutual Information), an evaluation metric for assessing language generation by Large Language Models (LLMs), particularly in generating informative judgments, without the need for a gold standard reference. GEM broadens the scenarios where we can benchmark LLM generation performance-from traditional ones, like machine translation and summarization, where gold standard references are readily available, to subjective tasks without clear gold standards, such as academic peer review. GEM uses a generative model to estimate mutual information between candidate and reference responses, without requiring the reference to be a gold standard. In experiments on a human-annotated dataset, GEM demonstrates competitive correlations with human scores compared to the state-of-the-art GPT-4o Examiner, and outperforms all other baselines. Additionally, GEM is more robust against strategic manipulations, such as rephrasing or elongation, which can artificially inflate scores under a GPT-4o Examiner. We also present GRE-bench (Generating Review Evaluation Benchmark) which evaluates LLMs based on how well they can generate high-quality peer reviews for academic research papers. Because GRE-bench is based upon GEM, it inherits its robustness properties. Additionally, GRE-bench circumvents data contamination problems (or data leakage) by using the continuous influx of new open-access research papers and peer reviews each year. We show GRE-bench results of various popular LLMs on their peer review capabilities using the ICLR2023 dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Incentive-Aligned Multi-Source LLM SummariesYanchen Jiang, Zhe Feng, Aranyak MehtaICLR 2026 · 被引用 2 次
- CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI ReviewersHexuan Deng, Xiaopeng Ke, Yichen Li, Ruina Hu 等ICML 2026
- Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM GenerationShuyao Xiao, Shengling Wang, Ke ChaoACL 2026
- PMIScore: An Unsupervised Approach to Quantify Dialogue EngagementYongkang Guo, Zhihuan Huang, Yuqing KongWWW 2026
它引用的顶会 Paper6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- Proving Test Set Contamination in Black-Box Language ModelsYonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak 等ICLR 2024 · 被引用 220 次
- Time Travel in LLMs: Tracing Data Contamination in Large Language ModelsShahriar Golchin, Mihai SurdeanuICLR 2024 · 被引用 165 次
相关 Paper
- AIR-Bench: Automated Heterogeneous Information Retrieval BenchmarkJianlyu Chen, Nan Wang, Chaofan Li, Bo Wang 等ACL 2025
- CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model GenerationPei Ke, Bosi Wen, Andrew Feng, Xiao Liu 等ACL 2024 · 被引用 9 次
- Foundational Autoraters: Taming Large Language Models for Better Automatic EvaluationTu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar 等EMNLP 2024 · 被引用 14 次
- WaterBench: Towards Holistic Evaluation of Watermarks for Large Language ModelsShangqing Tu, Yuliang Sun, Yushi Bai, Jifan Yu 等ACL 2024
- ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated AgentsZhuofeng Li, Yi Lu, Dongfu Jiang, Haoxiang Zhang 等ACL 2026 · 被引用 1 次
