Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluations
Jin Peng Zhou, Sébastien M. R. Arnold, Nan Ding, Kilian Q. Weinberger, Nan Hua, Fei Sha
摘要
Auto-evaluating language models (LMs), i.e., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and reduce the cost associated with it. But this presents a paradox: how can we trust the grader LM, which is presumably weaker than the candidate LM, to assess problems that are beyond the frontier of the capabilities of either model or both? For instance, today's LMs struggle on graduate-level physics and Olympiad-level math, making them unreliable graders in these domains. We show that providing privileged information -such as ground-truth solutions or problem-specific guidelines -improves automated evaluations on such frontier problems. This approach offers two key advantages. First, it expands the range of problems where LMs graders apply. Specifically, weaker models can now rate the predictions of stronger models. Second, privileged information can be used to devise easier variations of challenging problems which improves the separability of different LMs on tasks where their performance is generally low. With this approach, general-purpose LM graders match the state of the art performance on RewardBench, surpassing almost all the specially-tuned models. LM graders also outperform individual human raters on Vibe-Eval, and approach human expert graders on Olympiad-level math problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Gained in Translation: Privileged Pairwise Judges Enhance Multilingual ReasoningLintang Sutawika, Gokul Swamy, Steven Wu, Graham NeubigACL 2026 · 被引用 6 次
- Training AI Co-Scientists Using Rubric RewardsShashwat Goel, Rishi Hazra, Dulhan Jayalath, Timon Willi 等ICML 2026
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
相关 Paper
- Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation EvaluationJunjie Chen, Weihang Su, Zhumin Chu, Haitao Li 等AAAI 2026
- RocketEval: Efficient automated LLM evaluation via grading checklistTianjun Wei, Wei Wen, Ruizhi Qiao, Xing Sun 等ICLR 2025
- Reliable Fine-Grained Evaluation of Natural Language Math ProofsWenjie Ma, Andrei Cojocaru, Neel Kolhe, Haihan Zhang 等ICLR 2026 · 被引用 14 次
- QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical ProofsSantiago Gonzalez, Alireza Amiribavandpour, Peter Ye, Edward Zhang 等ICML 2026 · 被引用 1 次
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and JudgeZhilin Wang, Jaehun Jung, Ximing Lu, Shizhe Diao 等ICLR 2026 · 被引用 20 次
