Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
Austin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz, Shafiq Joty
摘要
The large language model (LLM)-as-judge paradigm has been used to meet the demand for a cheap, reliable, and fast evaluation of model outputs during AI system development and post-deployment monitoring. While judge models -- LLMs finetuned to specialize in assessing and critiquing model outputs -- have been touted as general purpose evaluators, they are typically evaluated only on non-contextual scenarios, such as instruction following. The omission of contextual settings -- those where external information is used as context to generate an output -- is surprising given the increasing prevalence of retrieval-augmented generation (RAG) and summarization use cases. Contextual assessment is uniquely challenging, as evaluation often depends on practitioner priorities, leading to conditional evaluation criteria (e.g., comparing responses based on factuality and then considering completeness if they are equally factual). To address the gap, we propose ContextualJudgeBench, a judge benchmark with 2,000 challenging response pairs across eight splits inspired by real-world contextual evaluation scenarios. We build our benchmark with a multi-pronged data construction pipeline that leverages both existing human annotations and model-based perturbations. Our comprehensive study across 11 judge models and 9 general purpose models, reveals that the contextual information and its assessment criteria present a significant challenge to even state-of-the-art models. For example, OpenAI's o1, the best-performing model, barely reaches 55% consistent accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the WildJiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen 等ICLR 2026 · 被引用 37 次
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric DomainsAustin Xu, Xuan-Phi Nguyen, Yilun Zhou, Chien-Sheng Wu 等ICLR 2026 · 被引用 8 次
- On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question GeneralizationJanvijay Singh, Austin Xu, Yilun Zhou, Yefan Zhou 等ICLR 2026 · 被引用 5 次
- Direct Judgement Preference OptimizationPeifeng Wang, Austin Xu, Yilun Zhou, Caiming Xiong 等EMNLP 2025 · 被引用 1 次
- RAGferee: Building Contextual Reward Models for Retrieval-Augmented GenerationAndrei Catalin Coman, Ionut-Teodor Sorodoc, Leonardo F. R. Ribeiro, Bill Byrne 等EMNLP 2025
它引用的顶会 Paper26
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri 等NeurIPS 2023 · 被引用 516 次
相关 Paper
- FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke 等ICLR 2025
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented GenerationZhishang Xiang, Chuanjie Wu, Qinggang Zhang, Shengyuan Chen 等ICLR 2026 · 被引用 56 次
- CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksHongchao Jiang, Yiming Chen, Yushi Cao, Hung-Yi Lee 等ACL 2026 · 被引用 33 次
- JudgeBench: A Benchmark for Evaluating LLM-Based JudgesSijun Tan, Siyuan Zhuang, Kyle Montgomery, William Yuan Tang 等ICLR 2025
- CUB: Benchmarking Context Utilisation Techniques for Language ModelsLovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee 等ACL 2026 · 被引用 5 次
