Validating LLM-as-a-Judge Systems under Rating Indeterminacy
Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach, Steven Z. Wu, Alexandra Chouldechova
Abstract
The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and standardizing GenAI evaluations. To validate such judge systems, evaluators assess human-judge agreement by first collecting multiple human ratings for each item in a validation corpus, then aggregating the ratings into a single, per-item gold label rating. For many items, however, rating criteria may admit multiple valid interpretations, so a human or LLM rater may deem multiple ratings "reasonable" or "correct". We call this condition rating indeterminacy. Problematically, many rating tasks that contain rating indeterminacy rely on forced-choice elicitation, whereby raters are instructed to select only one rating for each item. In this paper, we introduce a framework for validating LLM-as-a-judge systems under rating indeterminacy. We draw theoretical connections between different measures of judge system performance under different human-judge agreement metrics, and different rating elicitation and aggregation schemes. We demonstrate that differences in how humans and LLMs resolve rating indeterminacy when responding to forced-choice rating instructions can heavily bias LLM-as-a-judge validation. Through extensive experiments involving 11 real-world rating tasks and 9 commercial LLMs, we show that standard validation approaches that rely upon forced-choice ratings select judge systems that are highly suboptimal, performing as much as 31% worse than judge systems selected by our approach that uses multi-label "response set" ratings to account for rating indeterminacy. We conclude with concrete recommendations for more principled approaches to LLM-as-a-judge validation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b5128cb-aecb-4995-8e68-cdd82d398c3eCited by top-tier papers4
- Navigating Uncertainties: How GenAI Developers Document Their Models on Open-Source PlatformsNingjing Tang, Megan Li, Amy A. Winecoff, Michael Madaio et al.CHI 2026 · 3 citations
- Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-MakingShreya Chappidi, Jatinder Singh, Andra Valentina KrauzeCHI 2026 · 2 citations
- Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL GenerationTarfah Alrashed, Madhup Sukoon, David R. Karger, Natasha F. NoyVLDB 2026
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response TheoryJunhyuk Choi, Sohhyung Park, chanhee cho, Hyeonchu Park et al.ICML 2026
Builds on27
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
Related papers
- JuStRank: Benchmarking LLM Judges for System RankingAriel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim et al.ACL 2025
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development QualityChunyang Li, Yilun Zheng, Xinting Huang, Tianqing Fang et al.ICLR 2026 · 14 citations
- CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklistsYukyung Lee, JoongHoon Kim, Jaehee Kim, Hyowon Cho et al.EMNLP 2025 · 2 citations
- Quantifying Biases in LLM-as-a-Judge EvaluationsMagda Dubois, Harry Coppock, Mario Giulianelli, Ole Jorgensen et al.ICML 2026
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human EvaluationJiaju Chen, Yuxuan Lu, Xiaojie Wang, Huimin Zeng et al.ACL 2026 · 30 citations
