Learning to Rank with Multi-Criteria LLM-Judge Annotations
Naghmeh Farzi, Laura Dietz
摘要
Large Language Models (LLMs) are increasingly used as automated judges (LLM judges) to evaluate Information Retrieval (IR) systems, offering a cost-effective complement to human assessments. However, most prior work treats evaluation mainly as a tool for comparison rather than as a signal for improving the retrieval system. We study whether criterion grades from Multi-Criteria LLM-Judge relevance labeling, which decomposes relevance into Exactness, Coverage, Topicality, and Contextual Fit, can serve as effective ranking features for learning-to-rank (L2R). We evaluate this approach using manual relevance labels from TREC TREC DL 2019, DL 2020, and DL 2023. We then analyze how the Multi-Criteria LLM-Judge feature importance varies relative to retrieval scores across IR system performance levels.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
- Formalized Information Needs Improve Large-Language-Model Relevance JudgmentsJüri Keller, Maik Fröbe, Björn Engelmann, Fabian Haak 等SIGIR 2026 · 被引用 1 次
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme 等ACL 2024 · 被引用 27 次
- Self-Calibrated Listwise Reranking with Large Language ModelsRuiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao 等WWW 2025 · 被引用 12 次
- Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMsLukas Gienapp, Martin Potthast, Andrew Yates, Harrisen Scells 等SIGIR 2026
