Learning to Rank with Multi-Criteria LLM-Judge Annotations
Naghmeh Farzi, Laura Dietz
Abstract
Large Language Models (LLMs) are increasingly used as automated judges (LLM judges) to evaluate Information Retrieval (IR) systems, offering a cost-effective complement to human assessments. However, most prior work treats evaluation mainly as a tool for comparison rather than as a signal for improving the retrieval system. We study whether criterion grades from Multi-Criteria LLM-Judge relevance labeling, which decomposes relevance into Exactness, Coverage, Topicality, and Contextual Fit, can serve as effective ranking features for learning-to-rank (L2R). We evaluate this approach using manual relevance labels from TREC TREC DL 2019, DL 2020, and DL 2023. We then analyze how the Multi-Criteria LLM-Judge feature importance varies relative to retrieval scores across IR system performance levels.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9b09372e-e9e3-4fd3-9860-f28385ed90ecRelated papers
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
- Formalized Information Needs Improve Large-Language-Model Relevance JudgmentsJüri Keller, Maik Fröbe, Björn Engelmann, Fabian Haak et al.SIGIR 2026 · 1 citation
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- Self-Calibrated Listwise Reranking with Large Language ModelsRuiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao et al.WWW 2025 · 12 citations
- Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMsLukas Gienapp, Martin Potthast, Andrew Yates, Harrisen Scells et al.SIGIR 2026
