Formalized Information Needs Improve Large-Language-Model Relevance Judgments
Jüri Keller, Maik Fröbe, Björn Engelmann, Fabian Haak, Timo Breuer, Birger Larsen, Philipp Schaer
Abstract
Cranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In evaluations with human assessors, information needs are often formalized as retrieval topics to avoid an excessive number of relevant documents while maintaining good agreement. However, emerging evaluation setups that use Large Language Models (LLMs) as relevance assessors often use only queries, potentially decreasing the reliability. To study whether LLM relevance assessors benefit from formalized information needs, we synthetically formalize information needs with LLMs into topics that follow the established structure from previous human relevance assessments (i.e., descriptions and narratives). We compare assessors using synthetically formalized topics against the LLM-default query-only assessor on the 2019/2020 editions of TREC Deep Learning and Robust04. We find that assessors without formalization judge many more documents relevant and have a lower agreement, leading to reduced reliability in retrieval evaluations. Furthermore, we show that the formalized topics improve agreement between human and LLM relevance judgments, even when the topics are not highly similar to their human counterparts. Our findings indicate that LLM relevance assessors should use formalized information needs, as is standard for human assessment, and synthetically formalize topics when no human formalization exists to improve evaluation reliability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 972b07b1-b26d-4fdf-81e4-0d2aa350ad48Builds on3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Large Language Models can Accurately Predict Searcher PreferencesPaul Thomas, Seth Spielman, Nick Craswell, Bhaskar MitraSIGIR 2024 · 153 citations
- LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?Rikiya Takehi, Ellen M. Voorhees, Tetsuya Sakai, Ian SoboroffSIGIR 2025 · 10 citations
Related papers
- Learning to Rank with Multi-Criteria LLM-Judge AnnotationsNaghmeh Farzi, Laura DietzSIGIR 2026
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
- The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language ModelsRonak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos et al.SIGIR 2025 · 13 citations
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-JudgeMd. Tahmid Rahman Laskar, Israt Jahan, Elham Dolatabadi, Chun Peng et al.ACL 2025
