LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
Rikiya Takehi, Ellen M. Voorhees, Tetsuya Sakai, Ian Soboroff
Abstract
Test collections are information retrieval tools that allow researchers to quickly and easily evaluate ranking algorithms. While test collections have become an integral part of IR research, the process of data creation involves significant manual annotation effort, which often makes it very expensive and time-consuming. Consequently, test collections could become too small when the budget is limited, which may lead to unstable evaluations. As a cheaper alternative, recent studies have proposed the use of large language models (LLMs) to completely replace human assessors. However, while LLMs seem to somewhat correlate with human judgments, their predictions are not perfect and often show bias. Thus, a complete replacement with LLMs is argued to be too risky and not fully reliable.
In this paper, we propose LLM-Assisted Relevance Assessments (LARA) 1 , an effective method to balance manual annotations with LLM annotations, which helps to build a rich and reliable test collection even under a low budget. We use the LLM's predicted relevance probabilities to select the most profitable documents to manually annotate under a budget constraint. With theoretical reasoning, LARA effectively guides the human annotation process by actively learning to calibrate the LLM's predicted relevance probabilities. Then, using the calibration model learned from the limited manual annotations, LARA debiases the LLM predictions to annotate the remaining non-assessed data. Experiments on TREC-7 Ad Hoc, TREC-8 Ad Hoc, TREC Robust 2004, and TREC-COVID datasets show that LARA outperforms alternative solutions under almost any budget constraint. While the community debates humans vs. LLMs in relevance assessments, we contend that, given the same amount of human effort, it is reasonable to leverage LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3311f38-1550-460b-9238-9e069ec83c87Cited by top-tier papers5
- Validating LLM-as-a-Judge Systems under Rating IndeterminacyLuke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach et al.NeurIPS 2025 · 31 citations
- Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR BenchmarksMinjeong Ban, Jeonghwan Choi, Hyangsuk Min, Nicole Hee-Yeon Kim et al.ICLR 2026 · 1 citation
- Formalized Information Needs Improve Large-Language-Model Relevance JudgmentsJüri Keller, Maik Fröbe, Björn Engelmann, Fabian Haak et al.SIGIR 2026 · 1 citation
- Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMsLukas Gienapp, Martin Potthast, Andrew Yates, Harrisen Scells et al.SIGIR 2026
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
Builds on1
Related papers
- The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMsNitay Calderon, Roi Reichart, Rotem DrorACL 2025
- Learning to Rank with Multi-Criteria LLM-Judge AnnotationsNaghmeh Farzi, Laura DietzSIGIR 2026
- Human-LLM Collaborative Annotation Through Effective Verification of LLM LabelsXinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra et al.CHI 2024 · 127 citations
- Let the LLM Stick to Its Strengths: Learning to Route Economical LLMYi-Kai Zhang, Shiyin Lu, Qingguo Chen, Weihua Luo et al.NeurIPS 2025 · 3 citations
- Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.IHarrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang et al.KDD 2024 · 2 citations
