LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
Rikiya Takehi, Ellen M. Voorhees, Tetsuya Sakai, Ian Soboroff
摘要
Test collections are information retrieval tools that allow researchers to quickly and easily evaluate ranking algorithms. While test collections have become an integral part of IR research, the process of data creation involves significant manual annotation effort, which often makes it very expensive and time-consuming. Consequently, test collections could become too small when the budget is limited, which may lead to unstable evaluations. As a cheaper alternative, recent studies have proposed the use of large language models (LLMs) to completely replace human assessors. However, while LLMs seem to somewhat correlate with human judgments, their predictions are not perfect and often show bias. Thus, a complete replacement with LLMs is argued to be too risky and not fully reliable.
In this paper, we propose LLM-Assisted Relevance Assessments (LARA) 1 , an effective method to balance manual annotations with LLM annotations, which helps to build a rich and reliable test collection even under a low budget. We use the LLM's predicted relevance probabilities to select the most profitable documents to manually annotate under a budget constraint. With theoretical reasoning, LARA effectively guides the human annotation process by actively learning to calibrate the LLM's predicted relevance probabilities. Then, using the calibration model learned from the limited manual annotations, LARA debiases the LLM predictions to annotate the remaining non-assessed data. Experiments on TREC-7 Ad Hoc, TREC-8 Ad Hoc, TREC Robust 2004, and TREC-COVID datasets show that LARA outperforms alternative solutions under almost any budget constraint. While the community debates humans vs. LLMs in relevance assessments, we contend that, given the same amount of human effort, it is reasonable to leverage LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Validating LLM-as-a-Judge Systems under Rating IndeterminacyLuke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach 等NeurIPS 2025 · 被引用 31 次
- Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR BenchmarksMinjeong Ban, Jeonghwan Choi, Hyangsuk Min, Nicole Hee-Yeon Kim 等ICLR 2026 · 被引用 1 次
- Formalized Information Needs Improve Large-Language-Model Relevance JudgmentsJüri Keller, Maik Fröbe, Björn Engelmann, Fabian Haak 等SIGIR 2026 · 被引用 1 次
- Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMsLukas Gienapp, Martin Potthast, Andrew Yates, Harrisen Scells 等SIGIR 2026
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
它引用的顶会 Paper1
相关 Paper
- The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMsNitay Calderon, Roi Reichart, Rotem DrorACL 2025
- Learning to Rank with Multi-Criteria LLM-Judge AnnotationsNaghmeh Farzi, Laura DietzSIGIR 2026
- Human-LLM Collaborative Annotation Through Effective Verification of LLM LabelsXinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra 等CHI 2024 · 被引用 127 次
- Let the LLM Stick to Its Strengths: Learning to Route Economical LLMYi-Kai Zhang, Shiyin Lu, Qingguo Chen, Weihua Luo 等NeurIPS 2025 · 被引用 3 次
- Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.IHarrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang 等KDD 2024 · 被引用 2 次
