Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
Lukas Gienapp, Martin Potthast, Andrew Yates, Harrisen Scells, Eugene Yang
摘要
The unjudged document problem, where systems that did not contribute to the original judgement pool may retrieve documents without a relevance judgement, is a key obstacle to the reuseability of test collections in information retrieval. While the de facto standard to deal with the problem is to treat unjudged documents as non-relevant, many alternatives have been proposed, such as the use of large language models (LLMs) as a relevance judge (LLM-as-a-judge). However, this has been criticized, among other things, as circular, since the same LLM can be used as the ranker and the judge. We propose to train topic-specific relevance classifiers instead: By finetuning monoT5 with independent LoRA weight adaptation on the judgments of a single assessor for a single topic's pool, we align it to that assessor's notion of relevance for the topic. The system rankings obtained through our classifier's relevance judgments achieve a Spearmans' correlation of with ground truth system rankings. As little as 128 initial human judgments per topic suffice to improve the comparability of models, compared to treating unjudged documents as non-relevant, while achieving more reliability than existing LLM-as-a-judge approaches. Topic-specific relevance classifiers are thus a lightweight and straightforward way to tackle the unjudged document problem, while maintaining human judgments as the gold standard for retrieval evaluation. Code, models, and data are made openly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- A Survey of Active Learning for Natural Language ProcessingZhisong Zhang, Emma Strubell, Eduard H. HovyEMNLP 2022 · 被引用 60 次
- LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?Rikiya Takehi, Ellen M. Voorhees, Tetsuya Sakai, Ian SoboroffSIGIR 2025 · 被引用 10 次
- Justice or Prejudice? Quantifying Biases in LLM-as-a-JudgeJiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen 等ICLR 2025
相关 Paper
- Learning to Rank with Multi-Criteria LLM-Judge AnnotationsNaghmeh Farzi, Laura DietzSIGIR 2026
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
- Formalized Information Needs Improve Large-Language-Model Relevance JudgmentsJüri Keller, Maik Fröbe, Björn Engelmann, Fabian Haak 等SIGIR 2026 · 被引用 1 次
- Leveraging LLMs for Unsupervised Dense Retriever RankingEkaterina Khramtsova, Shengyao Zhuang, Mahsa Baktashmotlagh, Guido ZucconSIGIR 2024 · 被引用 21 次
- REALM: Recursive Relevance Modeling for LLM-based Document Re-RankingPinhuan Wang, Zhiqiu Xia, Chunhua Liao, Feiyi Wang 等EMNLP 2025
