Large Language Models can Accurately Predict Searcher Preferences
Paul Thomas, Seth Spielman, Nick Craswell, Bhaskar Mitra
Abstract
Much of the evaluation and tuning of a search system relies on relevance labels---annotations that say whether a document is useful for a given search and searcher. Ideally these come from real searchers, but it is hard to collect this data at scale, so typical experiments rely on third-party labellers who may or may not produce accurate annotations. Label quality is managed with ongoing auditing, training, and monitoring. We discuss an alternative approach. We take careful feedback from real searchers and use this to select a large language model (LLM), and prompt, that agrees with this feedback; the LLM can then produce labels at scale. Our experiments show LLMs are as accurate as human labellers and as useful for finding the best systems and hardest queries. LLM performance varies with prompt features, but also varies unpredictably with simple paraphrases. This unpredictability reinforces the need for high-quality "gold" labels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7ae9a28-8523-47ed-bdb9-38ad33b64048Cited by top-tier papers27
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Prediction-Powered Ranking of Large Language ModelsIvi Chatzi, Eleni Straitouri, Suhas Thejaswi, Manuel Gomez RodriguezNeurIPS 2024 · 34 citations
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- Leveraging LLMs for Unsupervised Dense Retriever RankingEkaterina Khramtsova, Shengyao Zhuang, Mahsa Baktashmotlagh, Guido ZucconSIGIR 2024 · 21 citations
- Small, Medium, Large? A Meta-Study of Effect Sizes at CHI to Aid Interpretation of Effect Sizes and Power CalculationAnna-Marie Ortloff, Florin Martius, Mischa Meier, Theo Raimbault et al.CHI 2025 · 17 citations
Builds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster et al.ICLR 2023 · 297 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
Related papers
- Consolidating Ranking and Relevance Predictions of Large Language Models through Post-ProcessingLe Yan, Zhen Qin, Honglei Zhuang, Rolf Jagerman et al.EMNLP 2024 · 6 citations
- LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?Rikiya Takehi, Ellen M. Voorhees, Tetsuya Sakai, Ian SoboroffSIGIR 2025 · 10 citations
- GLaPE: Gold Label-agnostic Prompt Evaluation for Large Language ModelsXuanchang Zhang, Zhuosheng Zhang, Hai ZhaoEMNLP 2024 · 3 citations
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
- Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation EvaluationJunjie Chen, Weihang Su, Zhumin Chu, Haitao Li et al.AAAI 2026
