Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations
Maria Movin, Claudia Hauff
Abstract
LLMs have been explored for their use in IR as end-to-end rankers, rerankers and assessors. Recently, the exploration of the prompt-and-predict paradigm for reranking in combination with highly performant LLMs have drawn the attention of researchers. Instead of training or fine-tuning a reranker, LLMs are prompted in a zero-shot manner to produce relevance scores, pairwise preferences, or reranked lists. Existing research, though, has been confined to unstructured text corpora, leaving a gap in our understanding: to what extent do the findings of zero-shot LLM rerankers established on plain text corpora hold for datasets containing predominantly precomputed ranking features as is common in industrial settings? We explore this question via an empirical study on one public learning-to-rank dataset (MSLR-WEB10K) and two datasets collected from an audio streaming platform's search logs. Our results paint a differentiated picture: On average, there remains a significant performance gap: prompting the high-capacity LLM GPT-4 results in up to 16% lower NDCG@10 compared to the traditional supervised learning-to-rank (LTR) approach LambdaMART on the public MSLR-WEB10K dataset. However, when focusing only on a subset of hard queries-i.e. queries where the LTR approach ranks a non-relevant document at the top-the zero-shot LLM reranking outperforms the LTR baseline. We confirm the same trends on two proprietary audio search datasets. We also provide insights into prompt design choices and their impact on LLM reranking. We show that LLMs remain brittle, with the same strategies sometimes helping or hurting depending on the model size and dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
Related papers
- PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document RetrievalShengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin et al.EMNLP 2024 · 26 citations
- REALM: Recursive Relevance Modeling for LLM-based Document Re-RankingPinhuan Wang, Zhiqiu Xia, Chunhua Liao, Feiyi Wang et al.EMNLP 2025
- UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian et al.EMNLP 2023 · 23 citations
- Large Language Models as Narrative-Driven RecommendersLukas Eberhard, Thorsten Ruprechter, Denis HelicWWW 2025 · 3 citations
- REARANK: Reasoning Re-ranking Agent via Reinforcement LearningLe Zhang, Bo Wang, Xipeng Qiu, Siva Reddy et al.EMNLP 2025 · 14 citations
