Attention in Large Language Models Yields Efficient Zero-Shot Re-Rankers
Shijie Chen, Bernal Jimenez Gutierrez, Yu Su
摘要
Information retrieval (IR) systems have played a vital role in modern digital life and have cemented their continued usefulness in this new era of generative AI via retrieval-augmented generation. With strong language processing capabilities and remarkable versatility, large language models (LLMs) have become popular choices for zero-shot re-ranking in IR systems. So far, LLM-based re-ranking methods rely on strong generative capabilities, which restricts their use to either specialized or powerful proprietary models. Given these restrictions, we ask: is autoregressive generation necessary and optimal for LLMs to perform re-ranking? We hypothesize that there are abundant signals relevant to re-ranking within LLMs that might not be used to their full potential via generation. To more directly leverage such signals, we propose in-context re-ranking (ICR), a novel method that leverages the change in attention pattern caused by the search query for accurate and efficient re-ranking. We assume that more relevant documents should receive more attention weights when an LLM is processing the query tokens, and leverage such signals for re-ranking. To mitigate the intrinsic biases in LLMs, we propose a calibration method using a content-free query. Due to the absence of generation, ICR only requires two (O(1)) forward passes to re-rank N documents, making it substantially more efficient than generative re-ranking methods that require at least O(N ) forward passes. Our novel design also enables ICR to be applied to any LLM without specialized training while guaranteeing a well-formed ranking. Extensive experiments with two popular open-weight LLMs on standard singlehop and multi-hop information retrieval benchmarks show that ICR outperforms RankGPT while cutting the latency by more than 60% in practice. Through detailed analyses, we show that ICR's performance is specially strong on tasks that require more complex re-ranking signals, such as handling contextualization and contradiction between the query and passages, as well as information integration across multiple passages. Our findings call for further exploration on novel ways of utilizing open-weight LLMs beyond text generation. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented GenerationJiashuo Sun, Xianrui Zhong, Sizhe Zhou, Jiawei HanNeurIPS 2025 · 被引用 19 次
- Scalable In-context Ranking with Generative ModelsNilesh Gupta, Chong You, Srinadh Bhojanapalli, Sanjiv Kumar 等NeurIPS 2025 · 被引用 10 次
- Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in RecommendationTianjun Wei, Huizhong Guo, Yingpeng Du, Zhu Sun 等ACL 2026 · 被引用 4 次
- Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM PersonalizationLinfeng Du, Ye Yuan, Zichen Zhao, Fuyuan Lyu 等ACL 2026 · 被引用 3 次
- Contextual Relevance and Adaptive Sampling for LLM-Based Document RerankingJerry Huang, Siddarth Madala, Cheng Niu, Julia Hockenmaier 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsBernal Jimenez Gutierrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga 等NeurIPS 2024 · 被引用 395 次
- Distilling Knowledge from Reader to Retriever for Question AnsweringGautier Izacard, Edouard GraveICLR 2021 · 被引用 317 次
相关 Paper
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang 等EMNLP 2023 · 被引用 182 次
- List-aware Reranking-Truncation Joint Model for Search and Retrieval-augmented GenerationShicheng Xu, Liang Pang, Jun Xu, Huawei Shen 等WWW 2024 · 被引用 13 次
- Attention Basin: Why Contextual Position Matters in Large Language ModelsZihao Yi, Zhenqing Ling, Delong Zeng, Haohao Luo 等ACL 2026 · 被引用 2 次
- REALM: Recursive Relevance Modeling for LLM-based Document Re-RankingPinhuan Wang, Zhiqiu Xia, Chunhua Liao, Feiyi Wang 等EMNLP 2025
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan 等EMNLP 2022 · 被引用 69 次
