Leveraging Passage Embeddings for Efficient Listwise Reranking with Large Language Models
Qi Liu, Bo Wang, Nan Wang, Jiaxin Mao
摘要
Recent studies have demonstrated the effectiveness of using large language language models (LLMs) in passage ranking. The listwise approaches, such as RankGPT, have become new state-of-the-art in this task. However, the efficiency of RankGPT models is limited by the maximum context length and relatively high latency of LLM inference. To address these issues, in this paper, we propose PE-Rank, leveraging the single passage embedding as a good context compression for efficient listwise passage reranking. By treating each passage as a special token, we can directly input passage embeddings into LLMs, thereby reducing input length. Additionally, we introduce an inference method that dynamically constrains the decoding space to these special tokens, accelerating the decoding process. For adapting the model to reranking, we employ listwise learning to rank loss for training. Evaluation results on multiple benchmarks demonstrate that PE-Rank significantly improves efficiency in both prefilling and decoding, while maintaining competitive ranking effectiveness. The code is available at https://github.com/liuqi6777/pe_rank
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- xRAG: Extreme Context Compression for Retrieval-augmented Generation with One TokenXin Cheng, Xun Wang, Xingxing Zhang, Tao Ge 等NeurIPS 2024 · 被引用 156 次
- Self-Calibrated Listwise Reranking with Large Language ModelsRuiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao 等WWW 2025 · 被引用 12 次
- Supervised Fine-Tuning or Contrastive Learning? Towards Better Multimodal LLM RerankingXin Zhang, Ziqi Dai, Mingxin Li, Yanzhao Zhang 等ICLR 2026 · 被引用 8 次
- DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG RerankersNavve Wasserman, Oliver Heinimann, Yuval Golbari, Tal Zimbalist 等EMNLP 2025 · 被引用 7 次
- Reason-to-Rank: Distilling Direct and Comparative Reasoning from Large Language Models for Document RerankingYuelyu Ji, Zhuochun Li, Rui Meng, Daqing HeSIGIR 2025 · 被引用 3 次
它引用的顶会 Paper14
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 被引用 488 次
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang 等EMNLP 2023 · 被引用 182 次
- xRAG: Extreme Context Compression for Retrieval-augmented Generation with One TokenXin Cheng, Xun Wang, Xingxing Zhang, Tao Ge 等NeurIPS 2024 · 被引用 156 次
相关 Paper
- Compress-then-Rank: Faster and Better Listwise Reranking with Large Language Models via Ranking-Aware Passage CompressionZhewei Zhi, Yingyi Zhang, Yizhen Jing, Xianneng Li 等AAAI 2026 · 被引用 1 次
- Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language ModelsWenhan Liu, Xinyu Ma, Yutao Zhu, Ziliang Zhao 等ACL 2025 · 被引用 10 次
- FIRST: Faster Improved Listwise Reranking with Single Token DecodingRevanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md. Arafat Sultan 等EMNLP 2024 · 被引用 14 次
- LongRanker: Efficient One-Pass Document Reranking with Long-Context Large Language ModelsChangjiang Zhou, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke 等WWW 2026
- ListT5: Listwise Reranking with Fusion-in-Decoder Improves Zero-shot RetrievalSoyoung Yoon, Eunbi Choi, Jiyeon Kim, Hyeongu Yun 等ACL 2024
