Compress-then-Rank: Faster and Better Listwise Reranking with Large Language Models via Ranking-Aware Passage Compression
Zhewei Zhi, Yingyi Zhang, Yizhen Jing, Xianneng Li, Jianing Liu, Huajie Liu, Yongliang Ding
摘要
Listwise reranking with Large Language Models (LLMs) has emerged as the state-of-the-art approach, consistently establishing new performance benchmarks in passage reranking. However, their practical application faces two critical hurdles: the prohibitive computational overhead and high latency of processing long token sequences, and the performance degradation caused by phenomena like "lost in the middle" in long contexts. To address these challenges, we introduce Compress-then-Rank (C2R), an efficient framework that performs listwise reranking not on original passages, but on their compact multi-vector surrogates. These surrogates can be pre-computed and cached for all passages in the corpus. The effectiveness of C2R hinges on three key innovations. First, the compressor model is pre-trained on a combination of text restoration and continuation objectives, enabling high-fidelity compressed vector sequences that mitigate the semantic loss common in single-vector methods. Second, a novel input scheme prepends embeddings of each ordinal index (e.g., [1]:) to its corresponding compressed vector sequence, which both delineates passage boundaries and guides the reranker LLM to generate a ranked list. Finally, the compressor and reranker are jointly optimized, making the compression ranking-aware for the ranking objective. Extensive experiments on major reranking benchmarks demonstrate that C2R provides substantial speedups while achieving competitive and even superior ranking performance compared to full-text reranking methods. The related code is provided in the supplementary materials.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Make Your LLM Fully Utilize the ContextShengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng 等NeurIPS 2024 · 被引用 212 次
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang 等EMNLP 2023 · 被引用 182 次
- In-context Autoencoder for Context Compression in a Large Language ModelTao Ge, Jing Hu, Lei Wang, Xun Wang 等ICLR 2024 · 被引用 158 次
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan 等EMNLP 2022 · 被引用 69 次
相关 Paper
- Leveraging Passage Embeddings for Efficient Listwise Reranking with Large Language ModelsQi Liu, Bo Wang, Nan Wang, Jiaxin MaoWWW 2025 · 被引用 26 次
- LongRanker: Efficient One-Pass Document Reranking with Long-Context Large Language ModelsChangjiang Zhou, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke 等WWW 2026
- Very Efficient Listwise Multimodal Reranking for Long DocumentsYiqun Sun, Pengfei Wei, Lawrence HsiehICML 2026 · 被引用 1 次
- FIRST: Faster Improved Listwise Reranking with Single Token DecodingRevanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md. Arafat Sultan 等EMNLP 2024 · 被引用 14 次
- Self-Calibrated Listwise Reranking with Large Language ModelsRuiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao 等WWW 2025 · 被引用 12 次
