Towards Lossless Token Pruning in Late-Interaction Retrieval Models
Yuxuan Zong, Benjamin Piwowarski
摘要
Late interaction neural IR models like ColBERT offer a competitive effectiveness-efficiency trade-off across many benchmarks. However, they require a huge memory space to store the contextual representation for all the document tokens. Some works have proposed using either heuristics or statistical-based techniques to prune tokens from each document. This however doesn't guarantee that the removed tokens have no impact on the retrieval score. Our work uses a principled approach to define how to prune tokens without impacting the score between a document and a query. We introduce three regularization losses, that induce a solution with high pruning ratios, as well as two pruning strategies. We study them experimentally (in and out-domain), showing that we can preserve ColBERT's performance while using only 30% of the tokens.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Multi-Vector Index Compression in Any ModalityHanxiang Qin, Alexander Martin, Rohan Jha, Chunsheng Zuo 等SIGIR 2026 · 被引用 11 次
- A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval ModelsYash Kankanampati, Yuxuan Zong, Nadi Tomeh, Benjamin Piwowarski 等SIGIR 2026
它引用的顶会 Paper9
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin 等SIGIR 2021 · 被引用 297 次
- RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-rankingRuiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao 等EMNLP 2021 · 被引用 147 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Multi-View Document Representation Learning for Open-Domain Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang 等ACL 2022 · 被引用 80 次
相关 Paper
- CrossQ: Task-Aligned Cross-Token Conditional Quantization for Late Interaction RetrievalRohit Kumar Salla, Manoj Saravanan, Ramya AmancherlaICML 2026
- Rethinking the Role of Token Retrieval in Multi-Vector RetrievalJinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei 等NeurIPS 2023 · 被引用 60 次
- Compact Token Representations with Contextual Quantization for Efficient Document Re-rankingYingrui Yang, Yifan Qiao, Tao YangACL 2022 · 被引用 8 次
- Incorporating Token Importance in Multi-Vector RetrievalArchish S, Ankit Garg, Kirankumar Shiragur, Neeraj KayalAAAI 2026 · 被引用 2 次
- Dynamic Context Pruning for Efficient and Interpretable Autoregressive TransformersSotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci 等NeurIPS 2023 · 被引用 95 次
