Towards Lossless Token Pruning in Late-Interaction Retrieval Models
Yuxuan Zong, Benjamin Piwowarski
Abstract
Late interaction neural IR models like ColBERT offer a competitive effectiveness-efficiency trade-off across many benchmarks. However, they require a huge memory space to store the contextual representation for all the document tokens. Some works have proposed using either heuristics or statistical-based techniques to prune tokens from each document. This however doesn't guarantee that the removed tokens have no impact on the retrieval score. Our work uses a principled approach to define how to prune tokens without impacting the score between a document and a query. We introduce three regularization losses, that induce a solution with high pruning ratios, as well as two pruning strategies. We study them experimentally (in and out-domain), showing that we can preserve ColBERT's performance while using only 30% of the tokens.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b45531a1-1bb4-4941-aedd-d4fb07eccb74Cited by top-tier papers2
- Multi-Vector Index Compression in Any ModalityHanxiang Qin, Alexander Martin, Rohan Jha, Chunsheng Zuo et al.SIGIR 2026 · 11 citations
- A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval ModelsYash Kankanampati, Yuxuan Zong, Nadi Tomeh, Benjamin Piwowarski et al.SIGIR 2026
Builds on9
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin et al.SIGIR 2021 · 297 citations
- RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-rankingRuiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao et al.EMNLP 2021 · 147 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Multi-View Document Representation Learning for Open-Domain Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang et al.ACL 2022 · 80 citations
Related papers
- CrossQ: Task-Aligned Cross-Token Conditional Quantization for Late Interaction RetrievalRohit Kumar Salla, Manoj Saravanan, Ramya AmancherlaICML 2026
- Rethinking the Role of Token Retrieval in Multi-Vector RetrievalJinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei et al.NeurIPS 2023 · 60 citations
- Compact Token Representations with Contextual Quantization for Efficient Document Re-rankingYingrui Yang, Yifan Qiao, Tao YangACL 2022 · 8 citations
- Incorporating Token Importance in Multi-Vector RetrievalArchish S, Ankit Garg, Kirankumar Shiragur, Neeraj KayalAAAI 2026 · 2 citations
- Dynamic Context Pruning for Efficient and Interpretable Autoregressive TransformersSotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci et al.NeurIPS 2023 · 95 citations
