CrossQ: Task-Aligned Cross-Token Conditional Quantization for Late Interaction Retrieval
Rohit Kumar Salla, Manoj Saravanan, Ramya Amancherla
Abstract
Late-interaction retrievers like ColBERT achieve high quality but suffer from large multi-vector indices. Standard compression minimizes token reconstruction error, while ranking depends critically on preserving scores of sparse "winner" tokens. We introduce CrossQ, which adaptively improves effective token fidelity within documents by conditioning token codes on lightweight document context computed at indexing time (but not stored). CrossQ is trained with ranking-aligned objectives that preserve candidate score distributions and protect hard-negative margins. At 2 B/token, CrossQ improves MRR@10 by +0.010 over the strongest strictly footprint-matched quantization baseline and by +0.012 over the strongest candidate-matched system reference. On the nine-dataset BEIR subset reported in Appendix G.1, CrossQ improves average nDCG@10 by +0.009 at 4 B/token over the strongest candidate-matched system reference. At 4 B/token, CrossQ achieves raw token-storage reduction, approximately including metadata and approximately under conservative padding/alignment accounting. At 8 B/token, CrossQ + light fine-tuning retains approximately 98% of full-precision ColBERT MRR@10, improving the footprint-quality tradeoff for memory-constrained late-interaction retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1d89526-52b1-4573-b7fd-22ed60460bfbBuilds on2
- Progressively Optimized Bi-Granular Document Representation for Scalable Embedding Based RetrievalShitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang et al.WWW 2022 · 19 citations
- Qinco2: Vector Compression and Search with Improved Implicit Neural CodebooksThéophane Vallaeys, Matthew J. Muckley, Jakob Verbeek, Matthijs DouzeICLR 2025
Related papers
- Rethinking the Role of Token Retrieval in Multi-Vector RetrievalJinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei et al.NeurIPS 2023 · 60 citations
- Towards Lossless Token Pruning in Late-Interaction Retrieval ModelsYuxuan Zong, Benjamin PiwowarskiSIGIR 2025 · 2 citations
- No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector RetrievalLixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng et al.ICML 2026
- Compact Token Representations with Contextual Quantization for Efficient Document Re-rankingYingrui Yang, Yifan Qiao, Tao YangACL 2022 · 8 citations
- Incorporating Token Importance in Multi-Vector RetrievalArchish S, Ankit Garg, Kirankumar Shiragur, Neeraj KayalAAAI 2026 · 2 citations
