Certified Error Control of Candidate Set Pruning for Two-Stage Relevance Ranking
Minghan Li, Xinyu Zhang, Ji Xin, Hongyang Zhang, Jimmy Lin
摘要
In information retrieval (IR), candidate set pruning has been commonly used to speed up two-stage relevance ranking. However, such an approach lacks accurate error control and often trades accuracy against computational efficiency in an empirical fashion, missing theoretical guarantees. In this paper, we propose the concept of certified error control of candidate set pruning for relevance ranking, which means that the test error after pruning is guaranteed to be controlled under a user-specified threshold with high probability. Both in-domain and out-of-domain experiments show that our method successfully prunes the first-stage retrieved candidate sets to improve the second-stage reranking speed while satisfying the pre-specified accuracy constraints in both settings. For example, on MS MARCO Passage v1, our method reduces the average candidate set size from 1000 to 27, increasing reranking speed by about 37 times, while keeping MRR@10 greater than a pre-specified value of 0.38 with about 90% empirical coverage. In contrast, empirical baselines fail to meet such requirements. Code and data are available at: https://github.com/alexlimh/CEC-Ranking.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- PAC Confidence Sets for Deep Neural Networks via Calibrated PredictionSangdon Park, Osbert Bastani, Nikolai Matni, Insup LeeICLR 2020 · 被引用 77 次
相关 Paper
- Threshold-driven Pruning with Segmented Maximum Term Weights for Approximate Cluster-based Sparse RetrievalYifan Qiao, Parker Carlson, Shanxiu He, Yingrui Yang 等EMNLP 2024 · 被引用 7 次
- FIRST: Faster Improved Listwise Reranking with Single Token DecodingRevanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md. Arafat Sultan 等EMNLP 2024 · 被引用 14 次
- Provence: efficient and robust context pruning for retrieval-augmented generationNadezhda Chirkova, Thibault Formal, Vassilina Nikoulina, Stéphane ClinchantICLR 2025 · 被引用 2 次
- Efficient Sparse Retrieval with Lightweight Superblock PruningParker Carlson, Wentai Xie, Rohil Shah, Tao YangSIGIR 2026 · 被引用 2 次
- CODER: An efficient framework for improving retrieval through COntextual Document Embedding RerankingGeorge Zerveas, Navid Rekabsaz, Daniel Cohen, Carsten EickhoffEMNLP 2022 · 被引用 10 次
