Intra-Document Cascading: Learning to Select Passages for Neural Document Ranking
Sebastian Hofstätter, Bhaskar Mitra, Hamed Zamani, Nick Craswell, Allan Hanbury
Abstract
An emerging recipe for achieving state-of-the-art effectiveness in neural document re-ranking involves utilizing large pre-trained language models - e.g., BERT - to evaluate all individual passages in the document and then aggregating the outputs by pooling or additional Transformer layers. A major drawback of this approach is high query latency due to the cost of evaluating every passage in the document with BERT. To make matters worse, this high inference cost and latency varies based on the length of the document, with longer documents requiring more time and computation. To address this challenge, we adopt an intra-document cascading strategy, which prunes passages of a candidate document using a less expensive model, called ESM, before running a scoring model that is more expensive and effective, called ETM. We found it best to train ESM (short for Efficient Student Model) via knowledge distillation from the ETM (short for Effective Teacher Model) e.g., BERT. This pruning allows us to only run the ETM model on a smaller set of passages whose size does not vary by document length. Our experiments on the MS MARCO and TREC Deep Learning Track benchmarks suggest that the proposed Intra-Document Cascaded Ranking Model (IDCM) leads to over 400% lower query latency by providing essentially the same effectiveness as the state-of-the-art BERT-based document ranking models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe01eeb9-3651-445b-95c4-1f48f058b7f2Cited by top-tier papers7
- Efficient Neural Ranking using Forward IndexesJurek Leonhardt, Koustav Rudra, Megha Khosla, Abhijit Anand et al.WWW 2022 · 16 citations
- Efficient Re-ranking with Cross-encoders via Early ExitFrancesco Busolin, Claudio Lucchese, Franco Maria Nardini, Salvatore Orlando et al.SIGIR 2025 · 9 citations
- Socialformer: Social Network Inspired Long Document Modeling for Document RankingYujia Zhou, Zhicheng Dou, Huaying Yuan, Zhengyi MaWWW 2022 · 7 citations
- RankingSHAP - Faithful Listwise Feature Attribution Explanations for Ranking ModelsMaria Heuss, Maarten de Rijke, Avishek AnandSIGIR 2025 · 6 citations
- Pseudo-Relevance for Enhancing Document RepresentationJihyuk Kim, Seung-won Hwang, Seoho Song, Hyeseon Ko et al.EMNLP 2022 · 1 citation
Builds on1
Related papers
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- TILDE: Term Independent Likelihood moDEl for Passage Re-rankingShengyao Zhuang, Guido ZucconSIGIR 2021 · 86 citations
- SDR: Efficient Neural Re-ranking using Succinct Document RepresentationNachshon Cohen, Amit Portnoy, Besnik Fetahu, Amir IngberACL 2022
- Efficient Document Re-Ranking for Transformers by Precomputing Term RepresentationsSean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto et al.SIGIR 2020 · 62 citations
- A Unified Pretraining Framework for Passage Ranking and ExpansionMing Yan, Chenliang Li, Bin Bi, Wei Wang et al.AAAI 2021 · 14 citations
