The Cascade Transformer: an Application for Efficient Answer Sentence Selection
Luca Soldaini, Alessandro Moschitti
Abstract
Large transformer-based language models have been shown to be very effective in many classification tasks. However, their computational complexity prevents their use in applications requiring the classification of a large set of candidates. While previous works have investigated approaches to reduce model size, relatively little attention has been paid to techniques to improve batch throughput during inference. In this paper, we introduce the Cascade Transformer, a simple yet effective technique to adapt transformer-based models into a cascade of rankers. Each ranker is used to prune a subset of candidates in a batch, thus dramatically increasing throughput at inference time. Partial encodings from the transformer model are shared among rerankers, providing further speed-up. When compared to a state-of-the-art transformer model, our approach reduces computation by 37% with almost no impact on accuracy, as measured on two English Question Answering datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4a27368-d52d-4679-be35-10841cf04122Cited by top-tier papers8
- Path Independent Equilibrium Models Can Better Exploit Test-Time ComputationCem Anil, Ashwini Pokle, Kaiqu Liang, Johannes Treutlein et al.NeurIPS 2022 · 43 citations
- DeeCap: Dynamic Early Exiting for Efficient Image CaptioningZhengcong Fei, Xu Yan, Shuhui Wang, Qi TianCVPR 2022 · 39 citations
- Model Cascading: Towards Jointly Improving Efficiency and Accuracy of NLP SystemsNeeraj Varshney, Chitta BaralEMNLP 2022 · 11 citations
- Certified Error Control of Candidate Set Pruning for Two-Stage Relevance RankingMinghan Li, Xinyu Zhang, Ji Xin, Hongyang Zhang et al.EMNLP 2022 · 3 citations
- Timing Channels in Adaptive Neural NetworksAyomide Akinsanya, Tegan BrennanNDSS 2024
Builds on2
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence SelectionSiddhant Garg, Thuy Vu, Alessandro MoschittiAAAI 2020 · 229 citations
Related papers
- Block-Skim: Efficient Question Answering for TransformerYue Guan, Zhengyi Li, Zhouhan Lin, Yuhao Zhu et al.AAAI 2022 · 33 citations
- Dynamic Context Pruning for Efficient and Interpretable Autoregressive TransformersSotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci et al.NeurIPS 2023 · 95 citations
- Efficient Document Re-Ranking for Transformers by Precomputing Term RepresentationsSean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto et al.SIGIR 2020 · 62 citations
- DeFormer: Decomposing Pre-trained Transformers for Faster Question AnsweringQingqing Cao, Harsh Trivedi, Aruna Balasubramanian, Niranjan BalasubramanianACL 2020 · 61 citations
- EEL: Efficiently Encoding Lattices for RerankingPrasann Singhal, Jiacheng Xu, Xi Ye, Greg DurrettACL 2023
