Consistent Accelerated Inference via Confident Adaptive Transformers
Tal Schuster, Adam Fisch, Tommi S. Jaakkola, Regina Barzilay
Abstract
We develop a novel approach for confidently accelerating inference in the large and expensive multilayer Transformers that are now ubiquitous in natural language processing (NLP). Amortized or approximate computational methods increase efficiency, but can come with unpredictable performance costs. In this work, we present CATs – Confident Adaptive Transformers – in which we simultaneously increase computational efficiency, while guaranteeing a specifiable degree of consistency with the original model with high confidence. Our method trains additional prediction heads on top of intermediate layers, and dynamically decides when to stop allocating computational effort to each input using a meta consistency classifier. To calibrate our early prediction stopping rule, we formulate a unique extension of conformal prediction. We demonstrate the effectiveness of this approach on four classification and regression tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1d7133b-98f3-454b-8fd6-77bd18816288Cited by top-tier papers24
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei et al.ICLR 2024 · 242 citations
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala et al.ICLR 2024 · 132 citations
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 96 citations
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
- PAC Confidence Predictions for Deep Neural Network ClassifiersSangdon Park, Shuo Li, Insup Lee, Osbert BastaniICLR 2021 · 27 citations
Related papers
- Few-Shot Conformal Prediction with Auxiliary TasksAdam Fisch, Tal Schuster, Tommi S. Jaakkola, Regina BarzilayICML 2021 · 66 citations
- The Art of Abstention: Selective Prediction and Error Regularization for Natural Language ProcessingJi Xin, Raphael Tang, Yaoliang Yu, Jimmy LinACL 2021
- Efficient Conformal Prediction via Cascaded Inference with Expanded AdmissionAdam Fisch, Tal Schuster, Tommi S. Jaakkola, Regina BarzilayICLR 2021 · 53 citations
- CaTS: Calibrated Test-Time Scaling for Efficient LLM ReasoningChengsong Huang, Langlin Huang, Jixuan Leng, Jiacheng Liu et al.ICLR 2026
- Memory Efficient Continual Learning with TransformersBeyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal et al.NeurIPS 2022 · 75 citations
