AdapLeR: Speeding up Inference by Adaptive Length Reduction
Ali Modarressi, Hosein Mohebbi, Mohammad Taher Pilehvar
摘要
Pre-trained language models have shown stellar performance in various downstream tasks. But, this usually comes at the cost of high latency and computation, hindering their usage in resource-limited settings. In this work, we propose a novel approach for reducing the computational cost of BERT with minimal loss in downstream performance. Our method dynamically eliminates less contributing tokens through layers, resulting in shorter lengths and consequently lower computational cost. To determine the importance of each token representation, we train a Contribution Predictor for each layer using a gradient-based saliency method. Our experiments on several diverse classification tasks show speedups up to 22x during inference time without much sacrifice in performance. We also validate the quality of the selected tokens in our method using human annotations in the ERASER benchmark. In comparison to other widely used strategies for selecting important tokens, such as saliency and attention, our proposed method has a significantly lower false positive rate in generating rationales. Our code is freely available at https://github.com/amodaresi/ AdapLeR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- C3oT: Generating Shorter Chain-of-Thought Without Compromising EffectivenessYu Kang, Xianghui Sun, Liangyu Chen, Wei ZouAAAI 2025 · 被引用 162 次
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language ModelsHuiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang 等EMNLP 2023 · 被引用 94 次
- LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt CompressionHuiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li 等ACL 2024 · 被引用 59 次
- TextFusion: Privacy-Preserving Pre-trained Model Inference via Token FusionXin Zhou, Jinzhu Lu, Tao Gui, Ruotian Ma 等EMNLP 2022 · 被引用 12 次
- Model Cascading: Towards Jointly Improving Efficiency and Accuracy of NLP SystemsNeeraj Varshney, Chitta BaralEMNLP 2022 · 被引用 11 次
它引用的顶会 Paper16
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma 等AAAI 2020 · 被引用 656 次
- PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector EliminationSaurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy 等ICML 2020 · 被引用 260 次
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
相关 Paper
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu 等ACL 2022
- Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer InferenceJunyan Li, Li Lyna Zhang, Jiahang Xu, Yujing Wang 等KDD 2023 · 被引用 12 次
- Pyramid-BERT: Reducing Complexity via Successive Core-set based Token SelectionXin Huang, Ashish Khetan, Rene Bidart, Zohar S. KarninACL 2022
- Fine- and Coarse-Granularity Hybrid Self-Attention for Efficient BERTJing Zhao, Yifan Wang, Junwei Bao, Youzheng Wu 等ACL 2022 · 被引用 7 次
- SkipBERT: Efficient Inference with Shallow Layer SkippingJue Wang, Ke Chen, Gang Chen, Lidan Shou 等ACL 2022
