Block-Skim: Efficient Question Answering for Transformer
Yue Guan, Zhengyi Li, Zhouhan Lin, Yuhao Zhu, Jingwen Leng, Minyi Guo
摘要
Transformer models have achieved promising results on natural language processing (NLP) tasks including extractive question answering (QA). Common Transformer encoders used in NLP tasks process the hidden states of all input tokens in the context paragraph throughout all layers. However, different from other tasks such as sequence classification, answering the raised question does not necessarily need all the tokens in the context paragraph. Following this motivation, we propose Block-skim, which learns to skim unnecessary context in higher hidden layers to improve and accelerate the Transformer performance. The key idea of Block-Skim is to identify the context that must be further processed and those that could be safely discarded early on during inference. Critically, we find that such information could be sufficiently derived from the self-attention weights inside the Transformer model. We further prune the hidden states corresponding to the unnecessary positions early in lower layers, achieving significant inference-time speedup. To our surprise, we observe that models pruned in this way outperform their full-size counterparts. Block-Skim improves QA models' accuracy on different datasets and achieves 3 times speedup on BERT-base model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair QuantizationCong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng 等ISCA 2023 · 被引用 151 次
- ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationCong Guo, Chen Zhang, Jingwen Leng, Zihan Liu 等MICRO 2022 · 被引用 109 次
- SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian ApproximationCong Guo, Yuxian Qiu, Jingwen Leng, Xiaotian Gao 等ICLR 2022 · 被引用 92 次
- Nimbus: Secure and Efficient Two-Party Inference for TransformersZhengyi Li, Kang Yang, Jin Tan, Wen-jie Lu 等NeurIPS 2024 · 被引用 34 次
- A Survey for Efficient Open Domain Question AnsweringQin Zhang, Shangsi Chen, Dongkuan Xu, Qingqing Cao 等ACL 2023 · 被引用 28 次
它引用的顶会 Paper8
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Lite Transformer with Long-Short Range AttentionZhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin 等ICLR 2020 · 被引用 379 次
- SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian ApproximationCong Guo, Yuxian Qiu, Jingwen Leng, Xiaotian Gao 等ICLR 2022 · 被引用 92 次
- Accelerating sparse DNN models without hardware-support via tile-wise sparsityCong Guo, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu 等SC 2020 · 被引用 65 次
相关 Paper
- Transkimmer: Transformer Learns to Layer-wise SkimYue Guan, Zhengyi Li, Jingwen Leng, Zhouhan Lin 等ACL 2022
- DeFormer: Decomposing Pre-trained Transformers for Faster Question AnsweringQingqing Cao, Harsh Trivedi, Aruna Balasubramanian, Niranjan BalasubramanianACL 2020 · 被引用 61 次
- SkipBERT: Efficient Inference with Shallow Layer SkippingJue Wang, Ke Chen, Gang Chen, Lidan Shou 等ACL 2022
- Dynamic Context Pruning for Efficient and Interpretable Autoregressive TransformersSotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci 等NeurIPS 2023 · 被引用 95 次
- Block Pruning For Faster TransformersFrançois Lagunas, Ella Charlaix, Victor Sanh, Alexander M. RushEMNLP 2021 · 被引用 2 次
