DeFormer: Decomposing Pre-trained Transformers for Faster Question Answering
Qingqing Cao, Harsh Trivedi, Aruna Balasubramanian, Niranjan Balasubramanian
摘要
Transformer-based QA models use input-wide self-attention -i.e. across both the question and the input passage -at all layers, causing them to be slow and memory-intensive. It turns out that we can get by without inputwide self-attention at all layers, especially in the lower layers. We introduce DeFormer, a decomposed transformer, which substitutes the full self-attention with question-wide and passage-wide self-attentions in the lower layers. This allows for question-independent processing of the input text representations, which in turn enables pre-computing passage representations reducing runtime compute drastically. Furthermore, because DeFormer is largely similar to the original model, we can initialize DeFormer with the pre-training weights of a standard transformer, and directly fine-tune on the target QA dataset. We show DeFormer versions of BERT and XLNet can be used to speed up QA by over 4.3x and with simple distillation-based losses they incur only a 1% drop in accuracy. We open source the code at https://github.com/ StonyBrookNLP/deformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Which *BERT? A Survey Organizing Contextualized EncodersPatrick Xia, Shijie Wu, Benjamin Van DurmeEMNLP 2020 · 被引用 44 次
- Block-Skim: Efficient Question Answering for TransformerYue Guan, Zhengyi Li, Zhouhan Lin, Yuhao Zhu 等AAAI 2022 · 被引用 33 次
- Exploring extreme parameter compression for pre-trained language modelsBenyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang 等ICLR 2022 · 被引用 23 次
- Distilled Dual-Encoder Model for Vision-Language UnderstandingZekun Wang, Wenhui Wang, Haichao Zhu, Ming Liu 等EMNLP 2022 · 被引用 22 次
- Differentially Private Model CompressionFatemehsadat Mireshghallah, Arturs Backurs, Huseyin A. Inan, Lukas Wutschitz 等NeurIPS 2022 · 被引用 18 次
相关 Paper
- Transformer-VQ: Linear-Time Transformers via Vector QuantizationLucas D. LingleICLR 2024 · 被引用 30 次
- You Only Need One Model for Open-domain Question AnsweringHaejun Lee, Akhil Kedia, Jongwon Lee, Ashwin Paranjape 等EMNLP 2022
- AQ-DETR: Low-Bit Quantized Detection Transformer with Auxiliary QueriesRunqi Wang, Huixin Sun, Linlin Yang, Shaohui Lin 等AAAI 2024 · 被引用 10 次
- ReadOnce Transformers: Reusable Representations of Text for TransformersShih-Ting Lin, Ashish Sabharwal, Tushar KhotACL 2021
- Sparse is Enough in Scaling TransformersSebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser 等NeurIPS 2021 · 被引用 127 次
