GhostBERT: Generate More Features with Cheap Operations for BERT
Zhiqi Huang, Lu Hou, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu
摘要
Transformer-based pre-trained language models like BERT, though powerful in many tasks, are expensive in both memory and computation, due to their large number of parameters. Previous works show that some parameters in these models can be pruned away without severe accuracy drop. However, these redundant features contribute to a comprehensive understanding of the training data and removing them weakens the model's representation ability. In this paper, we propose GhostBERT, which generates more features with very cheap operations from the remaining features. In this way, GhostBERT has similar memory and computational cost as the pruned model, but enjoys much larger representation power. The proposed ghost module can also be applied to unpruned BERT models to enhance their performance with negligible additional parameters and computation. Empirical results on the GLUE benchmark on three backbone models (i.e., BERT, RoBERTa and ELECTRA) verify the efficacy of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Efficient Post-training Quantization of Pre-trained Language ModelsHaoli Bai, Lu Hou, Lifeng Shang, Xin Jiang 等NeurIPS 2022 · 被引用 62 次
- Efficient Computation Sharing for Multi-Task Visual Scene UnderstandingSara Shoouri, Mingyu Yang, Zichen Fan, Hun-Seok KimICCV 2023 · 被引用 9 次
- E-LANG: Energy-Based Joint Inferencing of Super and Swift Language ModelsMohammad Akbari, Amin Banitalebi-Dehkordi, Yong ZhangACL 2022
它引用的顶会 Paper17
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma 等AAAI 2020 · 被引用 656 次
相关 Paper
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 被引用 126 次
- PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector EliminationSaurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy 等ICML 2020 · 被引用 260 次
- bert2BERT: Towards Reusable Pretrained Language ModelsCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang 等ACL 2022
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture SearchJin Xu, Xu Tan, Renqian Luo, Kaitao Song 等KDD 2021 · 被引用 49 次
