ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut
Abstract
Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases become harder due to GPU/TPU memory limitations and longer training times. To address these problems, we present two parameter-reduction techniques to lower memory consumption and increase the training speed of BERT. Comprehensive empirical evidence shows that our proposed methods lead to models that scale much better compared to the original BERT. We also use a self-supervised loss that focuses on modeling inter-sentence coherence, and show it consistently helps downstream tasks with multi-sentence inputs. As a result, our best model establishes new state-of-the-art results on the GLUE, RACE, and benchmarks while having fewer parameters compared to BERT-large. The code and the pretrained models are available at this https URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16431d55-c204-4bab-9bd4-89ddd6f00586Cited by top-tier papers846
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Self-supervised Graph Learning for RecommendationJiancan Wu, Xiang Wang, Fuli Feng, Xiangnan He et al.SIGIR 2021 · 1,476 citations
Builds on2
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu et al.ICLR 2020 · 297 citations
- DCMN+: Dual Co-Matching Network for Multi-Choice Reading ComprehensionShuailiang Zhang, Hai Zhao, Yuwei Wu, Zhuosheng Zhang et al.AAAI 2020 · 138 citations
Related papers
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture SearchJin Xu, Xu Tan, Renqian Luo, Kaitao Song et al.KDD 2021 · 49 citations
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 126 citations
- schuBERT: Optimizing Elements of BERTAshish Khetan, Zohar S. KarninACL 2020 · 2 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- BinaryBERT: Pushing the Limit of BERT QuantizationHaoli Bai, Wei Zhang, Lu Hou, Lifeng Shang et al.ACL 2021
