Distilling Knowledge Learned in BERT for Text Generation
Yen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu, Jingjing Liu
Abstract
Large-scale pre-trained language model such as BERT has achieved great success in language understanding tasks. However, it remains an open question how to utilize BERT for language generation. In this paper, we present a novel approach, Conditional Masked Language Modeling (C-MLM), to enable the finetuning of BERT on target generation tasks. The finetuned BERT (teacher) is exploited as extra supervision to improve conventional Seq2Seq models (student) for better text generation performance. By leveraging BERT's idiosyncratic bidirectional nature, distilling knowledge learned in BERT can encourage auto-regressive Seq2Seq models to plan ahead, imposing global sequence-level supervision for coherent text generation. Experiments show that the proposed approach significantly outperforms strong Transformer baselines on multiple language generation tasks such as machine translation and text summarization. Our proposed model also achieves new state of the art on IWSLT German-English and English-Vietnamese MT datasets. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b2d67c0-6823-4c20-bee9-c9960df39815Cited by top-tier papers23
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin et al.NeurIPS 2022 · 222 citations
- Contrastive Triple Extraction with Generative TransformerHongbin Ye, Ningyu Zhang, Shumin Deng, Mosha Chen et al.AAAI 2021 · 146 citations
- Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic ModelsAviv Bick, Kevin Y. Li, Eric P. Xing, J. Zico Kolter et al.NeurIPS 2024 · 78 citations
- Graph-less Collaborative FilteringLianghao Xia, Chao Huang, Jiao Shi, Yong XuWWW 2023 · 60 citations
- Prophet Attention: Predicting Attention with Future AttentionFenglin Liu, Xuancheng Ren, Xian Wu, Shen Ge et al.NeurIPS 2020 · 52 citations
Builds on1
Related papers
- Universal Conditional Masked Language Pre-training for Neural Machine TranslationPengfei Li, Liangyou Li, Meng Zhang, Minghao Wu et al.ACL 2022 · 32 citations
- AMOM: Adaptive Masking over Masking for Conditional Masked Language ModelYisheng Xiao, Ruiyang Xu, Lijun Wu, Juntao Li et al.AAAI 2023 · 14 citations
- Incorporating BERT into Parallel Sequence Decoding with AdaptersJunliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei et al.NeurIPS 2020 · 72 citations
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned GenerationBin Bi, Chenliang Li, Chen Wu, Ming Yan et al.EMNLP 2020 · 41 citations
- BERTGen: Multi-task Generation through BERTFaidon Mitzalis, Ozan Caglayan, Pranava Madhyastha, Lucia SpeciaACL 2021
