ELMER: A Non-Autoregressive Pre-trained Language Model for Efficient and Effective Text Generation
Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, Ji-Rong Wen
摘要
We study the text generation task under the approach of pre-trained language models (PLMs). Typically, an auto-regressive (AR) method is adopted for generating texts in a token-bytoken manner. Despite many advantages of AR generation, it usually suffers from inefficient inference. Therefore, non-autoregressive (NAR) models are proposed to generate all target tokens simultaneously. However, NAR models usually generate texts of lower quality due to the absence of token dependency in the output text. In this paper, we propose ELMER: an Efficient and effective PLM for NAR tExt geneRation to explicitly model the token dependency during NAR generation. By leveraging the early exit technique, ELMER enables the token generations at different layers, according to their prediction confidence (a more confident token will exit at a lower layer). Besides, we propose a novel pre-training objective, Layer Permutation Language Modeling, to pre-train ELMER by permuting the exit layer for each token in sequences. Experiments on three text generation tasks show that ELMER significantly outperforms NAR models and further narrows the performance gap with AR PLMs (e.g., ELMER (29.92) vs BART (30.61) ROUGE-L in XSUM) while achieving over 10 times inference speedup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- AR-Diffusion: Auto-Regressive Diffusion Model for Text GenerationTong Wu, Zhihao Fan, Xiao Liu, Hai-Tao Zheng 等NeurIPS 2023 · 被引用 170 次
- Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph DenoiseZhenghao Lin, Yeyun Gong, Yelong Shen, Tong Wu 等ICML 2023 · 被引用 107 次
- AMOM: Adaptive Masking over Masking for Conditional Masked Language ModelYisheng Xiao, Ruiyang Xu, Lijun Wu, Juntao Li 等AAAI 2023 · 被引用 14 次
- Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and BeyondSiyang Liu, Naihao Deng, Sahand Sabour, Yilin Jia 等EMNLP 2023 · 被引用 7 次
- Reflection-Window Decoding: Text Generation with Selective RefinementZeyu Tang, Zhenhao Chen, Xiangchen Song, Loka Li 等ICML 2025
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 被引用 264 次
- Non-autoregressive Translation with Layer-Wise Prediction and Deep SupervisionChenyang Huang, Hao Zhou, Osmar R. Zaïane, Lili Mou 等AAAI 2022 · 被引用 65 次
- BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale PretrainingWeizhen Qi, Yeyun Gong, Jian Jiao, Yu Yan 等ICML 2021 · 被引用 54 次
相关 Paper
- Dynamic and Efficient Inference for Text Generation via BERT FamilyXiaobo Liang, Juntao Li, Lijun Wu, Ziqiang Cao 等ACL 2023
- Open-ended Long Text Generation via Masked Language ModelingXiaobo Liang, Zecheng Tang, Juntao Li, Min ZhangACL 2023 · 被引用 10 次
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderYi Liao, Xin Jiang, Qun LiuACL 2020 · 被引用 28 次
- LeeBERT: Learned Early Exit for BERT with cross-level optimizationWei ZhuACL 2021
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
