Open-ended Long Text Generation via Masked Language Modeling
Xiaobo Liang, Zecheng Tang, Juntao Li, Min Zhang
摘要
Pre-trained autoregressive (AR) language models such as BART and GPTs have dominated Open-ended Long Text Generation (Open-LTG). However, the AR nature will decrease the inference efficiency along with the increase of generation length, which hinder their application in Open-LTG. To improve inference efficiency, we alternatively explore the potential of the pre-trained masked language models (MLMs) along with a representative iterative non-autoregressive (NAR) decoding strategy for Open-LTG. Our preliminary study shows that pre-trained MLMs can merely generate short text and will collapse for long text modeling. To enhance the long text generation capability of MLMs, we introduce two simple yet effective strategies for the iterative NAR model: dynamic sliding window attention (DSWA) and linear temperature decay (LTD). It can alleviate long-distance collapse problems and achieve longer text generation with a flexible trade-off between performance and inference speedup. Experiments on the storytelling and multi-paragraph opinionated article writing tasks show that pre-trained MLMs can achieve more than 3 × → 13 × speedup with better performance than strong AR models. Our code is available at GitHub * .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- LaViDa: A Large Diffusion Language Model for Multimodal UnderstandingShufan Li, Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul 等NeurIPS 2025 · 被引用 89 次
- Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed InteractionShiyan Hu, Jianxin Jin, Yang Shu, Peng Chen 等ICLR 2026 · 被引用 7 次
- FastLog: An End-to-End Method to Efficiently Generate and Insert Logging StatementsXiaoyuan Xie, Zhipeng Cai, Songqiang Chen, Jifeng XuanISSTA 2024 · 被引用 3 次
- Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language ModelsYi Zhou, Danushka Bollegala, José Camacho-ColladosEMNLP 2024 · 被引用 3 次
- Are Bert Family Good Instruction Followers? A Study on Their Potential And LimitationsYisheng Xiao, Juntao Li, Zechen Sun, Zechang Li 等ICLR 2024 · 被引用 2 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Content Planning for Neural Story Generation with Aristotelian RescoringSeraphina Goldfarb-Tarrant, Tuhin Chakrabarty, Ralph M. Weischedel, Nanyun PengEMNLP 2020 · 被引用 106 次
- Improving Non-Autoregressive Translation Models Without DistillationXiao Shi Huang, Felipe Pérez, Maksims VolkovsICLR 2022 · 被引用 60 次
- BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale PretrainingWeizhen Qi, Yeyun Gong, Jian Jiao, Yu Yan 等ICML 2021 · 被引用 54 次
相关 Paper
- Dynamic and Efficient Inference for Text Generation via BERT FamilyXiaobo Liang, Juntao Li, Lijun Wu, Ziqiang Cao 等ACL 2023
- XLM-D: Decorate Cross-lingual Pre-training Model as Non-Autoregressive Neural Machine TranslationYong Wang, Shilin He, Guanhua Chen, Yun Chen 等EMNLP 2022 · 被引用 4 次
- ELMER: A Non-Autoregressive Pre-trained Language Model for Efficient and Effective Text GenerationJunyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie 等EMNLP 2022 · 被引用 11 次
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderYi Liao, Xin Jiang, Qun LiuACL 2020 · 被引用 28 次
- Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical GuidelinesYuchen Li, Alexandre Kirchmeyer, Aashay Mehta, Yilong Qin 等ICML 2024 · 被引用 5 次
