Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word Order
Yi Liao, Xin Jiang, Qun Liu
Abstract
Masked language model and autoregressive language model are two types of language models. While pretrained masked language models such as BERT (Devlin et al., 2019) overwhelm the line of natural language understanding (NLU) tasks, autoregressive language models such as GPT (Radford et al., 2018) are especially capable in natural language generation (NLG). In this paper, we propose a probabilistic masking scheme for the masked language model, which we call probabilistically masked language model (PMLM). We implement a specific PMLM with a uniform prior distribution on the masking ratio named u-PMLM. We prove that u-PMLM is equivalent to an autoregressive permutated language model. One main advantage of the model is that it supports text generation in arbitrary order with surprisingly good quality, which could potentially enable new applications over traditional unidirectional generation. Besides, the pretrained u-PMLM also outperforms BERT on a set of downstream NLU tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8645cd72-15aa-4302-9091-2e1ec5bb7e19Cited by top-tier papers10
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal TransformersJaemin Cho, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi et al.EMNLP 2020 · 80 citations
- UFC-BERT: Unifying Multi-Modal Controls for Conditional Image SynthesisZhu Zhang, Jianxin Ma, Chang Zhou, Rui Men et al.NeurIPS 2021 · 43 citations
- BLens: Contrastive Captioning of Binary Functions using Ensemble EmbeddingTristan Benoit, Yunru Wang, Moritz Dannehl, Johannes KinderUSENIX Security 2025
- Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked DiffusionsJaeyeon Kim, Kulin Shah, Vasilis Kontonis, Sham M. Kakade et al.ICML 2025
Builds on3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu et al.ICLR 2020 · 297 citations
Related papers
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- DiffusionBERT: Improving Generative Masked Language Models with Diffusion ModelsZhengfu He, Tianxiang Sun, Qiong Tang, Kuanning Wang et al.ACL 2023 · 63 citations
- GLM: General Language Model Pretraining with Autoregressive Blank InfillingZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding et al.ACL 2022
- Distilling Knowledge Learned in BERT for Text GenerationYen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu et al.ACL 2020 · 116 citations
- XLM-D: Decorate Cross-lingual Pre-training Model as Non-Autoregressive Neural Machine TranslationYong Wang, Shilin He, Guanhua Chen, Yun Chen et al.EMNLP 2022 · 4 citations
