Masked Language Model Scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, Katrin Kirchhoff
Abstract
Pretrained masked language models (MLMs) require finetuning for most NLP tasks. Instead, we evaluate MLMs out of the box via their pseudo-log-likelihood scores (PLLs), which are computed by masking tokens one by one. We show that PLLs outperform scores from autoregressive language models like GPT-2 in a variety of tasks. By rescoring ASR and NMT hypotheses, RoBERTa reduces an endto-end LibriSpeech model's WER by 30% relative and adds up to +1.7 BLEU on state-of-theart baselines for low-resource translation pairs, with further gains from domain adaptation. We attribute this success to PLL's unsupervised expression of linguistic acceptability without a left-to-right bias, greatly improving on scores from GPT-2 (+10 points on island effects, NPI licensing in BLiMP). One can finetune MLMs to give scores without masking, enabling computation in a single inference pass. In all, PLLs and their associated pseudo-perplexities (PP-PLs) enable plug-and-play use of the growing number of pretrained MLMs; e.g., we use a single cross-lingual model to rescore translations in multiple languages. We release our library for language model scoring at https: //github.com/awslabs/mlm-scoring.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65d480e9-3e1a-482b-8f1a-5deadd725b03Cited by top-tier papers89
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.NeurIPS 2022 · 305 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau et al.EMNLP 2021 · 177 citations
- Lift Yourself Up: Retrieval-augmented Text Generation with Self-MemoryXin Cheng, Di Luo, Xiuying Chen, Lemao Liu et al.NeurIPS 2023 · 177 citations
- An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language ModelsNicholas Meade, Elinor Poole-Dayan, Siva ReddyACL 2022 · 160 citations
Builds on2
Related papers
- Improving Language Plasticity via Pretraining with Active ForgettingYihong Chen, Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani et al.NeurIPS 2023 · 49 citations
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderYi Liao, Xin Jiang, Qun LiuACL 2020 · 28 citations
- Simple and Effective Unsupervised Speech TranslationChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov et al.ACL 2023 · 9 citations
- Masking as an Efficient Alternative to Finetuning for Pretrained Language ModelsMengjie Zhao, Tao Lin, Fei Mi, Martin Jaggi et al.EMNLP 2020 · 61 citations
- When Do You Need Billions of Words of Pretraining Data?Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. BowmanACL 2021
