MLE-Guided Parameter Search for Task Loss Minimization in Neural Sequence Modeling
Sean Welleck, Kyunghyun Cho
Abstract
Neural autoregressive sequence models are used to generate sequences in a variety of natural language processing (NLP) tasks, where they are evaluated according to sequence-level task losses. These models are typically trained with maximum likelihood estimation, which ignores the task loss, yet empirically performs well as a surrogate objective. Typical approaches to directly optimizing the task loss such as policy gradient and minimum risk training are based around sampling in the sequence space to obtain candidate update directions that are scored based on the loss of a single sequence. In this paper, we develop an alternative method based on random search in the parameter space that leverages access to the maximum likelihood gradient. We propose maximum likelihood guided parameter search (MGS), which samples from a distribution over update directions that is a mixture of random search around the current parameters and around the maximum likelihood gradient, with each direction weighted by its improvement in the task loss. MGS shifts sampling to the parameter space, and scores candidates using losses that are pooled from multiple sequences. Our experiments show that MGS is capable of optimizing sequence-level losses, with substantial reductions in repetition and non-termination in sequence completion, and similar improvements to those of minimum risk training in machine translation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9728362-df26-481d-be19-086e44fb9958Cited by top-tier papers2
- ZARTS: On Zero-order Optimization for Neural Architecture SearchXiaoxing Wang, Wenxuan Guo, Jianlin Su, Xiaokang Yang et al.NeurIPS 2022 · 37 citations
- Learning to Learn Transferable AttackShuman Fang, Jie Li, Xianming Lin, Rongrong JiAAAI 2022 · 26 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
- Residual Energy-Based Models for Text GenerationYuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam et al.ICLR 2020 · 147 citations
Related papers
- On the Weaknesses of Reinforcement Learning for Neural Machine TranslationLeshem Choshen, Lior Fox, Zohar Aizenbud, Omri AbendICLR 2020 · 124 citations
- SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with BacktrackingChris Cundy, Stefano ErmonICLR 2024 · 17 citations
- Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--HastingsKartik Goyal, Chris Dyer, Taylor Berg-KirkpatrickICLR 2022 · 53 citations
- Learning Extrapolative Sequence Transformations from Markov ChainsSophia Hager, Aleem Khan, Andrew Wang, Nicholas AndrewsICML 2025
- Straight to the Gradient: Learning to Use Novel Tokens for Neural Text GenerationXiang Lin, Simeng Han, Shafiq R. JotyICML 2021 · 30 citations
