ColdGANs: Taming Language GANs with Cautious Sampling Strategies
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano
Abstract
Training regimes based on Maximum Likelihood Estimation (MLE) suffer from known limitations, often leading to poorly generated text sequences. At the root of these limitations is the mismatch between training and inference, i.e. the so-called exposure bias, exacerbated by considering only the reference texts as correct, while in practice several alternative formulations could be as good. Generative Adversarial Networks (GANs) can mitigate those limitations but the discrete nature of text has hindered their application to language generation: the approaches proposed so far, based on Reinforcement Learning, have been shown to underperform MLE. Departing from previous works, we analyze the exploration step in GANs applied to text generation, and show how classical sampling results in unstable training. We propose to consider alternative exploration strategies in a GAN framework that we name ColdGAN s, where we force the sampling to be close to the distribution modes to get smoother learning dynamics. For the first time, to the best of our knowledge, the proposed language GANs compare favorably to MLE, and obtain improvements over the state-of-the-art on three generative tasks, namely unconditional text generation, question generation, and abstractive summarization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e294554-5928-4862-b268-a5b63fbfea80Cited by top-tier papers4
- To Beam Or Not To Beam: That is a Question of Cooperation for Language GANsThomas Scialom, Paul-Alexis Dray, Jacopo Staiano, Sylvain Lamprier et al.NeurIPS 2021 · 23 citations
- Generative Cooperative Networks for Natural Language GenerationSylvain Lamprier, Thomas Scialom, Antoine Chaffin, Vincent Claveau et al.ICML 2022 · 13 citations
- Adaptive Bridge between Training and Inference for Dialogue GenerationHaoran Xu, Hainan Zhang, Yanyan Zou, Hongshen Chen et al.EMNLP 2021 · 5 citations
- Branch-GAN: Improving Text Generation with (not so) Large Language ModelsFredrik Carlsson, Johan Broberg, Erik Hillbom, Magnus Sahlgren et al.ICLR 2024 · 3 citations
Builds on6
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
- Discriminative Adversarial Search for Abstractive SummarizationThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski et al.ICML 2020 · 37 citations
Related papers
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 88 citations
- Improving Text Generation with Student-Forcing Optimal TransportJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu et al.EMNLP 2020 · 11 citations
- Improving GAN Training with Probability Ratio Clipping and Sample ReweightingYue Wu, Pan Zhou, Andrew Gordon Wilson, Eric P. Xing et al.NeurIPS 2020 · 39 citations
- On the Weaknesses of Reinforcement Learning for Neural Machine TranslationLeshem Choshen, Lior Fox, Zohar Aizenbud, Omri AbendICLR 2020 · 124 citations
- Meta-CoTGAN: A Meta Cooperative Training Paradigm for Improving Adversarial Text GenerationHaiyan Yin, Dingcheng Li, Xu Li, Ping LiAAAI 2020 · 24 citations
