Language GANs Falling Short
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau, Laurent Charlin
Abstract
Generating high-quality text with sufficient diversity is essential for a wide range of Natural Language Generation (NLG) tasks. Maximum-Likelihood (MLE) models trained with teacher forcing have consistently been reported as weak baselines, where poor performance is attributed to exposure bias (Bengio et al., 2015; Ranzato et al., 2015); at inference time, the model is fed its own prediction instead of a ground-truth token, which can lead to accumulating errors and poor samples. This line of reasoning has led to an outbreak of adversarial based approaches for NLG, on the account that GANs do not suffer from exposure bias. In this work, we make several surprising observations which contradict common beliefs. First, we revisit the canonical evaluation framework for NLG, and point out fundamental flaws with quality-only evaluation: we show that one can outperform such metrics using a simple, well-known temperature parameter to artificially reduce the entropy of the model's conditional distributions. Second, we leverage the control over the quality / diversity trade-off given by this parameter to evaluate models over the whole quality-diversity spectrum and find MLE models constantly outperform the proposed GAN variants over the whole quality-diversity space. Our results have several implications: 1) The impact of exposure bias on sample quality is less severe than previously thought, 2) temperature tuning provides a better quality / diversity trade-off than adversarial training while being easier to train, easier to cross-validate, and less computationally expensive. Code to reproduce the experiments is available at github.com/pclucas14/GansFallingShort
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c2d4b35-def3-4d07-9f22-2bc0c0aead95Cited by top-tier papers61
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Self-Evaluation Guided Beam Search for ReasoningYuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao et al.NeurIPS 2023 · 316 citations
- Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data HidingSahar Abdelnabi, Mario FritzS&P 2021 · 210 citations
Related papers
- ColdGANs: Taming Language GANs with Cautious Sampling StrategiesThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski et al.NeurIPS 2020 · 19 citations
- Meta-CoTGAN: A Meta Cooperative Training Paradigm for Improving Adversarial Text GenerationHaiyan Yin, Dingcheng Li, Xu Li, Ping LiAAAI 2020 · 24 citations
- Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation?Tianxing He, Jingzhao Zhang, Zhiming Zhou, James R. GlassEMNLP 2021 · 12 citations
- Are GANs overkill for NLP?David Alvarez-Melis, Vikas Garg, Adam KalaiNeurIPS 2022 · 2 citations
- G2: Guided Generation for Enhanced Output Diversity in LLMsZhiwen Ruan, Yixia Li, Yefeng Liu, Yun Chen et al.EMNLP 2025
