To Beam Or Not To Beam: That is a Question of Cooperation for Language GANs
Thomas Scialom, Paul-Alexis Dray, Jacopo Staiano, Sylvain Lamprier, Benjamin Piwowarski
Abstract
Due to the discrete nature of words, language GANs require to be optimized from rewards provided by discriminator networks, via reinforcement learning methods. This is a much harder setting than for continuous tasks, which enjoy gradient flows from discriminators to generators, usually leading to dramatic learning instabilities. However, we claim that this can be solved by making discriminator and generator networks cooperate to produce output sequences during training. These cooperative outputs, inherently built to obtain higher discrimination scores, not only provide denser rewards for training, but also form a more compact artificial set for discriminator training, hence improving its accuracy and stability. In this paper, we show that our SelfGAN framework, built on this cooperative principle, outperforms Teacher Forcing and obtains state-of-the-art results on two challenging tasks, Summarization and Question Generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02eeb74f-04ce-46ed-b08d-1f9e686fe774Cited by top-tier papers9
- Controlled Decoding from Language ModelsSidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li et al.ICML 2024 · 130 citations
- Transformer-based Planning for Symbolic RegressionParshin Shojaee, Kazem Meidani, Amir Barati Farimani, Chandan K. ReddyNeurIPS 2023 · 116 citations
- Solving Math Word Problems via Cooperative Reasoning induced Language ModelsXinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang et al.ACL 2023 · 16 citations
- Planning with Large Language Models for Code GenerationShun Zhang, Zhenfang Chen, Yikang Shen, Mingyu Ding et al.ICLR 2023 · 15 citations
- Taming Imperfect Process Verifiers: A Sampling Perspective on BacktrackingDhruv Rohatgi, Abhishek Shetty, Donya Saless, Yuchen Li et al.ICLR 2026 · 15 citations
Builds on8
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
- Self-Distillation as Instance-Specific Label SmoothingZhilu Zhang, Mert R. SabuncuNeurIPS 2020 · 155 citations
- Residual Energy-Based Models for Text GenerationYuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam et al.ICLR 2020 · 147 citations
- Discriminative Adversarial Search for Abstractive SummarizationThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski et al.ICML 2020 · 37 citations
Related papers
- Generative Cooperative Networks for Natural Language GenerationSylvain Lamprier, Thomas Scialom, Antoine Chaffin, Vincent Claveau et al.ICML 2022 · 13 citations
- ColdGANs: Taming Language GANs with Cautious Sampling StrategiesThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski et al.NeurIPS 2020 · 19 citations
- Improving GAN Training with Probability Ratio Clipping and Sample ReweightingYue Wu, Pan Zhou, Andrew Gordon Wilson, Eric P. Xing et al.NeurIPS 2020 · 39 citations
- Self-Adversarial Learning with Comparative Discrimination for Text GenerationWangchunshu Zhou, Tao Ge, Ke Xu, Furu Wei et al.ICLR 2020 · 20 citations
- Amalgamating Knowledge from Two Teachers for Task-oriented Dialogue System with Adversarial TrainingWanwei He, Min Yang, Rui Yan, Chengming Li et al.EMNLP 2020 · 22 citations
