TextGAIL: Generative Adversarial Imitation Learning for Text Generation
Qingyang Wu, Lei Li, Zhou Yu
Abstract
Generative Adversarial Networks (GANs) for text generation have recently received many criticisms, as they perform worse than their MLE counterparts (Caccia et al. 2020; Tevet et al. 2019; Semeniuta, Severyn, and Gelly 2018) . We suspect previous text GANs' inferior performance is due to the lack of a reliable guiding signal in their discriminators. To address this problem, we propose a generative adversarial imitation learning framework for text generation that uses large pre-trained language models to provide more reliable reward guidance. As previous text GANs suffer from high variance of gradients, we apply contrastive discriminator, and proximal policy optimization (PPO) to stabilize and improve text generation performance. For evaluation, we conduct experiments on a diverse set of unconditional and conditional text generation tasks. Experimental results show that TextGAIL achieves better performance in terms of both quality and diversity than the MLE baseline. We also validate our intuition that TextGAIL's discriminator demonstrates the capability of providing reasonable rewards with an additional task. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbdfe9ee-feb1-4248-b2dd-3a148b17ba72Cited by top-tier papers12
- Symbolic Music Generation with Transformer-GANsAashiq Muhamed, Liang Li, Xingjian Shi, Suri Yaddanapudi et al.AAAI 2021 · 78 citations
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy OptimizationRajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel et al.ICLR 2023 · 54 citations
- Imitating Language via Scalable Inverse Reinforcement LearningMarkus Wulfmeier, Michael Bloesch, Nino Vieillard, Arun Ahuja et al.NeurIPS 2024 · 26 citations
- Generative Data Augmentation with Contrastive Learning for Zero-Shot Stance DetectionYang Li, Jiawei YuanEMNLP 2022 · 18 citations
- Ess-InfoGAIL: Semi-supervised Imitation Learning from Imbalanced DemonstrationsHuiqiao Fu, Kaiqiang Tang, Yuanyang Lu, Yiming Qi et al.NeurIPS 2023 · 15 citations
Builds on6
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
- Making Efficient Use of Demonstrations to Solve Hard Exploration ProblemsÇaglar Gülçehre, Tom Le Paine, Bobak Shahriari, Misha Denil et al.ICLR 2020 · 97 citations
Related papers
- Improving GAN Training with Probability Ratio Clipping and Sample ReweightingYue Wu, Pan Zhou, Andrew Gordon Wilson, Eric P. Xing et al.NeurIPS 2020 · 39 citations
- Self-Adversarial Learning with Comparative Discrimination for Text GenerationWangchunshu Zhou, Tao Ge, Ke Xu, Furu Wei et al.ICLR 2020 · 20 citations
- Improving Adversarial Text Generation by Modeling the Distant FutureRuiyi Zhang, Changyou Chen, Zhe Gan, Wenlin Wang et al.ACL 2020 · 16 citations
- Generative Cooperative Networks for Natural Language GenerationSylvain Lamprier, Thomas Scialom, Antoine Chaffin, Vincent Claveau et al.ICML 2022 · 13 citations
- Branch-GAN: Improving Text Generation with (not so) Large Language ModelsFredrik Carlsson, Johan Broberg, Erik Hillbom, Magnus Sahlgren et al.ICLR 2024 · 3 citations
