GanLM: Encoder-Decoder Pre-training with an Auxiliary Discriminator
Jian Yang, Shuming Ma, Li Dong, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang, Liqun Yang, Furu Wei, Zhoujun Li
Abstract
Pre-trained models have achieved remarkable success in natural language processing (NLP). However, existing pre-training methods underutilize the benefits of language understanding for generation. Inspired by the idea of Generative Adversarial Networks (GANs), we propose a GAN-style model for encoder-decoder pretraining by introducing an auxiliary discriminator, unifying the ability of language understanding and generation in a single model. Our model, named as GANLM, is trained with two pre-training objectives: replaced token detection and replaced token denoising. Specifically, given masked source sentences, the generator outputs the target distribution and the discriminator predicts whether the target sampled tokens from distribution are incorrect. The target sentence is replaced with misclassified tokens to construct noisy previous context, which is used to generate the gold sentence. In general, both tasks improve the ability of language understanding and generation by selectively using the denoising data. Extensive experiments in language generation benchmarks show that GANLM with the powerful language understanding capability outperforms various strong pre-trained language models (PLMs) and achieves state-of-the-art performance. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47bbef42-9053-4728-8fc1-83ca20bbd839Cited by top-tier papers4
- XCOT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought ReasoningLinzheng Chai, Jian Yang, Tao Sun, Hongcheng Guo et al.AAAI 2025 · 70 citations
- OWL: A Large Language Model for IT OperationsHongcheng Guo, Jian Yang, Jiaheng Liu, Liqun Yang et al.ICLR 2024 · 64 citations
- Adaptive Neural Ranking Framework: Toward Maximized Business Goal for Cascade Ranking SystemsYunli Wang, Zhiqiang Wang, Jian Yang, Shiyang Wen et al.WWW 2024 · 16 citations
- Towards Real-world Scenario: Imbalanced New Intent DiscoveryShun Zhang, Chaoran Yan, Jian Yang, Jiaheng Liu et al.ACL 2024
Builds on7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 213 citations
- Cross-Lingual Natural Language Generation via Pre-TrainingZewen Chi, Li Dong, Furu Wei, Wenhui Wang et al.AAAI 2020 · 142 citations
Related papers
- GLM: General Language Model Pretraining with Autoregressive Blank InfillingZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding et al.ACL 2022
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderYi Liao, Xin Jiang, Qun LiuACL 2020 · 28 citations
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned GenerationBin Bi, Chenliang Li, Chen Wu, Ming Yan et al.EMNLP 2020 · 41 citations
- Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder NetworkYehao Li, Yingwei Pan, Ting Yao, Jingwen Chen et al.AAAI 2021 · 59 citations
- XLM-D: Decorate Cross-lingual Pre-training Model as Non-Autoregressive Neural Machine TranslationYong Wang, Shilin He, Guanhua Chen, Yun Chen et al.EMNLP 2022 · 4 citations
