Understand and Modularize Generator Optimization in ELECTRA-style Pretraining
Chengyu Dong, Liyuan Liu, Hao Cheng, Jingbo Shang, Jianfeng Gao, Xiaodong Liu
Abstract
Despite the effectiveness of ELECTRA-style pretraining, their performance is dependent on the careful selection of the model size for the auxiliary generator, leading to high trial-and-error costs. In this paper, we present the first systematic study of this problem. Our theoretical investigation highlights the importance of controlling the generator capacity in ELECTRA-style training. Meanwhile, we found it is not handled properly in the original ELECTRA design, leading to the sensitivity issue. Specifically, since adaptive optimizers like Adam will cripple the weighing of individual losses in the joint optimization, the original design fails to control the generator training effectively. To regain control over the generator, we modularize the generator optimization by decoupling the generator optimizer and discriminator optimizer completely, instead of simply relying on the weighted objective combination. Our simple technique reduced the sensitivity of ELECTRA training significantly and obtains considerable performance gain compared to the original design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui et al.NeurIPS 2020 · 755 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 358 citations
Related papers
- Fast-ELECTRA for Efficient Pre-trainingChengyu Dong, Liyuan Liu, Hao Cheng, Jingbo Shang et al.ICLR 2024 · 2 citations
- SAS: Self-Augmentation Strategy for Language Model Pre-trainingYifei Xu, Jingqiao Zhang, Ru He, Liangzhu Ge et al.AAAI 2022 · 2 citations
- Pretraining Text Encoders with Adversarial Mixture of Training Signal GeneratorsYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary et al.ICLR 2022 · 17 citations
- Lightweight Generative Adversarial Networks for Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Philip H. S. Torr, Thomas LukasiewiczNeurIPS 2020 · 76 citations
- Scalable GANs with TransformersSangeek Hyun, MinKyu Lee, Jae-Pil HeoICML 2026
