Understand and Modularize Generator Optimization in ELECTRA-style Pretraining
Chengyu Dong, Liyuan Liu, Hao Cheng, Jingbo Shang, Jianfeng Gao, Xiaodong Liu
摘要
Despite the effectiveness of ELECTRA-style pretraining, their performance is dependent on the careful selection of the model size for the auxiliary generator, leading to high trial-and-error costs. In this paper, we present the first systematic study of this problem. Our theoretical investigation highlights the importance of controlling the generator capacity in ELECTRA-style training. Meanwhile, we found it is not handled properly in the original ELECTRA design, leading to the sensitivity issue. Specifically, since adaptive optimizers like Adam will cripple the weighing of individual losses in the joint optimization, the original design fails to control the generator training effectively. To regain control over the generator, we modularize the generator optimization by decoupling the generator optimizer and discriminator optimizer completely, instead of simply relying on the weighted objective combination. Our simple technique reduced the sensitivity of ELECTRA training significantly and obtains considerable performance gain compared to the original design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui 等NeurIPS 2020 · 被引用 755 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 被引用 358 次
相关 Paper
- Fast-ELECTRA for Efficient Pre-trainingChengyu Dong, Liyuan Liu, Hao Cheng, Jingbo Shang 等ICLR 2024 · 被引用 2 次
- SAS: Self-Augmentation Strategy for Language Model Pre-trainingYifei Xu, Jingqiao Zhang, Ru He, Liangzhu Ge 等AAAI 2022 · 被引用 2 次
- Pretraining Text Encoders with Adversarial Mixture of Training Signal GeneratorsYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary 等ICLR 2022 · 被引用 17 次
- Lightweight Generative Adversarial Networks for Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Philip H. S. Torr, Thomas LukasiewiczNeurIPS 2020 · 被引用 76 次
- Scalable GANs with TransformersSangeek Hyun, MinKyu Lee, Jae-Pil HeoICML 2026
