R-Drop: Regularized Dropout for Neural Networks
Xiaobo Liang, Lijun Wu, Juntao Li, Yue Wang, Qi Meng, Tao Qin, Wei Chen, Min Zhang, Tie-Yan Liu
摘要
Dropout is a powerful and widely used technique to regularize the training of deep neural networks. Though effective and performing well, the randomness introduced by dropout causes unnegligible inconsistency between training and inference. In this paper, we introduce a simple consistency training strategy to regularize dropout, namely R-Drop, which forces the output distributions of different sub models generated by dropout to be consistent with each other. Specifically, for each training sample, R-Drop minimizes the bidirectional KL-divergence between the output distributions of two sub models sampled by dropout. Theoretical analysis reveals that R-Drop reduces the above inconsistency. Experiments on 5 widely used deep learning tasks (18 datasets in total), including neural machine translation, abstractive summarization, language understanding, language modeling, and image classification, show that R-Drop is universally effective. In particular, it yields substantial improvements when applied to fine-tune large-scale pre-trained models, e.g., ViT, RoBERTa-large, and BART, and achieves state-of-the-art (SOTA) performances with the vanilla Transformer model on WMT14 English→German translation (30.91 BLEU) and WMT14 English→French translation (43.95 BLEU), even surpassing models trained with extra large-scale data and expert-designed advanced variants of Transformer models. Our code is available at GitHub 2 . * Equal contribution (listed in alphabetical order).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper57
- RRHF: Rank Responses to Align Language Models with Human FeedbackHongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang 等NeurIPS 2023 · 被引用 515 次
- ST++: Make Self-trainingWork Better for Semi-supervised Semantic SegmentationLihe Yang, Wei Zhuo, Lei Qi, Yinghuan Shi 等CVPR 2022 · 被引用 467 次
- GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy InjectionWanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu 等AAAI 2022 · 被引用 181 次
- Frequency Enhanced Hybrid Attention Network for Sequential RecommendationXinyu Du, Huanhuan Yuan, Pengpeng Zhao, Jianfeng Qu 等SIGIR 2023 · 被引用 142 次
- CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language RecognitionPeiqi Jiao, Yuecong Min, Yanan Li, Xiaotao Wang 等ICCV 2023 · 被引用 52 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
相关 Paper
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 被引用 233 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language ModelsBiao Chen, Lin Zuo, Mengmeng Jing, Kunbin He 等AAAI 2026
- Consistency Regularization for Cross-Lingual Fine-TuningBo Zheng, Li Dong, Shaohan Huang, Wenhui Wang 等ACL 2021
- Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and InferenceMostafa Elhoushi, Alexander Pretko, Nolan Dey, Bin Zhang 等ICML 2026
