R-Drop: Regularized Dropout for Neural Networks
Xiaobo Liang, Lijun Wu, Juntao Li, Yue Wang, Qi Meng, Tao Qin, Wei Chen, Min Zhang, Tie-Yan Liu
Abstract
Dropout is a powerful and widely used technique to regularize the training of deep neural networks. Though effective and performing well, the randomness introduced by dropout causes unnegligible inconsistency between training and inference. In this paper, we introduce a simple consistency training strategy to regularize dropout, namely R-Drop, which forces the output distributions of different sub models generated by dropout to be consistent with each other. Specifically, for each training sample, R-Drop minimizes the bidirectional KL-divergence between the output distributions of two sub models sampled by dropout. Theoretical analysis reveals that R-Drop reduces the above inconsistency. Experiments on 5 widely used deep learning tasks (18 datasets in total), including neural machine translation, abstractive summarization, language understanding, language modeling, and image classification, show that R-Drop is universally effective. In particular, it yields substantial improvements when applied to fine-tune large-scale pre-trained models, e.g., ViT, RoBERTa-large, and BART, and achieves state-of-the-art (SOTA) performances with the vanilla Transformer model on WMT14 English→German translation (30.91 BLEU) and WMT14 English→French translation (43.95 BLEU), even surpassing models trained with extra large-scale data and expert-designed advanced variants of Transformer models. Our code is available at GitHub 2 . * Equal contribution (listed in alphabetical order).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa004fd0-85dd-4a7f-8475-d4c37da0af94Cited by top-tier papers57
- RRHF: Rank Responses to Align Language Models with Human FeedbackHongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang et al.NeurIPS 2023 · 515 citations
- ST++: Make Self-trainingWork Better for Semi-supervised Semantic SegmentationLihe Yang, Wei Zhuo, Lei Qi, Yinghuan Shi et al.CVPR 2022 · 467 citations
- GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy InjectionWanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu et al.AAAI 2022 · 181 citations
- Frequency Enhanced Hybrid Attention Network for Sequential RecommendationXinyu Du, Huanhuan Yuan, Pengpeng Zhao, Jianfeng Qu et al.SIGIR 2023 · 142 citations
- CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language RecognitionPeiqi Jiao, Yuecong Min, Yanan Li, Xiaotao Wang et al.ICCV 2023 · 52 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
Related papers
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 233 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language ModelsBiao Chen, Lin Zuo, Mengmeng Jing, Kunbin He et al.AAAI 2026
- Consistency Regularization for Cross-Lingual Fine-TuningBo Zheng, Li Dong, Shaohan Huang, Wenhui Wang et al.ACL 2021
- Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and InferenceMostafa Elhoushi, Alexander Pretko, Nolan Dey, Bin Zhang et al.ICML 2026
