AutoDropout: Learning Dropout Patterns to Regularize Deep Networks
Hieu Pham, Quoc V. Le
Abstract
Neural networks are often over-parameterized and hence benefit from aggressive regularization. Conventional regularization methods, such as Dropout (Srivastava et al. 2014) or weight decay, do not leverage the structures of the network's inputs and hidden states. As a result, these methods are less effective than recent methods that leverage the structures, such as SpatialDropout (Tompson et al. 2020) and DropBlock (Ghiasi, Lin, and Le 2018), which randomly drop the values at certain contiguous areas in the hidden states and setting them to zero. Although the locations of dropout areas are random, the patterns of SpatialDropout and DropBlock are manually designed and fixed. Here we propose AutoDropout, which automates the process of designing dropout patterns. In our method, a controller learns to generate a dropout pattern at every channel and layer of a target network, such as a Con-vNet or a Transformer. The target network is then trained with the dropout pattern, and its resulting validation performance is used as a signal for the controller to learn from. We show that this method works well for both image recognition on CIFAR-10 and ImageNet, as well as language modeling on Penn Treebank and WikiText-2. The learned dropout patterns also transfers to different tasks and datasets, such as from language model on Penn Treebank to Engligh-French translation on WMT 2014. Our code will be available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- AutoLoss-Zero: Searching Loss Functions from Scratch for Generic TasksHao Li, Tianwen Fu, Jifeng Dai, Hongsheng Li et al.CVPR 2022 · 21 citations
- Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction LearningYimeng Shan, Malu Zhang, Ruijie Zhu, Xuerui Qiu et al.AAAI 2025 · 14 citations
- RankNAS: Efficient Neural Architecture Search by Pairwise RankingChi Hu, Chenglong Wang, Xiangnan Ma, Xia Meng et al.EMNLP 2021 · 10 citations
- AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-TuningTao Yang, Jinghao Deng, Xiaojun Quan, Qifan Wang et al.NeurIPS 2022 · 7 citations
- HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better GeneralizabilityJiaao Chen, Dinghan Shen, Weizhu Chen, Diyi YangACL 2021
Builds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
Related papers
- Learning to Drop Out: An Adversarial Approach to Training Sequence VAEsDjordje Miladinovic, Kumar Shridhar, Kushal Jain, Max B. Paulus et al.NeurIPS 2022 · 5 citations
- Regularized Mask Tuning: Uncovering Hidden Knowledge in Pre-trained Vision-Language ModelsKecheng Zheng, Wei Wu, Ruili Feng, Kai Zhu et al.ICCV 2023 · 13 citations
- Domain Generalization Guided by Gradient Signal to Noise Ratio of ParametersMateusz Michalkiewicz, Masoud Faraki, Xiang Yu, Manmohan Chandraker et al.ICCV 2023 · 9 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient TrainingAnup Sarma, Sonali Singh, Huaipan Jiang, Rui Zhang et al.NeurIPS 2021 · 1 citation
