AutoDropout: Learning Dropout Patterns to Regularize Deep Networks
Hieu Pham, Quoc V. Le
摘要
Neural networks are often over-parameterized and hence benefit from aggressive regularization. Conventional regularization methods, such as Dropout (Srivastava et al. 2014) or weight decay, do not leverage the structures of the network's inputs and hidden states. As a result, these methods are less effective than recent methods that leverage the structures, such as SpatialDropout (Tompson et al. 2020) and DropBlock (Ghiasi, Lin, and Le 2018), which randomly drop the values at certain contiguous areas in the hidden states and setting them to zero. Although the locations of dropout areas are random, the patterns of SpatialDropout and DropBlock are manually designed and fixed. Here we propose AutoDropout, which automates the process of designing dropout patterns. In our method, a controller learns to generate a dropout pattern at every channel and layer of a target network, such as a Con-vNet or a Transformer. The target network is then trained with the dropout pattern, and its resulting validation performance is used as a signal for the controller to learn from. We show that this method works well for both image recognition on CIFAR-10 and ImageNet, as well as language modeling on Penn Treebank and WikiText-2. The learned dropout patterns also transfers to different tasks and datasets, such as from language model on Penn Treebank to Engligh-French translation on WMT 2014. Our code will be available. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- AutoLoss-Zero: Searching Loss Functions from Scratch for Generic TasksHao Li, Tianwen Fu, Jifeng Dai, Hongsheng Li 等CVPR 2022 · 被引用 21 次
- Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction LearningYimeng Shan, Malu Zhang, Ruijie Zhu, Xuerui Qiu 等AAAI 2025 · 被引用 14 次
- RankNAS: Efficient Neural Architecture Search by Pairwise RankingChi Hu, Chenglong Wang, Xiangnan Ma, Xia Meng 等EMNLP 2021 · 被引用 10 次
- AD-DROP: Attribution-Driven Dropout for Robust Language Model Fine-TuningTao Yang, Jinghao Deng, Xiaojun Quan, Qifan Wang 等NeurIPS 2022 · 被引用 7 次
- HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better GeneralizabilityJiaao Chen, Dinghan Shen, Weizhu Chen, Diyi YangACL 2021
它引用的顶会 Paper9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
相关 Paper
- Learning to Drop Out: An Adversarial Approach to Training Sequence VAEsDjordje Miladinovic, Kumar Shridhar, Kushal Jain, Max B. Paulus 等NeurIPS 2022 · 被引用 5 次
- Regularized Mask Tuning: Uncovering Hidden Knowledge in Pre-trained Vision-Language ModelsKecheng Zheng, Wei Wu, Ruili Feng, Kai Zhu 等ICCV 2023 · 被引用 13 次
- Domain Generalization Guided by Gradient Signal to Noise Ratio of ParametersMateusz Michalkiewicz, Masoud Faraki, Xiang Yu, Manmohan Chandraker 等ICCV 2023 · 被引用 9 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient TrainingAnup Sarma, Sonali Singh, Huaipan Jiang, Rui Zhang 等NeurIPS 2021 · 被引用 1 次
