Contextual Dropout: An Efficient Sample-Dependent Dropout Module
Xinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian, Mingyuan Zhou
Abstract
Dropout has been demonstrated as a simple and effective module to not only regularize the training process of deep neural networks, but also provide the uncertainty estimation for prediction. However, the quality of uncertainty estimation is highly dependent on the dropout probabilities. Most current models use the same dropout distributions across all data samples due to its simplicity. Despite the potential gains in the flexibility of modeling uncertainty, sample-dependent dropout, on the other hand, is less explored as it often encounters scalability issues or involves non-trivial model changes. In this paper, we propose contextual dropout with an efficient structural design as a simple and scalable sample-dependent dropout module, which can be applied to a wide range of models at the expense of only slightly increased memory and computational cost. We learn the dropout probabilities with a variational objective, compatible with both Bernoulli dropout and Gaussian dropout. We apply the contextual dropout module to various models with applications to image classification and visual question answering and demonstrate the scalability of the method with large-scale datasets, such as ImageNet and VQA 2.0. Our experimental results show that the proposed method outperforms baseline methods in terms of both accuracy and quality of uncertainty estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f469bfa-b0bd-4129-8449-5c2906ca5e0bCited by top-tier papers12
- CARD: Classification and Regression Diffusion ModelsXizewen Han, Huangjie Zheng, Mingyuan ZhouNeurIPS 2022 · 185 citations
- Bayesian Attention Belief NetworksShujian Zhang, Xinjie Fan, Bo Chen, Mingyuan ZhouICML 2021 · 38 citations
- NPCL: Neural Processes for Uncertainty-Aware Continual LearningSaurav Jha, Dong Gong, He Zhao, Lina YaoNeurIPS 2023 · 27 citations
- Alignment Attention by Matching Key and Query DistributionsShujian Zhang, Xinjie Fan, Huangjie Zheng, Korawat Tanwisuth et al.NeurIPS 2021 · 20 citations
- Deep Multitask Learning with Progressive Parameter SharingHaosen Shi, Shen Ren, Tianwei Zhang, Sinno Jialin PanICCV 2023 · 15 citations
Builds on3
- Bayesian Attention ModulesXinjie Fan, Shujian Zhang, Bo Chen, Mingyuan ZhouNeurIPS 2020 · 78 citations
- DisARM: An Antithetic Gradient Estimator for Binary Latent VariablesZhe Dong, Andriy Mnih, George TuckerNeurIPS 2020 · 43 citations
- Thompson Sampling via Local UncertaintyZhendong Wang, Mingyuan ZhouICML 2020 · 21 citations
Related papers
- Structured Dropout Variational Inference for Bayesian Neural NetworksSon Nguyen, Duong Nguyen, Khai Nguyen, Khoat Than et al.NeurIPS 2021 · 11 citations
- Bayesian Nested Neural Networks for Uncertainty Calibration and Adaptive CompressionYufei Cui, Ziquan Liu, Qiao Li, Antoni B. Chan et al.CVPR 2021
- GFlowOut: Dropout with Generative Flow NetworksDianbo Liu, Moksh Jain, Bonaventure F. P. Dossou, Qianli Shen et al.ICML 2023 · 27 citations
- Bayesian Posterior Approximation With Stochastic EnsemblesOleksandr Balabanov, Bernhard Mehlig, Hampus LinanderCVPR 2023
- Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty EstimationTal Zeevi, Ravid Shwartz-Ziv, Yann LeCun, Lawrence H. Staib et al.CVPR 2025
