Bayesian Nested Neural Networks for Uncertainty Calibration and Adaptive Compression
Yufei Cui, Ziquan Liu, Qiao Li, Antoni B. Chan, Chun Jason Xue
摘要
Nested networks or slimmable networks are neural networks whose architectures can be adjusted instantly during testing time, e.g., based on computational constraints. Recent studies have focused on a "nested dropout" layer, which is able to order the nodes of a layer by importance during training, thus generating a nested set of subnetworks that are optimal for different configurations of resources. However, the dropout rate is fixed as a hyperparameter over different layers during the whole training process. Therefore, when nodes are removed, the performance decays in a human-specified trajectory rather than in a trajectory learned from data. Another drawback is the generated sub-networks are deterministic networks without well-calibrated uncertainty. To address these two problems, we develop a Bayesian approach to nested neural networks. We propose a variational ordering unit that draws samples for nested dropout at a low cost, from a proposed Downhill distribution, which provides useful gradients to the parameters of nested dropout. Based on this approach, we design a Bayesian nested neural network that learns the order knowledge of the node distributions. In experiments, we show that the proposed approach outperforms the nested network in terms of accuracy, calibration, and out-of-domain detection in classification tasks. It also outperforms the related approach on uncertainty-critical tasks in computer vision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- The Pitfalls and Promise of Conformal Inference Under Adversarial AttacksZiquan Liu, Yufei Cui, Yan Yan, Yi Xu 等ICML 2024 · 被引用 9 次
- ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation FusionShangyu Wu, Ying Xiong, Yufei Cui, Xue Liu 等ICLR 2024 · 被引用 7 次
- Bayes-MIL: A New Probabilistic Perspective on Attention-based Multiple Instance Learning for Whole Slide ImagesYufei Cui, Ziquan Liu, Xiangyu Liu, Xue Liu 等ICLR 2023
它引用的顶会 Paper4
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 被引用 444 次
- Modeling Noisy Annotations for Crowd CountingJia Wan, Antoni B. ChanNeurIPS 2020 · 被引用 120 次
- Crowd Counting with Decomposed UncertaintyMin-hwan Oh, Peder A. Olsen, Karthikeyan Natesan RamamurthyAAAI 2020 · 被引用 118 次
相关 Paper
- Structured Dropout Variational Inference for Bayesian Neural NetworksSon Nguyen, Duong Nguyen, Khai Nguyen, Khoat Than 等NeurIPS 2021 · 被引用 11 次
- Variational Bayesian Last LayersJames Harrison, John Willes, Jasper SnoekICLR 2024 · 被引用 75 次
- Deep Hierarchical Learning with Nested Subspace Networks for Large Language ModelsPaulius Rauba, Mihaela van der SchaarICLR 2026 · 被引用 3 次
- GFlowOut: Dropout with Generative Flow NetworksDianbo Liu, Moksh Jain, Bonaventure F. P. Dossou, Qianli Shen 等ICML 2023 · 被引用 27 次
- On the Expressiveness of Approximate Inference in Bayesian Neural NetworksAndrew Y. K. Foong, David R. Burt, Yingzhen Li, Richard E. TurnerNeurIPS 2020 · 被引用 142 次
