DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
Alexandre Ramé, Matthieu Cord
Abstract
Deep ensembles perform better than a single network thanks to the diversity among their members. Recent approaches regularize predictions to increase diversity; however, they also drastically decrease individual members' performances. In this paper, we argue that learning strategies for deep ensembles need to tackle the trade-off between ensemble diversity and individual accuracies. Motivated by arguments from information theory and leveraging recent advances in neural estimation of conditional mutual information, we introduce a novel training criterion called DICE: it increases diversity by reducing spurious correlations among features. The main idea is that features extracted from pairs of members should only share information useful for target class prediction without being conditionally redundant. Therefore, besides the classification loss with information bottleneck, we adversarially prevent features from being conditionally predictable from each other. We manage to reduce simultaneous errors while protecting class information. We obtain state-of-the-art accuracy results on CIFAR-10/100: for example, an ensemble of 5 networks trained with DICE matches an ensemble of 7 networks trained independently. We further analyze the consequences on calibration, uncertainty estimation, out-of-distribution detection and online co-distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e7ab38a-80fa-48af-883a-2f8074fe33f6Cited by top-tier papers21
- Diverse Weight Averaging for Out-of-Distribution GeneralizationAlexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy et al.NeurIPS 2022 · 183 citations
- Reducing Information Bottleneck for Weakly Supervised Semantic SegmentationJungbeom Lee, Jooyoung Choi, Jisoo Mok, Sungroh YoonNeurIPS 2021 · 174 citations
- Repulsive Deep Ensembles are BayesianFrancesco D'Angelo, Vincent FortuinNeurIPS 2021 · 141 citations
- Model Ratatouille: Recycling Diverse Models for Out-of-Distribution GeneralizationAlexandre Ramé, Kartik Ahuja, Jianyu Zhang, Matthieu Cord et al.ICML 2023 · 108 citations
- MixMo: Mixing Multiple Inputs for Multiple Outputs via Deep SubnetworksAlexandre Ramé, Rémy Sun, Matthieu CordICCV 2021 · 64 citations
Builds on21
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng et al.AAAI 2020 · 354 citations
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen et al.ICLR 2020 · 292 citations
Related papers
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 50 citations
- Improving Accuracy and Calibration via Differentiated Deep Mutual LearningHan Liu, Peng Cui, Bingning Wang, Weipeng Chen et al.CVPR 2025
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
- A Rate-Distortion View of Uncertainty QuantificationIfigeneia Apostolopoulou, Benjamin Eysenbach, Frank Nielsen, Artur DubrawskiICML 2024 · 3 citations
- Going Beyond Feature Similarity: Effective Dataset distillation based on Class-aware Conditional Mutual InformationXinhao Zhong, Bin Chen, Hao Fang, Xulin Gu et al.ICLR 2025
