MixMo: Mixing Multiple Inputs for Multiple Outputs via Deep Subnetworks
Alexandre Ramé, Rémy Sun, Matthieu Cord
摘要
Recent strategies achieved ensembling "for free" by fitting concurrently diverse subnetworks inside a single base network. The main idea during training is that each sub-network learns to classify only one of the multiple inputs simultaneously provided. However, the question of how to best mix these multiple inputs has not been studied so farIn this paper, we introduce MixMo, a new generalized framework for learning multi-input multi-output deep subnetworks. Our key motivation is to replace the suboptimal summing operation hidden in previous approaches by a more appropriate mixing mechanism. For that purpose, we draw inspiration from successful mixed sample data augmentations. We show that binary mixing in features - particularly with rectangular patches from CutMix - enhances results by making subnetworks stronger and more diverse.We improve state of the art for image classification on CIFAR-100 and Tiny ImageNet datasets. Our easy to implement models notably outperform data augmented deep ensembles, without the inference and memory overheads. As we operate in features and simply better leverage the expressiveness of large networks, we open a new line of research complementary to previous works.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Diverse Weight Averaging for Out-of-Distribution GeneralizationAlexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy 等NeurIPS 2022 · 被引用 183 次
- FiLM-Ensemble: Probabilistic Deep Learning via Feature-wise Linear ModulationMehmet Ozgur Turkoglu, Alexander Becker, Hüseyin Anil Gündüz, Mina Rezaei 等NeurIPS 2022 · 被引用 46 次
- MIMONets: Multiple-Input-Multiple-Output Neural Networks Exploiting Computation in SuperpositionNicolas Menet, Michael Hersche, Geethan Karunaratne, Luca Benini 等NeurIPS 2023 · 被引用 32 次
- DataMUX: Data Multiplexing for Neural NetworksVishvak Murahari, Carlos E. Jimenez, Runzhe Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 27 次
- DART: Diversify-Aggregate-Repeat Training Improves Generalization of Neural NetworksSamyak Jain, Sravanti Addepalli, Pawan Kumar Sahu, Priyam Dey 等CVPR 2023
它引用的顶会 Paper20
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph 等ICLR 2020 · 被引用 1,572 次
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 被引用 457 次
相关 Paper
- GradAug: A New Regularization Method for Deep Neural NetworksTaojiannan Yang, Sijie Zhu, Chen ChenNeurIPS 2020 · 被引用 43 次
- Training independent subnetworks for robust predictionMarton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu 等ICLR 2021 · 被引用 235 次
- RecursiveMix: Mixed Learning with HistoryLingfeng Yang, Xiang Li, Borui Zhao, Renjie Song 等NeurIPS 2022 · 被引用 27 次
- StyleMix: Separating Content and Style for Enhanced Data AugmentationMinui Hong, Jinwoo Choi, Gunhee KimCVPR 2021
- Provably Learning Diverse Features in Multi-View Data with Midpoint MixupMuthu Chidambaram, Xiang Wang, Chenwei Wu, Rong GeICML 2023 · 被引用 13 次
