DART: Diversify-Aggregate-Repeat Training Improves Generalization of Neural Networks
Samyak Jain, Sravanti Addepalli, Pawan Kumar Sahu, Priyam Dey, R. Venkatesh Babu
摘要
Generalization of Neural Networks is crucial for deploying them safely in the real world. Common training strategies to improve generalization involve the use of data augmentations, ensembling and model averaging. In this work, we first establish a surprisingly simple but strong benchmark for generalization which utilizes diverse augmentations within a training minibatch, and show that this can learn a more balanced distribution of features. Further, we propose Diversify-Aggregate-Repeat Training (DART) strategy that first trains diverse models using different augmentations (or domains) to explore the loss basin, and further Aggregates their weights to combine their expertise and obtain improved generalization. We find that Repeating the step of Aggregation throughout training improves the overall optimization trajectory and also ensures that the individual models have sufficiently low loss barrier to obtain improved generalization on combining them. We theoretically justify the proposed approach and show that it indeed generalizes better. In addition to improvements in In-Domain generalization, we demonstrate SOTA performance on the Domain Generalization benchmarks in the popular DomainBed framework as well. Our method is generic and can easily be integrated with several base training algorithms to achieve performance gains. Our code is available here: https://github.com/val-iisc/DART.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- MIST: Defending Against Membership Inference Attacks Through Membership-Invariant Subspace TrainingJiacheng Li, Ninghui Li, Bruno RibeiroUSENIX Security 2024 · 被引用 16 次
- Sharpness-Aware Minimization Enhances Feature Quality via Balanced LearningJacob Mitchell Springer, Vaishnavh Nagarajan, Aditi RaghunathanICLR 2024 · 被引用 13 次
- Lookaround Optimizer: k steps around, 1 step averageJiangtao Zhang, Shunyu Liu, Jie Song, Tongtian Zhu 等NeurIPS 2023 · 被引用 12 次
- DIVE: Subgraph Disagreement for Graph Out-of-Distribution GeneralizationXin Sun, Liang Wang, Qiang Liu, Shu Wu 等KDD 2024 · 被引用 6 次
- A2XP: Towards Private Domain GeneralizationGeunhyeok Yu, Hyoseok HwangCVPR 2024 · 被引用 3 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
相关 Paper
- Ensemble of Averages: Improving Model Selection and Boosting Performance in Domain GeneralizationDevansh Arpit, Huan Wang, Yingbo Zhou, Caiming XiongNeurIPS 2022 · 被引用 232 次
- Diverse Weight Averaging for Out-of-Distribution GeneralizationAlexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy 等NeurIPS 2022 · 被引用 183 次
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho 等NeurIPS 2021 · 被引用 630 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- PGrad: Learning Principal Gradients For Domain GeneralizationZhe Wang, Jake Grigsby, Yanjun QiICLR 2023 · 被引用 1 次
