Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes Well
Vipul Gupta, Santiago Akle Serrano, Dennis DeCoste
Abstract
We propose Stochastic Weight Averaging in Parallel (SWAP), an algorithm to accelerate DNN training. Our algorithm uses large mini-batches to compute an approximate solution quickly and then refines it by averaging the weights of multiple models computed independently and in parallel. The resulting models generalize equally well as those trained with small mini-batches but are produced in a substantially shorter time. We demonstrate the reduction in training time and the good generalization performance of the resulting models on the computer vision datasets CIFAR10, CIFAR100, and ImageNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c918240d-c47d-4271-8bf5-a7b431d39eaaCited by top-tier papers23
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu et al.ICLR 2024 · 230 citations
- Diverse Weight Averaging for Out-of-Distribution GeneralizationAlexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy et al.NeurIPS 2022 · 183 citations
- Consensus Control for Decentralized Deep LearningLingjing Kong, Tao Lin, Anastasia Koloskova, Martin Jaggi et al.ICML 2021 · 100 citations
- Representation Surgery for Multi-Task Model MergingEnneng Yang, Li Shen, Zhenyi Wang, Guibing Guo et al.ICML 2024 · 96 citations
- Parameter Competition Balancing for Model MergingGuodong Du, Junlin Lee, Jing Li, Runhua Jiang et al.NeurIPS 2024 · 91 citations
Builds on1
Related papers
- Trainable Weight Averaging: Efficient Training by Optimizing Historical SolutionsTao Li, Zhehao Huang, Qinghua Tao, Yingwen Wu et al.ICLR 2023
- Sparse Weight Activation TrainingMd Aamir Raihan, Tor M. AamodtNeurIPS 2020 · 83 citations
- Augment Your Batch: Improving Generalization Through Instance RepetitionElad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi et al.CVPR 2020
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson et al.NSDI 2021
- Stochastic Whitening Batch NormalizationShengdong Zhang, Ehsan Nezhadarya, Homa Fashandi, Jiayi Liu et al.CVPR 2021
