Lookaround Optimizer: k steps around, 1 step average
Jiangtao Zhang, Shunyu Liu, Jie Song, Tongtian Zhu, Zhengqi Xu, Mingli Song
Abstract
Weight Average (WA) is an active research topic due to its simplicity in ensembling deep networks and the effectiveness in promoting generalization. Existing weight average approaches, however, are often carried out along only one training trajectory in a post-hoc manner (i.e., the weights are averaged after the entire training process is finished), which significantly degrades the diversity between networks and thus impairs the effectiveness. In this paper, inspired by weight average, we propose Lookaround, a straightforward yet effective SGD-based optimizer leading to flatter minima with better generalization. Specifically, Lookaround iterates two steps during the whole training period: the around step and the average step. In each iteration, 1) the around step starts from a common point and trains multiple networks simultaneously, each on transformed data by a different data augmentation, and 2) the average step averages these trained networks to get the averaged network, which serves as the starting point for the next iteration. The around step improves the functionality diversity while the average step guarantees the weight locality of these networks during the whole training, which is essential for WA to work. We theoretically explain the superiority of Lookaround by convergence analysis, and make extensive experiments to evaluate Lookaround on popular benchmarks including CIFAR and ImageNet with both CNNs and ViTs, demonstrating clear superiority over state-of-the-arts. Our code is available at https://github.com/Ardcy/Lookaround .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- WATT: Weight Average Test Time Adaptation of CLIPDavid Osowiechi, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah et al.NeurIPS 2024 · 46 citations
- HVAdam: A Full-Dimension Adaptive OptimizerYiheng Zhang, Shaowu Wu, Yuanzhuo Xu, Jiajun Wu et al.AAAI 2025 · 1 citation
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
Related papers
- When Do Flat Minima Optimizers Work?Jean Kaddour, Linqing Liu, Ricardo Silva, Matt J. KusnerNeurIPS 2022 · 102 citations
- Neural networks with late-phase weightsJohannes von Oswald, Seijin Kobayashi, João Sacramento, Alexander Meulemans et al.ICLR 2021 · 38 citations
- Random Sharpness-Aware MinimizationYong Liu, Siqi Mai, Minhao Cheng, Xiangning Chen et al.NeurIPS 2022 · 38 citations
- Towards Efficient and Scalable Sharpness-Aware MinimizationYong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh et al.CVPR 2022 · 61 citations
- Trainable Weight Averaging: Efficient Training by Optimizing Historical SolutionsTao Li, Zhehao Huang, Qinghua Tao, Yingwen Wu et al.ICLR 2023
