Concurrent Adversarial Learning for Large-Batch Training
Yong Liu, Xiangning Chen, Minhao Cheng, Cho-Jui Hsieh, Yang You
摘要
Large-batch training has become a commonly used technique when training neural networks with a large number of GPU/TPU processors. As batch size increases, stochastic optimizers tend to converge to sharp local minima, leading to degraded test performance. Current methods usually use extensive data augmentation to increase the batch size, but we found the performance gain with data augmentation decreases as batch size increases, and data augmentation will become insufficient after certain point. In this paper, we propose to use adversarial learning to increase the batch size in large-batch training. Despite being a natural choice for smoothing the decision surface and biasing towards a flat region, adversarial learning has not been successfully applied in large-batch training since it requires at least two sequential gradient computations at each step, which will at least double the running time compared with vanilla training even with a large number of processors. To overcome this issue, we propose a novel Concurrent Adversarial Learning (ConAdv) method that decouple the sequential gradient computations in adversarial learning by utilizing staled parameters. Experimental results demonstrate that ConAdv can successfully increase the batch size on ResNet-50 training on ImageNet while maintaining high accuracy. In particular, we show ConAdv along can achieve 75.3% top-1 accuracy on ImageNet ResNet-50 training with 96K batch size, and the accuracy can be further improved to 76.2% when combining ConAdv with data augmentation. This is the first work successfully scales ResNet-50 training batch size to 96K.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Make Sharpness-Aware Minimization Stronger: A Sparsified Perturbation ApproachPeng Mi, Li Shen, Tianhe Ren, Yiyi Zhou 等NeurIPS 2022 · 被引用 102 次
- Towards Efficient and Scalable Sharpness-Aware MinimizationYong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh 等CVPR 2022 · 被引用 61 次
- Deep Perturbation Learning: Enhancing the Network Performance via Image PerturbationsZifan Song, Xiao Gong, Guosheng Hu, Cairong ZhaoICML 2023 · 被引用 10 次
- Fast AdvPropJieru Mei, Yucheng Han, Yutong Bai, Yixiao Zhang 等ICLR 2022 · 被引用 10 次
- Large-batch Optimization for Dense Visual Predictions: Training Faster R-CNN in 4.2 MinutesZeyue Xue, Jianming Liang, Guanglu Song, Zhuofan Zong 等NeurIPS 2022
它引用的顶会 Paper4
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Understanding and Improving Fast Adversarial TrainingMaksym Andriushchenko, Nicolas FlammarionNeurIPS 2020 · 被引用 366 次
- Robust and Accurate Object Detection via Adversarial LearningXiangning Chen, Cihang Xie, Mingxing Tan, Li Zhang 等CVPR 2021
- Adversarial Examples Improve Image RecognitionCihang Xie, Mingxing Tan, Boqing Gong, Jiang Wang 等CVPR 2020
相关 Paper
- Augment Your Batch: Improving Generalization Through Instance RepetitionElad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi 等CVPR 2020
- Large Batch Optimization for Deep Learning Using New Complete Layer-Wise Adaptive Rate ScalingZhouyuan Huo, Bin Gu, Heng HuangAAAI 2021 · 被引用 35 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Not All Layers Are Equal: A Layer-Wise Adaptive Approach Toward Large-Scale DNN TrainingYun-Yong Ko, Dongwon Lee, Sang-Wook KimWWW 2022 · 被引用 11 次
- Adversarial AutoAugmentXinyu Zhang, Qiang Wang, Jian Zhang, Zhao ZhongICLR 2020 · 被引用 210 次
