Convergence-Aware Neural Network Training
Hyungjun Oh, Yongseung Yu, Giha Ryu, Gunjoo Ahn, Yuri Jeong, Yongjun Park, Jiwon Seo
摘要
Training a deep neural network(DNN) is expensive, requiring a large amount of computation time. While the training overhead is high, not all computation in DNN training is equal. Some parameters converge faster and thus their gradient computation may contribute little to the parameter update; in nearstationary points a subset of parameters may change very little. In this paper we exploit the parameter convergence to optimize gradient computation in DNN training. We design a light-weight monitoring technique to track the parameter convergence; we prune the gradient computation stochastically for a group of semantically related parameters, exploiting their convergence correlations. These techniques are efficiently implemented in existing GPU kernels. In our evaluation the optimization techniques substantially and robustly improve the training throughput for four DNN models on three public datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks TrainingPengcheng Dai, Jianlei Yang, Xucheng Ye, Xingzhou Cheng 等DAC 2020 · 被引用 27 次
- ADA-GP: Accelerating DNN Training By Adaptive Gradient PredictionVahid Janfaza, Shantanu Mandal, Farabi Mahmud, Abdullah MuzahidMICRO 2023 · 被引用 3 次
- Understanding the effects of data parallelism and sparsity on neural network trainingNamhoon Lee, Thalaiyasingam Ajanthan, Philip H. S. Torr, Martin JaggiICLR 2021 · 被引用 8 次
- Egeria: Efficient DNN Training with Knowledge-Guided Layer FreezingYiding Wang, Decang Sun, Kai Chen, Fan Lai 等EuroSys 2023 · 被引用 43 次
- DEPrune: Depth-wise Separable Convolution Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyunchan Moon, Myung Kuk Yoon 等NeurIPS 2024 · 被引用 10 次
