Accumulated Decoupled Learning with Gradient Staleness Mitigation for Convolutional Neural Networks
Huiping Zhuang, Zhenyu Weng, Fulin Luo, Kar-Ann Toj, Haizhou Li, Zhiping Lin
摘要
Gradient staleness is a major side effect in decoupled when training convolutional neural asynchronously. Existing methods that this effect might result in reduced generalization even divergence. In this paper, propose an accumulated decoupled learning (ADL), which includes a module-wise gradient in order to mitigate the gradient . Unlike prior arts ignoring the gradient , we quantify the staleness in such a way its mitigation can be quantitatively visualized. a new learning scheme, the proposed ADL is shown to converge to critical points spite of its asynchronism. Extensive experiments CIFAR-10 and ImageNet datasets are , demonstrating that ADL gives promising results while the state-of-theart experience reduced generalization divergence. In addition, our ADL is shown to the fastest training speed among the compared . The code will be ready soon https://github.com/ZHUANGHP/Accumulated- -Learning.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FedImpro: Measuring and Improving Client Update in Federated LearningZhenheng Tang, Yonggang Zhang, Shaohuai Shi, Xinmei Tian 等ICLR 2024 · 被引用 25 次
- PETRA: Parallel End-to-end Training with Reversible ArchitecturesStéphane Rivaud, Louis Fournier, Thomas Pumir, Eugene Belilovsky 等ICLR 2025
它引用的顶会 Paper2
相关 Paper
- Gap-Aware Mitigation of Gradient StalenessSaar Barkai, Ido Hakimi, Assaf SchusterICLR 2020 · 被引用 27 次
- Concurrent Adversarial Learning for Large-Batch TrainingYong Liu, Xiangning Chen, Minhao Cheng, Cho-Jui Hsieh 等ICLR 2022 · 被引用 14 次
- Accelerated Vertical Federated Adversarial Learning through Decoupling Layer-Wise DependenciesTianxing Man, Yu Bai, Ganyu Wang, Jinjie Fang 等NeurIPS 2025
- Nesterov Method for Asynchronous Pipeline Parallel OptimizationThalaiyasingam Ajanthan, Sameera Ramasinghe, Yan Zuo, Gil Avraham 等ICML 2025
- Taming unbalanced training workloads in deep learning with partial collective operationsShigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh 等PPoPP 2020 · 被引用 52 次
