Accumulated Decoupled Learning with Gradient Staleness Mitigation for Convolutional Neural Networks
Huiping Zhuang, Zhenyu Weng, Fulin Luo, Kar-Ann Toj, Haizhou Li, Zhiping Lin
Abstract
Gradient staleness is a major side effect in decoupled when training convolutional neural asynchronously. Existing methods that this effect might result in reduced generalization even divergence. In this paper, propose an accumulated decoupled learning (ADL), which includes a module-wise gradient in order to mitigate the gradient . Unlike prior arts ignoring the gradient , we quantify the staleness in such a way its mitigation can be quantitatively visualized. a new learning scheme, the proposed ADL is shown to converge to critical points spite of its asynchronism. Extensive experiments CIFAR-10 and ImageNet datasets are , demonstrating that ADL gives promising results while the state-of-theart experience reduced generalization divergence. In addition, our ADL is shown to the fastest training speed among the compared . The code will be ready soon https://github.com/ZHUANGHP/Accumulated- -Learning.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac7468fb-2b73-4c83-a7e9-a246426b03dfCited by top-tier papers2
- FedImpro: Measuring and Improving Client Update in Federated LearningZhenheng Tang, Yonggang Zhang, Shaohuai Shi, Xinmei Tian et al.ICLR 2024 · 25 citations
- PETRA: Parallel End-to-end Training with Reversible ArchitecturesStéphane Rivaud, Louis Fournier, Thomas Pumir, Eugene Belilovsky et al.ICLR 2025
Builds on2
Related papers
- Gap-Aware Mitigation of Gradient StalenessSaar Barkai, Ido Hakimi, Assaf SchusterICLR 2020 · 27 citations
- Concurrent Adversarial Learning for Large-Batch TrainingYong Liu, Xiangning Chen, Minhao Cheng, Cho-Jui Hsieh et al.ICLR 2022 · 14 citations
- Accelerated Vertical Federated Adversarial Learning through Decoupling Layer-Wise DependenciesTianxing Man, Yu Bai, Ganyu Wang, Jinjie Fang et al.NeurIPS 2025
- Nesterov Method for Asynchronous Pipeline Parallel OptimizationThalaiyasingam Ajanthan, Sameera Ramasinghe, Yan Zuo, Gil Avraham et al.ICML 2025
- Taming unbalanced training workloads in deep learning with partial collective operationsShigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh et al.PPoPP 2020 · 52 citations
