MAST: model-agnostic sparsified training
Yury Demidovich, Grigory Malinovsky, Egor Shulgin, Peter Richtárik
摘要
We introduce a novel optimization problem formulation that departs from the conventional way of minimizing machine learning model loss as a black-box function. Unlike traditional formulations, the proposed approach explicitly incorporates an initially pre-trained model and random sketch operators, allowing for sparsification of both the model and gradient during training. We establish the insightful properties of the proposed objective function and highlight its connections to the standard formulation. Furthermore, we present several variants of the Stochastic Gradient Descent (SGD) method adapted to the new problem formulation, including SGD with general sampling, a distributed version, and SGD with variance reduction techniques. We achieve tighter convergence rates and relax assumptions, bridging the gap between theoretical principles and practical applications, covering several important techniques such as Dropout and Sparse training. This work presents promising opportunities to enhance the theoretical understanding of model training through a sparsification-aware optimization approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered DropoutSamuel Horváth, Stefanos Laskaridis, Mário Almeida, Ilias Leontiadis 等NeurIPS 2021 · 被引用 390 次
- FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model ExtractionSamiul Alam, Luyang Liu, Ming Yan, Mi ZhangNeurIPS 2022 · 被引用 261 次
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev 等ICLR 2020 · 被引用 229 次
相关 Paper
- A Single-Step, Sharpness-Aware Minimization is All You Need to Achieve Efficient and Accurate Sparse TrainingJie Ji, Gen Li, Jingjing Fu, Fatemeh Afghah 等NeurIPS 2024 · 被引用 14 次
- Efficient Generalization with Distributionally Robust LearningSoumyadip Ghosh, Mark S. Squillante, Ebisa D. WollegaNeurIPS 2021 · 被引用 4 次
- Detached Error Feedback for Distributed SGD with Random SparsificationAn Xu, Heng HuangICML 2022 · 被引用 12 次
- Sketched Adaptive Distributed Deep Learning: A Sharp Convergence AnalysisZhijie Chen, Qiaobo Li, Arindam BanerjeeNeurIPS 2025 · 被引用 1 次
- Communication-Efficient Distributed Deep Learning with Merged Gradient Sparsification on GPUsShaohuai Shi, Qiang Wang, Xiaowen Chu, Bo Li 等INFOCOM 2020 · 被引用 66 次
