Variance Reduction With Sparse Gradients
Melih Elibol, Lihua Lei, Michael I. Jordan
摘要
Variance reduction methods such as SVRG and SpiderBoost use a mixture of large and small batch gradients to reduce the variance of stochastic gradients. Compared to SGD, these methods require at least double the number of operations per update to model parameters. To reduce the computational cost of these methods, we introduce a new sparsity operator: The random-top-k operator. Our operator reduces computational complexity by estimating gradient sparsity exhibited in a variety of applications by combining the top-k operator and the randomized coordinate descent operator. With this operator, large batch gradients offer an extra benefit beyond variance reduction: A reliable estimate of gradient sparsity. Theoretically, our algorithm is at least as good as the best algorithm (SpiderBoost), and further excels in performance whenever the random-top-k operator captures gradient sparsity. Empirically, our algorithm consistently outperforms SpiderBoost using various models on various tasks including image classification, natural language processing, and sparse matrix factorization. We also provide empirical evidence to support the intuition behind our algorithm via a simple gradient entropy computation, which serves to quantify gradient sparsity at every iteration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- A Better Alternative to Error Feedback for Communication-Efficient Distributed LearningSamuel Horváth, Peter RichtárikICLR 2021 · 被引用 66 次
- DeepReduce: A Sparse-tensor Communication Framework for Federated Deep LearningHang Xu, Kelly Kostopoulou, Aritra Dutta, Xin Li 等NeurIPS 2021 · 被引用 48 次
- Increasing ising machine capacity with multi-chip architecturesAnshujit Sharma, Richard Afoakwa, Zeljko Ignjatovic, Michael C. HuangISCA 2022 · 被引用 28 次
- Detached Error Feedback for Distributed SGD with Random SparsificationAn Xu, Heng HuangICML 2022 · 被引用 12 次
- JointSQ: Joint Sparsification-Quantization for Distributed LearningWeiying Xie, Haowei Li, Jitao Ma, Yunsong Li 等CVPR 2024 · 被引用 9 次
相关 Paper
- History-Gradient Aided Batch Size Adaptation for Variance Reduced AlgorithmsKaiyi Ji, Zhe Wang, Bowen Weng, Yi Zhou 等ICML 2020 · 被引用 19 次
- An Effective Hard Thresholding Method Based on Stochastic Variance Reduction for Nonconvex Sparse LearningGuannan Liang, Qianqian Tong, Chunjiang Zhu, Jinbo BiAAAI 2020 · 被引用 5 次
- A Coefficient Makes SVRG EffectiveYida Yin, Zhiqiu Xu, Zhiyuan Li, Trevor Darrell 等ICLR 2025
- Stochastic Reweighted Gradient DescentAyoub El Hanchi, David A. Stephens, Chris J. MaddisonICML 2022 · 被引用 10 次
- Near-optimal sparse allreduce for distributed deep learningShigang Li, Torsten HoeflerPPoPP 2022 · 被引用 57 次
