Stochastic Reweighted Gradient Descent
Ayoub El Hanchi, David A. Stephens, Chris J. Maddison
摘要
Despite the strong theoretical guarantees that variance-reduced finite-sum optimization algorithms enjoy, their applicability remains limited to cases where the memory overhead they introduce (SAG/SAGA), or the periodic full gradient computation they require (SVRG/SARAH) are manageable. A promising approach to achieving variance reduction while avoiding these drawbacks is the use of importance sampling instead of control variates. While many such methods have been proposed in the literature, directly proving that they improve the convergence of the resulting optimization algorithm has remained elusive. In this work, we propose an importance-sampling-based algorithm we call SRG (stochastic reweighted gradient). We analyze the convergence of SRG in the strongly-convex case and show that, while it does not recover the linear rate of control variates methods, it provably outperforms SGD. We pay particular attention to the time and memory overhead of our proposed method, and design a specialized red-black tree allowing its efficient implementation. Finally, we present empirical results to support our findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Conditional Mixture Path Guiding for Differentiable RenderingZhimin Fan, Pengcheng Shi, Mufan Guo, Ruoyu Fu 等SIGGRAPH 2024 · 被引用 7 次
- Changing the Training Data Distribution to Reduce Simplicity Bias Improves In-distribution GeneralizationDang Nguyen, Paymon Haddad, Eric Gan, Baharan MirzasoleimanNeurIPS 2024 · 被引用 4 次
- Global Perception Based Autoregressive Neural ProcessesJinyang TaiICCV 2023 · 被引用 1 次
它引用的顶会 Paper4
- SGD with shuffling: optimal rates without component convexity and large epoch requirementsKwangjun Ahn, Chulhee Yun, Suvrit SraNeurIPS 2020 · 被引用 83 次
- Variance Reduction via Accelerated Dual Averaging for Finite-Sum OptimizationChaobing Song, Yong Jiang, Yi MaNeurIPS 2020 · 被引用 25 次
- On Convergence-Diagnostic based Step Sizes for Stochastic Gradient DescentScott Pesme, Aymeric Dieuleveut, Nicolas FlammarionICML 2020 · 被引用 19 次
- Adaptive Importance Sampling for Finite-Sum Optimization and Sampling with Decreasing Step-SizesAyoub El Hanchi, David A. StephensNeurIPS 2020 · 被引用 18 次
相关 Paper
- A Short and Unified Convergence Analysis of the SAG, SAGA, and IAG AlgorithmsFeng Zhu, Robert Heath, Aritra MitraICML 2026
- On the Convergence of Hamiltonian Monte Carlo with Stochastic GradientsDifan Zou, Quanquan GuICML 2021 · 被引用 20 次
- Tackling Data Heterogeneity: A New Unified Framework for Decentralized SGD with Sample-induced TopologyYan Huang, Ying Sun, Zehan Zhu, Changzhi Yan 等ICML 2022 · 被引用 18 次
- History-Gradient Aided Batch Size Adaptation for Variance Reduced AlgorithmsKaiyi Ji, Zhe Wang, Bowen Weng, Yi Zhou 等ICML 2020 · 被引用 19 次
- Variance Reduction in Stochastic Particle-Optimization SamplingJianyi Zhang, Yang Zhao, Changyou ChenICML 2020 · 被引用 13 次
