Sketchy: Memory-efficient Adaptive Regularization with Frequent Directions
Vladimir Feinberg, Xinyi Chen, Y. Jennifer Sun, Rohan Anil, Elad Hazan
摘要
Adaptive regularization methods that exploit more than the diagonal entries exhibit state of the art performance for many tasks, but can be prohibitive in terms of memory and running time. We find the spectra of the Kronecker-factored gradient covariance matrix in deep learning (DL) training tasks are concentrated on a small leading eigenspace that changes throughout training, motivating a low-rank sketching approach. We describe a generic method for reducing memory and compute requirements of maintaining a matrix preconditioner using the Frequent Directions (FD) sketch. While previous approaches have explored applying FD for second-order optimization, we present a novel analysis which allows efficient interpolation between resource requirements and the degradation in regret guarantees with rank : in the online convex optimization (OCO) setting over dimension , we match full-matrix memory regret using only memory up to additive error in the bottom eigenvalues of the gradient covariance. Further, we show extensions of our work to Shampoo, resulting in a method competitive in quality with Shampoo and Adam, yet requiring only sub-linear memory for tracking second moments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Flora: Low-Rank Adapters Are Secretly Gradient CompressorsYongchang Hao, Yanshuai Cao, Lili MouICML 2024 · 被引用 113 次
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren 等NeurIPS 2025 · 被引用 58 次
- Memory-Efficient LLM Training with Online Subspace DescentKaizhao Liang, Bo Liu, Lizhang Chen, Qiang LiuNeurIPS 2024 · 被引用 46 次
- MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable ConvergenceIonut-Vlad Modoranu, Mher Safaryan, Grigory Malinovsky, Eldar Kurtic 等NeurIPS 2024 · 被引用 32 次
- Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its PreconditionerRuna Eschenhagen, Aaron Defazio, Tsung-Hsien Lee, Richard E. Turner 等NeurIPS 2025 · 被引用 24 次
它引用的顶会 Paper5
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- A Deeper Look at the Hessian Eigenspectrum of Deep Neural Networks and its Applications to RegularizationAdepu Ravi Sankar, Yash Khasbage, Rahul Vigneswaran, Vineeth N. BalasubramanianAAAI 2021 · 被引用 60 次
- Extreme Tensoring for Low-Memory PreconditioningXinyi Chen, Naman Agarwal, Elad Hazan, Cyril Zhang 等ICLR 2020 · 被引用 11 次
- When does preconditioning help or hurt generalization?Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li 等ICLR 2021 · 被引用 11 次
- Better Full-Matrix Regret via Parameter-Free Online LearningAshok CutkoskyNeurIPS 2020 · 被引用 7 次
相关 Paper
- Dimension-Free Adaptive Subgradient Methods with Frequent DirectionsSifan Yang, Yuanyu Wan, Peijia Li, Yibo Wang 等ICML 2025
- Combining Axes Preconditioners through Kronecker Approximation for Deep LearningSai Surya Duvvuri, Devvrit, Rohan Anil, Cho-Jui Hsieh 等ICLR 2024 · 被引用 16 次
- A Computationally Efficient Sparsified Online Newton MethodDevvrit, Sai Surya Duvvuri, Rohan Anil, Vineet Gupta 等NeurIPS 2023 · 被引用 1 次
- A New Perspective on Shampoo's PreconditionerDepen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach 等ICLR 2025
- Structured Preconditioners in Adaptive Optimization: A Unified AnalysisShuo Xie, Tianhao Wang, Sashank J. Reddi, Sanjiv Kumar 等ICML 2025
