Block Low-Rank Preconditioner with Shared Basis for Stochastic Optimization
Jui-Nan Yen, Sai Surya Duvvuri, Inderjit S. Dhillon, Cho-Jui Hsieh
摘要
Adaptive methods with non-diagonal preconditioning have shown state-of-the-art results on various tasks. However, their computational complexity and memory requirement make it challenging to scale these methods to modern neural network architectures. To address this challenge, some previous works have adopted blockdiagonal preconditioners. However, the memory cost of storing the block-diagonal matrix remains substantial, leading to the use of smaller block sizes that ultimately leads to suboptimal performance. To reduce the time and memory complexity without sacrificing performance, we propose approximating the diagonal blocks of the second moment matrix by low-rank matrices and enforcing the same basis for the blocks within each layer. We provide theoretical justification for such basis sharing and design an algorithm to efficiently maintain this shared-basis block low-rank approximation during training. Our results on a deep autoencoder and a Transformer benchmark demonstrate that the proposed method outperforms firstorder methods with slightly more time and memory usage, while also achieving competitive or superior performance compared to other second-order methods with less time and memory usage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- 4-bit Shampoo for Memory-Efficient Network TrainingSike Wang, Pan Zhou, Jia Li, Hua HuangNeurIPS 2024 · 被引用 19 次
- BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network InferenceChangwoo Lee, Soo Min Kwon, Qing Qu, Hun-Seok KimNeurIPS 2024 · 被引用 5 次
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan 等ICML 2026 · 被引用 1 次
- Asymmetric Learning for Spectral Graph Neural NetworksFangbing Liu, Qing WangAAAI 2025 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- IO-Adam: Rethinking Memory-Efficient Adaptive Optimizers from Gradient ComputationYiting Chen, Zongwei Huo, Junchi YanICML 2026
- RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based OptimizationShenyang Deng, Zhuoli Ouyang, Tianyu Pang, Zihang Liu 等ICML 2026 · 被引用 7 次
- Sketchy: Memory-efficient Adaptive Regularization with Frequent DirectionsVladimir Feinberg, Xinyi Chen, Y. Jennifer Sun, Rohan Anil 等NeurIPS 2023 · 被引用 21 次
- Combining Axes Preconditioners through Kronecker Approximation for Deep LearningSai Surya Duvvuri, Devvrit, Rohan Anil, Cho-Jui Hsieh 等ICLR 2024 · 被引用 16 次
- Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveWu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae 等ICML 2024 · 被引用 23 次
