Error Feedback Can Accurately Compress Preconditioners
Ionut-Vlad Modoranu, Aleksei Kalinov, Eldar Kurtic, Elias Frantar, Dan Alistarh
摘要
Leveraging second-order information about the loss at the scale of deep networks is one of the main lines of approach for improving the performance of current optimizers for deep learning. Yet, existing approaches for accurate full-matrix preconditioning, such as Full-Matrix Adagrad (GGT) or Matrix-Free Approximate Curvature (M-FAC) suffer from massive storage costs when applied even to small-scale models, as they must store a sliding window of gradients, whose memory requirements are multiplicative in the model dimension. In this paper, we address this issue via a novel and efficient error-feedback technique that can be applied to compress preconditioners by up to two orders of magnitude in practice, without loss of convergence. Specifically, our approach compresses the gradient information via sparsification or low-rank compression before it is fed into the preconditioner, feeding the compression error back into future iterations. Experiments on deep neural networks show that this approach can compress full-matrix preconditioners to up to 99% sparsity without accuracy loss, effectively removing the memory overhead of full-matrix preconditioners such as GGT and M-FAC. Our code is available at https://github.com/IST-DASLab/EFCP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable ConvergenceIonut-Vlad Modoranu, Mher Safaryan, Grigory Malinovsky, Eldar Kurtic 等NeurIPS 2024 · 被引用 32 次
- SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM TrainingYehonathan Refael, Guy Smorodinsky, Tom Tirer, Ofir LindenbaumNeurIPS 2025 · 被引用 17 次
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan 等ICML 2026 · 被引用 1 次
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel 等ICLR 2025
它引用的顶会 Paper3
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa 等AAAI 2021 · 被引用 358 次
- On Distributed Adaptive Optimization with Gradient CompressionXiaoyun Li, Belhal Karimi, Ping LiICLR 2022 · 被引用 34 次
相关 Paper
- M-FAC: Efficient Matrix-Free Approximations of Second-Order InformationElias Frantar, Eldar Kurtic, Dan AlistarhNeurIPS 2021 · 被引用 69 次
- Sketchy: Memory-efficient Adaptive Regularization with Frequent DirectionsVladimir Feinberg, Xinyi Chen, Y. Jennifer Sun, Rohan Anil 等NeurIPS 2023 · 被引用 21 次
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
- Step-Ahead Error Feedback for Distributed Training with Compressed GradientAn Xu, Zhouyuan Huo, Heng HuangAAAI 2021 · 被引用 17 次
- Memory-Efficient 4-bit Preconditioned Stochastic OptimizationJingyang Li, Kuangyu Ding, Kim-Chuan Toh, Pan ZhouICCV 2025 · 被引用 1 次
