Extreme Tensoring for Low-Memory Preconditioning
Xinyi Chen, Naman Agarwal, Elad Hazan, Cyril Zhang, Yi Zhang
Abstract
State-of-the-art models are now trained with billions of parameters, reaching hardware limits in terms of memory consumption. This has created a recent demand for memory-efficient optimizers. To this end, we investigate the limits and performance tradeoffs of memory-efficient adaptively preconditioned gradient methods. We propose extreme tensoring for high-dimensional stochastic optimization, showing that an optimizer needs very little memory to benefit from adaptive preconditioning. Our technique applies to arbitrary models (not necessarily with tensor-shaped parameters), and is accompanied by regret and convergence guarantees, which shed light on the tradeoffs between preconditioner quality and expressivity. On a large-scale NLP model, we reduce the optimizer memory overhead by three orders of magnitude, without degrading performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a8cd16e-3d41-41a9-a127-e614e9f403d4Cited by top-tier papers6
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 457 citations
- Memory Efficient Optimizers with 4-bit StatesBingrui Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 72 citations
- Sketchy: Memory-efficient Adaptive Regularization with Frequent DirectionsVladimir Feinberg, Xinyi Chen, Y. Jennifer Sun, Rohan Anil et al.NeurIPS 2023 · 21 citations
- Stochastic Optimization with Laggard Data PipelinesNaman Agarwal, Rohan Anil, Tomer Koren, Kunal Talwar et al.NeurIPS 2020 · 14 citations
- Better Full-Matrix Regret via Parameter-Free Online LearningAshok CutkoskyNeurIPS 2020 · 7 citations
Related papers
- LoQT: Low-Rank Adapters for Quantized PretrainingSebastian Loeschcke, Mads Toftrup, Michael J. Kastoryano, Serge J. Belongie et al.NeurIPS 2024 · 14 citations
- FOAM: Blocked State Folding for Memory-Efficient LLM TrainingZiqing Wen, Jiahuan Wang, ping luo, Dongsheng Li et al.ICML 2026 · 2 citations
- On the Duality between Gradient Transformations and AdaptersLucas Torroba Hennigen, Hunter Lang, Han Guo, Yoon KimICML 2025
- SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM TrainingYehonathan Refael, Guy Smorodinsky, Tom Tirer, Ofir LindenbaumNeurIPS 2025 · 17 citations
- Memory-Efficient 4-bit Preconditioned Stochastic OptimizationJingyang Li, Kuangyu Ding, Kim-Chuan Toh, Pan ZhouICCV 2025 · 1 citation
