Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization
Zitao Song, Cedar Site Bai, Zhe Zhang, Brian Bullins, David Gleich
摘要
Adaptive methods like Adam have become the de facto standard for large-scale vector and Euclidean optimization due to their coordinate-wise adaptation with a second-order nature. More recently, matrix-based spectral optimizers like Muon (Jordan et al., 2024b) show the power of treating weight matrices as matrices rather than long vectors. Linking these is hard because many natural generalizations are not feasible to implement, and we also cannot simply move the Adam adaptation to the matrix spectrum. To address this, we reformulate the AdaGrad update and decompose it into a variance adaptation term and a scale-invariant term. This decoupling produces DeVA (Decoupled Variance Adaptation), a framework that bridges between vector-based variance adaptation and matrix spectral optimization, enabling a seamless transition from Adam to adaptive spectral descent. Extensive experiments across language modeling and image classification demonstrate that DeVA consistently outperforms state-of-the-art methods such as Muon and SOAP (Vyas et al., 2024) , reducing token usage by around 6.6%. Theoretically, we show that the variance adaptation term effectively improves the blockwise smoothness, facilitating faster convergence. Our implementation is available at https://github.com/Tsedao/ Decoupled-Variance-Adaptation
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li 等NeurIPS 2024 · 被引用 149 次
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 被引用 92 次
相关 Paper
- Delving into Muon and Beyond: Deep Analysis and ExtensionsXianbiao Qi, Marco Chen, Jiaquan Ye, Yelin He 等ICML 2026 · 被引用 6 次
- How Muon’s Spectral Design Benefits Generalization: A Study on Imbalanced DataBhavya Vasudeva, Puneesh Deora, Yize Zhao, Vatsal Sharan 等ICLR 2026 · 被引用 16 次
- Enhancing LLM Training via Spectral ClippingXiaowen Jiang, Andrei Semenov, Sebastian StichICML 2026 · 被引用 4 次
- Domain-Independent Dominance of Adaptive MethodsPedro Savarese, David McAllester, Sudarshan Babu, Michael MaireCVPR 2021
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large ModelsJunkang Liu, Fanhua Shang, Hongying Liu, Yuxuan Tian 等AAAI 2026 · 被引用 12 次
