Modular Duality in Deep Learning
Jeremy Bernstein, Laker Newhouse
摘要
An old idea in optimization theory says that since the gradient is a dual vector it may not be subtracted from the weights without first being mapped to the primal space where the weights reside. We take this idea seriously in this paper and construct such a duality map for general neural networks. Our map, which we call modular dualization, forms a unifying theoretical basis for training algorithms that are a) fast and b) scalable. Modular dualization involves first assigning operator norms to layers based on the semantics of each layer, and then using these layerwise norms to recursively induce a duality map on the weight space of the full neural architecture. We conclude by deriving GPU-friendly algorithms for dualizing Embed, Linear and Conv2D layers-the latter two methods are based on a rectangular Newton-Schulz iteration (Kovarik, 1970; Björck & Bowie, 1971) . A variant of our methods was used to set speed records for training NanoGPT. Overall, we hope that our theory of modular duality will yield a next generation of fast and scalable optimizers for general neural architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- The Polar Express: Optimal Matrix Sign Methods and their Application to the Muon AlgorithmNoah Amsel, David Persson, Christopher Musco, Robert M. GowerICLR 2026 · 被引用 115 次
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 被引用 92 次
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren 等NeurIPS 2025 · 被引用 58 次
- MuonBP: Faster Muon via Block-Periodic OrthogonalizationAhmed Khaled, Kaan Ozkara, Tao Yu, Mingyi Hong 等ICLR 2026 · 被引用 35 次
- COSMOS: A Hybrid Adaptive Optimizer for Efficient Training of Large Language ModelsLiming Liu, Zhenghao Xu, Zixuan Zhang, Hao Kang 等ICLR 2026 · 被引用 27 次
它引用的顶会 Paper4
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng 等NeurIPS 2024 · 被引用 70 次
- Scaling Exponents Across Parameterizations and OptimizersKatie E. Everett, Lechao Xiao, Mitchell Wortsman, Alexander A. Alemi 等ICML 2024 · 被引用 59 次
- Sketchy: Memory-efficient Adaptive Regularization with Frequent DirectionsVladimir Feinberg, Xinyi Chen, Y. Jennifer Sun, Rohan Anil 等NeurIPS 2023 · 被引用 21 次
相关 Paper
- Does Preprocessing Help Training Over-parameterized Neural Networks?Zhao Song, Shuo Yang, Ruizhe ZhangNeurIPS 2021 · 被引用 52 次
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan 等ICML 2026 · 被引用 1 次
- PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network TrainingShenghao Yang, Zhichao Wang, Oleg Balabanov, N. Benjamin Erichson 等ICML 2026 · 被引用 3 次
- Training Deep Learning Models with Norm-Constrained LMOsThomas Pethick, Wanyun Xie, Kimon Antonakopoulos, Zhenyu Zhu 等ICML 2025
- Demystifying Batch Normalization in ReLU Networks: Equivalent Convex Optimization Models and Implicit RegularizationTolga Ergen, Arda Sahiner, Batu Ozturkler, John M. Pauly 等ICLR 2022 · 被引用 34 次
