Enhancing Neural Training via a Correlated Dynamics Model
Jonathan Brokman, Roy Betser, Rotem Turjeman, Tom Berkov, Ido Cohen, Guy Gilboa
Abstract
As neural networks grow in scale, their training becomes both computationally demanding and rich in dynamics. Amidst the flourishing interest in these training dynamics, we present a novel observation: Parameters during training exhibit intrinsic correlations over time. Capitalizing on this, we introduce Correlation Mode Decomposition (CMD). This algorithm clusters the parameter space into groups, termed modes, that display synchronized behavior across epochs. This enables CMD to efficiently represent the training dynamics of complex networks, like ResNets and Transformers, using only a few modes. Moreover, test set generalization is enhanced. We introduce an efficient CMD variant, designed to run concurrently with training. Our experiments indicate that CMD surpasses the state-of-the-art method for compactly modeled dynamics on image classification. Our modeling can improve training efficiency and lower communication overhead, as shown by our preliminary experiments in the context of federated learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 997d6fef-d84c-44c5-bbd0-5e4807d54bc9Cited by top-tier papers1
Ask how each one uses itBuilds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Optimizing Neural Networks via Koopman Operator TheoryAkshunna S. Dogra, William T. RedmanNeurIPS 2020 · 65 citations
- Improving Neural Network Training in Low Dimensional Random BasesFrithjof Gressmann, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2020 · 35 citations
- An Operator Theoretic View On Pruning Deep Neural NetworksWilliam T. Redman, Maria Fonoberova, Ryan Mohr, Yannis G. Kevrekidis et al.ICLR 2022 · 21 citations
- StarGAN v2: Diverse Image Synthesis for Multiple DomainsYunjey Choi, Youngjung Uh, Jaejun Yoo, Jung-Woo HaCVPR 2020
Related papers
- FedBCGD: Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated LearningJunkang Liu, Fanhua Shang, Yuanyuan Liu, Hongying Liu et al.ACM MM 2024 · 6 citations
- Distributed Learning of Fully Connected Neural Networks using Independent Subnet TrainingBinhang Yuan, Cameron R. Wolfe, Chen Dun, Yuxin Tang et al.VLDB 2022 · 42 citations
- Federated Dynamic Sparse Training: Computing Less, Communicating Less, Yet Learning BetterSameer Bibikar, Haris Vikalo, Zhangyang Wang, Xiaohan ChenAAAI 2022 · 133 citations
- The Panaceas for Improving Low-Rank Decomposition in Communication-Efficient Federated LearningShiwei Li, Xiandi Luo, Haozhao Wang, Xing Tang et al.ICML 2025
- HODEC: Towards Efficient High-Order DEcomposed Convolutional Neural NetworksMiao Yin, Yang Sui, Wanzhao Yang, Xiao Zang et al.CVPR 2022 · 17 citations
