Enhancing Neural Training via a Correlated Dynamics Model
Jonathan Brokman, Roy Betser, Rotem Turjeman, Tom Berkov, Ido Cohen, Guy Gilboa
摘要
As neural networks grow in scale, their training becomes both computationally demanding and rich in dynamics. Amidst the flourishing interest in these training dynamics, we present a novel observation: Parameters during training exhibit intrinsic correlations over time. Capitalizing on this, we introduce Correlation Mode Decomposition (CMD). This algorithm clusters the parameter space into groups, termed modes, that display synchronized behavior across epochs. This enables CMD to efficiently represent the training dynamics of complex networks, like ResNets and Transformers, using only a few modes. Moreover, test set generalization is enhanced. We introduce an efficient CMD variant, designed to run concurrently with training. Our experiments indicate that CMD surpasses the state-of-the-art method for compactly modeled dynamics on image classification. Our modeling can improve training efficiency and lower communication overhead, as shown by our preliminary experiments in the context of federated learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Optimizing Neural Networks via Koopman Operator TheoryAkshunna S. Dogra, William T. RedmanNeurIPS 2020 · 被引用 65 次
- Improving Neural Network Training in Low Dimensional Random BasesFrithjof Gressmann, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2020 · 被引用 35 次
- An Operator Theoretic View On Pruning Deep Neural NetworksWilliam T. Redman, Maria Fonoberova, Ryan Mohr, Yannis G. Kevrekidis 等ICLR 2022 · 被引用 21 次
- StarGAN v2: Diverse Image Synthesis for Multiple DomainsYunjey Choi, Youngjung Uh, Jaejun Yoo, Jung-Woo HaCVPR 2020
相关 Paper
- FedBCGD: Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated LearningJunkang Liu, Fanhua Shang, Yuanyuan Liu, Hongying Liu 等ACM MM 2024 · 被引用 6 次
- Distributed Learning of Fully Connected Neural Networks using Independent Subnet TrainingBinhang Yuan, Cameron R. Wolfe, Chen Dun, Yuxin Tang 等VLDB 2022 · 被引用 42 次
- Federated Dynamic Sparse Training: Computing Less, Communicating Less, Yet Learning BetterSameer Bibikar, Haris Vikalo, Zhangyang Wang, Xiaohan ChenAAAI 2022 · 被引用 133 次
- The Panaceas for Improving Low-Rank Decomposition in Communication-Efficient Federated LearningShiwei Li, Xiandi Luo, Haozhao Wang, Xing Tang 等ICML 2025
- HODEC: Towards Efficient High-Order DEcomposed Convolutional Neural NetworksMiao Yin, Yang Sui, Wanzhao Yang, Xiao Zang 等CVPR 2022 · 被引用 17 次
