Train-by-Reconnect: Decoupling Locations of Weights from Their Values
Yushi Qiu, Reiji Suda
Abstract
What makes untrained deep neural networks (DNNs) different from the trained performant ones? By zooming into the weights in well-trained DNNs, we found that it is the location of weights that holds most of the information encoded by the training. Motivated by this observation, we hypothesized that weights in DNNs trained using stochastic gradient-based methods can be separated into two dimensions: the location of weights, and their exact values. To assess our hypothesis, we propose a novel method called lookahead permutation (LaPerm) to train DNNs by reconnecting the weights. We empirically demonstrate LaPerm's versatility while producing extensive evidence to support our hypothesis: when the initial weights are random and dense, our method demonstrates speed and performance similar to or better than that of regular optimizers, e.g., Adam. When the initial weights are random and sparse (many zeros), our method changes the way neurons connect, achieving accuracy comparable to that of a well-trained dense network. When the initial weights share a single value, our method finds a weight agnostic neural network with far-better-than-chance accuracy. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- A Signal Propagation Perspective for Pruning Neural Networks at InitializationNamhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, Philip H. S. TorrICLR 2020 · 174 citations
Related papers
- Slot Machines: Discovering Winning Combinations of Random Weights in Neural NetworksMaxwell Mbabilla Aladago, Lorenzo TorresaniICML 2021 · 11 citations
- Learning with RetrospectionXiang Deng, Zhongfei ZhangAAAI 2021 · 20 citations
- Does Preprocessing Help Training Over-parameterized Neural Networks?Zhao Song, Shuo Yang, Ruizhe ZhangNeurIPS 2021 · 52 citations
- AdaSTE: An Adaptive Straight-Through Estimator to Train Binary Neural NetworksHuu Le, Rasmus Kjær Høier, Che-Tsung Lin, Christopher ZachCVPR 2022
- What's Hidden in a Randomly Weighted Neural Network?Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi et al.CVPR 2020
