Adaptive Width Neural Networks
Federico Errica, Henrik Christiansen, Viktor Zaverkin, Mathias Niepert, Francesco Alesiani
Abstract
For almost 70 years, researchers have typically selected the width of neural networks’ layers either manually or through automated hyperparameter tuning methods such as grid search and, more recently, neural architecture search. This paper challenges the status quo by introducing an easy-to-use technique to learn an unbounded width of a neural network's layer during training. The method jointly optimizes the width and the parameters of each layer via standard backpropagation. We apply the technique to a broad range of data domains such as tables, images, text, sequences, and graphs, showing how the width adapts to the task's difficulty. A by product of our width learning approach is the easy truncation of the trained network at virtually zero cost, achieving a smooth trade-off between performance and compute resources. Alternatively, one can dynamically compress the network until performances do not degrade. In light of recent foundation models trained on large datasets, requiring billions of parameters and where hyper-parameter tuning is unfeasible due to huge training costs, our approach introduces a viable alternative for width learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09163011-622e-43db-a973-9acb4dc3644aBuilds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- A Fair Comparison of Graph Neural Networks for Graph ClassificationFederico Errica, Marco Podda, Davide Bacciu, Alessio MicheliICLR 2020 · 508 citations
- Firefly Neural Architecture Descent: a General Approach for Growing Neural NetworksLemeng Wu, Bo Liu, Peter Stone, Qiang LiuNeurIPS 2020 · 79 citations
Related papers
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked LayersJunjie Liu, Zhe Xu, Runbin Shi, Ray C. C. Cheung et al.ICLR 2020 · 136 citations
- Unveiling The Matthew Effect Across Channels: Assessing Layer Width Sufficiency via Weight Norm VarianceYiting Chen, Jiazi Bu, Junchi YanNeurIPS 2024 · 4 citations
- pscaling small models: Principled warm starts and hyperparameter transferYuxin Ma, Nan Chen, Mateo D Diaz, Soufiane Hayou et al.ICML 2026 · 1 citation
- Variational Inference for Infinitely Deep Neural NetworksAchille Nazaret, David M. BleiICML 2022 · 13 citations
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 111 citations
