The Limitations of Large Width in Neural Networks: A Deep Gaussian Process Perspective
Geoff Pleiss, John P. Cunningham
摘要
Large width limits have been a recent focus of deep learning research: modulo computational practicalities, do wider networks outperform narrower ones? Answering this question has been challenging, as conventional networks gain representational power with width, potentially masking any negative effects. Our analysis in this paper decouples capacity and width via the generalization of neural networks to Deep Gaussian Processes (Deep GP), a class of nonparametric hierarchical models that subsume neural nets. In doing so, we aim to understand how width affects (standard) neural networks once they have sufficient capacity for a given modeling task. Our theoretical and empirical results on Deep GP suggest that large width can be detrimental to hierarchical models. Surprisingly, we prove that even nonparametric Deep GP converge to Gaussian processes, effectively becoming shallower without any increase in representational power. The posterior, which corresponds to a mixture of data-adaptable basis functions, becomes less data-dependent with width. Our tail analysis demonstrates that width and depth have opposite effects: depth accentuates a model's non-Gaussianity, while width makes models increasingly Gaussian. We find there is a"sweet spot"that maximizes test performance before the limiting GP behavior prevents adaptability, occurring at width = 1 or width = 2 for nonparametric Deep GP. These results make strong predictions about the same phenomenon in conventional neural networks trained with L2 regularization (analogous to a Gaussian prior on parameters): we show that such neural networks may need up to 500 - 1000 hidden units for sufficient capacity - depending on the dataset - but further width degrades performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Exact marginal prior distributions of finite Bayesian neural networksJacob A. Zavatone-Veth, Cengiz PehlevanNeurIPS 2021 · 被引用 19 次
- A theory of representation learning gives a deep generalisation of kernel methodsAdam X. Yang, Maxime Robeyns, Edward Milsom, Ben Anson 等ICML 2023 · 被引用 15 次
- Variational Inference for Infinitely Deep Neural NetworksAchille Nazaret, David M. BleiICML 2022 · 被引用 13 次
- Convolutional Deep Kernel MachinesEdward Milsom, Ben Anson, Laurence AitchisonICLR 2024 · 被引用 6 次
- On permutation symmetries in Bayesian neural network posteriors: a variational perspectiveSimone Rossi, Ankit Singh, Thomas HannaganNeurIPS 2023 · 被引用 5 次
它引用的顶会 Paper17
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 被引用 458 次
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
相关 Paper
- Critical feature learning in deep neural networksKirsten Fischer, Javed Lindner, David Dahmen, Zohar Ringel 等ICML 2024 · 被引用 15 次
- Characterizing Deep Gaussian Processes via Nonlinear Recurrence SystemsAnh Tong, Jaesik ChoiAAAI 2021 · 被引用 2 次
- Bayesian Deep Ensembles via the Neural Tangent KernelBobby He, Balaji Lakshminarayanan, Yee Whye TehNeurIPS 2020 · 被引用 136 次
- Scale Mixtures of Neural Network Gaussian ProcessesHyungi Lee, Eunggu Yun, Hongseok Yang, Juho LeeICLR 2022 · 被引用 7 次
- A self consistent theory of Gaussian Processes captures feature learning effects in finite CNNsGadi Naveh, Zohar RingelNeurIPS 2021 · 被引用 38 次
