Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth
Thao Nguyen, Maithra Raghu, Simon Kornblith
Abstract
A key factor in the success of deep neural networks is the ability to scale models to improve performance by varying the architecture depth and width. This simple property of neural network design has resulted in highly effective architectures for a variety of tasks. Nevertheless, there is limited understanding of effects of depth and width on the learned representations. In this paper, we study this fundamental question. We begin by investigating how varying depth and width affects model hidden representations, finding a characteristic block structure in the hidden representations of larger capacity (wider or deeper) models. We demonstrate that this block structure arises when model capacity is large relative to the size of the training set, and is indicative of the underlying layers preserving and propagating the dominant principal component of their representations. This discovery has important ramifications for features learned by different models, namely, representations outside the block structure are often similar across architectures with varying widths and depths, but the block structure is unique to each model. We analyze the output predictions of different model architectures, finding that even when the overall accuracy is similar, wide and deep models exhibit distinctive error patterns and variations across classes. * Work done as a member of the Google AI Residency program.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c026b3d9-9532-4a2c-8938-4d0d1b988ba2Cited by top-tier papers93
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc et al.ICML 2023 · 260 citations
- Revisiting Model Stitching to Compare Neural RepresentationsYamini Bansal, Preetum Nakkiran, Boaz BarakNeurIPS 2021 · 253 citations
Builds on6
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 654 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
- What shapes feature representations? Exploring datasets, architectures, and trainingKatherine L. Hermann, Andrew K. LampinenNeurIPS 2020 · 186 citations
- Deep neuroethology of a virtual rodentJosh Merel, Diego Aldarondo, Jesse Marshall, Yuval Tassa et al.ICLR 2020 · 77 citations
- Let's Agree to Agree: Neural Networks Share Classification Order on Real DatasetsGuy Hacohen, Leshem Choshen, Daphna WeinshallICML 2020 · 63 citations
Related papers
- Understanding Robust Learning through the Lens of Representation SimilaritiesChristian Cianfarani, Arjun Nitin Bhagoji, Vikash Sehwag, Ben Y. Zhao et al.NeurIPS 2022 · 20 citations
- Feature-Learning Networks Are Consistent Across Widths At Realistic ScalesNikhil Vyas, Alexander B. Atanasov, Blake Bordelon, Depen Morwani et al.NeurIPS 2023 · 47 citations
- Why bigger is not always better: on finite and infinite neural networksLaurence AitchisonICML 2020 · 59 citations
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- Redundant representations help generalization in wide neural networksDiego Doimo, Aldo Glielmo, Sebastian Goldt, Alessandro LaioNeurIPS 2022 · 13 citations
