Can Neural Nets Learn the Same Model Twice? Investigating Reproducibility and Double Descent from the Decision Boundary Perspective
Gowthami Somepalli, Liam Fowl, Arpit Bansal, Ping-Yeh Chiang, Yehuda Dar, Richard G. Baraniuk, Micah Goldblum, Tom Goldstein
Abstract
We discuss methods for visualizing neural network decision boundaries and decision regions. We use these visual-izations to investigate issues related to reproducibility and generalization in neural network training. We observe that changes in model architecture (and its associate inductive bias) cause visible changes in decision boundaries, while multiple runs with the same architecture yield results with strong similarities, especially in the case of wide architectures. We also use decision boundary methods to visualize double descent phenomena. We see that decision boundary reproducibility depends strongly on model width. Near the threshold of interpolation, neural network decision bound-aries become fragmented into many small decision regions, and these regions are non-reproducible. Meanwhile, very narrows and very wide networks have high levels of re-producibility in their decision boundaries with relatively few decision regions. We discuss how our observations re-late to the theory of double descent phenomena in convex models. Code is available at https://github.com/somepago/dbViz.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e5c32b8-f6cb-4cc7-8e63-e85463a01063Cited by top-tier papers34
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin et al.NeurIPS 2022 · 222 citations
- Rashomon Capacity: A Metric for Predictive Multiplicity in ClassificationHsiang Hsu, Flávio P. CalmonNeurIPS 2022 · 65 citations
- Latent Space Translation via Semantic AlignmentValentino Maiorca, Luca Moschella, Antonio Norelli, Marco Fumero et al.NeurIPS 2023 · 59 citations
- What Knowledge Gets Distilled in Knowledge Distillation?Utkarsh Ojha, Yuheng Li, Anirudh Sundara Rajan, Yingyu Liang et al.NeurIPS 2023 · 58 citations
- Retrospective Adversarial Replay for Continual LearningLilly Kumari, Shengjie Wang, Tianyi Zhou, Jeff A. BilmesNeurIPS 2022 · 57 citations
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi et al.NeurIPS 2021 · 318 citations
Related papers
- On the geometry of generalization and memorization in deep neural networksCory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui et al.ICLR 2021 · 95 citations
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- Taxonomizing local versus global structure in neural network loss landscapesYaoqing Yang, Liam Hodgkinson, Ryan Theisen, Joe Zou et al.NeurIPS 2021 · 51 citations
- Phenomenology of Double Descent in Finite-Width Neural NetworksSidak Pal Singh, Aurélien Lucchi, Thomas Hofmann, Bernhard SchölkopfICLR 2022 · 12 citations
- Understanding Nonlinear Implicit Bias via Region Counts in Input SpaceJingwei Li, Jing Xu, Zifan Wang, Huishuai Zhang et al.ICML 2025
